We need open frameworks to evaluate model behavior. Discussions need to be grounded in auditable measurements rather than "us vs them" vibes.
@cyrilgorlla and the team at CTGT are doing important work in this space
François Chollet呼吁建立开源框架评估模型行为,主张以可审计的测量取代“我们vs他们”的立场之争。其引用推文指出,在8k token预算下,CTGT的120B模型在FinanceReasoning上得分83.61%,超过Kimi K3(81.93%)和Inkling(65.13%),且单次查询成本低62至160倍。
We need open frameworks to evaluate model behavior. Discussions need to be grounded in auditable measurements rather than "us vs them" vibes.
@cyrilgorlla and the team at CTGT are doing important work in this space
François Chollet呼吁建立开源框架评估模型行为,主张以可审计的测量取代“我们vs他们”的立场之争。其引用推文指出,在8k token预算下,CTGT的120B模型在FinanceReasoning上得分83.61%,超过Kimi K3(81.93%)和Inkling(65.13%),且单次查询成本低62至160倍。
We need open frameworks to evaluate model behavior. Discussions need to be grounded in auditable measurements rather than "us vs them" vibes.
@cyrilgorlla and the team at CTGT are doing important work in this space