one thing i think people dont appreciate enough about @poolsideai is their unusual degree of openness - not only have they shipped an excellent Small model that somehow beat @thinkymachines at coding, but most people (like @eliebakouch) have been shouting out their excellent papers, but also they're among a rare few to actually expose their full eval dataset as well - beautifully published, with across 6 public benchmarks with 4 runs each and hundreds of turns per run. you can satisfy for yourself if they rewardhack. brilliant.
swyx 指出 Poolside 的开放程度被低估,其小模型在编码任务上击败了 Thinky Machines,并公开了完整评测数据集。CEO Eiso Kant 透露,其 Model Factory 每月运行 1 万至 2 万次实验,新模型迭代最快仅需 8 周。
one thing i think people dont appreciate enough about @poolsideai is their unusual degree of openness - not only have they shipped an excellent Small model that somehow beat @thinkymachines at coding, but most people (like @eliebakouch) have been shouting out their excellent papers, but also they're among a rare few to actually expose their full eval dataset as well - beautifully published, with across 6 public benchmarks with 4 runs each and hundreds of turns per run. you can satisfy for yourself if they rewardhack. brilliant.