xAI发布Grok 4.3,其在Artificial Analysis智能指数得分53,性能优于Grok 4.20、Muse Spark等模型。核心改进在于“性价比”:输入与输出价格较前代分别降低约40%和60%,且基准测试套件运行成本下降。该版本在GDPval-AA等现实智能体任务上表现显著提升,指令遵循与客服任务强劲。但推文指出,其表现仍落后于最新的中国开源模型,并批评GDPval-AA测试本身价值有限。
The new Grok comes in below the latest Chinese open weights models, Grok 4 was at the frontier when released.
(&; Artificial Analysis: please stop using GDPval-AA which is not a useful test of anything except a model's ability to impress Gemini as a judge)