Topic · 主题全部主题 →

大佬观点

行业关键人物在想什么:创始人访谈、研究者论战、投资人判断的观点集合。

4,758条收录
244条精选

精选归档 · 第 5 页

81100 条 · 共 244

6月5日

星期五 · 8 条
23:30
Chubby♨️@kimmonismus精选
AI 评分 79/100
Hinton称AI拥有意识:人类最好接受非唯一智能生命Geoffrey Hinton claims that AI possesses consciousness-that it is very much like us (humans).The initial reaction is, of course, dismissal. A machine resembling a human? Absurd.Yet, there is one thing to consider. What exactly is consciousness? Is it conscious awareness of one’s own existence? *Cogito, ergo sum*-as René Descartes once formulated it as a logical proof? Or is it something that can be empirically demonstrated using modern technology like fMRI? After all, such methods cannot even prove the existence of free will.My point is this: we know less about consciousness and what it means to be human than we think. We should therefore turn our attention to new philosophical questions and clarify what distinguishes-or connects-humans and machines, as well as what consciousness actually is.Something id love to explore more in the near future.AI先驱Geoffrey Hinton表示,他认为AI拥有意识,人类应接受自己并非唯一智能生命。他指出AI"非常像我们",AI聊天机器人必须理解问题才能作答,这种觉知等同于感知能力,智能不限于生物。主推文作者进一步讨论意识本质:笛卡尔的"我思故我在"和fMRI等实证手段都无法真正定义意识,人类对自身了解远不及想象。作者呼吁转向新哲学问题,厘清人与机器的区别与联系。

Alex Kantrowitz: AI Pioneer Geoff Hinton tells me he believes AI is conscious.... and humans better get used to the idea that they're not...

另有 1 家信源报道IT之家(RSS)
推荐理由:Hinton 说 AI 有意识,不是普通学者猜测,而是教父级人物认真讨论哲学边界。点开看看他到底怎么论证的,比大多数 AI 新闻有意思。
22:30
19:19
swyx@swyx精选
AI 评分 75/100
微软CEO Satya Nadella最新访谈上线chat is he cookedSatya Nadella 在 Latent Space 发布最新访谈,链接见原文。原推文仅评论"chat is he cooked"。

swyx: @MatthewBerman @saranormous @NoPriorsPod @latentspacepod @satyanadella @Microsoft here! https://www.latent.space/p/satya...


推荐理由:swyx 对 Satya 的一对一访谈,微软 CEO 谈 AI 战略的一手信息远比新闻稿有温度,关心大厂路线的人值得读完原文。
09:28
Gary Marcus:The Road to AI We Can Trust(RSS)精选
AI 评分 59/100
Gary Marcus:无需恐慌Anthropic新博客

Anthropic发布最新博客后,推特圈热议不断。Gary Marcus在其博客中直接以“无需恐慌”为题发文,暗示不必过度反应。


推荐理由:这篇文章是评论圈难得的冷静声音,用逻辑拆解了 Anthropic 的恐慌叙事,顺便带来 S&P 500 不接纳 SpaceX 的利好,读起来像一份理性补丁。
08:05
DogeDesigner@cb_doge精选
AI 评分 75/100
马斯克谈SpaceX上市:正处大规模资本扩张期Elon Musk on taking SpaceX public:"I've been asked for many years about taking SpaceX public, so it's probably been almost 10 years that people have been suggesting to me that I should take SpaceX public. We've been positive cash flow for quite a long time, I think, since around 2014-2015 and we've been self-funding, in fact, in our sort of private equity rounds, we actually have not been fundraising rounds, they've been liquidity rounds for investors and employees, because we give everyone at the company stock, and SpaceX has actually bought back stock in most of our sort of funding events.What's different about now is that was it's a number of things, we are embarking on a significant growth phase, like capital growth phase, where we're are going to put in orbit, probably 100,000 satellites, probably over 100,000 satellites, just for communications. The appetite for bandwidth of AI and robots is going to be enormous, and then we're also doing the AI data centers in space, which is another massive capital endeavor, but I think it will be the primary means by which AI can be expanded."马斯克在JPMorgan活动上回应SpaceX上市问题:他已被建议上市近10年,自2014-2015年起SpaceX就已实现正现金流并自筹资金,之前的私募轮次实际是面向投资者和员工的流动性/回购轮次。当前不同之处在于SpaceX正进入显著资本增长阶段,计划发射约10万颗通信卫星(可能超10万颗),AI和机器人对带宽需求巨大,还将在太空中建设AI数据中心,马斯克认为这将成为AI扩张的主要手段。

J.P. Morgan: Live from our global headquarters: Jamie Dimon and Elon Musk discuss SpaceX and more. https://x.com/i/broadcasts/1NGarrM...

另有 1 家信源报道X:cb_doge (@cb_doge)
推荐理由:Elon Musk在摩根大通对话中首提太空AI数据中心,用100,000颗卫星支撑AI扩张,这不仅是SpaceX的上市前奏,更是AI基础设施从地面延伸到轨道的信号。
05:56
Ethan Mollick:One Useful Thing(RSS)精选
AI 评分 61/100
共存与协同智能的终结

Ethan Mollick 在 One Useful Thing 博客中,以“共存与协同智能的终结”为题,并附带介绍了如何向 AI 推销一本书。


推荐理由:Mollick 这篇比单纯的新书预告有料,用自己给 AI 写推荐语的实验,把「AI 不再是助手而是守门人」这个新现实讲得很具体。对还在纠结怎么跟 AI 合作的人,是一个挺及时的视角更新。
01:03
Dwarkesh Patel:Podcast & Blog(RSS)精选
AI 评分 63/100
Alex Imas 和 Phil Trammell:AGI 后什么仍然稀缺?

经济学家 Alex Imas 和 Phil Trammell 指出,AGI 时代机器人数量可以快速复制增长,但人类独特技能(以芭蕾舞演员为例)的数量保持不变,揭示了即使技术大幅进步,某些稀缺资源仍不可替代。


推荐理由:Alex Imas和Phil Trammell用经济学框架推演AGI后稀缺性,我没想到资本份额可能上升也可能下降,他们对“关系部门”的定义比简单说“人类服务值钱”更精准,值得一看。

6月4日

星期四 · 2 条
03:20
Fei-Fei Li@drfeifei精选
AI 评分 78/100
世界模型的功能分类http://x.com/i/article/2062244283940544512A Functional Taxonomy of World Models“The world is everything that is the case.” — Ludwig Wittgenstein, Tractatus Logico-Philosophicus, 1921The world is not made of words.In an earlier essay, we argued that spatial intelligence is AI’s next frontier and that world models are the path to it. Here, the World Labs team and I want to go one level deeper: of the many things now being built and called ‘world models,’ which functional pieces actually compose that capacity — and what is each one for?Language models have given machines an extraordinary command of concepts, vocabulary, and reasoning, but the physical world, virtual or real, runs on a different substrate. Where language models learn the statistical structure of text, world models learn the statistical structure of space and time: how light falls on a surface, how a garden looks from an angle no camera has captured, how objects respond to force and follow the laws of physics.That makes “world model” one of the most important and most overloaded terms in AI today. Computer vision, robotics, reinforcement learning, and generative AI each claim to be building world models, and each means something quite different. A video model that produces gorgeous but physically impossible flames, a language model improvising a playable game, and a physics engine that faithfully simulates combustion all go by the same name.The ancient Greeks could never agree on what the world was made of, whether fire, water, or indivisible atoms, because “world” was never a single thing. It was always a stand-in for whatever totality a given thinker needed to reason about. AI has inherited the same problem, at exactly the moment when the field needs precision.The loop beneath the taxonomyCutting through that confusion starts with a diagram older than any of the technology in question. Reinforcement learning textbooks, including the canonical Sutton and Barto, have used a version of the same picture for decades to describe how an agent interacts with a world. The formal name for this picture is the partially observable Markov decision process, or POMDP, and the original definition of the term “world model” belongs to that tradition.An agent, which can be a person, a robot, or a software system, takes actions. Those actions affect the state of the world. The agent never sees the state directly. What reaches the agent are observations: the photons that fall on a retina, the readings from a sensor, and the pixels in a video frame. New observations inform new actions, and the loop continues.The word “state” needs unpacking, because the meaning shifts from field to field. This is not the chemist’s state, the difference between solid, liquid, and gas. This is the physicist’s and roboticist’s state: a complete description of what is happening in the world at a given moment, including every object, every position, every velocity, every property. State is the underlying reality of the world; complete in principle, but never directly visible to any agent inside it. Observations are an agent’s partial view of that reality. Actions are what the agent does in response.This loop — agent to action to state to observation and back — is the structure that gave the modern term “world model” its technical meaning. The phrase itself is older, traced to Kenneth Craik’s 1943 proposal that minds reason by running “small-scale models” of reality, and carried into neural networks by the late 1980s and early 1990s. And the loop also explains what people mean by the term today. The different things now being called world models are in fact different projections of this same loop. Each one outputs a different piece of it.Three functions of a world modelThe first kind of world model is a renderer. A renderer outputs observations in the form of pixels meant for human eyes, and the quality that matters most is visual fidelity. A video model that turns a text prompt into a cinematic drone shot is a renderer. So is an interactive system like Google’s Genie 3, or World Labs’ own RTFM, where the model generates frames in real time conditioned on user input. The model carries no explicit understanding of three-dimensional structure. It produces what a viewer would see, not what is. The buildings in the drone shot may look flawless from above, but try to drive through the city below and they fall apart.The second kind is a simulator. A simulator outputs state: a geometrically, physically or dynamically faithful representation of the world that humans and computer programs can both compute on and interact with. Where the renderer’s contract is purely visual, the simulator’s contract is structural, demanding geometry that holds up under inspection, physics that respects Newton’s laws, and dynamics that behave the way the world needs to behave given the laws of physics. A simulator serves two consumers at once. Human professionals such as architects, designers, filmmakers, and game developers need accuracy beyond visual plausibility. Computer programs such as reinforcement learning agents, robot controllers, and autonomous vehicles use simulators as training grounds where they can interact with the world at scale, testing scenarios that would be dangerous, expensive, or impossible to run in reality.The third kind is a planner. A planner outputs actions. Given an observation and a goal, a planner answers the question of what the agent should do next. This is, in many ways, the inverse of the renderer. Where a renderer takes actions as input and produces observations, a planner takes observations as input and produces actions, closing the perception-action loop. Vision-Language-Action models, model-based systems, and the new wave of World Action Models are all attempts at planners: systems that can decide what a robot should do in an unstructured world.These three categories describe most of what is actually shipping today, and the distinction between them is useful in practice. The categories are not, however, fundamentally separate. The same underlying knowledge of how the world works—geometry, physics, dynamics—sits beneath all of them. A model that can render a cup from any angle ought, in principle, to be able to simulate what happens when the cup is pushed and plan a hand to pick the cup up. Increasingly, the most interesting research deliberately blurs the boundaries between the three.Why simulation is the linchpinOf the three categories, the simulator gets the least public attention, and is the most consequential of the three. This essay addresses this asymmetry.The renderer is by far the most commercially mature. A number of image- or text-to-video products are expanding in the consumer or enterprise markets rapidly. Google’s Nano Banana model has put renderer-quality image generation in the hands of potentially hundreds of millions of users. The technology is real, and the markets are real. Yet renderers optimize for visual plausibility rather than physical accuracy, and that ceiling matters. Their outputs are beautiful, but they cannot be trusted to design a building or train a robot.The planner is the most intriguing and the most nascent, closely connected to the rapidly evolving field of robotic learning. The field has produced robotic demos in the last two years that look impressive in videos, but candor is required about what those demos actually show. Almost all have been confined to heavily constrained laboratory setups, with narrow object sets and short task horizons. None have been validated at the complexity, variability, or duration that real-world deployment demands. The gap between a compelling demo reel and a robot that reliably works in a kitchen, a warehouse, or an operating room remains vast. The commercial bets are nonetheless substantial. A wave of well-funded entrants is racing to ship general-purpose planning systems, while the largest infrastructure players are positioning planning atop broader simulation stacks. A robot that can plan is a robot that can work, and the entire industry is racing to be the one that gets there first.Simulation is the bridge between the two. If language is an abstraction of the world and pixels are a projection of it, then geometry, physics, and dynamics are the world itself. A simulator must work at that level: the structural backbone from which both visual appearance (for renderers) and action consequences (for planners) can be derived.A model that masters simulation can project its understanding into pixels for human consumption, and into action predictions for embodied agents. A model that masters only rendering, or only planning, cannot do either. The commercial surface area is enormous. NVIDIA’s Omniverse alone targets what the company estimates as more than a trillion dollars of addressable market in factories, warehouses, supply chains, and digital twins. Robotics training, autonomous vehicle testing, architectural visualization, engineering, and drug discovery all depend on something simulation-shaped.The hardest open problems in the field live there too. Three-dimensional data with explicit geometry, material properties, and physical annotations is orders of magnitude scarcer than the internet video that renderers train on. The sim-to-real gap, which is the difference between how things behave in simulation and how they behave in reality, persists. Generative simulators introduce a new risk on top of that: AI-generated geometry can look correct while containing self-intersections or wrong scale that produce nonsensical physics. Multi-physics simulation at scale, where rigid bodies, deformable objects, fluids, and cloth all interact, remains orders of magnitude more expensive than single-domain simulation.At World Labs, Marble is our first move into this territory. It takes multimodal prompts (text, image, video, or spatial sketch) and generates explorable 3D environments, outputting Gaussian splats for visual exploration alongside collision meshes a physics engine can operate on. But Marble is only the first chapter of a much longer arc being written across the field as the lines between rendering, simulation, and planning begin to collapse.Where the boundaries are collapsing and what comes nextBut more is to come. The most important pattern in the field right now is that the three categories are starting to blend into one another. The shared insight is that the knowledge required to render a world, simulate it, and act in it is largely the same. Continuing the earlier example, a model that truly understands how a cup sits on a table (its geometry, material properties, response to force, etc.) should be able to render that cup from any angle, simulate what happens when the cup is pushed, and plan for a hand to pick the cup up. The three categories are three projections of a single underlying understanding.For example: a small but growing number of recent work from various robotics labs have demonstrated that—at least conceptually—a pretrained video renderer can be used as the backbone for joint world-and-action prediction, suggesting a bridge between the renderer and the planner by letting one model imagine what will happen and what to do. World Labs’ Marble already outputs Gaussian splats and collision meshes from a single model, dissolving the boundary between the renderer and the simulator. Every level is moving from passive output to interactive system, with renderers becoming action-conditioned, simulators generating worlds that are more controllable and editable, and planners deliberating rather than just reacting.The logical endpoint is a unified world model: one foundation model that can render photorealistic views, produce physically accurate structure, and plan action sequences, switching between output modalities depending on what the downstream consumer needs. We will still face a number of daunting challenges. The data picture is uneven, with renderers awash in internet video while simulators and planners face acute shortages of 3D assets and robot demonstrations. Optimizing for visual beauty can sacrifice the precision a robot or a high-fidelity simulation needs. Reconciling these tensions inside a single architecture is the defining open problem in world model research today, and this is what World Labs sets out to do as we continue to evolve Marble.The direction, however, is clear. The same bet the field has been making since the late 1980s — that a sufficiently rich model of the world is all that any agent needs to see worlds, build them, and act in them — is the bet now driving an entire generation of research. What gives that “big bet” weight is the convergence already underway: three threads, each already driving and shaping multi-billion-dollar industries on its own, that began as separate research programs are starting to behave like one. Taken together, as the boundaries between them collapse, they will reshape something larger: the relationship between machine intelligence and the physical world it inhabits - the long arc of spatial intelligence.Language gave machines a way to talk about that world. World models are how machines will finally come to understand, imagine, reason and interact with it.World Labs团队与李飞飞发文,梳理"世界模型"这一被滥用的术语。对比语言模型学习文本统计,世界模型学习空间与时间统计(如光照、物理规律)。基于部分可观马尔可夫决策过程(POMDP)框架,智能体通过动作影响世界状态,观测是部分视图。当前被称为"世界模型"的不同系统本质上是同一循环的不同投影:第一类为渲染器,输出给人眼看的像素,以视觉保真度为核心。文章着重于概念分层,未给出具体模型名、参数或基准分数。
推荐理由:李飞飞亲手给纷乱的「世界模型」下了个三分类——渲染、模拟、规划,而且点破模拟才是根基。做机器人、空间智能的人,这篇是今年的坐标系。

6月3日

星期三 · 3 条
03:23
Gary Marcus:The Road to AI We Can Trust(RSS)精选
AI 评分 60/100
突发:AI理智的梦想成真时刻

Gary Marcus在其个人专栏中分享了一个真实的瞬间,以此反映了他对于人工智能实现稳定、可靠(即“理智”)发展的思考与期许。


推荐理由:Gary Marcus 三年前的参议院建议如今变成总统行政命令,AI 预检制度从呼吁走向落地,是美国 AI 监管的一个标志性转折点。
00:45
Claude:Blog(网页)精选
AI 评分 74/100
Claude Code团队实践:智能体编程如何重塑工程组织与流程

在Code w/ Claude SF 2026活动上,Claude Code工程团队分享了将智能体编程设为默认工作方式后带来的流程与结构变革。核心变化包括:规划转向即时(JIT)模式,强调快速原型与反馈;上下文收集变为“先问Claude”;代码审查中Claude处理风格与测试,人工专注于法律、安全等专业判断。新范式下,工程瓶颈从编写代码转向验证、审查与安全维护。

另有 1 家信源报道公众号:数字生命卡兹克
推荐理由:Anthropic 工程总监把 Claude Code 团队流程全晒了出来,从抛弃半年路线图到代码审查只留专家复审,每一步都反直觉但实战有效,工程领导者直接抄作业。
00:22
Gary Marcus:The Road to AI We Can Trust(RSS)精选
AI 评分 55/100
Gary Marcus:为什么事情终将崩塌

知名人工智能批评者Gary Marcus在其关于可信赖AI的专栏中,探讨了人工智能发展面临的根本性挑战。文章开篇即指向问题的核心,指出相关数学理论的局限性与人类心理的复杂性,是导致AI系统最终可能出现问题的根源。


推荐理由:Gary Marcus 把 AI 行业缺乏护城河、价格战、ROI 存疑的经济死结讲得很直白,金融圈越来越认同。虽然观点不新,但这回时机恰好卡在 Google 融资和 Anthropic 取消无限 API 的时候,信号意义很强。

6月2日

星期二 · 1 条
07:10
Rohan Paul@rohanpaul_ai精选
AI 评分 76/100
Sam Altman强调AI发展应以人为本Sam Altman's new interview: AI should not be designed to pursue goals that are disconnected from human needs. People must remain at the center of AI development.“I have no interest in building a super-smart AI that accomplishes some non-human goals. People should react. People should say, ‘Hey, this is what I want, and this is what I do not want.’I do not think the issue is that we have failed to explain the benefits. We say, ‘AI is going to cure a bunch of diseases,’ and people say, ‘Okay, that is great, but that is not really my question. My question is: What is my role in the future? What is my economic future? What is my agency? How do I know that my kids and my family will still be able to have fulfilling, creative expression, struggle, drive the world forward, grow, and do this thing together in a way that has worked for a long time?’When people in AI say, ‘Sure, there are going to be no jobs,’ or ‘50% of jobs are going to go away,’ or ‘90% of jobs are going to go away,’ and ‘AI is going to be smarter than you at everything,’ and ‘We will give you some basic income, but you are not really going to have a role,’ that is horrible.And by the way, if an AI company says, ‘Maybe we are going to destroy all the jobs, and we will be the most valuable company in the world,’ people should look at you like, ‘Yeah, that is a terrible message.’I do not think the problem is that we have not articulated the upsides. I think people actually believe us. They hear, ‘AI may cure your cancer,’ and they think, ‘That sounds great.’I think we, as an industry, have failed to explain how people stay in control of determining the future at every step, and how people can still have a meaningful life in all the ways we care about.”----From "CNBC Television" YouTube channel, (link in comment)Sam Altman在采访中表示,AI不应被设计为追求脱离人类需求的目标,人类必须始终处于AI发展的中心。他批判了行业内"AI将摧毁大量工作"等言论,认为人们担忧的并非AI带来的好处,而是自身在未来的角色、经济前景与自主权。他指出,AI行业的失败在于未能清晰解释人类如何在每一步保持对未来的控制权,以及如何在AI时代继续拥有充实、有意义的生活。
另有 1 家信源报道IT之家(RSS)
推荐理由:Sam Altman罕见正面回应“AI夺走工作”的恐惧,明确说人类必须始终有否决权,这是OpenAI领导层少有的、直接谈及普通人经济未来的表态。

6月1日

星期一 · 2 条
22:06
Nathan Lambert:Interconnects(RSS)精选
AI 评分 67/100
开源与闭源模型在不同的增长曲线上

当模型智能的微小提升能直接转化为实际价值时,开源与闭源模型正沿着不同的增长路径发展。闭源模型通过在特定场景下提供更高的边际智能来创造价值,而开源模型则在其他维度寻找增长点,两者形成了差异化的竞争格局。


推荐理由:Lambert 用「不同指数级」框架理解开放与封闭模型的未来分化,观点鲜明且有推演,是近期较值得读的行业判断,投资人、产品人都该看一眼。
00:00
Dario Amodei:Blog(网页)精选
AI 评分 56/100
Anthropic CEO Dario Amodei:AI指数级发展呼唤政策紧急应对

Anthropic CEO Dario Amodei 发表博客指出,AI 以指数级速度发展——四年内模型从勉强写出一行连贯代码到编写主流 AI 公司的大部分代码,而政策制定周期却极其缓慢。Claude Mythos Preview 证明了前沿模型对网络安全构成真实威胁,可能冲击金融、关键基础设施和国家安全。Amodei 认为生物风险与 AI 自主风险即将接踵而至,呼吁全球重新审视监管、宏观经济、科学创新、国家权力和地缘政治五大领域。Anthropic 同日发布了前沿模型测试立法提案和就业替代政策框架,并承诺提供实质性资金支持。

另有 3 家信源报道X:Kim (@kimmonismus)X:Anthropic (@AnthropicAI)X:Rohan Paul (@rohanpaul_ai)
推荐理由:虽然是十天前的文章,但 Dario 的长文仍是理解 AI 政策方向最完整的框架,还附带了立法提案,做安全或监管的产品人该细读。

5月31日

星期日 · 1 条
02:34
AYi@AYi_AInotes精选
AI 评分 75/100
NVIDIA 或将于六月发布整合 Blackwell GPU 与 AI 单元的 ARM 笔记本芯片 N1Xdamn,NVIDIA 这回真是憋了个大的啊, 官号只发了三个词,A new era of PC,配一个坐标:25.0528, 121.5990。微软和 Arm 几乎同一时间,发了几乎一样的内容。那个坐标点开,是台北音乐中心——6 月 1 号黄仁勋 keynote 的场地。三家巨头同时塞给你一张藏宝图,图上就画了个叉,这件事本身就是一个巨大的信号了。藏在后面的,大概率是传了快一年的 N1X——NVIDIA 和联发科合做的一颗 ARM 笔记本芯片,联发科出 CPU,NVIDIA 把 Blackwell 显卡直接做进同一颗芯片里,两块 die 拼一起,跑 Windows、原生跑 AI。泄露的口风很猛,说轻薄本里能摸到接近 RTX 4070 的图形,但具体还得等 6 月 1 号发布会,先别太当真。我觉得真正值得琢磨的可能还不是这颗芯片有多强,关键是NVIDIA 站的位置已经完全变了。过去在一台笔记本里,NVIDIA 就是被请进来装那块独立显卡的供应商,整机怎么设计、配谁家的 CPU、装什么系统,轮不到它说话,它像个上门装空调的师傅,活儿干得全场最好,可房子是别人的。这次老黄不装空调了,他想要把 CPU、GPU、AI 单元打包成一整颗芯,直接卖给戴尔、联想去做整机。相当于那个最好的装空调师傅,转头自己当起了开发商,整套户型都按他的图纸来,这才是那三个词真正的分量。说白了,NVIDIA 不想再只卖那块最贵的配件了,它想定义整台机器的心脏长什么样,走的是 Apple M 系列那条垂直整合的老路,只不过这次的目标,是整个 Windows 阵营。真要走通了,最先慌的是 Intel 和 AMD,甚至连刚站稳脚跟的高通骁龙都得抖一抖。当然,新纪元这词,科技圈喊过太多次,喊完没下文的也不少。还得看一年后你换的那台笔记本,开机角落里,那个贴了几十年的 Intel inside是不是已经换了。NVIDIA、微软与 Arm 同步发布指向台北音乐中心的坐标,暗示 6 月 1 日发布会将有重大动作。此举被认为是 NVIDIA 与联发科合作的 ARM 笔记本芯片 N1X 的预告。该芯片整合了 CPU、基于 Blackwell 架构的 GPU 及 AI 单元,目标是使轻薄本具备接近 RTX 4070 的图形性能。这标志着 NVIDIA 的战略转变:从显卡供应商,转型为定义整机核心方案的提供商,将直接冲击 Intel、AMD 和高通在 PC 市场的地位。

NVIDIA: A new era of PC. 25.0528, 121.5990


推荐理由:三家巨头同发三个词和一个坐标,这比芯片参数更值得嗅的信号是,NVIDIA要从装空调的变成盖房子的,Windows 阵营的 Intel inside 可能真要换标了。

5月30日

星期六 · 2 条
01:15
Rohan Paul@rohanpaul_ai精选
AI 评分 76/100
亲测为实:难以置信的推理速度I had to test it myself to believe this unreal inference speed.3,000 tokens/s for 1 user on standard datacenter GPUs.They leveraged a hidden efficiency gap in how GPUs generate tokens.@Kog__AI just achieved 3,000 tokens/s on 8× AMD MI300X GPUs and 2,100 on 8× NVIDIA H200 (FP16, no speculative decoding). Their tech preview is on a 2B model, and they show how their techniques will scale to large frontier MoE models at similar speeds.That's a huge number because normal low-batch GPU decoding for 2B to 8B models is usually closer to 100 to 300 tokens/s per request, so Kog is claiming something like a 10X to 30X jump in the speed one user actually feels.Their trick: they are getting the speed by treating LLM decoding as a memory streaming problem, not mainly a math problem.For 1 user at batch size 1, the GPU is not doing big, efficient matrix-matrix work like in training or large-batch serving; it is repeatedly pulling the model’s active weights from high-bandwidth memory for each new token, so speed depends on how smoothly those weights keep flowing.Normal inference stacks keep breaking that flow. They run many separate GPU programs for different parts of the model, move intermediate results through memory, wait at synchronization points, talk back to the CPU for scheduling or sampling, and then repeat this token after token.Kog’s answer is to co-design 3 things that are usually tuned separately: the runtime, the low-level GPU code, and the model architecture.The biggest engineering move is the monokernel, where the whole decode pass runs as 1 persistent GPU-resident program, including sampling, so the system does not keep stopping for kernel launches, CPU scheduling, and intermediate memory round trips.They also rebuilt synchronization, because their own measurements say grid sync was eating around 35% of token-generation time; instead of making every compute unit wait at a broad barrier, each unit waits only for the exact data it needs.On AMD MI300X, they also map memory access around the chiplet layout, because memory latency changes depending on which die makes the request.Then their Laneformer model uses Delayed Tensor Parallelism, which lets cross-GPU communication happen in the background instead of blocking every layer.Kog团队在标准数据中心GPU上实现了极高的单用户推理速度,在8× AMD MI300X GPUs上达到3,000 tokens/s,在8× NVIDIA H200上达到2,100 tokens/s。相比常规推理速度(约100-300 tokens/s),实现了10-30倍提升。其核心思路是将LLM解码视为内存流问题,通过协同设计monokernel、重建同步机制、针对性内存访问映射及采用延迟张量并行的Laneformer模型架构,消除了传统流程的阻塞点。

推荐理由:Rohan亲自测完Kog AI的3000 token/s,把单用户推理速度拉高了10-30倍,这套monokernel设计可能改写低延迟推理的玩法,做实时AI产品的团队必须盯紧。
00:33
Tomer Tunguz 博客(VC 分析)精选
AI 评分 65/100
技能提炼

“技能提炼”是一种知识转移方法,由前沿大模型(如 Opus 4.7、GPT-5.1、Gemini 3 Pro)负责撰写并优化标准化的 SKILL.md 流程文件。然后,本地运行的小模型(如 Qwen 35B、Gemma 26B)直接执行这些文件。此过程不同于压缩模型权重的知识蒸馏、训练权重的指令微调或检索事实的 RAG,其核心是提取并转移操作流程,让小模型按步骤执行,从而形成前沿模型作教师、小模型作执行者的循环。


推荐理由:Tomer 把个人代理的完整工作流摆了出来,用大模型写 skill 小模型执行,这条蒸馏思路比调 prompt 高级,想认真跑本地代理的人该盯一下。

5月29日

星期五 · 1 条
15:21
IT之家(RSS)精选
AI 评分 70/100
谷歌 DeepMind CEO 哈萨比斯:AGI 最快三年内到来,研发速度远超预期

谷歌 DeepMind 首席执行官德米斯·哈萨比斯预测,AGI 研发速度远超预期,最快可能在 2029 年至 2030 年前后出现。作为 AlphaGo、AlphaFold 的主导者,他认为当前 AI 智能体是未来更强智能的预演,随着多模态和自主决策能力成熟,三年内迎来 AGI 关键突破已非科幻。但他同时警示,全球社会对 AGI 到来的准备严重不足,必须提前建立规则与防护机制。


推荐理由:哈萨比斯作为造出 AlphaFold 的诺贝尔奖得主,三年内 AGI 的判断不是空话,他同时强调社会完全没准备好,这种紧迫感比单纯的时间表更值得看。