Robotics is more and more getting close to get its version of in-context learning.
Skild just released S1, a robot foundation model that uses video demonstrations to define tasks instead of language instructions.
Give S1 1 human video showing a long, multi-step task, and the robot executes it straight away. No retraining. No fine-tuning.
If this scales, you pay the enormous data bill once during foundation-model training, then amortize it across thousands of new tasks through prompting. That could change the economics of robot learning completely.