This is a neat feature.
I wrote an article a few weeks back about how I built this into my agent orchestrator: https://x.com/omarsar0/status/2073404610501329247?s=20
But I made it multimodal from the ground up. Text, screenshots, audio, video, and annotations can all be wired up as a reusable skill. Magical!