Karpathy's point about long voice rambles goes further when you add more modalities.
My favorite way to prompt agents lately is a bigger unit I've been calling a task.
A task bundles a long voice note, the current screen, annotations, and exact text into a single turn. The agent reconstructs my intent from all of those signals, so most of the correction loops disappear and I can hand off larger pieces of work.