You can bootstrap your agent quickly with the Omni API using the skill we published:
https://github.com/google-gemini/gemini-skills
It includes:
• video editing • text to video • video generation with image references • first frame to video
But it also has some helper tools for:
• prepping input videos for editing (10s, 720p) • audio stripping if you want to generate new audio • video inspection