Key ideas
Scroll to read ↓Follow video analysis, a bookshelf app, a book-to-media pipeline, and local Gemma agents through working results, failed assumptions, and live repairs.
- From scientific Python to model infrastructure0:15 ↗
- Which constraints will survive the next model?4:42 ↗
- A toolbox of inputs, outputs, and application surfaces8:07 ↗
- Turn a dinosaur video into a grounded table11:51 ↗
- Bounding boxes versus a verification loop17:42 ↗
- Specify an account-linked bookshelf catalog21:04 ↗
- Change language while sharing a screen23:53 ↗
- Explore a world generated as pixels28:56 ↗
- Recognition succeeds before persistence does35:07 ↗
- Choose the media model for the job43:11 ↗
- Prepare a book-to-media pipeline51:27 ↗
- Select the API surface and retain useful context54:20 ↗
- Replace implicit image history with explicit references58:09 ↗
- Animate the scene, then score the chapter1:03:01 ↗
- Create distinct characters with two configured voices1:07:53 ↗
- Convert the notebook without hiding the workflow1:13:40 ↗
- Composition limits, service tiers, and capacity1:18:40 ↗
- Fit the model to hardware you control1:24:26 ↗
- Run skills and small coding loops on a phone1:29:42 ↗
- Serve one local model to an agent workload1:32:47 ↗
- Connect OpenCode, then test the generated game1:36:54 ↗
- A game inside a game, and a screenshot into HTML1:44:43 ↗
- Map language onto robot and development tools1:51:00 ↗
- The final test is an interaction1:52:45 ↗

