Storytime Archives
How We Built a Full Cinematic YouTube Short Every 2.5 Hours
Powered by Claude + Higgsfield MCP + Kling 3.0
What used to take a traditional video production team days — research, scripting, storyboarding, shooting, editing — we compressed into a single automated pipeline that produces a complete, cinematic, narrated short film in under 2.5 hours from a blank page.
Here's exactly how it works.
The Pipeline
Step 1 — Research & Script (Claude)
Claude researches the historical story, fact-checks it, and writes a tight Mr. Ballen-style narrative script — 350 to 420 spoken words, six dramatic beats, a cold open that builds tension before the reveal. The whole script is designed to run exactly 2.5 to 3 minutes. No filler, no padding, every sentence earning its place.
Step 2 — Voiceover (ElevenLabs)
The script goes straight into ElevenLabs for narration. The director gets back a finished audio file with a timestamped transcript — every line of narration mapped to an exact second.
Step 3 — Shot List (Claude)
Claude takes the timestamped transcript and builds a complete shot list against it — every second of narration gets footage coverage. No gaps, no reshoots. 30 to 32 shots per episode, each one with a specific duration, a camera direction, a dramatic purpose, and a fully written cinematic image prompt ready to generate.
Step 4 — Image Generation (Higgsfield MCP → Nano Banana Pro)
Here's where the automation lives. Higgsfield is connected directly to Claude via MCP — no copy-pasting, no switching apps, no tab juggling. Claude fires image generations in batches of 4 directly through the Higgsfield connector. The director reviews a batch, approves it, and the next batch is already firing. Each image costs 2 credits and takes roughly 30 seconds to render. Nano Banana Pro delivers gritty, cinematic, photoreal stills that look like frames pulled from a prestige TV production.
Step 5 — Video Generation (Higgsfield MCP → Kling 3.0)
The moment a batch of images is approved, Claude simultaneously fires the matching videos AND the next batch of images in the same turn — 8 generations running in parallel. Kling 3.0 takes each confirmed image and animates it into a 3 to 6 second cinematic shot. Camera moves, actor performances, dust in the air, fire reflected in glass, crowds moving through tunnel entrances — all of it rendered at a quality level that would have required a full crew and a location shoot 5 years ago.
Step 6 — Assembly (CapCut)
The director takes the approved image and video files into CapCut, drops them against the ElevenLabs VO track, adds auto-captions, keyword highlights, a music bed, and the branded end card. Done. Picture locked. Ready to publish.
The Numbers
- Total pipeline time: ~2.5 hours, blank page to picture lock
- Shots per episode: 30–32
- Images generated: ~32 (Nano Banana Pro, 2cr each)
- Videos generated: ~35–40 (Kling 3.0, 6–12cr each)
- Credits per episode: ~500–520 credits
- Final runtime: 2.5–3 minutes
- Human decisions required: Script approval, image batch approvals, CapCut assembly
Why Kling 3.0 Specifically
Kling 3.0 is the reason this pipeline works at the quality level it does.
Every shot in a Storytime Archives episode starts as a still image — a precisely directed, fully prompted cinematic frame. What Kling 3.0 does is take that frame and breathe life into it with a level of coherence and cinematic weight that earlier video models couldn't deliver. Dust clouds billow correctly. Crowds move with natural energy. A man wiping silica dust from his eyes does exactly that — he doesn't suddenly turn toward the camera or dissolve into something unrecognizable.
The key technical decisions that made Kling 3.0 reliable at scale:
- Every call uses a confirmed start image — we never generate video blind
- Shot durations are matched to content (3s inserts, 6s hero shots)
- A mandatory declined preset ID is passed on every call to prevent the model from overriding our style with its own preset aesthetic
- Video prompts are written in a precise SHOT FLOW format — 0:00 to 0:03, then 0:03 to 0:06 — so the model knows exactly what happens when
The result is a consistent cinematic language across 30+ shots — the same gritty, raw, 35mm aesthetic, the same low-key lighting, the same motivated performances — that holds from the first frame of the episode to the last.
Why the Higgsfield MCP Connection Is the Key
Before MCP, AI video production looked like this: prompt in one window, generate in another tab, download the file, re-upload it somewhere else, repeat 60 times. Each handoff between tools costs time and breaks creative flow.
With Higgsfield connected directly to Claude via MCP, the entire generation layer lives inside the conversation. Claude writes the prompt, fires the generation, receives the job ID, and uses that same ID as the start image for the video — all without the director touching a single file. The director's only job is to look at the result and say "good" or "redo it."
That's what compresses 2+ days of traditional production into 2.5 hours. Not just faster generation — the elimination of all the dead time between steps.
The Creative Result
Storytime Archives produces gritty cinematic reenactment documentaries about historical injustices — the companies that killed their workers and covered it up, the disasters that never made it into the history books, the people erased by money and power. The content is serious, the production is cinematic, and the pipeline runs on roughly $5 worth of AI credits per episode.
Three episodes completed. Each one better than the last. The workflow is locked, the quality bar is set, and the pipeline scales.
Built with Claude (Anthropic) + Higgsfield AI (MCP) + Kling 3.0 + ElevenLabs + CapCut.
Storytime Archives — storytimearchives.com








