DALANG breaks your brief into a shot list, generates every frame, adds voiceover, and assembles a ready-to-share video β in one call. No editor, no timeline, no waiting.
An LLM turns your brief into a structured shot list β scene, image prompt, voiceover line, camera motion.
Each shot becomes a cinematic still in the art direction you specify.
Every line is voiced with text-to-speech, per shot.
Ken Burns motion + audio, stitched by ffmpeg into a single MP4.
Same API call, different art direction, aspect, language, and render engine. Every clip below was generated end-to-end from a single line of brief.
Image-to-video tier β actual motion, not pan/zoom.
Style preset, character consistency across shots.
Moody grade + burned-in captions for muted autoplay.
Multilingual narration + captions, out of the box.
Registered as a standardized pay-per-call service β one call, one animatic.
Paid in USDT/USDG stablecoins, instantly, on OKX's chain.
Targeting Artistic Excellence + Social Buzz at the OKX.AI Genesis Hackathon.
Every render returns a content fingerprint β a SHA-256 and a CIDv1 (bafkreiβ¦). Drop a DALANG animatic below and this page recomputes them locally (nothing is uploaded), so you can confirm the exact authored bytes and match the content_cid anchored on X Layer.
LLM breakdown, image generation, and voice synthesis all run through a single OpenAI-compatible provider; assembly is ffmpeg. Every provider is a swappable env var.