A Vietnamese article in. A 9:16 short out.
One command · zero editing · deterministic renders.
🌐 English · Tiếng Việt
Quick Start · How It Works · Usage · TemplatesThe split that makes it reliable: AI handles content (the script + template choices), deterministic code handles production (the pixels). The same
script.jsonalways renders the same video — no surprises, no manual editing.
You supply the text. The templates own all the design, layout, and motion. The pipeline does TTS, sound design, rendering, and the final mux — and hands you three files ready for CapCut / TikTok / Shorts / Reels:
| File | What it's for |
|---|---|
video.mp4 |
Final 9:16 video with voice + SFX baked in |
voice.mp3 |
Narration track — drop into CapCut |
script.txt |
Plain text — CapCut auto-caption |
Vibe Coding Thực Chiến với Claude Code: Từ Zero đến Hero
Senior AI Engineer @ AI Coding
Setup · Permission Modes · Memory · Hooks · Skills · MCP Servers · Subagents · GitHub
Từ zero đến hero — đúng cách build agent & tự động hoá như repo này.
Làm chủ Claude Cowork - Tự động hoá các công việc hàng ngày
Senior AI Engineer @ AI Coding
Setup · Tổng hợp email · Lên lịch chạy tự động · Xử lý file · Tạo report · Tạo interactive dashboard · Claude in chrome ·
📺 Detailed guide: Watch the video walkthrough on YouTube
git clone https://github.com/huytranvan2010/AI-auto-generate-video.git
cd AI-auto-generate-video
npm install
# start your local OmniVoice server, then generate video|
With Claude Code — recommended Claude fetches the article, writes |
Manual — bring your own npm run pipeline -- output/my-video/script.jsonFull control over every scene and template. |
A few minutes later → output/<slug>/video.mp4 (1080×1920).
Note: The .agent directory has been added, so all these instructions work not only with Claude Code but can also work with other Coding Assistants, since everything here is a skill.
flowchart LR
A["📰 URL / .txt"] -->|/create-template-video| B[Claude Code]
B -->|fetch + write text| C["script.json<br/>renderer: hyperframes"]
C -->|Zod validate| D[Template Pipeline]
D -->|TTS per scene| E[OmniVoice]
E -->|concat + SFX mix| F[voice.mp3]
D -->|render each template| G["HyperFrames<br/>Chromium"]
G -->|fit clip to narration| H["clips/scene-*.mp4"]
F --> I[mux audio]
H --> I
I -->|🎬| J["video.mp4<br/>1080×1920"]
style A fill:#0f172a,color:#fff,stroke:#334155
style B fill:#6366f1,color:#fff,stroke:#6366f1
style E fill:#f59e0b,color:#fff,stroke:#f59e0b
style G fill:#ec4899,color:#fff,stroke:#ec4899
style J fill:#10b981,color:#fff,stroke:#10b981
Eight deterministic steps in src/render/template-pipeline.ts:
| # | Step | Output |
|---|---|---|
| 1 | Validate | script.json checked against the Zod schema |
| 2 | Caption text | script.txt — all voiceText joined (CapCut auto-caption) |
| 3 | TTS / scene | voice/scene-<id>.mp3 via OmniVoice (idempotent) |
| 4 | Concat voice | voice-raw.mp3 with 0.3s gaps + per-scene start times |
| 5 | SFX mix | voice.mp3 — sound effects layered onto the narration |
| 6 | Render clips | clips/scene-<id>-fit.mp4 — template → MP4, fit to narration |
| 7 | Concat + mux | video-silent.mp4 → video.mp4 (voice muxed in) |
| 8 | Done | prints result paths + total duration |
Prerequisites
| Item | Need | Notes |
|---|---|---|
| Node.js | ≥ 22 | node --version |
| FFmpeg + ffprobe | any modern | must be in PATH (ffmpeg -version) |
| Chrome / Chromium | any | used by HyperFrames to render each template |
| OmniVoice server | running | local TTS at OMNIVOICE_ENDPOINT (default http://127.0.0.1:8123) |
| Claude Code CLI | optional | only for the /create-template-video skill |
Install FFmpeg:
- Windows —
winget install Gyan.FFmpeg - macOS —
brew install ffmpeg - Linux —
sudo apt install ffmpeg
Configuration — .env.local
OmniVoice is the only TTS provider, and it's local — no API keys.
TTS_PROVIDER=omnivoice
OMNIVOICE_ENDPOINT=http://127.0.0.1:8123The server must accept POST /tts with { text } and return audio/mpeg bytes.
Inside Claude Code (recommended) — pass a URL or a local .txt:
/create-template-video https://aicodingvn.vercel.app/iphone-17-200mp
/create-template-video news/my-article.txt
The skill reads the content, writes script.json, and runs the pipeline. Authoring rules
(template mapping + Vietnamese TTS number handling) live in the
skill spec.
Or run the pipeline directly on an existing script.json:
npm run pipeline -- output/<slug>/script.json📄 script.json shape (template mode)
{
"version": "1.0",
"renderer": "hyperframes",
"aspect": "9:16",
"metadata": {
"title": "Apple ra mắt iPhone 17 camera 200MP",
"source": {
"url": "https://...",
"domain": "aicodingvn.vercel.app",
"image": null
},
"channel": "AI Coding"
},
"voice": { "provider": "omnivoice", "speed": 1.0 },
"scenes": [
{
"id": "hook",
"type": "hook",
"voiceText": "Apple vừa ra mắt iPhone mười bảy với camera hai trăm megapixel.",
"templateId": "frame-liquid-bg-hero",
"inputs": {
"kicker": "🔥 Tin nóng",
"headline": "iPhone 17",
"subheadline": "Camera 200MP",
"cta": "Theo dõi ngay",
"brand": "AI Coding"
}
},
{
"id": "body-1",
"type": "body",
"voiceText": "Cảm biến mới thu nhiều ánh sáng hơn, ảnh đêm sắc nét hơn rõ rệt.",
"templateId": "frame-pentagram-stat",
"inputs": {
"label": "Camera",
"headline": "200MP",
"subtitle": "Cảm biến lớn nhất từ trước tới nay",
"anchor": "200"
}
},
{
"id": "outro",
"type": "outro",
"voiceText": "Theo dõi AI Coding để xem bản tin công nghệ mới mỗi ngày.",
"templateId": "frame-logo-outro",
"inputs": {
"brand_name": "AI Coding",
"tagline": "Tin công nghệ mỗi ngày",
"primary_url": "https://aicodingvn.vercel.app/"
}
}
]
}Schema rules: 3–12 scenes · scenes[0].type === "hook" · last scene type === "outro" ·
every templateId must exist under templates/.
📁 Output structure
output/<slug>-<timestamp>/
├── script.json # input (skill-generated or hand-written)
├── script.txt # all voiceText joined — CapCut auto-caption
├── voice/
│ ├── scene-hook.mp3 # TTS per scene (idempotent)
│ └── scene-*.mp3
├── voice-raw.mp3 # concatenated voices, no SFX (intermediate)
├── voice.mp3 # final audio with SFX mixed in
├── clips/
│ ├── scene-hook.mp4 # rendered template clip (idempotent)
│ └── scene-hook-fit.mp4 # fitted to the scene's narration length
├── video-silent.mp4 # concatenated clips, no audio (intermediate)
└── video.mp4 # 🎉 final — 1080×1920 + voice + SFX
Idempotent. Delete
voice/scene-<id>.mp3to force re-TTS, orclips/scene-<id>.mp4to re-render just that scene, then re-run the pipeline.
Every visual is a self-contained HyperFrames project under templates/ — index.html (16:9)
and compositions/portrait.html (9:16). You fill the text inputs; the template owns the design.
Full slot reference: templates/CATALOG.md.
| Template | Role | Best for |
|---|---|---|
frame-liquid-bg-hero |
hook | Opening hook — aurora hero with headline + CTA pill |
frame-vignelli |
body | A single striking stat — dark charcoal + red accent |
frame-pentagram-stat |
body | A hero number / benchmark — dark neon + bar chart |
frame-bold-poster |
body | A punchy multi-line statement + giant figure |
frame-build-minimal |
body | One bold word revealed letter-by-letter — dark/amber |
frame-creative-voltage |
body | A creative slogan — electric-blue split + handwriting |
frame-glitch-title |
body | Breaking / tech news — cyberpunk RGB-split glitch |
frame-aicoding-list |
body | A list of 2–5 items (icon + level tag) |
frame-aicoding-comparison |
body | A head-to-head comparison of two things |
frame-logo-outro |
outro | Default brand end-card — logo glow + name + tagline + URL |
frame-statement-outro |
outro | Alternative outro — red statement card on paper |
Add your own: drop
templates/<id>/withindex.html,compositions/portrait.html,hyperframes.json,meta.json(+NOTICE.mdif vendored), then add a row toCATALOG.md. Use a Vietnamese-capable font stack.
SFX live in assets/sfx/<category>/<name>.mp3. Per scene, the picker
(src/assets/sfx-selector.ts) resolves in three tiers:
1. scene.sfx override → exact file, or { "name": "none" } to mute
2. semantic match → voiceText keywords (cảnh báo→alert, kỷ lục→success, ra mắt→reveal …)
3. scene-type default → hook→hook · body→callout · outro→outro
Within a category the file is chosen deterministically by hashing the scene id — same script gives the same SFX, different scenes get different files. The library is large and not committed:
npm run sfx:download # fetch the SFX library
npm run sfx:filter # prune / filter itNo assets/sfx/? The pipeline just renders without SFX.
| Layer | Technology |
|---|---|
| Runtime | Node ≥22 · TypeScript 6 · ESM · tsx |
| Render | HyperFrames 0.6.94 (HTML→MP4 via Chromium) |
| TTS | OmniVoice (local) |
| Schema | Zod ^4 |
| HTTP | axios + nock |
| Concurrency | p-limit |
| A/V | FFmpeg + ffprobe |
| Tests | Vitest ^4 |
| Orchestration | Claude Code skill |
- HyperFrames — the HTML-to-video engine behind the templates
- OmniVoice — local Vietnamese text-to-speech
- html-video — HTML-to-video approach this project builds on
- Auto-Create-Video — the original project this is based on
If this project saved you time, please consider:
- ⭐ Star this repo — it really helps with discoverability
- 🎓 Check out AI Coding's courses on Udemy
- 📱 Follow AI Coding on Facebook · TikTok · YouTube
- 💬 Tell a friend who creates content
- 🐛 Report bugs or request features
