API Reference
API Credit Costs
Use this page as the source of truth for Subclip API credit usage.
| API | Cost |
|---|---|
| AI Image Generation | Model credits per output x model-specific output-tier multiplier x output count. Minimum 1 credit; rounded to 2 decimals. |
| AI Video Generation | Model credits per second x billable duration x resolution multiplier x generated-audio multiplier. One output per job; minimum 1 credit; rounded to 2 decimals. |
| Viral Captions | 1 credit per rendered minute + 1 transcription credit + 1 caption analysis credit. Face tracking adds 1 credit minimum, or rendered minutes x 1 credit, whichever is higher. Minimum charge is 2 credits. |
| AI Dubbing | Duration preflight: transcription minutes when no transcript is supplied + dubbing minutes + 5 credits for voice clone/multi-speaker + 3 credits for video background preservation. |
| AI Clipping | Existing AI clipping pipeline credits plus 5 API orchestration credits for import/render orchestration. Viral Captions analysis cost is added when dynamicCaptions is enabled. YouTube transcript fetch adds 1 credit when fetched. |
| Text to Speech | 5 credits per generated audio minute, minimum 1 credit. Voice clone adds 5 credits. |
| Social Media Transcript | 1 credit per source media minute after import and duration probe, rounded up to a whole minute. |
| Video Studio API without AI analysis | 2 credits minimum, or rendered minutes x 1 credit, whichever is higher. Optional BGM/SFX adds 1 credit total. |
| Video Studio API with AI analysis | Base render cost + 2 AI planning credits. Vision analysis adds 0.25 credits per image and 1 credit per video when enabled. |
| Video Studio API voiceover | Adds 1 credit minimum, or rendered minutes x 5 credits, whichever is higher. |
| Enhance Audio API | 1 credit minimum, or duration seconds x 0.10816 credits, whichever is higher. Rounded up to a whole credit. |
Viral Captions API
Workflow: Viral Captions API docs.
Base render is 1 credit per rendered minute. Transcription adds 1 credit. Caption analysis adds 1 credit.
Face tracking adds 1 credit minimum, or rendered minutes x 1 credit, whichever is higher. Minimum charge is 2 credits.
The same formula applies whether captions are generated from ASR or supplied with SRT.
| Job | Credits |
|---|---|
| 1 minute, no face tracking | 3 |
| 1 minute, face tracking | 4 |
| 3 minutes, no face tracking | 5 |
| 3 minutes, face tracking | 8 |
AI Dubbing API
Workflow: AI Dubbing API docs.
Every AI Dubbing API request must include durationSeconds. Subclip checks the estimated credit requirement before creating the job.
The preflight formula is transcription minutes when no transcript is supplied, plus dubbing minutes, plus voice clone and background-preservation add-ons when requested.
| Job | Credits |
|---|---|
| 3 minutes, transcript supplied, catalog voice | 3 |
| 3 minutes, no transcript, catalog voice | 6 |
| 3 minutes, no transcript, clone voice, video background preserved | 14 |
| 3 minutes, multi-speaker | 11 |
AI Clipping API
Workflow: AI Clipping API docs.
AI Clipping API jobs add 5 API orchestration credits for source import, API packaging, and final render orchestration.
When you omit segments, the existing AI clipping pipeline costs still apply. When dynamicCaptions is true, Viral Captions analysis costs also apply.
When youtubeTranscript is true and the YouTube transcript is fetched, the completed job adds 1 credit. If transcript fetch fails and the job falls back to ASR, that extra credit is not counted.
| Job | Credits |
|---|---|
| Selected segments, dynamicCaptions false | +5 API credits |
| AI-selected clips | AI clipping pipeline credits + 5 API credits |
| AI-selected clips with dynamicCaptions true | AI clipping pipeline credits + Viral Captions analysis + 5 API credits |
| YouTube transcript fetched | Any AI Clipping job + 1 credit |
AI Image Generation API
Workflow: AI Image Generation API docs.
Cost is the model's base credits per output x its model-specific output-tier multiplier x output count. The result has a 1-credit minimum and is rounded to 2 decimal places.
Send the exact output tier returned by the selected model in GET /models. GPT Image 2 uses quality tiers (Low, Medium, High); other models may use resolution tiers (1K, 2K, 4K). These values are not aliases and do not imply physical pixel equivalence.
Billing groups are 0.7x for Low; 1x for 1K or Medium; 1.4x for 2K or High; and 2.5x for 4K.
For runtime planning, use the selected model's live estimatedCreditsExamples from GET /models and the estimatedCredits returned when the job starts.
Output count is 1 unless the model supports batch output. Batch models accept 1 to 4 outputs.
| Model | Base credits per output | Allowed outputs |
|---|---|---|
| GPT Image 2 | 6 | 1β4 |
| Nano Banana 2 | 2.5 | 1 |
| Nano Banana Pro | 6 | 1 |
| FLUX.2 Pro | 2.5 | 1 |
| Seedream 5 Pro | 3.5 | 1β4 |
| Ideogram 4 | 5 | 1 |
| Recraft V4.1 | 2 | 1 |
| Qwen Image 2 Pro | 3.75 | 1 |
| Krea 2 Large | 3 | 1 |
| FLUX 1.1 Ultra | 3 | 1 |
| Job | Credits |
|---|---|
| GPT Image 2, Medium, 1 output | 6 |
| GPT Image 2, High, 4 outputs | 33.6 |
| Nano Banana 2, 1K, 1 output | 2.5 |
| Nano Banana 2, 4K, 1 output | 6.25 |
| Seedream 5 Pro, 2K, 4 outputs | 19.6 |
| Recraft V4.1, 1K, 1 output | 2 |
AI Video Generation API
Workflow: AI Video Generation API docs.
Cost is the model's base credits per second x billable duration x the resolution multiplier x the generated-audio multiplier. Video jobs return one output. The result has a 1-credit minimum and is rounded to 2 decimal places.
Resolution multipliers are 1x for 720p, 768p, or source; 1.4x for 1080p or 1440p; and 2.5x for 2160p.
Generated audio adds a 1.2x multiplier when enabled on a supported model. Uploaded audio references do not add this multiplier. Wan 2.7 requires generated audio, so its estimate always includes 1.2x.
Most models use one of their fixed native duration choices below. Gemini Omni Flash is source-video editing only and bills the uploaded source video's matching duration from 3 to 10 seconds. Veo 3.1's three-reference-image mode uses 8 seconds, and Hailuo 2.3 at 1080p uses 6 seconds.
No configured text-to-video or image-to-video model accepts a native 3-second request. For an exact 3.000-second deliverable, trim after download; billing uses the native generated duration before that downstream trim.
For runtime planning, use the selected model's live estimatedCreditsExamples from GET /models and the estimatedCredits returned when the job starts.
| Model | Base credits per second | Billable duration | Generated audio |
|---|---|---|---|
| Gemini Omni Flash | 10 | 3β10s source video | Not generated |
| Kling 3 | 6 | 5s, 10s, 15s | Optional (+20% when enabled) |
| Veo 3.1 | 15 | 4s, 6s, 8s | Optional (+20% when enabled) |
| Runway Gen-4.5 | 6 | 5s, 10s | Not generated |
| Seedance 2 | 12 | 5s, 10s, 15s | Optional (+20% when enabled) |
| Pika 2.2 | 2.5 | 5s, 10s | Not generated |
| Luma Ray 2 | 5 | 5s, 9s | Not generated |
| Hailuo 2.3 | 3 | 6s, 10s | Not generated |
| PixVerse V6 | 6 | 5s, 10s, 15s | Optional (+20% when enabled) |
| Wan 2.7 | 5 | 5s, 10s, 15s | Required (+20%) |
| LTX 2.3 Pro | 4 | 6s, 8s, 10s | Optional (+20% when enabled) |
| Job | Credits |
|---|---|
| Pika 2.2, 5s, 720p, no generated audio | 12.5 |
| Kling 3, 5s, 720p, no generated audio | 30 |
| Kling 3, 5s, 1080p, generated audio | 50.4 |
| Wan 2.7, 5s, 720p, required generated audio | 30 |
| Gemini Omni Flash, 3s source video | 30 |
| Gemini Omni Flash, 10s source video | 100 |
| LTX 2.3 Pro, 6s, 2160p, no generated audio | 60 |
Text to Speech API
Workflow: Text to Speech API docs.
Credits are estimated before the job starts. Text to Speech costs 5 credits per generated audio minute with a 1 credit minimum. Voice clone adds 5 credits.
| Job | Credits |
|---|---|
| 30 seconds generated speech | 3 |
| 1 minute generated speech | 5 |
| 1 minute generated speech with voice clone | 10 |
| 3 minutes generated speech | 15 |
Video Studio API
Workflow: Video Studio API docs.
Without AI analysis or add-ons, cost is 2 credits minimum, or rendered minutes x 1 credit, whichever is higher.
BGM or SFX adds 1 credit total. bgmQuery has no separate search charge.
AI planning adds 2 credits when aiAnalysis is true. Vision analysis adds 0.25 credits per image and 1 credit per video. Audio analysis adds 1 credit per uploaded audio file.
Voiceover adds 1 credit minimum, or rendered minutes x 5 credits, whichever is higher. Stock B-roll search is not exposed in the Video Studio API right now.
| Job | Credits |
|---|---|
| 1 minute render, no AI or add-ons | 2 |
| 3 minute render, no AI or add-ons | 3 |
| 3 minute render with AI planning only | 5 |
| 3 minute render with AI planning, 4 images, 1 video | 7 |
| 3 minute render with BGM or SFX | 4 |
| 3 minute render with voiceover | 18 |
Enhance Audio API
Workflow: Enhance Audio API docs.
API background music separation is disabled. Provider choice does not change the credit formula.
Cost is 1 credit minimum, or duration seconds x 0.10816 credits, whichever is higher. The result is rounded up to the next whole credit.
That works out to about 6.49 credits per minute.
| Job | Credits |
|---|---|
| 1 minute | 7 |
| 3 minutes | 20 |
| 10 minutes | 65 |
Social Media Transcript API
Workflow: Social Media Transcript API docs.
Subclip imports the media first, probes duration, then checks credits before transcription starts.
Cost is 1 credit per source media minute, rounded up to a whole minute.