WizCut API
Programmatically upload, process, and render multicam podcast edits with the WizCut API.
The WizCut API lets you automate multicam podcast editing. Upload your camera angles, let WizCut detect speakers, work out who’s on which camera and generate cuts, optionally review them in a hosted editor, and get a rendered video back via webhook.
Authentication
Create an API key at wizcut.com/settings. Include it in every request:
Authorization: Bearer wc_live_your_key_here
Keys can be revoked at any time from the settings page.
Create a job
POST /api/jobs
{
"sources": [
{ "label": "Camera 1", "kind": "video", "ext": "mp4", "fileSize": 2147483648 },
{ "label": "Camera 2", "kind": "video", "ext": "mp4", "fileSize": 1932735283 },
{ "label": "Mic mix", "kind": "audio", "ext": "wav" }
],
"callbackUrl": "https://your-server.com/webhook",
"review": true,
"autoMap": "confident",
"silence": { "mode": "tighten" }
}
Fields:
| Field | Type | Description |
|---|---|---|
sources | array | Camera angles or audio files, up to 12 per job. A job needs at least two video sources before it can be processed. Each has label, optional kind ("video" or "audio", default "video"), optional ext (mp4, mov, webm, mkv, wav, mp3, m4a or aac; default "mp4" for video, "wav" for audio) and optional fileSize in bytes. |
callbackUrl | string | Optional. URL to receive webhook notifications — see Webhooks. |
review | boolean | Default true. When true, the job pauses at “ready” for human review of cuts before rendering. When false, rendering starts automatically as soon as the speaker mapping is submitted — by a person, your code, or WizCut itself. |
autoMap | string | Default "confident". Whether WizCut may decide on its own who’s on which camera: "off", "confident" or "always" — see Automatic camera mapping. |
silence | object | Optional. Automatic pause removal — see Silence removal. Omitted means off. |
You can’t pass the speaker mapping (tracks) here: source IDs only exist once the job is created, so there’s nothing to point a speaker at yet. A tracks list that names speakers is rejected with 400 and "code": "TRACKS_ON_CREATE". Use autoMap, or map speakers via the API once speaker detection is done.
Pass fileSize for every file you know the size of. Files over 100 MB then get a multipart upload, which lets you upload parts in parallel, retry a single part, and is the only way to upload files larger than 5 GB. Without fileSize, every file gets a single presigned PUT.
Response:
{
"jobId": "uuid",
"uploads": {
"source-uuid-1": {
"method": "multipart",
"uploadId": "...",
"key": "sources/uuid/source-uuid-1.mp4",
"partSize": 52428800,
"partCount": 41,
"urls": ["https://presigned-part-1-url...", "https://presigned-part-2-url...", "..."]
},
"source-uuid-3": {
"method": "put",
"url": "https://presigned-upload-url..."
}
},
"audioUploads": { "...": "..." }
}
uploads is keyed by the source IDs WizCut assigned, in the same order as your sources array. You can ignore audioUploads — it’s used by the WizCut apps, which upload an extracted audio track ahead of the video.
Upload files
No auth header is needed for the presigned URLs — they’re self-authenticating.
"method": "put" — PUT the whole file to url within an hour:
curl -X PUT -T mic-mix.wav "https://presigned-upload-url..."
"method": "multipart" — split the file into partSize-byte chunks (the last one is shorter) and PUT chunk n to the URL for part n:
-
urlsholds the URLs for the first parts (urls[0]is part 1). Request more as you go — up to 16 per request:POST /api/jobs/{jobId}/sources/{sourceId}/parts{ "uploadId": "...", "partNumbers": [5, 6, 7, 8] }The response maps part numbers to URLs:
{ "urls": { "5": "https://...", "6": "https://..." } }. Part URLs are valid for two hours, so fetch them shortly before you use them. -
Once every part is uploaded, finish the upload:
POST /api/uploads/complete{ "key": "sources/uuid/source-uuid-1.mp4", "uploadId": "..." }
To give up on an upload, send the same body to POST /api/uploads/abort.
Start processing
POST /api/jobs/{jobId}/process
Call this once every file is uploaded. It kicks off the pipeline: audio sync, speaker detection, proxy rendering, camera mapping and cut generation. The request returns immediately — processing happens asynchronously.
Optional body:
{ "diarizeSourceIds": ["source-uuid-3"] }
diarizeSourceIds picks which sources WizCut listens to for speaker detection. If you leave it out, WizCut uses your audio sources, or the first source if there are none. Pick the sources with the cleanest audio of everyone talking — a camera with a muted or distant mic leads to a job with no cuts.
Response:
{ "jobId": "uuid", "status": "syncing" }
List jobs
GET /api/jobs
Returns { "jobs": [...] } with your 50 most recent jobs, newest first. Each has id, title, status, created_at and updated_at.
Get job status
GET /api/jobs/{jobId}
Response:
{
"job": {
"id": "uuid",
"status": "complete",
"sources": [...],
"tracks": [...],
"turns": [...],
"cuts": [...],
"removed_ranges": [...],
"camera_map": {...},
"output_url": "https://presigned-download-url...",
"error_message": null
},
"sourceUrls": { "source-uuid-1": "https://..." },
"audioUrl": "https://..."
}
The job itself sits under job. Signed URLs in the response (output_url, sourceUrls, audioUrl) are freshly signed on every request, so fetch the job again rather than storing them.
camera_map is WizCut’s proposal for who’s on which camera — see The camera map. It’s null until it has been computed.
Status values:
| Status | Meaning |
|---|---|
created | Job created, waiting for uploads |
uploading | Waiting for files to finish uploading (e.g. after a failed job is reopened) |
syncing | Aligning audio across sources |
diarizing | Detecting speakers |
mapping | Speakers detected. WizCut is still working out who’s on which camera, or a person (or your agent) needs to confirm it |
ready | Cuts generated, ready for review or rendering |
rendering | Final render in progress |
complete | Render done, output_url available |
approved | Human approved the output |
failed | Something went wrong, see error_message |
Delete a job
DELETE /api/jobs/{jobId}
Deletes the job and every file it stored: uploads, proxies, speaker-detection audio and the render. Works in any status; if a stage is still running, its output is cleaned up when it finishes. Deleting a job does not give back used minutes. Returns { "ok": true }, or 404 if the job doesn’t exist or isn’t yours.
Webhooks
When you provide a callbackUrl, WizCut sends a JSON POST request when a job reaches one of the statuses below. WizCut tries each delivery up to three times and waits up to 10 seconds for your endpoint to answer.
Webhooks aren’t signed yet, so treat one as a signal to fetch the job with GET /api/jobs/{jobId} rather than trusting its contents. Poll as a fallback, too, in case a delivery gets lost.
“mapping” webhook (a person needs to confirm who’s on which camera):
{
"jobId": "uuid",
"status": "mapping",
"reviewUrl": "https://wizcut.com/jobs/uuid/edit?reviewToken=...",
"speakers": ["SPEAKER_00", "SPEAKER_01"],
"cameraMap": {
"version": 1,
"speakers": ["SPEAKER_00", "SPEAKER_01"],
"unassignedSpeakers": ["SPEAKER_01"],
"droppedSpeakers": [],
"cameras": [
{ "sourceId": "source-uuid-1", "kind": "closeup", "confidence": "high", "speakers": ["SPEAKER_00"], "...": "..." },
{ "sourceId": "source-uuid-2", "kind": "closeup", "confidence": "low", "speakers": ["SPEAKER_01"], "...": "..." }
]
}
}
Send a human to the reviewUrl to map speakers to camera sources and review cuts, or map them via the API. Review links are valid for 72 hours. speakers lists the detected speakers, most talkative first.
cameraMap is WizCut’s proposal, in the same shape as camera_map on the job — see The camera map. It’s left out when there’s no proposal yet.
When this webhook arrives depends on autoMap:
"confident"or"always"— the job waits in “mapping” without a webhook while WizCut works out the cameras, which takes a few minutes once the proxies are built. Then it either maps the cameras itself, and the next webhook you get is “ready”, or it sends “mapping” with the proposal attached. If WizCut can’t compute a proposal (for example because a proxy failed), “mapping” is sent withoutcameraMap— at the latest about 30 minutes after speaker detection."off"— “mapping” is sent right after speaker detection, usually withoutcameraMap. The proposal shows up inGET /api/jobs/{jobId}a few minutes later.
“ready” webhook (speaker mapping submitted — by a person, your code or WizCut — and cuts generated):
{
"jobId": "uuid",
"status": "ready",
"reviewUrl": "https://wizcut.com/jobs/uuid/edit?reviewToken=..."
}
“complete” webhook (render finished):
{
"jobId": "uuid",
"status": "complete",
"outputUrl": "https://presigned-download-url..."
}
“approved” webhook (human approved via review UI):
{
"jobId": "uuid",
"status": "approved",
"outputUrl": "https://presigned-download-url..."
}
“failed” webhook (the job failed, at any stage):
{
"jobId": "uuid",
"status": "failed",
"error": "description of what went wrong",
"failedAtStatus": "syncing"
}
error is the same message GET /api/jobs/{jobId} returns as error_message, and failedAtStatus names the stage that failed, such as syncing, diarizing or rendering. For most failures in sync, speaker detection and rendering, WizCut first retries once on its own — so when this webhook arrives, the retry didn’t work either.
The outputUrl in the “complete” and “approved” webhooks is valid for seven days after the render finishes. GET /api/jobs/{jobId} always returns a freshly signed output_url.
Speaker mapping and review (human-in-the-loop)
After speaker detection, the job enters “mapping” status. The detected speakers have to be mapped to camera sources — speaker detection identifies when someone speaks but not which camera they’re on.
WizCut tries to work that out for you. People move when they talk — they gesture, nod, lean in — so WizCut compares each speaker’s talking with the movement on each camera and proposes who’s on which one. How much it decides on its own is up to you, with autoMap.
When a person is needed, WizCut sends the “mapping” webhook. It includes a reviewUrl — a signed link to the WizCut editor. Send a human there. In the editor, they:
- Check which speaker is on which camera (the cameras WizCut is sure about are already selected)
- Preview and edit the auto-generated cuts
- Click “Render” when satisfied (if
review: true, the default) - Optionally click “Approve” after reviewing the rendered output
With review: false, rendering starts automatically as soon as the speaker mapping is submitted — the human doesn’t get to review cuts before rendering.
Automatic camera mapping
Set autoMap when you create the job:
| Mode | What WizCut does | Pick it when |
|---|---|---|
"off" | Never maps the cameras on its own. Sends “mapping” right after speaker detection; the proposal shows up on the job a few minutes later for a person or your agent to use. | A person always reviews the mapping anyway. |
"confident" | Default. Maps the cameras itself when it’s sure it has found a close-up of every speaker, then moves on to “ready”. Otherwise sends “mapping” with the proposal attached. A wide or two-shot is used in the edit but never counts as a speaker’s own camera. | Most integrations. You only hear from WizCut when it needs help. |
"always" | Doesn’t wait for a person. Uses the answers it’s sure about, then its likely answers, and places anyone left over on the remaining cameras by talk time. Only if it can’t compute a proposal at all (for example because a proxy failed) does it send “mapping”. | Fully unattended pipelines that would rather get a guessed edit than a stalled job. |
“Sure about every speaker” means: everyone with at least a minute of speech is on a camera WizCut answered with high confidence. Cameras it isn’t sure about get no speakers, so the edit doesn’t cut to them — fix that in the editor if you want them used.
Be ready for “confident” to ask a person fairly often. It holds back on footage where the movement doesn’t follow one speaker: TV studios with switched feeds, events, handheld or operated cameras. It also needs about 20 minutes of conversation to be sure, so short clips usually end up in “mapping”. In our tests so far, WizCut hasn’t put a speaker on the wrong camera when it said it was sure — but the sample is still small.
Combined with review: false, "confident" and "always" mean a job can render without anyone looking at it. That’s the point for unattended pipelines; if you want a person to see every edit, keep review: true.
The camera map
The proposal is on the job as camera_map (and in the “mapping” webhook as cameraMap). It’s null until computed — a few minutes after the proxies are built. Here’s one for a two-person show with a close-up of each person and a two-shot:
{
"version": 1,
"computedAt": "2026-10-01T09:42:17.000Z",
"speakers": ["SPEAKER_00", "SPEAKER_01"],
"unassignedSpeakers": [],
"droppedSpeakers": ["SPEAKER_02"],
"cameras": [
{
"sourceId": "source-uuid-1",
"kind": "closeup",
"confidence": "high",
"speakers": ["SPEAKER_00"],
"z": { "SPEAKER_00": 11.5, "SPEAKER_01": 0.9 },
"lagErrorS": { "SPEAKER_00": 0.1, "SPEAKER_01": 41.6 }
},
{
"sourceId": "source-uuid-2",
"kind": "closeup",
"confidence": "high",
"speakers": ["SPEAKER_01"],
"z": { "SPEAKER_01": 12.4, "SPEAKER_00": 1.3 },
"lagErrorS": { "SPEAKER_01": 0.2, "SPEAKER_00": 63.0 }
},
{
"sourceId": "source-uuid-3",
"kind": "twoshot",
"confidence": "high",
"speakers": ["SPEAKER_01", "SPEAKER_00"],
"positions": { "left": "SPEAKER_01", "right": "SPEAKER_00" },
"z": { "SPEAKER_00": 6.8, "SPEAKER_01": 5.2 },
"lagErrorS": { "SPEAKER_00": 0.3, "SPEAKER_01": 0.4 }
}
]
}
Per camera:
| Field | Description |
|---|---|
sourceId | One of the job’s video sources. |
kind | What the camera shows: closeup (one person), twoshot (two people side by side), wide (three people), multi (several people, one after another — usually an operated or switched camera) or unknown. |
confidence | high (WizCut is sure), low (a likely answer, not certain) or none (no answer). |
speakers | Who the camera shows. Empty when confidence is none. For two-shots and wides, from left to right. |
positions | Two-shots and wides only: who’s on the left, middle and right of the frame. |
reason | Set to no_motion when a camera couldn’t be analysed at all. |
z, lagErrorS | Diagnostic numbers per speaker: how strongly the camera’s movement follows that person’s talking, and how many seconds off the best match landed. Handy for debugging, not something to build on. |
Top level:
| Field | Description |
|---|---|
speakers | The speakers WizCut judged: everyone with at least a minute of speech. |
unassignedSpeakers | Judged speakers who aren’t on any high camera. If this isn’t empty, “confident” leaves the decision to a person. |
droppedSpeakers | Speakers with too little speech to judge — usually a stray label or a short interjection. WizCut doesn’t place them on a camera. |
To accept the proposal as it is, turn each camera you trust into a track — its sourceId and speakers — and submit it.
Map speakers via the API
If you already know who sits in front of which camera, or your agent has checked the proposal, submit the mapping yourself while the job is in “mapping” status:
POST /api/jobs/{jobId}/tracks
{
"tracks": [
{ "sourceId": "source-uuid-1", "speakers": ["SPEAKER_00"] },
{ "sourceId": "source-uuid-2", "speakers": ["SPEAKER_01"] }
]
}
Use the speaker names from the “mapping” webhook (or the turns in GET /api/jobs/{jobId}) and your video source IDs. A camera can show more than one speaker. WizCut then generates the cuts, moves the job to “ready”, sends the “ready” webhook, and — with review: false — starts the render. The response contains the new cuts.
Every sourceId has to be one of the job’s video sources; anything else gets a 400 with "code": "UNKNOWN_SOURCE". You can submit while WizCut is still working out the cameras — whichever comes first wins. If the mapping was already submitted, for example automatically, you get a 409 with "code": "INVALID_STATUS".
Advanced: correcting a mapping after cuts exist. Once the job is “ready” or “complete”, the same endpoint can still change the mapping, but it no longer generates cuts for you: send the corrected cuts along with tracks (and optionally removedRanges), or you get a 400 with "code": "CUTS_REQUIRED". WizCut saves both together, so a later recut uses the corrected mapping. It doesn’t send a webhook or start a render — trigger one if you need a new output. The WizCut editor does all of this for you when someone fixes a mapping there.
Silence removal
WizCut can automatically shorten or remove long pauses before generating cuts. This happens first, so camera switching is decided on the tightened conversation — you never get a cut to a camera that immediately loses most of its screen time to a removed pause.
Enable it with the silence field when creating a job:
{ "silence": { "mode": "tighten" } }
Modes:
| Mode | Behavior |
|---|---|
off | Keep all pauses as recorded (default when silence is omitted). |
tighten | Shorten long pauses to a natural beat (~0.7s). Recommended — keeps the conversation feeling human. |
remove | Cut long pauses out almost entirely, leaving just enough padding to avoid clipped words. |
Optional tuning knobs (defaults work well for most recordings):
| Field | Default | Range | Description |
|---|---|---|---|
minPauseSec | 1.5 | 0.3–10 | Only pauses longer than this are touched. |
keepPauseSec | 0.7 | 0–3 | The beat left in place of a tightened pause (tighten mode only). |
paddingSec | 0.25 | 0–1 | Safety margin kept around speech at every removed range. |
Detection is conservative: a pause is only removed where speaker detection and an audio energy analysis both agree it’s quiet, so laughter and other non-speech moments survive.
The job object (GET /api/jobs/{jobId}) includes removed_ranges — the removed spans in source time as { startMs, endMs } — alongside cuts. The final render skips these ranges; the review editor shows them as restorable hatched sections.
Change it after processing with the recut endpoint:
POST /api/jobs/{jobId}/recut
{ "silence": { "mode": "remove", "minPauseSec": 1.0 } }
Recut regenerates cuts and removed ranges from the already-detected speakers — no reprocessing, so it returns immediately with the new cuts, removedRanges, and savedMs (how much shorter the output gets). Pass { "mode": "off" } to turn silence removal off again; omit silence entirely to keep the job’s stored setting.
Billing note: usage is measured in input minutes (the footage you upload), so silence removal doesn’t change what a job costs — it just makes the output tighter.
Advanced: programmatic cut control
For fully automated pipelines that skip the UI entirely, let WizCut map the cameras or map speakers via the API, then manage cuts and rendering yourself. These endpoints work while the job is “ready” or “complete”:
Update cuts:
PATCH /api/jobs/{jobId}/cuts
{
"cuts": [
{ "startMs": 0, "endMs": 5000, "sourceId": "source-uuid-1" },
{ "startMs": 5000, "endMs": 12000, "sourceId": "source-uuid-2" }
]
}
The same endpoint also accepts removedRanges (array of { startMs, endMs }) to directly edit the silence-removal spans — for example to restore a pause the automatic pass removed. You can send cuts, removedRanges, or both.
Trigger render:
POST /api/jobs/{jobId}/render
Works from “ready”, or from “complete” to render again after editing cuts. Returns { "jobId": "uuid", "status": "rendering" }, or a 409 with "code": "NO_CUTS" if the job has no cuts to render.
Approve output:
POST /api/jobs/{jobId}/approve
Only works once the job is “complete”. Moves it to “approved” and sends the “approved” webhook.
Error responses
All endpoints return errors in this format:
{ "error": "Description of what went wrong" }
Some errors also carry a machine-readable code:
| Code | HTTP | Meaning |
|---|---|---|
TRACKS_ON_CREATE | 400 | POST /api/jobs got a speaker mapping. Leave tracks out and use autoMap, or set the mapping later. |
UNKNOWN_SOURCE | 400 | A track names a sourceId that isn’t one of the job’s video sources. |
CUTS_REQUIRED | 400 | Changing the mapping after cuts exist needs the updated cuts too. |
NOT_ENOUGH_VIDEOS | 400 | The job needs at least two video sources to be processed. |
QUOTA_EXCEEDED | 403 | You’ve used up your minutes. |
INVALID_STATUS | 409 | The job isn’t in a status that allows this — for example, its mapping was already submitted. |
NO_CUTS | 409 | There are no cuts to render. |
Common HTTP status codes:
| Code | Meaning |
|---|---|
| 400 | Invalid request (e.g. fewer than two video sources when processing) |
| 401 | Missing or invalid API key |
| 403 | Quota exceeded |
| 404 | Job not found or not owned by you |
| 409 | Invalid status transition (e.g. rendering a job that isn’t ready) |
| 500 | Server error |
| 502 | A processing service rejected the job or failed to start |
| 503 | Service not available |
Complete flow
Here’s the full happy path for an AI agent or automation:
# 1. Create a job with two cameras (files up to 5 GB each — pass
# fileSize to get multipart uploads for larger files)
JOB=$(curl -s -X POST https://wizcut.com/api/jobs \
-H "Authorization: Bearer wc_live_your_key" \
-H "Content-Type: application/json" \
-d '{
"sources": [
{"label": "Camera 1", "kind": "video"},
{"label": "Camera 2", "kind": "video"}
],
"callbackUrl": "https://your-server.com/webhook"
}')
JOB_ID=$(echo $JOB | jq -r '.jobId')
# 2. Upload video files to the presigned URLs
curl -X PUT -T camera1.mp4 "$(echo $JOB | jq -r '.uploads | to_entries[0].value.url')"
curl -X PUT -T camera2.mp4 "$(echo $JOB | jq -r '.uploads | to_entries[1].value.url')"
# 3. Start processing
curl -s -X POST "https://wizcut.com/api/jobs/$JOB_ID/process" \
-H "Authorization: Bearer wc_live_your_key"
# 4. Wait for the next webhook. With the default autoMap ("confident"):
# - { status: "ready", reviewUrl: "..." } — WizCut mapped the cameras itself
# - { status: "mapping", reviewUrl: "...", speakers: [...], cameraMap: {...} }
# — it wasn't sure. Send a human to reviewUrl to confirm the cameras
# (or POST the mapping to /api/jobs/$JOB_ID/tracks yourself)
# 5. Human reviews and edits cuts → clicks Render in the editor
# Wait for webhook: { status: "complete", outputUrl: "..." }
# 6. Download the rendered video from outputUrl
Fully unattended (autoMap: "always", review: false):
# Same as above, but add "autoMap": "always" and "review": false to step 1
# WizCut maps the cameras itself, even when it isn't sure,
# and starts rendering right after — no person involved
# You'll get: { status: "ready", ... }, then { status: "complete", outputUrl: "..." }
Integrations
Prefer no-code? You can drive the WizCut API from automation platforms without writing any of the requests above.
n8n
We maintain an official n8n integration — the verified community node @wizcut/n8n-nodes-wizcut. Install it from Settings → Community Nodes in n8n (it’s verified, so it’s also available on n8n Cloud).
It ships two nodes:
- WizCut — actions for Create Job, Get Job, Start Processing, Start Render, and Approve.
- WizCut Trigger — a webhook trigger that starts your workflow whenever a job changes status (mapping, ready, complete, approved).
Point the trigger’s webhook URL at your job’s callbackUrl and you get a fully event-driven pipeline — for example, ping Slack when WizCut needs someone to confirm who’s on which camera, auto-start the render when cuts are ready, and post the download link when the episode is done. Both nodes also work as tools for n8n AI agents.
Want a head start? Grab the ready-made template: Notify and manage podcast edit status with WizCut and Slack.
More platforms
Make.com and Zapier integrations are on the roadmap. Building on a different platform, or want a workflow template for your exact setup? Reach out to WizCut support — feature requests are welcome.