How to use WizCut

WizCut API

Programmatically upload, process, and render multicam podcast edits with the WizCut API.

The WizCut API lets you automate multicam podcast editing. Upload your camera angles, let WizCut detect speakers, work out who’s on which camera and generate cuts, optionally review them in a hosted editor, and get a rendered video back via webhook.

Authentication

Create an API key at wizcut.com/settings. Include it in every request:

Authorization: Bearer wc_live_your_key_here

Keys can be revoked at any time from the settings page.

Create a job

POST /api/jobs
{
  "sources": [
    { "label": "Camera 1", "kind": "video", "ext": "mp4", "fileSize": 2147483648 },
    { "label": "Camera 2", "kind": "video", "ext": "mp4", "fileSize": 1932735283 },
    { "label": "Mic mix", "kind": "audio", "ext": "wav" }
  ],
  "callbackUrl": "https://your-server.com/webhook",
  "review": true,
  "autoMap": "confident",
  "silence": { "mode": "tighten" }
}

Fields:

FieldTypeDescription
sourcesarrayCamera angles or audio files, up to 12 per job. A job needs at least two video sources before it can be processed. Each has label, optional kind ("video" or "audio", default "video"), optional ext (mp4, mov, webm, mkv, wav, mp3, m4a or aac; default "mp4" for video, "wav" for audio) and optional fileSize in bytes.
callbackUrlstringOptional. URL to receive webhook notifications — see Webhooks.
reviewbooleanDefault true. When true, the job pauses at “ready” for human review of cuts before rendering. When false, rendering starts automatically as soon as the speaker mapping is submitted — by a person, your code, or WizCut itself.
autoMapstringDefault "confident". Whether WizCut may decide on its own who’s on which camera: "off", "confident" or "always" — see Automatic camera mapping.
silenceobjectOptional. Automatic pause removal — see Silence removal. Omitted means off.

You can’t pass the speaker mapping (tracks) here: source IDs only exist once the job is created, so there’s nothing to point a speaker at yet. A tracks list that names speakers is rejected with 400 and "code": "TRACKS_ON_CREATE". Use autoMap, or map speakers via the API once speaker detection is done.

Pass fileSize for every file you know the size of. Files over 100 MB then get a multipart upload, which lets you upload parts in parallel, retry a single part, and is the only way to upload files larger than 5 GB. Without fileSize, every file gets a single presigned PUT.

Response:

{
  "jobId": "uuid",
  "uploads": {
    "source-uuid-1": {
      "method": "multipart",
      "uploadId": "...",
      "key": "sources/uuid/source-uuid-1.mp4",
      "partSize": 52428800,
      "partCount": 41,
      "urls": ["https://presigned-part-1-url...", "https://presigned-part-2-url...", "..."]
    },
    "source-uuid-3": {
      "method": "put",
      "url": "https://presigned-upload-url..."
    }
  },
  "audioUploads": { "...": "..." }
}

uploads is keyed by the source IDs WizCut assigned, in the same order as your sources array. You can ignore audioUploads — it’s used by the WizCut apps, which upload an extracted audio track ahead of the video.

Upload files

No auth header is needed for the presigned URLs — they’re self-authenticating.

"method": "put" — PUT the whole file to url within an hour:

curl -X PUT -T mic-mix.wav "https://presigned-upload-url..."

"method": "multipart" — split the file into partSize-byte chunks (the last one is shorter) and PUT chunk n to the URL for part n:

  1. urls holds the URLs for the first parts (urls[0] is part 1). Request more as you go — up to 16 per request:

    POST /api/jobs/{jobId}/sources/{sourceId}/parts
    
    { "uploadId": "...", "partNumbers": [5, 6, 7, 8] }
    

    The response maps part numbers to URLs: { "urls": { "5": "https://...", "6": "https://..." } }. Part URLs are valid for two hours, so fetch them shortly before you use them.

  2. Once every part is uploaded, finish the upload:

    POST /api/uploads/complete
    
    { "key": "sources/uuid/source-uuid-1.mp4", "uploadId": "..." }
    

To give up on an upload, send the same body to POST /api/uploads/abort.

Start processing

POST /api/jobs/{jobId}/process

Call this once every file is uploaded. It kicks off the pipeline: audio sync, speaker detection, proxy rendering, camera mapping and cut generation. The request returns immediately — processing happens asynchronously.

Optional body:

{ "diarizeSourceIds": ["source-uuid-3"] }

diarizeSourceIds picks which sources WizCut listens to for speaker detection. If you leave it out, WizCut uses your audio sources, or the first source if there are none. Pick the sources with the cleanest audio of everyone talking — a camera with a muted or distant mic leads to a job with no cuts.

Response:

{ "jobId": "uuid", "status": "syncing" }

List jobs

GET /api/jobs

Returns { "jobs": [...] } with your 50 most recent jobs, newest first. Each has id, title, status, created_at and updated_at.

Get job status

GET /api/jobs/{jobId}

Response:

{
  "job": {
    "id": "uuid",
    "status": "complete",
    "sources": [...],
    "tracks": [...],
    "turns": [...],
    "cuts": [...],
    "removed_ranges": [...],
    "camera_map": {...},
    "output_url": "https://presigned-download-url...",
    "error_message": null
  },
  "sourceUrls": { "source-uuid-1": "https://..." },
  "audioUrl": "https://..."
}

The job itself sits under job. Signed URLs in the response (output_url, sourceUrls, audioUrl) are freshly signed on every request, so fetch the job again rather than storing them.

camera_map is WizCut’s proposal for who’s on which camera — see The camera map. It’s null until it has been computed.

Status values:

StatusMeaning
createdJob created, waiting for uploads
uploadingWaiting for files to finish uploading (e.g. after a failed job is reopened)
syncingAligning audio across sources
diarizingDetecting speakers
mappingSpeakers detected. WizCut is still working out who’s on which camera, or a person (or your agent) needs to confirm it
readyCuts generated, ready for review or rendering
renderingFinal render in progress
completeRender done, output_url available
approvedHuman approved the output
failedSomething went wrong, see error_message

Delete a job

DELETE /api/jobs/{jobId}

Deletes the job and every file it stored: uploads, proxies, speaker-detection audio and the render. Works in any status; if a stage is still running, its output is cleaned up when it finishes. Deleting a job does not give back used minutes. Returns { "ok": true }, or 404 if the job doesn’t exist or isn’t yours.

Webhooks

When you provide a callbackUrl, WizCut sends a JSON POST request when a job reaches one of the statuses below. WizCut tries each delivery up to three times and waits up to 10 seconds for your endpoint to answer.

Webhooks aren’t signed yet, so treat one as a signal to fetch the job with GET /api/jobs/{jobId} rather than trusting its contents. Poll as a fallback, too, in case a delivery gets lost.

“mapping” webhook (a person needs to confirm who’s on which camera):

{
  "jobId": "uuid",
  "status": "mapping",
  "reviewUrl": "https://wizcut.com/jobs/uuid/edit?reviewToken=...",
  "speakers": ["SPEAKER_00", "SPEAKER_01"],
  "cameraMap": {
    "version": 1,
    "speakers": ["SPEAKER_00", "SPEAKER_01"],
    "unassignedSpeakers": ["SPEAKER_01"],
    "droppedSpeakers": [],
    "cameras": [
      { "sourceId": "source-uuid-1", "kind": "closeup", "confidence": "high", "speakers": ["SPEAKER_00"], "...": "..." },
      { "sourceId": "source-uuid-2", "kind": "closeup", "confidence": "low", "speakers": ["SPEAKER_01"], "...": "..." }
    ]
  }
}

Send a human to the reviewUrl to map speakers to camera sources and review cuts, or map them via the API. Review links are valid for 72 hours. speakers lists the detected speakers, most talkative first.

cameraMap is WizCut’s proposal, in the same shape as camera_map on the job — see The camera map. It’s left out when there’s no proposal yet.

When this webhook arrives depends on autoMap:

  • "confident" or "always" — the job waits in “mapping” without a webhook while WizCut works out the cameras, which takes a few minutes once the proxies are built. Then it either maps the cameras itself, and the next webhook you get is “ready”, or it sends “mapping” with the proposal attached. If WizCut can’t compute a proposal (for example because a proxy failed), “mapping” is sent without cameraMap — at the latest about 30 minutes after speaker detection.
  • "off" — “mapping” is sent right after speaker detection, usually without cameraMap. The proposal shows up in GET /api/jobs/{jobId} a few minutes later.

“ready” webhook (speaker mapping submitted — by a person, your code or WizCut — and cuts generated):

{
  "jobId": "uuid",
  "status": "ready",
  "reviewUrl": "https://wizcut.com/jobs/uuid/edit?reviewToken=..."
}

“complete” webhook (render finished):

{
  "jobId": "uuid",
  "status": "complete",
  "outputUrl": "https://presigned-download-url..."
}

“approved” webhook (human approved via review UI):

{
  "jobId": "uuid",
  "status": "approved",
  "outputUrl": "https://presigned-download-url..."
}

“failed” webhook (the job failed, at any stage):

{
  "jobId": "uuid",
  "status": "failed",
  "error": "description of what went wrong",
  "failedAtStatus": "syncing"
}

error is the same message GET /api/jobs/{jobId} returns as error_message, and failedAtStatus names the stage that failed, such as syncing, diarizing or rendering. For most failures in sync, speaker detection and rendering, WizCut first retries once on its own — so when this webhook arrives, the retry didn’t work either.

The outputUrl in the “complete” and “approved” webhooks is valid for seven days after the render finishes. GET /api/jobs/{jobId} always returns a freshly signed output_url.

Speaker mapping and review (human-in-the-loop)

After speaker detection, the job enters “mapping” status. The detected speakers have to be mapped to camera sources — speaker detection identifies when someone speaks but not which camera they’re on.

WizCut tries to work that out for you. People move when they talk — they gesture, nod, lean in — so WizCut compares each speaker’s talking with the movement on each camera and proposes who’s on which one. How much it decides on its own is up to you, with autoMap.

When a person is needed, WizCut sends the “mapping” webhook. It includes a reviewUrl — a signed link to the WizCut editor. Send a human there. In the editor, they:

  1. Check which speaker is on which camera (the cameras WizCut is sure about are already selected)
  2. Preview and edit the auto-generated cuts
  3. Click “Render” when satisfied (if review: true, the default)
  4. Optionally click “Approve” after reviewing the rendered output

With review: false, rendering starts automatically as soon as the speaker mapping is submitted — the human doesn’t get to review cuts before rendering.

Automatic camera mapping

Set autoMap when you create the job:

ModeWhat WizCut doesPick it when
"off"Never maps the cameras on its own. Sends “mapping” right after speaker detection; the proposal shows up on the job a few minutes later for a person or your agent to use.A person always reviews the mapping anyway.
"confident"Default. Maps the cameras itself when it’s sure it has found a close-up of every speaker, then moves on to “ready”. Otherwise sends “mapping” with the proposal attached. A wide or two-shot is used in the edit but never counts as a speaker’s own camera.Most integrations. You only hear from WizCut when it needs help.
"always"Doesn’t wait for a person. Uses the answers it’s sure about, then its likely answers, and places anyone left over on the remaining cameras by talk time. Only if it can’t compute a proposal at all (for example because a proxy failed) does it send “mapping”.Fully unattended pipelines that would rather get a guessed edit than a stalled job.

“Sure about every speaker” means: everyone with at least a minute of speech is on a camera WizCut answered with high confidence. Cameras it isn’t sure about get no speakers, so the edit doesn’t cut to them — fix that in the editor if you want them used.

Be ready for “confident” to ask a person fairly often. It holds back on footage where the movement doesn’t follow one speaker: TV studios with switched feeds, events, handheld or operated cameras. It also needs about 20 minutes of conversation to be sure, so short clips usually end up in “mapping”. In our tests so far, WizCut hasn’t put a speaker on the wrong camera when it said it was sure — but the sample is still small.

Combined with review: false, "confident" and "always" mean a job can render without anyone looking at it. That’s the point for unattended pipelines; if you want a person to see every edit, keep review: true.

The camera map

The proposal is on the job as camera_map (and in the “mapping” webhook as cameraMap). It’s null until computed — a few minutes after the proxies are built. Here’s one for a two-person show with a close-up of each person and a two-shot:

{
  "version": 1,
  "computedAt": "2026-10-01T09:42:17.000Z",
  "speakers": ["SPEAKER_00", "SPEAKER_01"],
  "unassignedSpeakers": [],
  "droppedSpeakers": ["SPEAKER_02"],
  "cameras": [
    {
      "sourceId": "source-uuid-1",
      "kind": "closeup",
      "confidence": "high",
      "speakers": ["SPEAKER_00"],
      "z": { "SPEAKER_00": 11.5, "SPEAKER_01": 0.9 },
      "lagErrorS": { "SPEAKER_00": 0.1, "SPEAKER_01": 41.6 }
    },
    {
      "sourceId": "source-uuid-2",
      "kind": "closeup",
      "confidence": "high",
      "speakers": ["SPEAKER_01"],
      "z": { "SPEAKER_01": 12.4, "SPEAKER_00": 1.3 },
      "lagErrorS": { "SPEAKER_01": 0.2, "SPEAKER_00": 63.0 }
    },
    {
      "sourceId": "source-uuid-3",
      "kind": "twoshot",
      "confidence": "high",
      "speakers": ["SPEAKER_01", "SPEAKER_00"],
      "positions": { "left": "SPEAKER_01", "right": "SPEAKER_00" },
      "z": { "SPEAKER_00": 6.8, "SPEAKER_01": 5.2 },
      "lagErrorS": { "SPEAKER_00": 0.3, "SPEAKER_01": 0.4 }
    }
  ]
}

Per camera:

FieldDescription
sourceIdOne of the job’s video sources.
kindWhat the camera shows: closeup (one person), twoshot (two people side by side), wide (three people), multi (several people, one after another — usually an operated or switched camera) or unknown.
confidencehigh (WizCut is sure), low (a likely answer, not certain) or none (no answer).
speakersWho the camera shows. Empty when confidence is none. For two-shots and wides, from left to right.
positionsTwo-shots and wides only: who’s on the left, middle and right of the frame.
reasonSet to no_motion when a camera couldn’t be analysed at all.
z, lagErrorSDiagnostic numbers per speaker: how strongly the camera’s movement follows that person’s talking, and how many seconds off the best match landed. Handy for debugging, not something to build on.

Top level:

FieldDescription
speakersThe speakers WizCut judged: everyone with at least a minute of speech.
unassignedSpeakersJudged speakers who aren’t on any high camera. If this isn’t empty, “confident” leaves the decision to a person.
droppedSpeakersSpeakers with too little speech to judge — usually a stray label or a short interjection. WizCut doesn’t place them on a camera.

To accept the proposal as it is, turn each camera you trust into a track — its sourceId and speakers — and submit it.

Map speakers via the API

If you already know who sits in front of which camera, or your agent has checked the proposal, submit the mapping yourself while the job is in “mapping” status:

POST /api/jobs/{jobId}/tracks
{
  "tracks": [
    { "sourceId": "source-uuid-1", "speakers": ["SPEAKER_00"] },
    { "sourceId": "source-uuid-2", "speakers": ["SPEAKER_01"] }
  ]
}

Use the speaker names from the “mapping” webhook (or the turns in GET /api/jobs/{jobId}) and your video source IDs. A camera can show more than one speaker. WizCut then generates the cuts, moves the job to “ready”, sends the “ready” webhook, and — with review: false — starts the render. The response contains the new cuts.

Every sourceId has to be one of the job’s video sources; anything else gets a 400 with "code": "UNKNOWN_SOURCE". You can submit while WizCut is still working out the cameras — whichever comes first wins. If the mapping was already submitted, for example automatically, you get a 409 with "code": "INVALID_STATUS".

Advanced: correcting a mapping after cuts exist. Once the job is “ready” or “complete”, the same endpoint can still change the mapping, but it no longer generates cuts for you: send the corrected cuts along with tracks (and optionally removedRanges), or you get a 400 with "code": "CUTS_REQUIRED". WizCut saves both together, so a later recut uses the corrected mapping. It doesn’t send a webhook or start a render — trigger one if you need a new output. The WizCut editor does all of this for you when someone fixes a mapping there.

Silence removal

WizCut can automatically shorten or remove long pauses before generating cuts. This happens first, so camera switching is decided on the tightened conversation — you never get a cut to a camera that immediately loses most of its screen time to a removed pause.

Enable it with the silence field when creating a job:

{ "silence": { "mode": "tighten" } }

Modes:

ModeBehavior
offKeep all pauses as recorded (default when silence is omitted).
tightenShorten long pauses to a natural beat (~0.7s). Recommended — keeps the conversation feeling human.
removeCut long pauses out almost entirely, leaving just enough padding to avoid clipped words.

Optional tuning knobs (defaults work well for most recordings):

FieldDefaultRangeDescription
minPauseSec1.50.3–10Only pauses longer than this are touched.
keepPauseSec0.70–3The beat left in place of a tightened pause (tighten mode only).
paddingSec0.250–1Safety margin kept around speech at every removed range.

Detection is conservative: a pause is only removed where speaker detection and an audio energy analysis both agree it’s quiet, so laughter and other non-speech moments survive.

The job object (GET /api/jobs/{jobId}) includes removed_ranges — the removed spans in source time as { startMs, endMs } — alongside cuts. The final render skips these ranges; the review editor shows them as restorable hatched sections.

Change it after processing with the recut endpoint:

POST /api/jobs/{jobId}/recut
{ "silence": { "mode": "remove", "minPauseSec": 1.0 } }

Recut regenerates cuts and removed ranges from the already-detected speakers — no reprocessing, so it returns immediately with the new cuts, removedRanges, and savedMs (how much shorter the output gets). Pass { "mode": "off" } to turn silence removal off again; omit silence entirely to keep the job’s stored setting.

Billing note: usage is measured in input minutes (the footage you upload), so silence removal doesn’t change what a job costs — it just makes the output tighter.

Advanced: programmatic cut control

For fully automated pipelines that skip the UI entirely, let WizCut map the cameras or map speakers via the API, then manage cuts and rendering yourself. These endpoints work while the job is “ready” or “complete”:

Update cuts:

PATCH /api/jobs/{jobId}/cuts
{
  "cuts": [
    { "startMs": 0, "endMs": 5000, "sourceId": "source-uuid-1" },
    { "startMs": 5000, "endMs": 12000, "sourceId": "source-uuid-2" }
  ]
}

The same endpoint also accepts removedRanges (array of { startMs, endMs }) to directly edit the silence-removal spans — for example to restore a pause the automatic pass removed. You can send cuts, removedRanges, or both.

Trigger render:

POST /api/jobs/{jobId}/render

Works from “ready”, or from “complete” to render again after editing cuts. Returns { "jobId": "uuid", "status": "rendering" }, or a 409 with "code": "NO_CUTS" if the job has no cuts to render.

Approve output:

POST /api/jobs/{jobId}/approve

Only works once the job is “complete”. Moves it to “approved” and sends the “approved” webhook.

Error responses

All endpoints return errors in this format:

{ "error": "Description of what went wrong" }

Some errors also carry a machine-readable code:

CodeHTTPMeaning
TRACKS_ON_CREATE400POST /api/jobs got a speaker mapping. Leave tracks out and use autoMap, or set the mapping later.
UNKNOWN_SOURCE400A track names a sourceId that isn’t one of the job’s video sources.
CUTS_REQUIRED400Changing the mapping after cuts exist needs the updated cuts too.
NOT_ENOUGH_VIDEOS400The job needs at least two video sources to be processed.
QUOTA_EXCEEDED403You’ve used up your minutes.
INVALID_STATUS409The job isn’t in a status that allows this — for example, its mapping was already submitted.
NO_CUTS409There are no cuts to render.

Common HTTP status codes:

CodeMeaning
400Invalid request (e.g. fewer than two video sources when processing)
401Missing or invalid API key
403Quota exceeded
404Job not found or not owned by you
409Invalid status transition (e.g. rendering a job that isn’t ready)
500Server error
502A processing service rejected the job or failed to start
503Service not available

Complete flow

Here’s the full happy path for an AI agent or automation:

# 1. Create a job with two cameras (files up to 5 GB each — pass
#    fileSize to get multipart uploads for larger files)
JOB=$(curl -s -X POST https://wizcut.com/api/jobs \
  -H "Authorization: Bearer wc_live_your_key" \
  -H "Content-Type: application/json" \
  -d '{
    "sources": [
      {"label": "Camera 1", "kind": "video"},
      {"label": "Camera 2", "kind": "video"}
    ],
    "callbackUrl": "https://your-server.com/webhook"
  }')

JOB_ID=$(echo $JOB | jq -r '.jobId')

# 2. Upload video files to the presigned URLs
curl -X PUT -T camera1.mp4 "$(echo $JOB | jq -r '.uploads | to_entries[0].value.url')"
curl -X PUT -T camera2.mp4 "$(echo $JOB | jq -r '.uploads | to_entries[1].value.url')"

# 3. Start processing
curl -s -X POST "https://wizcut.com/api/jobs/$JOB_ID/process" \
  -H "Authorization: Bearer wc_live_your_key"

# 4. Wait for the next webhook. With the default autoMap ("confident"):
#    - { status: "ready", reviewUrl: "..." } — WizCut mapped the cameras itself
#    - { status: "mapping", reviewUrl: "...", speakers: [...], cameraMap: {...} }
#      — it wasn't sure. Send a human to reviewUrl to confirm the cameras
#      (or POST the mapping to /api/jobs/$JOB_ID/tracks yourself)

# 5. Human reviews and edits cuts → clicks Render in the editor
#    Wait for webhook: { status: "complete", outputUrl: "..." }

# 6. Download the rendered video from outputUrl

Fully unattended (autoMap: "always", review: false):

# Same as above, but add "autoMap": "always" and "review": false to step 1
# WizCut maps the cameras itself, even when it isn't sure,
# and starts rendering right after — no person involved
# You'll get: { status: "ready", ... }, then { status: "complete", outputUrl: "..." }

Integrations

Prefer no-code? You can drive the WizCut API from automation platforms without writing any of the requests above.

n8n

We maintain an official n8n integration — the verified community node @wizcut/n8n-nodes-wizcut. Install it from Settings → Community Nodes in n8n (it’s verified, so it’s also available on n8n Cloud).

It ships two nodes:

  • WizCut — actions for Create Job, Get Job, Start Processing, Start Render, and Approve.
  • WizCut Trigger — a webhook trigger that starts your workflow whenever a job changes status (mapping, ready, complete, approved).

Point the trigger’s webhook URL at your job’s callbackUrl and you get a fully event-driven pipeline — for example, ping Slack when WizCut needs someone to confirm who’s on which camera, auto-start the render when cuts are ready, and post the download link when the episode is done. Both nodes also work as tools for n8n AI agents.

Want a head start? Grab the ready-made template: Notify and manage podcast edit status with WizCut and Slack.

More platforms

Make.com and Zapier integrations are on the roadmap. Building on a different platform, or want a workflow template for your exact setup? Reach out to WizCut support — feature requests are welcome.

WizCut API – WizCut Docs