← Research Notes中文
Technology · 2026-07-24 · 21 min read · Close read

Can Fine Cutting Be Handed to AI? The Boundary Isn't Model Precision — It's What Defines the Cut

"Can AI edit video" is the wrong question. Once you lay the boundary out against the evidence, the dividing line isn't model strength — it's what defines the position of a given cut. Cuts defined by a signal (word boundaries, silence, shot changes, beats) have been automated at frame accuracy for years, and that frame accuracy comes entirely from deterministic signal processing, not from AI. Cuts defined by meaning and emotion still can't be automated — and the blocker isn't resolution. It's the absence of counterfactual data, of an objective function, and of any way to grade the result.

Research statement

This piece reviews the capability boundary of "can an AI with tools complete video fine cutting." 182 research agents searched in parallel; 56 load-bearing conclusions each passed 3 independent adversarial verification perspectives (technical mechanism / product reality / recency), with verifiers instructed to default to "this doesn't hold unless you can confirm it."

The outcome is worth stating up front: not one of the 56 passed all three perspectives unamended, and not one was refuted — every single one landed on "directionally right, wording needs correction." Part of that 0% clean-pass rate is an artifact of my own "assume it fails" prior, so it does not mean the original conclusions were wrong. The real output is the 56 corrections, and what follows is the corrected version. Grades: [Solid] = survives verification (possibly narrowed) · [Narrowed] = the draft claim was overturned or heavily qualified, and what appears here is the corrected form · [Inference] = analytical derivation.

Recency: this round exhausted its web-search budget and had to rely on model knowledge (cutoff around early 2026) plus fetches of known authoritative URLs. February–July 2026 is an explicit blind spot; treat every product-version-dependent claim as possibly stale.

The verdict up front: "can fine cutting be automated" is not one question but two. Once you separate cuts by what defines their position — signal or meaning — the first has been shipping at frame accuracy for years and the second has zero products, with the blocker sitting nowhere near the model. More importantly: there is still no public way to score whether a cut is good. So this isn't a battle being slowly won. It's a battle with no scoreboard.

1. First, pin down "fine cut" versus "rough cut"

Without pinning the definition, this topic immediately slides into "AI can auto-edit video," which carries no information. The industry sequence is: assembly → rough cut (structure and selection) → fine cut (frame level) → picture lock.

Almost all "AI editing" marketing conflates these two layers. And the exit condition of fine cutting has only been partially formalized [Narrowed]: picture lock itself is not a machine-verifiable fact — it's a declaration made by a director or producer under scheduling and contractual pressure. What has been formalized is its consequence: once locked, downstream work is scheduled and billed against frozen frame numbers (VFX pull lists with head/tail handles, composer spotting sessions writing cue sheets against timecode, audio turnover carrying 2–10 second handles). Any change after lock must produce a change list so downstream can auto-conform.

One correction I would have gotten wrong myself: this diff is not an O(n) per-clip comparison. Insert one segment and every downstream timecode shifts, so you must first do sequence alignment (LCS/Myers class, keyed on reel + source timecode). Worse, the alignment is semantically ambiguous — "trim 12 frames off the head and slip" versus "delete the old shot and insert a different take" can produce a frame-identical new timeline while implying completely different conform instructions. This is exactly why commercial conform tools ship with interactive adjudication UIs and why assistant editors still QC change lists by hand. And "empty diff" is neither sufficient for, nor equivalent to, "the picture didn't change": VFX version swaps, reframes, timewarps, an equal-length dissolve changing duration, titles, and color all move zero timecode.

2. The real dividing line: what defines this cut

This is the highest-information finding of the round, and the one all 14 dimensions' corrections independently converged on [Solid]:

Whether something can be automated depends not on how "fine" the operation is, but on whether the position of the cut is defined by a detectable physical event, or by meaning, emotion, and style.

Signal-defined boundaries — word onsets and offsets, silence, shot changes, beats — already have frame-level or sub-frame annotations, metrics, off-the-shelf tools, and leaderboards. Forced alignment is evaluated with millisecond boundary error as its primary metric. Shot boundary detection is per-frame binary classification. Beat tracking's standard tolerance is ±70ms. All of these are at or below the size of one frame.

Meaning-defined boundaries — "the moment he starts to hesitate," "right as the emotion lands," "cut when the punchline finishes" — have no loss, no leaderboard, and more fundamentally no way to construct reliable frame-level ground truth: inter-annotator agreement on such boundaries sits at the ±1 second scale. This is not "nobody has trained it." Evaluation and supervision fail first, in principle.

This line cuts across the fine-cut checklist rather than running along an easy/hard axis:

Fine-cut operationVerdictWhat blocks it (mechanism)
Remove breaths / filler words / silenceAutomatedBoundary defined by words and silence; forced alignment's 10–20ms grid is finer than a frame
Music beat sync (cut on the beat)AutomatedBoundary defined by the beat — but only holds for steady-tempo produced music
Recover existing cut pointsAutomatedPer-frame detection; but only recovers cuts that already exist, never proposes new ones
Multicam switching by speakerAutomatedRule-based: cut to whoever is talking + minimum shot duration; an externally observable proxy variable
How long to hold a breathNot possibleEvery product exposes a global constant; no per-instance decision, no paired corpus, no acceptance test
How many frames to offset a J/L cutNot possibleWriting it into the format is trivial; what's missing is an objective for "how much"
Match cutNot possibleThe representation simply lacks the quantity "cross-shot formal correspondence" (see the model ladder below)
Shot-scale rhythm / emotional curveNot possibleThe criterion is non-local; valence is unreliable, arousal is reliable but isn't the axis you want

3. The part that is automated gets its frame accuracy from somewhere other than AI

This is the most easily misread point [Solid]. For all four rows marked "Automated" above, none of the frame-level precision comes from a video LLM. It all comes from deterministic detectors:

The LLM's place in this chain is selection and ranking at the symbolic layer — it looks at a handful to a few dozen already-textualized candidates per second, not at 24 arbitrary frames. An LLM never needs to "say frame 37," and shouldn't be the thing that does. Every shipping frame-accurate AI edit today has this two-layer shape: second-level semantic selection, frame-level deterministic snapping.

Two discounts on that "automated," though:

Discount one: what's called "pacing" is actually a constant. [Narrowed] The open-source silence remover's breath control is a single scalar defaulting to 0.2 seconds; the commercial "shorten word gaps" feature is one slider for the whole file. What they expose is g' = f(g) — a memoryless scalar map whose target depends only on the gap's own length, with exactly zero conditional dependence on semantic position, syntactic boundary, or emotion. And the result isn't a "flattened" gap distribution — it's one quantized to a single spike: anything under threshold survives untouched, everything else collapses to the same value. That is, by definition, not pacing.

Discount two: the domain gap is bigger than you'd think. Shot boundary detection scores about 0.96 F1 on professional footage but only 77.9 on a UGC short-video dataset — precisely the domain where AI editing actually gets deployed. The measurement primitive itself loses roughly 18 points before you start.

4. Three layers of blockage, none of which is "the model isn't strong enough"

4.1 Missing counterfactual data [Solid — the hardest finding of the round]

Every public corpus records only the version that shipped. Not one dataset records the in/out points that were rejected for the same footage, which means no model has ever seen a paired preference label for "this cut 6 frames earlier versus 6 frames later — which is better." Public work can only do contrastive learning between real cuts and randomly displaced fake ones — and "real versus random" is vastly easier than "good versus plausible-but-worse."

There's a directly citable number here: the representative work on learning cut placement from finished films achieves top-1 accuracy of 8.18% at a ±1 snippet (roughly 8 frames) tolerance, against a 0.60% random baseline. At a tolerance far looser than fine cutting requires, top-1 is 8%.

Notably this is not "unobtainable": that same work's human study had 10 professional editors do pairwise "which cut is better" comparisons — the protocol is validated, it just stopped at dozens of examples. Scaling to the 105 range needs roughly 600–800 editor-hours, on the order of US$40–80k — cheaper than a mid-sized preference dataset. So the accurate statement is that nobody wants to fund it for a small market, not that it's physically impossible.

4.2 No objective function, and no verifier [Solid]

As of 2026, there exists no public benchmark that takes "is this editing craft good" as the thing being measured. Every existing asset substitutes a different question — imitation of the released cut, i.e. "did you pick the same frame as that editor did."

How fatal that substitution is has already been demonstrated in an adjacent field: video summarization used multi-reference evaluation with several human-annotated versions of the same footage, and the result was that randomly generated summaries scored comparably to published methods. Multiple references didn't rescue the metric — they exposed that it had degenerated.

Sharper evidence comes from the one public benchmark that lists J cut, L cut, match cut and friends as explicit categories. On the easy version of the task — cut position already given, only the type needs classifying — the best baseline reaches mean AP of just 47.9%, with match cut at 2.43%.

Two corrections are mandatory here, or you'll cite it wrong:

4.3 Sampling economics kills reinforcement learning [Inference, arithmetic verified]

If you can't write the objective analytically, can you use real engagement as the reward? Run the numbers. To detect a 1 percentage point difference in retention via A/B, you need roughly 3.9×10⁴ plays per arm; for 0.5 points, 1.6×10⁵. The true effect of a 3-frame trim is almost certainly far below 0.5 points, and a 10-minute talking-head piece has 200–500 cuts. One round of per-cut credit assignment across a whole piece would need 10⁷–10⁸ plays — more than the lifetime view count of most target channels. This is sampling economics, not compute.

4.4 The criterion itself is non-local [Solid, with one common misquote to remove]

The most-cited weighting in editing is Walter Murch's Rule of Six: emotion 51% / story 23% / rhythm 10% / eye-trace 7% / two-dimensional plane 5% / three-dimensional space 4%. It's routinely used to argue that "only the bottom 16% is formalizable."

[Narrowed] That usage is a category error. Murch gives a lexicographic priority ordering (his own phrasing is to sacrifice your way up, item by item, from the bottom), not additive shares. Summing 7+5+4 into a "16% capability ceiling" is mathematically meaningless. What it actually demonstrates is something else — that no hand-written, decomposable objective function for edit quality exists, because the top two entries were never defined as measurable quantities within that vocabulary.

And it has a direct consequence: an A/B on ±6 frames, conducted outside the context of the whole scene, mostly measures noise. Which means that even if you obtained every EDL diff from v1 through v40 of some film, the labels would be specific to that editor and that scene, and might not transfer.

5. The claim verification overturned: "1fps is a hard wall" is wrong

This one deserves its own section, because it's the claim I was most likely to get wrong myself [Narrowed].

The usual argument runs: mainstream video LLMs ingest at 1 frame per second, so 96% of the frames in 24fps footage never enter the context, and frame-level fine cutting is therefore walled off at the level of input representation. That argument doesn't hold. 1 fps is a default, not a ceiling — the mainstream API's video metadata carries a configurable sampling-rate field (valid up to 24), and on the open-source side the sampling rate is entirely the caller's choice. 24fps footage can go into context at 100%.

Cost isn't the wall either: scanning an hour at 1fps is roughly 900k tokens, in the small-change range; even maxed to 24fps at low resolution, an hour is about 5.7M tokens. What stops it is neither money nor "pixels can't get in."

The blockers that actually survive are three, and they're more interesting:

  1. Output contract: the API returns no frame numbers, the model has no frame-index handle, and the documented convention is to write times as MM:SS. The sampling-rate ceiling of 24 also means 25/30/50/60fps footage can't be fully covered even maxed out. Getting frame numbers requires the caller to maintain their own extraction-to-frame mapping.
  2. Temporal localization accuracy: the vendors' own stated ceiling has consistently been "second-level." The strictest public evaluation tier in the entire temporal-grounding field stops at IoU=0.95, while frame-level accuracy corresponds to IoU≈0.99–0.997 — no benchmark, loss, or leaderboard points at that range. For scale: on segments averaging 8 seconds, IoU=0.7 permits a 34-frame displacement.
  3. Label grid: mainstream datasets quantize ground truth to 2-second clips, and human annotators disagree on boundaries at the second scale anyway. Evaluating IoU≈0.99 against labels with ±1s of noise is measuring the noise.

So the correct formulation is: a bare model emitting frame-level in/out points = can't (mechanism: second-level localization accuracy plus no frame-number output contract); a tool-equipped agent performing frame-level trims = can, but the frame accuracy is supplied by deterministic tools. Attributing the mechanism to "1fps sampling" invites the wrong rebuttal ("just raise the frame rate, then") — when what actually seals the road is labels and objective functions.

One counterexample is especially persuasive: back in 2017, a SIGGRAPH paper used hand-written film-idiom cost functions and dynamic programming to produce fully automatic, frame-level dialogue-scene editing decisions, with zero neural video understanding anywhere in the loop. Nine years on, the decision layer in shipping products is still hand-written heuristics. The bottleneck was never model intelligence.

6. "AI can't drive editing software" is an expired claim

The most counterintuitive finding of the round is at the execution layer [Narrowed — the draft badly underestimated this]. As of mid-2026, frame-level timeline manipulation is already open to agents:

In other words: AI can already emit a frame-accurate timeline containing J/L cuts today. It just doesn't know whether to offset, or by how much. The bottleneck has moved wholesale to judgment.

And the strongest corroboration is silence: the flagship NLE's latest release lists more than a dozen AI features on its own site — footage search, speech generation, focus adjustment, face age and retouching, slate identification, sharpening, motion deblur — and not one of them touches cut timing. The entire AI budget for that cycle went into retrieval, restoration, beautification, and generation.

Which means that at the fine-cut layer there is currently no gap between marketing and reality — because nobody is marketing it. The vendors aren't even overclaiming. That says more than any benchmark.

7. If you want to build this yourself: a pipeline that runs today

Given all of the above, there's only one viable architecture — and notably, it doesn't require waiting for any model to improve:

  1. Deterministic detection layer (the source of frame accuracy): word-level forced alignment, shot boundary detection, VAD/silence, beat grid, speaker diarization. All open source, mature, cheap.
  2. Symbolic selection layer (the LLM's correct position): turn the previous layer's output into timecoded text candidates and let the model do nothing but select, rank, and segment. An hour of footage as "word-level transcript + shot table" is roughly 12–20k tokens, about 70× cheaper than 1.08M video tokens — and breath, filler words, and talking-head pacing are exactly the decisions that can be made frame-accurately on that cheap path.
  3. Write-back layer: emit an interchange format, or drive the NLE scripting API directly.

Format choice has a counterintuitive conclusion [Narrowed]: don't use OTIO as your delivery format. Its core schema has no audio level automation, no keyframes, no A/V link semantics, and no speed ramps. Anything reducible to a point in time survives losslessly (frame-level in/out points, the J/L offset itself, beat sync), but the audio envelopes and ramps that fine cutting depends on are silently dropped at the adapter boundary. The right use is OTIO as internal intermediate representation (rational time base, small schema, JSON that's easy to generate and diff), with FCPXML generated directly at the delivery layer.

Avoid EDL more strongly still: it does not store the frame rate in the file at all (not an implementation gap — the spec has no such field), and it can't carry a second video track, keyframes, or speed ramps, with only 3-digit event numbers.

The ranking of real-world failure modes is also the reverse of intuition [Narrowed]. By damage, worst first:

  1. Rate semantics errors: treating 23.976 as 24 costs 86 frames per hour; treating drop-frame as non-drop costs 108 frames per hour.
  2. Interval semantics errors: some formats' out is a closed interval, others' end is open — a systematic one-frame error at every single cut. This is the classic cross-format off-by-one.
  3. Variable frame rate footage: phone and screen recordings have no constant frame duration, so "frame accurate" is undefined before conform.
  4. Numeric representation precisiondead last. Microsecond representation error is about 1.5×10⁻⁵ frames; the claim that "accumulating a few hundred cuts drifts to a whole frame" is off by roughly 100×. The real risk is a rounding-mode mismatch when the renderer converts back, not exhausted precision.

Put differently: time-base incompatibility is a deterministic, unit-testable bookkeeping problem the industry solved twenty years ago. It is not the killer. The perception layer's error floor is 100–1000× larger.

8. A few details worth remembering on their own

9. What this research did not cover

Honest labeling matters, especially with the search budget exhausted:

Main sources (all passed three-perspective verification): craft and workflow — Murch, In the Blink of an Eye (lexicographic priority, not additive weights), and the downstream contracts of picture lock, change lists, and conform; computational editing — Leake et al., SIGGRAPH 2017 (frame-level automatic dialogue-scene editing with zero neural video understanding), Berthouzoz et al., SIGGRAPH 2012; learning cut placement — Pardo et al., ICCV 2021, Learning to Cut by Watching Movies (255K cuts; top-1 8.18% at ~8-frame tolerance), MovieCuts, ECCV 2022 (173,967 clips / 10 classes, mean AP 47.9, match cut 2.43, model ladder 1.1→2.3), AVE, ECCV 2022, Netflix Match Cutting, WACV 2023 (candidate retrieval, explicitly leaving frame choice to the editor); evaluation methodology — Otani et al., CVPR 2019 (random summaries score comparably to published methods under multi-reference F1), Otani et al., BMVC 2020 (strong priors in temporal-localization datasets); temporal grounding — QVHighlights (ground truth quantized to 2-second clips), Charades-STA, ActivityNet Captions, the Qwen2.5-VL technical report (second-level event localization); signal layer — CTC forced alignment and phoneme-level aligners (10–20ms), TransNetV2 (BBC 96.2 / RAI 93.9 / ClipShots 77.9), beat tracking (mir_eval ±70ms tolerance), PodcastFillers (non-lexical filler event F1 92.8); sync thresholds — EBU R37 and ITU-R BT.1359-1 (asymmetric audiovisual thresholds); formats and interfaces — the FCPXML DTD (audioStart/audioDuration expressing split edits), the OTIO core schema, the CMX3600 spec, NLE scripting interfaces and the MCP ecosystem. 56 load-bearing conclusions entered three-perspective adversarial verification; all landed on "directionally sound, wording requires correction," and this piece presents the corrected versions, with grades and recency risk labeled throughout.

Amos · research.xishe.ai · Please credit when sharing