Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
7 changes: 7 additions & 0 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -93,3 +93,10 @@ also has `camkit <command> --help`.
kept ranges with `camkit silences` before finalizing a cut.
- After a rebuild the project is already cut and `.bak` holds the original;
to recut, restore the `.bak` first or you'll back up the cut file.

### Separate-media rough cuts

For editable screen/camera/audio cuts using converted MP4/WAV media, see
[the separate-media workflow](docs/separate-media.md). Includes synchronized
source groups, conversion into a new project, and validation that refuses
unsupported existing edits.
83 changes: 83 additions & 0 deletions docs/separate-media.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,83 @@
# Editable cuts with separate screen, camera and audio sources

`convert-media` creates a new `.cmproj` with one editable MP4/WAV track for
each placed recording stream. It never edits the input project, Camtasia
preferences, presets or proxy cache. Keep an independent backup of the original
bundle before a production edit.

```sh
camkit convert-media --project original.cmproj --out working.cmproj --dry-run
camkit convert-media --project original.cmproj --out working.cmproj --acknowledge-cursor-loss
```

Requirements: ffmpeg and ffprobe, H.264 encoding, and a decoder for the source
recording (including tscc2 for screen recordings). Conversion can take minutes
and uses additional disk space. MP4 conversion loses native cursor metadata,
including editable cursor enlargement/highlighting; the decoded screen video
may omit the cursor itself. It retains static media scale, crop, position,
audio gain and track visibility/mute settings. It includes only the recording
placed on the timeline, not unused TREC entries in the media bin.

The initial supported input is a single uncut recording with synchronized clips
starting at zero. Effects, animations, speed changes, mixed-source UnifiedMedia,
multiple clips of a source on a track and other unsupported structures fail
before conversion. Source streams are checked against the probe's types,
dimensions and durations. Nonzero audio start times are padded with silence;
nonzero video start times are currently refused. The last decoded video frame
may be held for up to one second to cover a TREC tail whose declared duration
exceeds its decoded frames; output is trimmed to the original timeline duration. All output streams are checked
for zero start time and duration before the project is written. Output must not
already exist. A failed conversion removes only the output directory it created.

## One cut plan for multiple sources

Conversion writes `working.cmproj/sync-groups.json`, for example:

```json
[{"sources":[{"src":10,"offset":0},{"src":12,"offset":0},{"src":14,"offset":0}]}]
```

The first source is the reference clock. `offset` means **member source time
minus reference source time**, in seconds. For manually synchronized files, a
member with `offset: 0.25` uses source time 10.25 when the reference uses 10.
The group's sources must occupy distinct tracks and each source can belong to
only one group. Supply offsets explicitly; camkit does not infer synchronization.

Use the first source ID in the keep list (use the IDs in the generated file,
not the example numbers):

```json
{"keep":[{"src":10,"start":1,"end":12},{"src":10,"start":14,"end":20}]}
```

```sh
camkit rebuild --project working.cmproj --from keep.json \
--sync-groups working.cmproj/sync-groups.json --dry-run
camkit rebuild --project working.cmproj --from keep.json \
--sync-groups working.cmproj/sync-groups.json
```

Cut boundaries are rounded once to reference video frames; offsets are rounded
to project units. Every member gets exactly the same timeline start and duration,
with globally unique clip IDs. All member source bounds are checked both before
and after rounding. Dry-run entries report the actual rounded reference times.

Rebuild remains a source-time rough-cut operation, not a general retiming engine
for finished projects. It now refuses repeated source clips on a track and
unsupported effects/keyframes instead of cloning only the first clip and silently
losing later edits. Use the uncut original for a new cut plan. Existing lock and
`.bak` protections still apply; `--force` does not bypass timeline validation.

Transcribe separately, select complete takes, check real silences, then supply an
explicit keep plan. Local Whisper uses full JSON for token offsets and omits
reversed token timestamps from word results while retaining segment text.
Transcripts can still have inaccurate timestamps and hallucinations during long
silences; do not treat every inferred word boundary as an exact audio cut point.

## Validation limits

This workflow avoids native TREC references in the generated project, motivated
by [issue #17](https://github.com/Orva-Studio/camkit/issues/17). It does not fix or
prove the absence of a Camtasia decoder bug. Validate playback, save/reopen and
native export on a project copy before adopting it for production. Keep the
original TREC for cursor features and further editing.
21 changes: 19 additions & 2 deletions packages/cli/src/camkit.ts
Original file line number Diff line number Diff line change
Expand Up @@ -38,6 +38,7 @@ import { camtasiaDocPaths, closeProject, exportVideo, openProject, projectStatus
import { exportAudio, runSilencedetect, transcribeRecording } from "./media.ts";
import { mediaProxiesDir, planPrune, proxyKey, type ProxyEntry } from "./proxies.ts";
import { listPresets, resolvePreset } from "./presets.ts";
import { convertMedia } from "./convert.ts";
import { version } from "../package.json";

/** Load for read-only commands: --project, else the ./search.cmproj default,
Expand Down Expand Up @@ -91,8 +92,12 @@ const HELP: Record<string, { usage: string; about: string[] }> = {
"including unplaced takes).",
],
},
"convert-media": {
usage: "camkit convert-media --project PATH --out NEW.cmproj [--dry-run] [--acknowledge-cursor-loss]",
about: ["Convert one uncut TREC into separate editable MP4/WAV tracks in a new bundle.", "Requires ffmpeg/ffprobe with a decoder for the recording. Preserves static framing and gain.", "Drops native cursor data and unused media-bin items; refuses unsupported existing edits.", "Writes sync-groups.json for rebuild --sync-groups; never changes the input project or app settings."],
},
rebuild: {
usage: 'camkit rebuild [--project PATH] --keep "SRC:start-end ..." | --from FILE [--dry-run] [--force]',
usage: 'camkit rebuild [--project PATH] --keep "SRC:start-end ..." | --from FILE [--sync-groups FILE] [--dry-run] [--force]',
about: [
"The core rough-cut op. Rewrites the timeline to keep only the listed",
"source segments, in order, ripple-laid with no gaps (seconds → editRate",
Expand All @@ -101,6 +106,9 @@ const HELP: Record<string, { usage: string; about: string[] }> = {
"",
' --keep "1:159.8-179.2 2:46.3-60.0" keep src-1 159.8–179.2s, then src-2 46.3–60.0s',
" --from FILE JSON [{src,start,end}] or {keep:[...]}",
" --sync-groups FILE JSON [{sources:[{src,offset}]}]; keep uses each group’s first source",
"Offsets are source seconds minus reference seconds. All members share cut boundaries.",
"Refuses already-cut clips, effects/keyframes and unsupported media to protect existing edits.",
" --dry-run print the plan, write nothing (ALWAYS do this first)",
" --force override a stale ~project.tscproj lock / overwrite an existing .bak",
"",
Expand Down Expand Up @@ -376,12 +384,16 @@ function cmdRebuild(argv: string[]) {
segs = parseKeep(keep);
}

const plan = planRebuild(doc, segs);
const syncFile = flag(argv, "--sync-groups");
const plan = planRebuild(doc, segs, { syncGroups: syncFile ? JSON.parse(readFileSync(resolve(syncFile), "utf8")) : undefined });
console.log(`rebuild plan (${plan.segmentCount} segments, ${plan.totalSeconds.toFixed(1)}s total):`);
for (const e of plan.entries) {
console.log(
` src ${e.src} ${e.sourceStart.toFixed(2)}-${e.sourceEnd.toFixed(2)}s → timeline ${e.timelineStart.toFixed(2)}-${e.timelineEnd.toFixed(2)}s (${e.trackCount} track[s])`,
);
for (const member of e.sources.slice(1)) {
console.log(` synced src ${member.src}: ${member.sourceStart.toFixed(3)}-${member.sourceEnd.toFixed(3)}s`);
}
}
if (dryRun) {
console.log("\n--dry-run: no files written.");
Expand Down Expand Up @@ -751,6 +763,11 @@ const COMMANDS: Record<string, (argv: string[]) => void | Promise<void>> = {
clips: cmdClips,
sources: cmdSources,
rebuild: cmdRebuild,
"convert-media": async (argv) => {
const project = flag(argv, "--project"), out = flag(argv, "--out");
if (!project || !out) throw new Error("convert-media needs --project PATH --out NEW.cmproj");
console.log(JSON.stringify(await convertMedia({ project, out, dryRun: has(argv, "--dry-run"), acknowledgeCursorLoss: has(argv, "--acknowledge-cursor-loss") }), null, 2));
},
"export-audio": cmdExportAudio,
"export-video": cmdExportVideo,
captions: cmdCaptions,
Expand Down
213 changes: 213 additions & 0 deletions packages/cli/src/convert.ts
Original file line number Diff line number Diff line change
@@ -0,0 +1,213 @@
import { spawn } from "node:child_process";
import { mkdirSync, existsSync, writeFileSync, rmSync } from "node:fs";
import { dirname, resolve, join } from "node:path";
import {
loadProject,
planSeparateMedia,
type ConvertedStream,
} from "@camkit/core";

function run(command: string, args: string[]): Promise<string> {
return new Promise((res, rej) => {
const p = spawn(command, args, { stdio: ["ignore", "pipe", "pipe"] });
let out = "",
err = "";
p.stdout.on("data", (d) => (out += d));
p.stderr.on("data", (d) => (err = (err + d).slice(-12000)));
p.on("error", rej);
p.on("close", (code) =>
code === 0
? res(out)
: rej(new Error(`${command} exited ${code}: ${err}`)),
);
});
}
export async function convertMedia(opts: {
project: string;
out: string;
dryRun?: boolean;
acknowledgeCursorLoss?: boolean;
}) {
const { path, doc } = loadProject(opts.project);
const out = resolve(opts.out);
if (!out.endsWith(".cmproj"))
throw new Error("--out must be a new .cmproj directory.");
if (existsSync(out)) throw new Error(`Output already exists: ${out}`);
// First validate the project without invoking ffmpeg or creating any output.
// Only stream metadata belonging to the placed source is used (bin can contain unused recordings).
const clips =
doc.timeline?.sceneTrack?.scenes?.[0]?.csml?.tracks?.flatMap(
(t: any) => t.medias ?? [],
) ?? [];
const first = clips[0];
const src = first?._type === "UnifiedMedia" ? first.video?.src : first?.src;
const bin = (doc.sourceBin ?? []).find((s: any) => s.id === src);
const preliminary = planSeparateMedia(
doc,
(bin?.sourceTracks ?? []).map((t: any, n: number) => ({
trackNumber: n,
file: `./media/stream-${n}.${t.type === 2 ? "wav" : "mp4"}`,
kind: t.type === 2 ? "audio" : "video",
width: t.trackRect?.[2],
height: t.trackRect?.[3],
sampleRate: Number(t.sampleRate),
channels: t.numChannels,
})),
);
const input = resolve(dirname(path), preliminary.sourcePath);
const probe = JSON.parse(
await run("ffprobe", [
"-v",
"error",
"-show_streams",
"-of",
"json",
input,
]),
);
const converted: ConvertedStream[] = [];
const args = [
"-hide_banner",
"-loglevel",
"warning",
"-nostdin",
"-n",
"-copyts",
"-i",
input,
];
const used = preliminary.doc.sourceBin.map((s: any) =>
Number(s.src.match(/stream-(\d+)/)[1]),
);
for (const n of used) {
const st = probe.streams?.find((s: any) => s.index === n);
const sourceTrack = bin.sourceTracks[n];
const audio = sourceTrack.type === 2;
if (!st || st.codec_type !== (audio ? "audio" : "video"))
throw new Error(`Cannot map TREC track ${n} to ffmpeg stream safely.`);
const start = Number(st.start_time),
duration = Number(st.duration);
if (
!Number.isFinite(start) ||
!Number.isFinite(duration) ||
start < 0 ||
start >= preliminary.durationSeconds ||
duration + start < preliminary.durationSeconds
)
throw new Error(`Stream ${n} does not cover the requested duration.`);
if (
!audio &&
(start !== 0 ||
st.width !== sourceTrack.trackRect?.[2] ||
st.height !== sourceTrack.trackRect?.[3])
)
throw new Error(`Unsupported video timing or dimensions on stream ${n}.`);
const s: ConvertedStream = {
trackNumber: n,
file: `./media/stream-${n}.${audio ? "wav" : "mp4"}`,
kind: audio ? "audio" : "video",
width: st.width,
height: st.height,
sampleRate: Number(st.sample_rate),
channels: st.channels,
};
converted.push(s);
args.push("-map", `0:${n}`, "-t", String(preliminary.durationSeconds));
if (audio)
args.push("-af", "aresample=async=1:first_pts=0", "-c:a", "pcm_s16le");
else
args.push(
"-vf",
`fps=${doc.videoFormatFrameRate}:start_time=0,tpad=stop_mode=clone:stop_duration=1`,
"-c:v",
"libx264",
"-preset",
"ultrafast",
"-crf",
"18",
"-pix_fmt",
"yuv420p",
"-an",
"-movflags",
"+faststart",
);
args.push(resolve(out, s.file));
}
const plan = planSeparateMedia(doc, converted);
if (opts.dryRun)
return {
out,
durationSeconds: plan.durationSeconds,
syncGroups: plan.syncGroups,
warnings: plan.warnings,
ffmpeg: args,
};
if (!opts.acknowledgeCursorLoss)
throw new Error(
"Conversion loses native cursor data; pass --acknowledge-cursor-loss after reviewing --dry-run.",
);
// Exclusive mkdir: never overwrite an existing bundle, even when two invocations race.
mkdirSync(out);
try {
mkdirSync(join(out, "media"));
await run("ffmpeg", args);
for (const s of converted) {
const check = JSON.parse(
await run("ffprobe", [
"-v",
"error",
"-show_streams",
"-of",
"json",
resolve(out, s.file),
]),
);
const st = check.streams?.[0];
const tolerance =
s.kind === "audio" ? 1 / s.sampleRate! : 1 / doc.videoFormatFrameRate;
if (
check.streams?.length !== 1 ||
st.codec_type !== s.kind ||
!Number.isFinite(Number(st.duration)) ||
Math.abs(Number(st.duration) - plan.durationSeconds) >
tolerance + 1e-6 ||
Math.abs(Number(st.start_time ?? 0)) > 1e-6
)
throw new Error(
`Converted stream failed duration/start validation: ${s.file} (start=${st?.start_time}, duration=${st?.duration}, expected=${plan.durationSeconds})`,
);
}
writeFileSync(
join(out, "conversion.json"),
JSON.stringify(
{
sourceProject: path,
sourceMedia: input,
durationSeconds: plan.durationSeconds,
syncGroups: plan.syncGroups,
warnings: plan.warnings,
},
null,
2,
),
);
writeFileSync(
join(out, "sync-groups.json"),
JSON.stringify(plan.syncGroups, null, 2),
);
// Write the project last so a failed conversion never looks like a usable project.
writeFileSync(
join(out, "project.tscproj"),
JSON.stringify(plan.doc, null, 2),
);
return {
out,
durationSeconds: plan.durationSeconds,
syncGroups: plan.syncGroups,
warnings: plan.warnings,
};
} catch (e) {
rmSync(out, { recursive: true, force: true }); // only the directory exclusively created by this invocation
throw e;
}
}
9 changes: 6 additions & 3 deletions packages/cli/src/media.ts
Original file line number Diff line number Diff line change
Expand Up @@ -105,8 +105,8 @@ async function callWhisper(audioPath: string, model: string): Promise<any> {
*/
async function callWhisperCpp(audioPath: string, modelPath: string): Promise<any> {
const outBase = audioPath.replace(/\.wav$/, "");
// -oj writes <outBase>.json with token offsets; -np keeps stdout quiet.
await run(WHISPER_BIN, ["-m", modelPath, "-f", audioPath, "-oj", "-of", outBase, "-np"]);
// Full JSON is required: ordinary -oj omits the per-token word offsets.
await run(WHISPER_BIN, ["-m", modelPath, "-f", audioPath, "-ojf", "-of", outBase, "-np"]);
const jsonPath = `${outBase}.json`;
const data = JSON.parse(await readFile(jsonPath, "utf8"));
await unlink(jsonPath).catch(() => {});
Expand Down Expand Up @@ -148,7 +148,10 @@ export function shapeWhisperCpp(data: any): any {
}
const lastEnd = segments.length ? segments[segments.length - 1].end : null;
const text = segments.map((s) => s.text).join("").trim();
return { duration: lastEnd, text, words, segments };
// Some whisper.cpp versions emit reversed token offsets after long pauses.
// Keep the segment text, but never expose these as usable word cut points.
const validWords = words.filter(w => Number.isFinite(w.start) && Number.isFinite(w.end) && w.start >= 0 && w.end >= w.start);
return { duration: lastEnd, text, words: validWords, segments };
}

export type Engine = "openai" | "whisper-cpp" | "replicate";
Expand Down
Loading