Skip to content

feat(ai-models): add opt-in TwelveLabs Pegasus caption provider - #246

Open
mohit-twelvelabs wants to merge 1 commit into
ncounterspecialist:mainfrom
mohit-twelvelabs:feat/twelvelabs-integration
Open

mohit-twelvelabs wants to merge 1 commit into
ncounterspecialist:mainfrom
mohit-twelvelabs:feat/twelvelabs-integration

Conversation

@mohit-twelvelabs

Copy link
Copy Markdown

Hi! I'm Mohit, I work at TwelveLabs (@mohit-twelvelabs).

What this adds

An opt-in TwelveLabsAdapter in @twick/ai-models that generates time-aligned captions directly from a video URL using TwelveLabs Pegasus video understanding — no separate transcription step. It implements the existing ProviderAdapter contract (the package README already invites provider implementations like Sora/HeyGen/ElevenLabs), so it slots into ProviderRegistry / GenerationOrchestrator exactly like any other provider and emits the standard TimelineCaptionPatch.

import { ProviderRegistry, GenerationOrchestrator, TwelveLabsAdapter } from "@twick/ai-models";

const registry = new ProviderRegistry();
registry.registerAdapter(new TwelveLabsAdapter());
registry.setProviderConfig({ provider: "twelvelabs", apiKey: process.env.TWELVELABS_API_KEY });

const job = await new GenerationOrchestrator(registry).createJob({
  type: "caption",
  provider: "twelvelabs",
  input: { videoUrl: "https://example.com/clip.mp4", language: "en" },
});

Under the hood it creates a Pegasus analyze task (POST /analyze/tasks, time_based_metadata mode) on startJob, and the orchestrator polls getJobStatus (GET /analyze/tasks/{id}) until the task is ready, then maps each segment to a TimedTextSegment.

Why it helps this project

Captioning is a core editor workflow. Pegasus produces semantically-aware, time-aligned captions straight from a video URL, giving Twick apps an AI caption option behind the same provider-agnostic seam — and the orchestrator's built-in fallback means it can sit alongside other providers without lock-in.

Opt-in / non-breaking

  • Adds "twelvelabs" to the AIModelProvider union and a new src/providers/twelvelabs-adapter.ts; touches no existing provider, default, or runtime path.
  • No new runtime dependencies — uses native fetch (Node 20+), consistent with the rest of the repo.

How it was tested

  • No-network unit test (twelvelabs-adapter.test.ts, node:test, zero new deps) covering caption parsing, metadata fallback, and malformed input — passing (3/3).
  • pnpm --filter @twick/ai-models build (tsc typecheck + vite lib build + dts rollup) — passing.
  • Live end-to-end against the real API: created a real Pegasus analyze task and polled it to ready. The returned result.data ({"caption":[{"start_time":0,"end_time":10,"metadata":{"text":"..."}}]}) parses correctly into { text, startMs: 0, endMs: 10000 }. Auth, both routes, request payload, and status transitions (queued → processing → ready) all verified.

A changeset (minor for @twick/ai-models) is included.

Per CONTRIBUTING, significant changes are encouraged to start with an issue — happy to open one or adjust the approach (segmentation prompt, sync vs async, etc.) to match maintainer preferences. You can grab a free API key at https://twelvelabs.io — there's a generous free tier.

Adds TwelveLabsAdapter implementing the ProviderAdapter contract to
generate time-aligned captions directly from a video URL via the
TwelveLabs Pegasus analyze API (time_based_metadata mode). Registered
like any other provider, so defaults and existing behavior are
unchanged. Includes a no-network unit test for caption parsing.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant