feat(ai-models): add opt-in TwelveLabs Pegasus caption provider - #246
Open
mohit-twelvelabs wants to merge 1 commit into
Open
mohit-twelvelabs wants to merge 1 commit into
mohit-twelvelabs wants to merge 1 commit into
Conversation
Adds TwelveLabsAdapter implementing the ProviderAdapter contract to generate time-aligned captions directly from a video URL via the TwelveLabs Pegasus analyze API (time_based_metadata mode). Registered like any other provider, so defaults and existing behavior are unchanged. Includes a no-network unit test for caption parsing.
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Hi! I'm Mohit, I work at TwelveLabs (@mohit-twelvelabs).
What this adds
An opt-in
TwelveLabsAdapterin@twick/ai-modelsthat generates time-aligned captions directly from a video URL using TwelveLabs Pegasus video understanding — no separate transcription step. It implements the existingProviderAdaptercontract (the package README already invites provider implementations like Sora/HeyGen/ElevenLabs), so it slots intoProviderRegistry/GenerationOrchestratorexactly like any other provider and emits the standardTimelineCaptionPatch.Under the hood it creates a Pegasus analyze task (
POST /analyze/tasks,time_based_metadatamode) onstartJob, and the orchestrator pollsgetJobStatus(GET /analyze/tasks/{id}) until the task isready, then maps each segment to aTimedTextSegment.Why it helps this project
Captioning is a core editor workflow. Pegasus produces semantically-aware, time-aligned captions straight from a video URL, giving Twick apps an AI caption option behind the same provider-agnostic seam — and the orchestrator's built-in fallback means it can sit alongside other providers without lock-in.
Opt-in / non-breaking
"twelvelabs"to theAIModelProviderunion and a newsrc/providers/twelvelabs-adapter.ts; touches no existing provider, default, or runtime path.fetch(Node 20+), consistent with the rest of the repo.How it was tested
twelvelabs-adapter.test.ts,node:test, zero new deps) covering caption parsing, metadata fallback, and malformed input — passing (3/3).pnpm --filter @twick/ai-models build(tsc typecheck + vite lib build + dts rollup) — passing.ready. The returnedresult.data({"caption":[{"start_time":0,"end_time":10,"metadata":{"text":"..."}}]}) parses correctly into{ text, startMs: 0, endMs: 10000 }. Auth, both routes, request payload, and status transitions (queued → processing → ready) all verified.A changeset (
minorfor@twick/ai-models) is included.Per CONTRIBUTING, significant changes are encouraged to start with an issue — happy to open one or adjust the approach (segmentation prompt, sync vs async, etc.) to match maintainer preferences. You can grab a free API key at https://twelvelabs.io — there's a generous free tier.