All notable changes to this project will be documented in this file.
The format is based on Keep a Changelog, and this project adheres to Semantic Versioning.
- Thinking works with any model that exposes plaintext reasoning, not just
one delta shape: recognized delta fields are now
reasoning,reasoning_text, andthinking, and thinking streamed viamessage.part.updatedsnapshot deltas is forwarded too (with a dedupe guard so backends sending both shapes never emit it twice). Verified live againstmimo-v2.6-flash-free(non-streamingreasoning_content, opt-in streaming deltas, cleancontentby default). - Backend per-model failures that arrive as HTTP 200 with
info.error(e.g.402 Insufficient account funds, found live on non-free models) no longer return an empty 200: non-streaming endpoints answer500 backend_errorwith the upstream message, and streams end with an error frame instead of a clean-but-empty[DONE]. - Long agentic runs no longer look hung: streams now send an SSE heartbeat
comment every
STREAM_HEARTBEAT_MS(default 15000,0disables), so client/proxy idle timeouts don't kill runs while the backend works tools with no text deltas. Verified live (heartbeat interleaved mid-run, clean[DONE]). - Streaming master switch:
ENABLE_STREAMING=0(or--no-streaming;--streamingre-enables) servesstream: truerequests blocking as regular JSON instead of SSE, so strict clients keep working. Forcing streaming onto non-streaming requests is deliberately not offered. - Ops config gaps closed:
HOSTbind address (--host, default keeps0.0.0.0for Docker; use127.0.0.1for localhost-only no-auth use),CORS_ORIGIN(--cors, off by default; preflights bypass auth),VISION_MODEL(--vision_model, was hardcoded),INCLUDE_REASONING(--include_reasoning, default-on thinking for clients that can't send the flag),LOG_LEVEL(--log_level,error|warn|info|debug, request lines without bodies/keys), and a--version/-vflag. - Help for humans and agents: bare
helpsubcommand (node server.js help) plus richer--helpwith endpoints, request-field hints, copy-paste examples, and exit codes.
- Some models (observed:
muse-spark-1.3-contributor-free) keep thinking encrypted server-side (reasoningEncryptedContent, empty text, zero deltas). For those onlyusage.completion_tokens_details.reasoning_tokenscomes through — there is no plaintext for the adapter to forward.
- Thinking output: backend reasoning parts are exposed OpenAI-style as
reasoning_content(DeepSeek/OpenRouter/vLLM convention)- Streaming chat completions forward thinking as
delta.reasoning_contentchunks when the client opts in (include_reasoning: true; also acceptsinclude_thinking,stream_options.include_reasoning,reasoning,reasoning_effort,thinking/enable_thinking). Default still drops thinking socontentstays clean. - Reasoning parts are tracked via
message.part.updatedsnapshots, so thinking never leaks intocontenteven when the backend streams it withfield: "text"on a reasoning part. - Non-streaming chat completions include
message.reasoning_contentwhenever the backend produced thinking (unless explicitly disabled, e.g.reasoning: { effort: "none" }). usage.completion_tokens_details.reasoning_tokenswhen the backend reports reasoning tokens.
- Streaming chat completions forward thinking as
- Thinking request aliases are accepted without error but never forwarded
to the backend: opencode tunes effort via model variants
(
opencode.json), the adapter only gates reasoning output. Effort tuning viatools-style per-message params is not supported byPOST /session/:id/message({ messageID?, model?, agent?, noReply?, system?, tools?, parts }).
- Plain
node server.jsnow reads.env(zero-dependency built-in loader), soPORT(andOPENCODE_URL,API_KEY, timeouts, retries) from.envare honored instead of silently falling back to port 80. Precedence: CLI flags > shell env >.envfile > defaults.
- True token streaming: chat and text completions forward backend
message.part.deltaevents live viaprompt_async+ the/eventbus (replaces buffered char-by-char replay);stream_options.include_usageadds ausageblock to the final chunk; client disconnect aborts the backend run - Conversation memory: opt-in
session_idrequest field pins turns to one backend session (no history replay); responses echo it back; unknown ids return404 session_not_found - Resilience: retries with backoff on network failures and 5xx GETs
(
OPENCODE_RETRIES), configurableOPENCODE_TIMEOUT/OPENCODE_MESSAGE_TIMEOUT - OpenAI-shaped errors everywhere on
/v1/*(401 invalid_api_key,400 invalid_json,404 session_not_found,502 backend_unreachable) - npm binary: global install exposes the
opencode-api-nodecommand
- OpenAI-compatible proxy for
opencode serve(server.js, Express, zero runtime deps besidesexpress) - Endpoints:
GET /,GET /health,GET /v1/models,POST /v1/chat/completions,POST /v1/completions - SSE character-by-character streaming for chat and text completions
- Multi-turn context via
noReplyhistory replay,systemprompt mapping,provider/modelsplitting, vision-model forcing withimage_urlinlining - Optional bearer auth (
API_KEY), backend selection (OPENCODE_URL), port selection (PORT/--port) --session [<id>]CLI action to dump backend session data as JSON- Docker image + compose setup with healthcheck
- Test suite with mock backend (
npm test) - MIT license