You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Status: Draft — April 2026 Scope: Sufficient to direct a clean-room implementation of the GitHub Copilot inference protocol in any language. Precondition: The implementer has working Anthropic Messages API and OpenAI Chat Completions + Responses API client implementations available. This document covers only what is specific to GitHub Copilot: authentication, identity headers, model discovery, protocol routing, and known Copilot-layer quirks. Wire format, streaming event shapes, and request body construction for each protocol are out of scope — refer to the respective vendor specifications.
A developer wishes to call GitHub Copilot's inference backend from a program that is not an officially supported Copilot client (VS Code, JetBrains, Copilot CLI). The GitHub Copilot inference API is not publicly documented by GitHub. Its protocol has been established through reverse engineering by the open-source community and is used by tools such as pi-agent (badlogic/pi-mono), ericc-ch/copilot-api, LiteLLM, and opencode.
This document consolidates that prior art into a single implementer-facing reference, scoped to the Copilot-specific layer on top of the underlying vendor APIs.
2. Architecture Overview
flowchart TD
PRE["PRECONDITION\nOAuth implementation available\n(RFC 8628 device flow or PAT — not specified here)"]
subgraph DOC ["THIS DOCUMENT"]
subgraph AUTH ["Authentication Layer"]
A1["Phase 1: OAuth Device Flow\ngithub.com/login/device/code"]
A2["Yields: gho_* OAuth token"]
A3["Phase 2: Copilot Token Exchange\napi.github.com/copilot_internal/v2/token"]
A4["Yields: short-lived JWT"]
A1 --> A2 --> A3 --> A4
end
subgraph ROUTE ["Routing Layer"]
R1["GET /models\n→ inspect supported_endpoints\n→ select protocol\n→ inject identity headers\n→ set base URL"]
end
A4 -->|"Bearer copilot-jwt"| R1
end
SDK_A["Anthropic SDK\n/v1/messages\n(Claude)"]
SDK_O["OpenAI SDK\n/v1/chat/completions\n/v1/responses\n(GPT, Codex)"]
PRE --> AUTH
R1 --> SDK_A
R1 --> SDK_O
Loading
Both the authentication layer and the routing layer are specified by this document. Below the routing layer, standard vendor SDK implementations are used without modification — except for base URL override and header injection.
All inference requests carry identity headers that identify the client as a VSCode instance. These are required by the Copilot API gateway; requests lacking them are rejected. [REF-5, REF-6]
3. Authentication Protocol
3.1 OAuth Device Flow — Phase 1
The Copilot API only accepts OAuth tokens issued to the VSCode OAuth app. The VSCode app's public client ID is Iv1.b507a08c87ecfe98. This ID is hardcoded in official Copilot extensions and is the same for all users. [REF-4, REF-13]
Some model families (GPT-5.x Codex variants) are gated to the VSCode client ID specifically, so using this ID rather than any other app registration is important for full model coverage. [REF-13]
[REF-11] (live example of this exact response shape)
Display verification_uri and user_code to the user. The device code expires in expires_in seconds (≈ 15 minutes).
Step 2 — Poll for the access token
Poll at the interval rate (default 5 seconds). GitHub may return a slow_down error, in which case the polling interval must be increased by 5 seconds and the next poll must wait for the new interval before firing. [REF-22] (pi-mono CHANGELOG: "Fixed GitHub Copilot device-code login polling to respect OAuth slow-down intervals, wait before the first token poll")
POST https://github.com/login/oauth/access_token
Content-Type: application/json
Accept: application/json
{
"client_id": "Iv1.b507a08c87ecfe98",
"device_code": "<device_code from step 1>",
"grant_type": "urn:ietf:params:oauth:grant-type:device_code"
}
While pending:{"error": "authorization_pending"} Slow-down signal:{"error": "slow_down"} — increase interval by 5s before next poll On approval:
Key fields [REF-12] (VS Studio debug log output), [REF-6]:
Field
Type
Description
token
string
Short-lived JWT. Use as Authorization: Bearer <token> on all inference and model discovery requests
expires_at
ISO 8601
Absolute expiry timestamp
refresh_in
integer (seconds)
Suggested time before proactive refresh; typically 1500 (25 min)
sku (encoded in token)
string
Subscription tier, e.g. monthly_subscriber_quota, free_educational_quota
chat_enabled (encoded in token)
bool
Whether chat is enabled for this account
The endpoints object may also be present in GHE configurations; see §3.5 and §8.
3.3 Token Lifecycle and Refresh
sequenceDiagram
participant App
participant GHOAuth as GitHub OAuth
participant CopilotAPI as Copilot API
participant VendorSDK as Vendor SDK
App->>GHOAuth: POST /login/device/code
GHOAuth-->>App: {device_code, user_code}
Note over App: Display user_code to user
loop Poll until approved
App->>GHOAuth: POST /login/oauth/access_token
GHOAuth-->>App: authorization_pending / slow_down / gho_token
end
Note over App: Persist gho_token
App->>CopilotAPI: GET /copilot_internal/v2/token
CopilotAPI-->>App: {token, expires_at, refresh_in}
App->>VendorSDK: Inference request (via vendor SDK)
VendorSDK-->>App: Response / Stream
Note over App: After refresh_in seconds...
App->>CopilotAPI: GET /copilot_internal/v2/token
CopilotAPI-->>App: {new_token, expires_at, refresh_in}
Loading
Expiry tolerance: Treat tokens as expired 60 seconds before their stated expires_at. In WSL2 and virtual machine environments, clock drift can cause premature expiry errors — the 60-second buffer mitigates this. [REF-22]
On 401 during inference: Refresh the Copilot token immediately using the stored gho_ token and retry the request exactly once.
3.4 Alternative: PAT Authentication
Fine-grained Personal Access Tokens (PATs) with the Copilot Requests permission can substitute for the gho_ OAuth token in the Phase 2 exchange. [REF-10] (GitHub Copilot CLI docs)
PATs are suitable for non-interactive and CI/CD environments where device flow is not feasible. The token format is github_pat_ rather than gho_.
Known limitation: PAT support in third-party implementations is inconsistent. Pi-mono's provider only accepts gho_-prefixed OAuth tokens [REF-14]. LiteLLM and the official Copilot CLI both accept PATs via environment variable. [REF-10]
3.5 GitHub Enterprise Server (GHE) Variants
For GitHub Enterprise Server instances, the auth endpoints change [REF-15]:
Model ID to pass in the model field of inference requests
supported_endpoints
Determines which protocol client to use — see §5
model_picker_enabled
false if the current plan SKU cannot use this model; filter these out before presenting to users
capabilities.limits
Context window and output token limits; use when configuring SDK parameters
Fetch timing: Fetch on login and on each Copilot token refresh. Do not hardcode model IDs; the available set changes as GitHub rolls out new models. [REF-16]
Model activation prerequisite: Many models must be explicitly enabled by the user in VS Code before they become usable via the API, even with a valid subscription and model_picker_enabled: true. A model not supported error for such a model means the user needs to activate it via the VS Code model picker. There is no programmatic activation path. [REF-16]
5. Protocol Routing
This is the central decision the Copilot routing layer makes. Based on the supported_endpoints field for the selected model, the routing layer configures the appropriate vendor SDK client with the Copilot base URL and identity headers, then delegates.
supported_endpoints includes "/v1/messages"
→ Use Anthropic SDK
→ base_url: https://api.githubcopilot.com (or endpoints.proxy)
→ extra headers: see §6
supported_endpoints includes "/v1/chat/completions" (but not "/v1/messages")
→ Use OpenAI SDK, chat completions mode
→ base_url: https://api.githubcopilot.com
supported_endpoints includes "/v1/responses"
→ Use OpenAI SDK, responses mode
→ base_url: https://api.githubcopilot.com
[REF-16]: "Infer API type from supported_endpoints — /v1/messages maps to Anthropic protocol, /responses maps to OpenAI Responses API"
Priority when multiple endpoints are listed: Prefer /v1/messages over /v1/chat/completions for Claude models. Routing Claude via Chat Completions causes a known conversation history replay bug where the model re-answers all previous prompts with each new message. [REF-18] (pi-mono issue #209)
Routing decision pseudocode:
func protocol_for(model):
eps = model.supported_endpoints
if "/v1/messages" in eps:
return ANTHROPIC
if "/v1/responses" in eps:
return OPENAI_RESPONSES
return OPENAI_CHAT_COMPLETIONS
6. Request Headers Reference
These headers must be injected into every request to a Copilot base URL — including GET /models and all inference calls. They identify the client to GitHub's API gateway. Missing or incorrect headers result in 401 or 403 responses. [REF-4, REF-5]
Header
Value
Notes
Authorization
Bearer <copilot-token>
The Phase 2 JWT — not the gho_ token
Copilot-Integration-Id
vscode-chat
Required. [REF-4, REF-5]
Editor-Version
vscode/1.104.1
Should reflect a real recent VSCode release
Editor-Plugin-Version
copilot-chat/0.26.7
Copilot Chat plugin version
User-Agent
GitHubCopilotChat/0.26.7
Must not be a generic HTTP client string
For the Anthropic path only, also include:
Header
Value
Notes
Anthropic-Version
2023-06-01
Required by the Anthropic Messages protocol
On header freshness: The VSCode version string should track a real recent release. GitHub has previously rejected requests from very stale version strings, though enforcement is inconsistent. The values above reflect current production usage across community implementations.
Most vendor SDKs support injecting arbitrary extra headers and overriding the base URL via constructor options. Use those mechanisms rather than reimplementing the HTTP layer.
7. Error Handling
HTTP Status
Meaning
Recovery
401 Unauthorized
Copilot token expired
Refresh Phase 2 token using stored gho_ token; retry request once
403 Forbidden
No subscription, model disabled for plan, or transient outage
Check model_picker_enabled; on outage, retry with backoff
429 Too Many Requests
Rate limit
Exponential backoff; respect Retry-After header if present
400 Bad Request
Malformed request
Check model-specific constraints (see §9)
503 Service Unavailable
Transient backend issue
Retry with backoff
[REF-7] (GitHub community 403 discussion), [REF-12] (VS Studio logs distinguishing client vs. server 403)
Distinguishing 401 types: A 401 on inference after a recently issued token is transient. A 401 immediately after a token refresh attempt means the gho_ token has been revoked — re-run device flow.
Distinguishing 403 types: The token exchange response body carries HasToken, ChatEnabled, and ExpiresAt fields. A 403 with HasToken: False indicates no subscription. A 403 with HasToken: True is likely a transient outage. [REF-12]
8. Per-Subscription Base URLs
The inference base URL varies by Copilot plan [REF-1]:
Plan
Base URL
Individual / Pro / Pro+
https://api.githubcopilot.com
Business
https://api.business.githubcopilot.com
Enterprise
https://api.enterprise.githubcopilot.com
When endpoints.proxy is present in the token exchange response (GHE or certain enterprise configurations), that value takes precedence. [REF-15]
9. Known Behavioural Quirks
These are Copilot-layer constraints that affect how you configure the underlying vendor SDK client, not the protocol itself.
Opus extended thinking: effort: "medium" only
Claude Opus 4.7 via Copilot currently only accepts effort: "medium" for extended thinking. Sending effort: "high" causes an error. When configuring the Anthropic SDK with extended thinking for an Opus model routed through Copilot, cap effort at "medium". [REF-19] (pi-mono issues #3291, #3438)
Codex call_id length limit
The Responses API via Copilot enforces a 64-character maximum on call_id fields in tool results. When replaying tool call history into a Codex model session (e.g. after a model switch mid-session), truncate call_id values to 64 characters before dispatch. Failure produces 400 errors. [REF-20] (claude-code-router gist)
Cross-protocol history translation on model switch
When switching a mid-session conversation from a Claude model (Anthropic format) to a GPT model (OpenAI format) or vice versa, the message history must be translated between formats. Tool call and tool result objects have different shapes in each protocol and cannot be passed to the wrong endpoint. This is a session management concern but is a common source of 400 errors in Copilot-backed agents. [REF-20]
Clock drift on token expiry
The Copilot JWT contains an exp field (Unix epoch). In WSL2 and VM environments, clock drift causes premature 401 errors. Apply a 60-second buffer: treat tokens as expired at expires_at - 60s. [REF-22]
10. Module Design
The Copilot-specific implementation separates into three focused modules. The vendor SDK clients (Anthropic, OpenAI) are dependencies, not internal modules.
supported_endpoints inspection for protocol selection
Expose capabilities.limits for caller use in SDK configuration
copilot_client module
This is the thin routing and integration layer — it is not a protocol implementation. It configures the vendor SDK clients with Copilot-specific settings and delegates all protocol work to them.
Interface:
complete(model, request) → Response | Stream
Responsibilities:
Look up protocol via models.protocol_for(model)
Select base URL (§8) or endpoints.proxy from token response
Inject identity headers into the chosen vendor SDK client
Dispatch to the appropriate vendor SDK
On 401: token_manager.invalidate(), refresh, retry once
The auth module shall exchange a valid gho_ token for a short-lived Copilot JWT by calling GET /copilot_internal/v2/token.
When polling for an OAuth access token, the auth module shall respect the slow_down error by increasing the polling interval by five seconds before the next poll.
While a Copilot token is active, the token manager shall schedule a proactive refresh at refresh_in seconds after issuance.
If the Copilot token exchange returns 403, then the auth module shall not retry automatically and shall surface the error to the caller.
When storing OAuth credentials to disk, the auth module shall set file permissions to 0600.
If an inference request returns 401, then the copilot_client module shall refresh the Copilot token via the token manager and retry the request exactly once before surfacing the error.
Model Discovery
The models module shall fetch GET /models with all required identity headers on first use and on each Copilot token refresh.
The models module shall exclude models where model_picker_enabled is false from the set returned by usable_models().
When determining the inference protocol, the models module shall return ANTHROPIC for any model whose supported_endpoints includes /v1/messages.
When determining the inference protocol, the models module shall return OPENAI_RESPONSES for any model whose supported_endpoints includes /v1/responses and does not include /v1/messages.
Routing and Header Injection
The copilot_client module shall include Copilot-Integration-Id: vscode-chat, Editor-Version, Editor-Plugin-Version, and User-Agent on every request to a Copilot base URL.
Where the Copilot token exchange response includes an endpoints.proxy field, the copilot_client module shall use that value as the inference base URL.
Where a GitHub Enterprise Server host is configured, the auth module shall use https://{ghe_host}/api/v3/copilot_internal/v2/token for the token exchange.
Model-Specific Constraints
Where an Opus model is configured with extended thinking, the copilot_client module shall set effort to "medium" regardless of any higher value requested by the caller.
Where a Codex model is used and tool call history is present, the copilot_client module shall truncate any call_id field to a maximum of 64 characters before dispatching to the OpenAI SDK.
Token Lifecycle
While a Copilot token is cached, the token manager shall treat the token as expired at expires_at minus sixty seconds.
12. Out of Scope
Wire format, streaming event shapes, and request body construction — refer to Anthropic and OpenAI vendor specifications.
The Copilot inline completions API (/v1/engines/copilot-codex/completions) — legacy tab-completion endpoint, not chat.
The GitHub Copilot Web thread API (/github/chat/threads) — used by GitHub.com chat interface, not the CLI/IDE inference path.
The Copilot REST management APIs (/orgs/{org}/copilot/*) — seat management and billing, officially documented by GitHub.
Content filter bypass or abuse-pattern usage.
Telemetry endpoints.
13. References
All prior art cited in this document. Sources are labelled REF-N inline.
Ref
Source
URL
REF-1
ericc-ch/copilot-api — README and proxy implementation
This is a variant of OpenAI, with some added routing. Is this something you'd like to see implemented?
Github Copilot protocol spec
GitHub Copilot Inference Protocol — Technical Reference PRD
Table of Contents
1. Problem Statement
A developer wishes to call GitHub Copilot's inference backend from a program that is not an officially supported Copilot client (VS Code, JetBrains, Copilot CLI). The GitHub Copilot inference API is not publicly documented by GitHub. Its protocol has been established through reverse engineering by the open-source community and is used by tools such as
pi-agent(badlogic/pi-mono),ericc-ch/copilot-api,LiteLLM, andopencode.This document consolidates that prior art into a single implementer-facing reference, scoped to the Copilot-specific layer on top of the underlying vendor APIs.
2. Architecture Overview
flowchart TD PRE["PRECONDITION\nOAuth implementation available\n(RFC 8628 device flow or PAT — not specified here)"] subgraph DOC ["THIS DOCUMENT"] subgraph AUTH ["Authentication Layer"] A1["Phase 1: OAuth Device Flow\ngithub.com/login/device/code"] A2["Yields: gho_* OAuth token"] A3["Phase 2: Copilot Token Exchange\napi.github.com/copilot_internal/v2/token"] A4["Yields: short-lived JWT"] A1 --> A2 --> A3 --> A4 end subgraph ROUTE ["Routing Layer"] R1["GET /models\n→ inspect supported_endpoints\n→ select protocol\n→ inject identity headers\n→ set base URL"] end A4 -->|"Bearer copilot-jwt"| R1 end SDK_A["Anthropic SDK\n/v1/messages\n(Claude)"] SDK_O["OpenAI SDK\n/v1/chat/completions\n/v1/responses\n(GPT, Codex)"] PRE --> AUTH R1 --> SDK_A R1 --> SDK_OBoth the authentication layer and the routing layer are specified by this document. Below the routing layer, standard vendor SDK implementations are used without modification — except for base URL override and header injection.
All inference requests carry identity headers that identify the client as a VSCode instance. These are required by the Copilot API gateway; requests lacking them are rejected. [REF-5, REF-6]
3. Authentication Protocol
3.1 OAuth Device Flow — Phase 1
The Copilot API only accepts OAuth tokens issued to the VSCode OAuth app. The VSCode app's public client ID is
Iv1.b507a08c87ecfe98. This ID is hardcoded in official Copilot extensions and is the same for all users. [REF-4, REF-13]Some model families (GPT-5.x Codex variants) are gated to the VSCode client ID specifically, so using this ID rather than any other app registration is important for full model coverage. [REF-13]
Step 1 — Request a device code
Response:
{ "device_code": "794ff3a7cbcc9c45280f1a4c8cae67938e5e6787", "user_code": "31B0-2B8C", "verification_uri": "https://github.com/login/device", "expires_in": 899, "interval": 5 }[REF-11] (live example of this exact response shape)
Display
verification_urianduser_codeto the user. The device code expires inexpires_inseconds (≈ 15 minutes).Step 2 — Poll for the access token
Poll at the
intervalrate (default 5 seconds). GitHub may return aslow_downerror, in which case the polling interval must be increased by 5 seconds and the next poll must wait for the new interval before firing. [REF-22] (pi-mono CHANGELOG: "Fixed GitHub Copilot device-code login polling to respect OAuth slow-down intervals, wait before the first token poll")While pending:
{"error": "authorization_pending"}Slow-down signal:
{"error": "slow_down"}— increase interval by 5s before next pollOn approval:
{ "access_token": "gho_xxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxx", "token_type": "bearer", "scope": "read:user" }[REF-4] (Alorse/copilot-to-api full curl walkthrough)
Store the
gho_token persistently. It does not expire unless the user revokes the OAuth app authorization.3.2 Copilot Token Exchange — Phase 2
The
gho_token cannot be used directly for inference. It must be exchanged for a short-lived Copilot JWT.Response:
{ "token": "tid=<...>;exp=1754130011;sku=monthly_subscriber_quota;...", "expires_at": "2025-08-01T19:40:11.000Z", "refresh_in": 1500 }Key fields [REF-12] (VS Studio debug log output), [REF-6]:
tokenAuthorization: Bearer <token>on all inference and model discovery requestsexpires_atrefresh_insku(encoded in token)monthly_subscriber_quota,free_educational_quotachat_enabled(encoded in token)The
endpointsobject may also be present in GHE configurations; see §3.5 and §8.3.3 Token Lifecycle and Refresh
sequenceDiagram participant App participant GHOAuth as GitHub OAuth participant CopilotAPI as Copilot API participant VendorSDK as Vendor SDK App->>GHOAuth: POST /login/device/code GHOAuth-->>App: {device_code, user_code} Note over App: Display user_code to user loop Poll until approved App->>GHOAuth: POST /login/oauth/access_token GHOAuth-->>App: authorization_pending / slow_down / gho_token end Note over App: Persist gho_token App->>CopilotAPI: GET /copilot_internal/v2/token CopilotAPI-->>App: {token, expires_at, refresh_in} App->>VendorSDK: Inference request (via vendor SDK) VendorSDK-->>App: Response / Stream Note over App: After refresh_in seconds... App->>CopilotAPI: GET /copilot_internal/v2/token CopilotAPI-->>App: {new_token, expires_at, refresh_in}Expiry tolerance: Treat tokens as expired 60 seconds before their stated
expires_at. In WSL2 and virtual machine environments, clock drift can cause premature expiry errors — the 60-second buffer mitigates this. [REF-22]On 401 during inference: Refresh the Copilot token immediately using the stored
gho_token and retry the request exactly once.3.4 Alternative: PAT Authentication
Fine-grained Personal Access Tokens (PATs) with the
Copilot Requestspermission can substitute for thegho_OAuth token in the Phase 2 exchange. [REF-10] (GitHub Copilot CLI docs)PATs are suitable for non-interactive and CI/CD environments where device flow is not feasible. The token format is
github_pat_rather thangho_.3.5 GitHub Enterprise Server (GHE) Variants
For GitHub Enterprise Server instances, the auth endpoints change [REF-15]:
https://github.com/login/device/codehttps://{ghe_host}/login/device/codehttps://github.com/login/oauth/access_tokenhttps://{ghe_host}/login/oauth/access_tokenhttps://api.github.com/copilot_internal/v2/tokenhttps://{ghe_host}/api/v3/copilot_internal/v2/tokenAdditionally, the token exchange response for GHE includes an
endpointsobject:{ "token": "...", "expires_at": "...", "refresh_in": 1500, "endpoints": { "proxy": "https://copilot-proxy.your-ghe-instance.com", "api": "https://api.your-ghe-instance.com" } }When
endpoints.proxyis present, it must be used as the inference base URL rather thanhttps://api.githubcopilot.com. [REF-15]4. Model Discovery
Response (abridged):
{ "data": [ { "id": "claude-sonnet-4-5", "name": "Claude Sonnet 4.5", "object": "model", "vendor": "anthropic", "model_picker_enabled": true, "supported_endpoints": ["/v1/messages", "/v1/chat/completions"], "capabilities": { "limits": { "max_prompt_tokens": 128000, "max_output_tokens": 8192 }, "supports": { "tool_calls": true, "streaming": true } } }, { "id": "gpt-5.2-codex", "name": "GPT-5.2 Codex", "object": "model", "vendor": "openai", "model_picker_enabled": true, "supported_endpoints": ["/v1/responses"] }, { "id": "claude-opus-4-7", "name": "Claude Opus 4.7", "object": "model", "vendor": "anthropic", "model_picker_enabled": false, "supported_endpoints": ["/v1/messages"] } ] }Key fields [REF-16] (pi-mono issue #2891), [REF-17] (pi-mono issue #2678):
idmodelfield of inference requestssupported_endpointsmodel_picker_enabledfalseif the current plan SKU cannot use this model; filter these out before presenting to userscapabilities.limitsFetch timing: Fetch on login and on each Copilot token refresh. Do not hardcode model IDs; the available set changes as GitHub rolls out new models. [REF-16]
5. Protocol Routing
This is the central decision the Copilot routing layer makes. Based on the
supported_endpointsfield for the selected model, the routing layer configures the appropriate vendor SDK client with the Copilot base URL and identity headers, then delegates.[REF-16]: "Infer API type from
supported_endpoints—/v1/messagesmaps to Anthropic protocol,/responsesmaps to OpenAI Responses API"Priority when multiple endpoints are listed: Prefer
/v1/messagesover/v1/chat/completionsfor Claude models. Routing Claude via Chat Completions causes a known conversation history replay bug where the model re-answers all previous prompts with each new message. [REF-18] (pi-mono issue #209)Routing decision pseudocode:
6. Request Headers Reference
These headers must be injected into every request to a Copilot base URL — including
GET /modelsand all inference calls. They identify the client to GitHub's API gateway. Missing or incorrect headers result in 401 or 403 responses. [REF-4, REF-5]AuthorizationBearer <copilot-token>gho_tokenCopilot-Integration-Idvscode-chatEditor-Versionvscode/1.104.1Editor-Plugin-Versioncopilot-chat/0.26.7User-AgentGitHubCopilotChat/0.26.7For the Anthropic path only, also include:
Anthropic-Version2023-06-01Most vendor SDKs support injecting arbitrary extra headers and overriding the base URL via constructor options. Use those mechanisms rather than reimplementing the HTTP layer.
7. Error Handling
401 Unauthorizedgho_token; retry request once403 Forbiddenmodel_picker_enabled; on outage, retry with backoff429 Too Many RequestsRetry-Afterheader if present400 Bad Request503 Service Unavailable[REF-7] (GitHub community 403 discussion), [REF-12] (VS Studio logs distinguishing client vs. server 403)
Distinguishing 401 types: A 401 on inference after a recently issued token is transient. A 401 immediately after a token refresh attempt means the
gho_token has been revoked — re-run device flow.Distinguishing 403 types: The token exchange response body carries
HasToken,ChatEnabled, andExpiresAtfields. A 403 withHasToken: Falseindicates no subscription. A 403 withHasToken: Trueis likely a transient outage. [REF-12]8. Per-Subscription Base URLs
The inference base URL varies by Copilot plan [REF-1]:
https://api.githubcopilot.comhttps://api.business.githubcopilot.comhttps://api.enterprise.githubcopilot.comWhen
endpoints.proxyis present in the token exchange response (GHE or certain enterprise configurations), that value takes precedence. [REF-15]9. Known Behavioural Quirks
These are Copilot-layer constraints that affect how you configure the underlying vendor SDK client, not the protocol itself.
Opus extended thinking:
effort: "medium"onlyClaude Opus 4.7 via Copilot currently only accepts
effort: "medium"for extended thinking. Sendingeffort: "high"causes an error. When configuring the Anthropic SDK with extended thinking for an Opus model routed through Copilot, cap effort at"medium". [REF-19] (pi-mono issues #3291, #3438)Codex
call_idlength limitThe Responses API via Copilot enforces a 64-character maximum on
call_idfields in tool results. When replaying tool call history into a Codex model session (e.g. after a model switch mid-session), truncatecall_idvalues to 64 characters before dispatch. Failure produces 400 errors. [REF-20] (claude-code-router gist)Cross-protocol history translation on model switch
When switching a mid-session conversation from a Claude model (Anthropic format) to a GPT model (OpenAI format) or vice versa, the message history must be translated between formats. Tool call and tool result objects have different shapes in each protocol and cannot be passed to the wrong endpoint. This is a session management concern but is a common source of 400 errors in Copilot-backed agents. [REF-20]
Clock drift on token expiry
The Copilot JWT contains an
expfield (Unix epoch). In WSL2 and VM environments, clock drift causes premature 401 errors. Apply a 60-second buffer: treat tokens as expired atexpires_at - 60s. [REF-22]10. Module Design
The Copilot-specific implementation separates into three focused modules. The vendor SDK clients (Anthropic, OpenAI) are dependencies, not internal modules.
authmoduleInterface:
Responsibilities:
State:
gho_token(or PAT)copilot_token,expires_at,refresh_intoken_managermoduleInterface:
Responsibilities:
refresh_inseconds after issuanceexpires_at - 60sinvalidate(), refresh immediately, signal caller to retry oncemodelsmoduleInterface:
Responsibilities:
supported_endpointsinspection for protocol selectioncapabilities.limitsfor caller use in SDK configurationcopilot_clientmoduleThis is the thin routing and integration layer — it is not a protocol implementation. It configures the vendor SDK clients with Copilot-specific settings and delegates all protocol work to them.
Interface:
Responsibilities:
models.protocol_for(model)endpoints.proxyfrom token responsetoken_manager.invalidate(), refresh, retry once11. EARS Requirements
Authentication
gho_token for a short-lived Copilot JWT by callingGET /copilot_internal/v2/token.slow_downerror by increasing the polling interval by five seconds before the next poll.refresh_inseconds after issuance.Model Discovery
GET /modelswith all required identity headers on first use and on each Copilot token refresh.model_picker_enabledisfalsefrom the set returned byusable_models().ANTHROPICfor any model whosesupported_endpointsincludes/v1/messages.OPENAI_RESPONSESfor any model whosesupported_endpointsincludes/v1/responsesand does not include/v1/messages.Routing and Header Injection
Copilot-Integration-Id: vscode-chat,Editor-Version,Editor-Plugin-Version, andUser-Agenton every request to a Copilot base URL.endpoints.proxyfield, the copilot_client module shall use that value as the inference base URL.https://{ghe_host}/api/v3/copilot_internal/v2/tokenfor the token exchange.Model-Specific Constraints
effortto"medium"regardless of any higher value requested by the caller.call_idfield to a maximum of 64 characters before dispatching to the OpenAI SDK.Token Lifecycle
expires_atminus sixty seconds.12. Out of Scope
/v1/engines/copilot-codex/completions) — legacy tab-completion endpoint, not chat./github/chat/threads) — used by GitHub.com chat interface, not the CLI/IDE inference path./orgs/{org}/copilot/*) — seat management and billing, officially documented by GitHub.13. References
All prior art cited in this document. Sources are labelled REF-N inline.
endpointsobjectsupported_endpointsroutingmodel_picker_enabledper plan SKUeffort: "medium"constraint via Copilotcall_id64-char limit, cross-protocol history