Skip to content

Github Copilot API support #23

Description

@jamestelfer

This is a variant of OpenAI, with some added routing. Is this something you'd like to see implemented?

Github Copilot protocol spec

GitHub Copilot Inference Protocol — Technical Reference PRD

Status: Draft — April 2026
Scope: Sufficient to direct a clean-room implementation of the GitHub Copilot inference protocol in any language.
Precondition: The implementer has working Anthropic Messages API and OpenAI Chat Completions + Responses API client implementations available. This document covers only what is specific to GitHub Copilot: authentication, identity headers, model discovery, protocol routing, and known Copilot-layer quirks. Wire format, streaming event shapes, and request body construction for each protocol are out of scope — refer to the respective vendor specifications.


Table of Contents

  1. Problem Statement
  2. Architecture Overview
  3. Authentication Protocol
  4. Model Discovery
  5. Protocol Routing
  6. Request Headers Reference
  7. Error Handling
  8. Per-Subscription Base URLs
  9. Known Behavioural Quirks
  10. Module Design
  11. EARS Requirements
  12. Out of Scope
  13. References

1. Problem Statement

A developer wishes to call GitHub Copilot's inference backend from a program that is not an officially supported Copilot client (VS Code, JetBrains, Copilot CLI). The GitHub Copilot inference API is not publicly documented by GitHub. Its protocol has been established through reverse engineering by the open-source community and is used by tools such as pi-agent (badlogic/pi-mono), ericc-ch/copilot-api, LiteLLM, and opencode.

This document consolidates that prior art into a single implementer-facing reference, scoped to the Copilot-specific layer on top of the underlying vendor APIs.


2. Architecture Overview

flowchart TD
    PRE["PRECONDITION\nOAuth implementation available\n(RFC 8628 device flow or PAT — not specified here)"]

    subgraph DOC ["THIS DOCUMENT"]
        subgraph AUTH ["Authentication Layer"]
            A1["Phase 1: OAuth Device Flow\ngithub.com/login/device/code"]
            A2["Yields: gho_* OAuth token"]
            A3["Phase 2: Copilot Token Exchange\napi.github.com/copilot_internal/v2/token"]
            A4["Yields: short-lived JWT"]
            A1 --> A2 --> A3 --> A4
        end

        subgraph ROUTE ["Routing Layer"]
            R1["GET /models\n→ inspect supported_endpoints\n→ select protocol\n→ inject identity headers\n→ set base URL"]
        end

        A4 -->|"Bearer copilot-jwt"| R1
    end

    SDK_A["Anthropic SDK\n/v1/messages\n(Claude)"]
    SDK_O["OpenAI SDK\n/v1/chat/completions\n/v1/responses\n(GPT, Codex)"]

    PRE --> AUTH
    R1 --> SDK_A
    R1 --> SDK_O
Loading

Both the authentication layer and the routing layer are specified by this document. Below the routing layer, standard vendor SDK implementations are used without modification — except for base URL override and header injection.

All inference requests carry identity headers that identify the client as a VSCode instance. These are required by the Copilot API gateway; requests lacking them are rejected. [REF-5, REF-6]


3. Authentication Protocol

3.1 OAuth Device Flow — Phase 1

The Copilot API only accepts OAuth tokens issued to the VSCode OAuth app. The VSCode app's public client ID is Iv1.b507a08c87ecfe98. This ID is hardcoded in official Copilot extensions and is the same for all users. [REF-4, REF-13]

Some model families (GPT-5.x Codex variants) are gated to the VSCode client ID specifically, so using this ID rather than any other app registration is important for full model coverage. [REF-13]

Step 1 — Request a device code

POST https://github.com/login/device/code
Content-Type: application/json
Accept: application/json

{
  "client_id": "Iv1.b507a08c87ecfe98",
  "scope": "read:user"
}

Response:

{
  "device_code": "794ff3a7cbcc9c45280f1a4c8cae67938e5e6787",
  "user_code": "31B0-2B8C",
  "verification_uri": "https://github.com/login/device",
  "expires_in": 899,
  "interval": 5
}

[REF-11] (live example of this exact response shape)

Display verification_uri and user_code to the user. The device code expires in expires_in seconds (≈ 15 minutes).

Step 2 — Poll for the access token

Poll at the interval rate (default 5 seconds). GitHub may return a slow_down error, in which case the polling interval must be increased by 5 seconds and the next poll must wait for the new interval before firing. [REF-22] (pi-mono CHANGELOG: "Fixed GitHub Copilot device-code login polling to respect OAuth slow-down intervals, wait before the first token poll")

POST https://github.com/login/oauth/access_token
Content-Type: application/json
Accept: application/json

{
  "client_id": "Iv1.b507a08c87ecfe98",
  "device_code": "<device_code from step 1>",
  "grant_type": "urn:ietf:params:oauth:grant-type:device_code"
}

While pending: {"error": "authorization_pending"}
Slow-down signal: {"error": "slow_down"} — increase interval by 5s before next poll
On approval:

{
  "access_token": "gho_xxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxx",
  "token_type": "bearer",
  "scope": "read:user"
}

[REF-4] (Alorse/copilot-to-api full curl walkthrough)

Store the gho_ token persistently. It does not expire unless the user revokes the OAuth app authorization.


3.2 Copilot Token Exchange — Phase 2

The gho_ token cannot be used directly for inference. It must be exchanged for a short-lived Copilot JWT.

GET https://api.github.com/copilot_internal/v2/token
Authorization: token <gho_token>

Response:

{
  "token": "tid=<...>;exp=1754130011;sku=monthly_subscriber_quota;...",
  "expires_at": "2025-08-01T19:40:11.000Z",
  "refresh_in": 1500
}

Key fields [REF-12] (VS Studio debug log output), [REF-6]:

Field Type Description
token string Short-lived JWT. Use as Authorization: Bearer <token> on all inference and model discovery requests
expires_at ISO 8601 Absolute expiry timestamp
refresh_in integer (seconds) Suggested time before proactive refresh; typically 1500 (25 min)
sku (encoded in token) string Subscription tier, e.g. monthly_subscriber_quota, free_educational_quota
chat_enabled (encoded in token) bool Whether chat is enabled for this account

The endpoints object may also be present in GHE configurations; see §3.5 and §8.


3.3 Token Lifecycle and Refresh

sequenceDiagram
    participant App
    participant GHOAuth as GitHub OAuth
    participant CopilotAPI as Copilot API
    participant VendorSDK as Vendor SDK

    App->>GHOAuth: POST /login/device/code
    GHOAuth-->>App: {device_code, user_code}
    Note over App: Display user_code to user

    loop Poll until approved
        App->>GHOAuth: POST /login/oauth/access_token
        GHOAuth-->>App: authorization_pending / slow_down / gho_token
    end
    Note over App: Persist gho_token

    App->>CopilotAPI: GET /copilot_internal/v2/token
    CopilotAPI-->>App: {token, expires_at, refresh_in}

    App->>VendorSDK: Inference request (via vendor SDK)
    VendorSDK-->>App: Response / Stream

    Note over App: After refresh_in seconds...

    App->>CopilotAPI: GET /copilot_internal/v2/token
    CopilotAPI-->>App: {new_token, expires_at, refresh_in}
Loading

Expiry tolerance: Treat tokens as expired 60 seconds before their stated expires_at. In WSL2 and virtual machine environments, clock drift can cause premature expiry errors — the 60-second buffer mitigates this. [REF-22]

On 401 during inference: Refresh the Copilot token immediately using the stored gho_ token and retry the request exactly once.


3.4 Alternative: PAT Authentication

Fine-grained Personal Access Tokens (PATs) with the Copilot Requests permission can substitute for the gho_ OAuth token in the Phase 2 exchange. [REF-10] (GitHub Copilot CLI docs)

export COPILOT_GITHUB_TOKEN=github_pat_xxxxx
# or
export GH_TOKEN=github_pat_xxxxx

PATs are suitable for non-interactive and CI/CD environments where device flow is not feasible. The token format is github_pat_ rather than gho_.

Known limitation: PAT support in third-party implementations is inconsistent. Pi-mono's provider only accepts gho_-prefixed OAuth tokens [REF-14]. LiteLLM and the official Copilot CLI both accept PATs via environment variable. [REF-10]


3.5 GitHub Enterprise Server (GHE) Variants

For GitHub Enterprise Server instances, the auth endpoints change [REF-15]:

Step github.com GHE
Device code https://github.com/login/device/code https://{ghe_host}/login/device/code
Token poll https://github.com/login/oauth/access_token https://{ghe_host}/login/oauth/access_token
Copilot token exchange https://api.github.com/copilot_internal/v2/token https://{ghe_host}/api/v3/copilot_internal/v2/token

Additionally, the token exchange response for GHE includes an endpoints object:

{
  "token": "...",
  "expires_at": "...",
  "refresh_in": 1500,
  "endpoints": {
    "proxy": "https://copilot-proxy.your-ghe-instance.com",
    "api": "https://api.your-ghe-instance.com"
  }
}

When endpoints.proxy is present, it must be used as the inference base URL rather than https://api.githubcopilot.com. [REF-15]


4. Model Discovery

GET https://api.githubcopilot.com/models
Authorization: Bearer <copilot-token>
Copilot-Integration-Id: vscode-chat
Editor-Version: vscode/1.104.1
Editor-Plugin-Version: copilot-chat/0.26.7
User-Agent: GitHubCopilotChat/0.26.7

Response (abridged):

{
  "data": [
    {
      "id": "claude-sonnet-4-5",
      "name": "Claude Sonnet 4.5",
      "object": "model",
      "vendor": "anthropic",
      "model_picker_enabled": true,
      "supported_endpoints": ["/v1/messages", "/v1/chat/completions"],
      "capabilities": {
        "limits": {
          "max_prompt_tokens": 128000,
          "max_output_tokens": 8192
        },
        "supports": {
          "tool_calls": true,
          "streaming": true
        }
      }
    },
    {
      "id": "gpt-5.2-codex",
      "name": "GPT-5.2 Codex",
      "object": "model",
      "vendor": "openai",
      "model_picker_enabled": true,
      "supported_endpoints": ["/v1/responses"]
    },
    {
      "id": "claude-opus-4-7",
      "name": "Claude Opus 4.7",
      "object": "model",
      "vendor": "anthropic",
      "model_picker_enabled": false,
      "supported_endpoints": ["/v1/messages"]
    }
  ]
}

Key fields [REF-16] (pi-mono issue #2891), [REF-17] (pi-mono issue #2678):

Field Meaning
id Model ID to pass in the model field of inference requests
supported_endpoints Determines which protocol client to use — see §5
model_picker_enabled false if the current plan SKU cannot use this model; filter these out before presenting to users
capabilities.limits Context window and output token limits; use when configuring SDK parameters

Fetch timing: Fetch on login and on each Copilot token refresh. Do not hardcode model IDs; the available set changes as GitHub rolls out new models. [REF-16]

Model activation prerequisite: Many models must be explicitly enabled by the user in VS Code before they become usable via the API, even with a valid subscription and model_picker_enabled: true. A model not supported error for such a model means the user needs to activate it via the VS Code model picker. There is no programmatic activation path. [REF-16]


5. Protocol Routing

This is the central decision the Copilot routing layer makes. Based on the supported_endpoints field for the selected model, the routing layer configures the appropriate vendor SDK client with the Copilot base URL and identity headers, then delegates.

supported_endpoints includes "/v1/messages"
    → Use Anthropic SDK
    → base_url: https://api.githubcopilot.com (or endpoints.proxy)
    → extra headers: see §6

supported_endpoints includes "/v1/chat/completions" (but not "/v1/messages")
    → Use OpenAI SDK, chat completions mode
    → base_url: https://api.githubcopilot.com

supported_endpoints includes "/v1/responses"
    → Use OpenAI SDK, responses mode
    → base_url: https://api.githubcopilot.com

[REF-16]: "Infer API type from supported_endpoints — /v1/messages maps to Anthropic protocol, /responses maps to OpenAI Responses API"

Priority when multiple endpoints are listed: Prefer /v1/messages over /v1/chat/completions for Claude models. Routing Claude via Chat Completions causes a known conversation history replay bug where the model re-answers all previous prompts with each new message. [REF-18] (pi-mono issue #209)

Routing decision pseudocode:

func protocol_for(model):
    eps = model.supported_endpoints
    if "/v1/messages" in eps:
        return ANTHROPIC
    if "/v1/responses" in eps:
        return OPENAI_RESPONSES
    return OPENAI_CHAT_COMPLETIONS

6. Request Headers Reference

These headers must be injected into every request to a Copilot base URL — including GET /models and all inference calls. They identify the client to GitHub's API gateway. Missing or incorrect headers result in 401 or 403 responses. [REF-4, REF-5]

Header Value Notes
Authorization Bearer <copilot-token> The Phase 2 JWT — not the gho_ token
Copilot-Integration-Id vscode-chat Required. [REF-4, REF-5]
Editor-Version vscode/1.104.1 Should reflect a real recent VSCode release
Editor-Plugin-Version copilot-chat/0.26.7 Copilot Chat plugin version
User-Agent GitHubCopilotChat/0.26.7 Must not be a generic HTTP client string

For the Anthropic path only, also include:

Header Value Notes
Anthropic-Version 2023-06-01 Required by the Anthropic Messages protocol

On header freshness: The VSCode version string should track a real recent release. GitHub has previously rejected requests from very stale version strings, though enforcement is inconsistent. The values above reflect current production usage across community implementations.

Most vendor SDKs support injecting arbitrary extra headers and overriding the base URL via constructor options. Use those mechanisms rather than reimplementing the HTTP layer.


7. Error Handling

HTTP Status Meaning Recovery
401 Unauthorized Copilot token expired Refresh Phase 2 token using stored gho_ token; retry request once
403 Forbidden No subscription, model disabled for plan, or transient outage Check model_picker_enabled; on outage, retry with backoff
429 Too Many Requests Rate limit Exponential backoff; respect Retry-After header if present
400 Bad Request Malformed request Check model-specific constraints (see §9)
503 Service Unavailable Transient backend issue Retry with backoff

[REF-7] (GitHub community 403 discussion), [REF-12] (VS Studio logs distinguishing client vs. server 403)

Distinguishing 401 types: A 401 on inference after a recently issued token is transient. A 401 immediately after a token refresh attempt means the gho_ token has been revoked — re-run device flow.

Distinguishing 403 types: The token exchange response body carries HasToken, ChatEnabled, and ExpiresAt fields. A 403 with HasToken: False indicates no subscription. A 403 with HasToken: True is likely a transient outage. [REF-12]


8. Per-Subscription Base URLs

The inference base URL varies by Copilot plan [REF-1]:

Plan Base URL
Individual / Pro / Pro+ https://api.githubcopilot.com
Business https://api.business.githubcopilot.com
Enterprise https://api.enterprise.githubcopilot.com

When endpoints.proxy is present in the token exchange response (GHE or certain enterprise configurations), that value takes precedence. [REF-15]


9. Known Behavioural Quirks

These are Copilot-layer constraints that affect how you configure the underlying vendor SDK client, not the protocol itself.

Opus extended thinking: effort: "medium" only

Claude Opus 4.7 via Copilot currently only accepts effort: "medium" for extended thinking. Sending effort: "high" causes an error. When configuring the Anthropic SDK with extended thinking for an Opus model routed through Copilot, cap effort at "medium". [REF-19] (pi-mono issues #3291, #3438)

Codex call_id length limit

The Responses API via Copilot enforces a 64-character maximum on call_id fields in tool results. When replaying tool call history into a Codex model session (e.g. after a model switch mid-session), truncate call_id values to 64 characters before dispatch. Failure produces 400 errors. [REF-20] (claude-code-router gist)

Cross-protocol history translation on model switch

When switching a mid-session conversation from a Claude model (Anthropic format) to a GPT model (OpenAI format) or vice versa, the message history must be translated between formats. Tool call and tool result objects have different shapes in each protocol and cannot be passed to the wrong endpoint. This is a session management concern but is a common source of 400 errors in Copilot-backed agents. [REF-20]

Clock drift on token expiry

The Copilot JWT contains an exp field (Unix epoch). In WSL2 and VM environments, clock drift causes premature 401 errors. Apply a 60-second buffer: treat tokens as expired at expires_at - 60s. [REF-22]


10. Module Design

The Copilot-specific implementation separates into three focused modules. The vendor SDK clients (Anthropic, OpenAI) are dependencies, not internal modules.

auth module

Interface:

login_device_flow() → Credentials
refresh_copilot_token(credentials) → CopilotToken
load_credentials(path) → Credentials
save_credentials(path, credentials)

Responsibilities:

  • Phase 1: POST device/code, display user_code, poll access_token with slow-down handling
  • Phase 2: GET /copilot_internal/v2/token, parse token and refresh_in
  • Credential persistence (file permissions: 0600)
  • GHE host configuration (parameterise the endpoint URLs, not hardcoded)

State:

  • Persisted: gho_token (or PAT)
  • Ephemeral: copilot_token, expires_at, refresh_in

token_manager module

Interface:

get_valid_token() → CopilotToken   // blocks if refresh in progress
invalidate()                        // force refresh on next call

Responsibilities:

  • Proactive background refresh triggered at refresh_in seconds after issuance
  • Mutex/lock to prevent concurrent refresh races
  • Expiry tolerance: treat token as expired at expires_at - 60s
  • On 401 from inference: invalidate(), refresh immediately, signal caller to retry once

models module

Interface:

list_models(token) → []Model
usable_models(models) → []Model          // filter model_picker_enabled == true
protocol_for(model) → Protocol           // ANTHROPIC | OPENAI_CHAT | OPENAI_RESPONSES

Responsibilities:

  • GET /models with full identity headers
  • Caching (invalidate on token refresh)
  • supported_endpoints inspection for protocol selection
  • Expose capabilities.limits for caller use in SDK configuration

copilot_client module

This is the thin routing and integration layer — it is not a protocol implementation. It configures the vendor SDK clients with Copilot-specific settings and delegates all protocol work to them.

Interface:

complete(model, request) → Response | Stream

Responsibilities:

  • Look up protocol via models.protocol_for(model)
  • Select base URL (§8) or endpoints.proxy from token response
  • Inject identity headers into the chosen vendor SDK client
  • Dispatch to the appropriate vendor SDK
  • On 401: token_manager.invalidate(), refresh, retry once
  • Apply model-specific constraints before dispatch (Opus effort cap, Codex call_id truncation)

11. EARS Requirements

Authentication

  1. The auth module shall exchange a valid gho_ token for a short-lived Copilot JWT by calling GET /copilot_internal/v2/token.
  2. When polling for an OAuth access token, the auth module shall respect the slow_down error by increasing the polling interval by five seconds before the next poll.
  3. While a Copilot token is active, the token manager shall schedule a proactive refresh at refresh_in seconds after issuance.
  4. If the Copilot token exchange returns 403, then the auth module shall not retry automatically and shall surface the error to the caller.
  5. When storing OAuth credentials to disk, the auth module shall set file permissions to 0600.
  6. If an inference request returns 401, then the copilot_client module shall refresh the Copilot token via the token manager and retry the request exactly once before surfacing the error.

Model Discovery

  1. The models module shall fetch GET /models with all required identity headers on first use and on each Copilot token refresh.
  2. The models module shall exclude models where model_picker_enabled is false from the set returned by usable_models().
  3. When determining the inference protocol, the models module shall return ANTHROPIC for any model whose supported_endpoints includes /v1/messages.
  4. When determining the inference protocol, the models module shall return OPENAI_RESPONSES for any model whose supported_endpoints includes /v1/responses and does not include /v1/messages.

Routing and Header Injection

  1. The copilot_client module shall include Copilot-Integration-Id: vscode-chat, Editor-Version, Editor-Plugin-Version, and User-Agent on every request to a Copilot base URL.
  2. Where the Copilot token exchange response includes an endpoints.proxy field, the copilot_client module shall use that value as the inference base URL.
  3. Where a GitHub Enterprise Server host is configured, the auth module shall use https://{ghe_host}/api/v3/copilot_internal/v2/token for the token exchange.

Model-Specific Constraints

  1. Where an Opus model is configured with extended thinking, the copilot_client module shall set effort to "medium" regardless of any higher value requested by the caller.
  2. Where a Codex model is used and tool call history is present, the copilot_client module shall truncate any call_id field to a maximum of 64 characters before dispatching to the OpenAI SDK.

Token Lifecycle

  1. While a Copilot token is cached, the token manager shall treat the token as expired at expires_at minus sixty seconds.

12. Out of Scope

  • Wire format, streaming event shapes, and request body construction — refer to Anthropic and OpenAI vendor specifications.
  • The Copilot inline completions API (/v1/engines/copilot-codex/completions) — legacy tab-completion endpoint, not chat.
  • The GitHub Copilot Web thread API (/github/chat/threads) — used by GitHub.com chat interface, not the CLI/IDE inference path.
  • The Copilot REST management APIs (/orgs/{org}/copilot/*) — seat management and billing, officially documented by GitHub.
  • Content filter bypass or abuse-pattern usage.
  • Telemetry endpoints.

13. References

All prior art cited in this document. Sources are labelled REF-N inline.

Ref Source URL
REF-1 ericc-ch/copilot-api — README and proxy implementation https://github.com/ericc-ch/copilot-api
REF-4 Alorse/copilot-to-api — complete curl walkthrough https://github.com/Alorse/copilot-to-api
REF-5 LiteLLM — GitHub Copilot provider documentation https://docs.litellm.ai/docs/providers/github_copilot
REF-6 Den Delimarsky — "Using GitHub Copilot From Inside GitHub Actions" https://den.dev/blog/github-copilot-inside-github-actions/
REF-7 GitHub Community — "GitHub Copilot: 403 token expired or invalid" https://github.com/orgs/community/discussions/165646
REF-10 GitHub Docs — Authenticating GitHub Copilot CLI https://docs.github.com/en/copilot/how-tos/copilot-cli/set-up-copilot-cli/authenticate-copilot-cli
REF-11 CherryHQ/cherry-studio issue — live device_code response example CherryHQ/cherry-studio#11905
REF-12 GitHub Community — VS Studio log output showing Copilot token fields https://github.com/orgs/community/discussions/168431
REF-13 cavanaug/opencode-copilot-vscode — VSCode client ID and model access gating https://github.com/cavanaug/opencode-copilot-vscode
REF-14 NousResearch/hermes-agent issue #11442 — PAT vs OAuth token support NousResearch/hermes-agent#11442
REF-15 NousResearch/hermes-agent issue #11442 — GHE endpoint variants and endpoints object NousResearch/hermes-agent#11442
REF-16 badlogic/pi-mono issue #2891 — dynamic model discovery and supported_endpoints routing earendil-works/pi#2891
REF-17 badlogic/pi-mono issue #2678 — model_picker_enabled per plan SKU earendil-works/pi#2678
REF-18 badlogic/pi-mono issue #209 — Claude replay bug when routed via Chat Completions earendil-works/pi#209
REF-19 badlogic/pi-mono issues #3291, #3438 — Opus 4.7 effort: "medium" constraint via Copilot earendil-works/pi#3291
REF-20 dpearson2699 gist — claude-code-router: Codex call_id 64-char limit, cross-protocol history https://gist.github.com/dpearson2699/d7e797a85b4286a822dcb9d00f2bebe8
REF-22 badlogic/pi-mono CHANGELOG — device code slow_down fix and clock drift notes https://github.com/badlogic/pi-mono/blob/main/packages/coding-agent/CHANGELOG.md

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions