This hobby, personal project powers a highly configurable, local-first LangChain DeepAgent. We also provide interface to the agent though Chainlit, CLI and TUI.
Implemented features include:
- Sub-agents — delegate tasks to agents with isolated context windows
- Filesystem — read, write, edit, or search over pluggable local, sandboxed, or remote backends
- Context management — summarize long threads and offload tool outputs to disk
- Persistent memory — pluggable state and store backends for cross-session recall
- Skills — reusable behaviors the agent can load on demand
- Tools — bring your own functions or any MCP server
ChatOllama,ChatOpenAI, orChatAnthropicwith configurable model backends- native Chainlit streaming for reasoning, tool calls, and final response
- optional Langfuse tracing through the LangChain callback handler
- Chainlit image uploads sent to vision-capable models as photo attachments for OCR or image analysis
- Chainlit OCR/image uploads accept PNG, JPEG, WEBP, and GIF files
- config-driven synchronous and async DeepAgents subagents
- DeepAgents
==0.7.17with explicit todo planning and safe filesystem defaults - per-response download buttons for Markdown and PDF exports
- configurable Chainlit response actions that ask the agent without showing a user prompt
- Postgres-backed LangGraph checkpoints and durable
/memories/whenDATABASE_URLis set - repo files mounted for the agent under
/workspace/ - Chainlit Modes support for per-message reasoning selection (
Low,Medium,High)
Set these variables before starting the app if you want environment-based overrides:
export DATABASE_URL="postgresql://USER:PASSWORD@HOST:5432/DBNAME?sslmode=disable"
export DEEPAGENT_MODEL_PROVIDER="ollama"
export DEEPAGENT_MODEL_BASE_URL="http://127.0.0.1:11434"
# export DEEPAGENT_MODEL_ENDPOINT_URL="https://api.example.test/custom/v1/messages"
# export DEEPAGENT_MODEL_DISABLE_STREAMING_FOR_TOOL_CALLS="true"
export DEEPAGENT_MODEL_NAME="gpt-oss:20b"
export DEEPAGENT_MODEL_REASONING="medium"
export DEEPAGENT_RECURSION_LIMIT="200"
# export DEEPAGENT_MODEL_API_KEY="optional-for-secured-openai-compatible-servers"
# export ANTHROPIC_API_KEY="required-for-provider-anthropic-unless-DEEPAGENT_MODEL_API_KEY-is-set"
# export SNOWFLAKE_PAT="required-for-provider-snowflake_cortex-unless-a-CLI-or-generic-key-is-set"
# export AWS_REGION="us-east-1" # provider = "bedrock" uses the standard AWS credential chain
export DEEPAGENT_CONFIG="deepagent.toml"
export CHAINLIT_AUTH_SECRET="replace-with-a-long-random-string"
export CHAINLIT_AUTH_USERS='{"admin":"change-me","alice":"alice-password"}'
# export LANGFUSE_PUBLIC_KEY="pk-lf-..."
# export LANGFUSE_SECRET_KEY="sk-lf-..."
# export LANGFUSE_BASE_URL="https://cloud.langfuse.com"DATABASE_URL is optional now:
- when set, LangGraph checkpoints and
/memories/are persisted in Postgres - when unset, the app falls back to in-memory persistence for the current process only
- if
[agent].state = "stateless", LangGraph checkpoint and store handles are not opened or passed to the agent graph even whenDATABASE_URLis set
DEEPAGENT_CONFIG is optional:
- defaults to
deepagent.tomlin the project root - if the file is missing, the app falls back to built-in model defaults and runs without extra skills, MCP servers, or custom subagents
DEEPAGENT_MODEL_* variables are optional:
- they override the matching
[model]values indeepagent.toml DEEPAGENT_MODEL_API_KEYis used for secured OpenAI-compatible servers and can also supply the Anthropic API key whenANTHROPIC_API_KEYis unsetANTHROPIC_API_KEYis read first whenprovider = "anthropic"orprovider = "claude", so stale generic keys do not override the Claude credentialSNOWFLAKE_PATis read first whenprovider = "snowflake_cortex"; Cortex still requires a key, resolved in this order:--api-key,SNOWFLAKE_PAT,DEEPAGENT_MODEL_API_KEY, then[model].api_keyprovider = "bedrock"andprovider = "anthropic_bedrock"ignore API-key variables and uses the standard AWS credential chain (AWS_ACCESS_KEY_ID/AWS_SECRET_ACCESS_KEY,AWS_PROFILE, SSO, or an instance/task role); set the region withAWS_REGIONorAWS_DEFAULT_REGION- when switching to Anthropic with
DEEPAGENT_MODEL_PROVIDER, unset staleDEEPAGENT_MODEL_BASE_URL; useDEEPAGENT_MODEL_ENDPOINT_URLwith the/v1/messagespath for env-based Anthropic proxy switches, or pass--base-urlexplicitly from the CLI DEEPAGENT_MODEL_DISABLE_STREAMINGacceptstrue,false, ortool_calling;DEEPAGENT_MODEL_DISABLE_STREAMING_FOR_TOOL_CALLS=trueis a convenience alias fortool_callingOLLAMA_BASE_URL,OLLAMA_MODEL, andOLLAMA_REASONINGremain supported as Ollama-only compatibility aliases
DEEPAGENT_RECURSION_LIMIT is optional:
- it overrides
[agent].recursion_limitindeepagent.toml - it controls the maximum LangGraph steps for a single agent run
- raise it when long tool-heavy Deep Agent runs hit
GraphRecursionError
CHAINLIT_AUTH_SECRET and Chainlit user credentials are optional:
- when
CHAINLIT_AUTH_SECRETandCHAINLIT_AUTH_USERSare set, the app enables Chainlit password authentication for each configured user CHAINLIT_AUTH_USERSmust be a JSON object mapping usernames to passwords, e.g.{"admin":"change-me","alice":"alice-password"}- the legacy
CHAINLIT_AUTH_USERNAMEandCHAINLIT_AUTH_PASSWORDpair still works for a single user whenCHAINLIT_AUTH_USERSis unset - together with
DATABASE_URL, that unlocks the native Chainlit history bar and chat resume UI - when auth credentials are unset, the app stays unauthenticated and the history bar remains unavailable
You only need Postgres if you want durable LangGraph checkpoints and /memories/.
If DATABASE_URL is unset, the app runs fully in memory for the current process.
This repo includes a Compose file for a local Postgres instance:
docker compose up -d postgresPoint the app at that database:
export DATABASE_URL="postgresql://chainagents:chainagents@127.0.0.1:5432/chainagents?sslmode=disable"Optional verification:
docker compose exec postgres psql -U chainagents -d chainagents -c "select 1;"Notes:
- The Compose file lives at compose.yaml and creates a persistent
postgres-datavolume. - If you already have Postgres installed locally, create an empty database and set
DATABASE_URLto that instance instead. - No separate migration step is required for this app. On startup it calls the LangGraph Postgres store/checkpointer
setup()routines and creates any missing Chainlit persistence tables ("User","Thread","Step","Feedback", and"Element") automatically. - If you manage the Chainlit schema externally, set
CHAINLIT_SCHEMA_BOOTSTRAP=falsebefore launching the app to skip the automatic Chainlit table bootstrap.
Chainlit only shows its built-in history sidebar when both persistence and authentication are enabled.
This app includes a simple password-based auth callback driven by environment variables:
export CHAINLIT_AUTH_SECRET="replace-with-a-long-random-string"
export CHAINLIT_AUTH_USERS='{"admin":"change-me","alice":"alice-password"}'For compatibility, a single user can still be configured with:
export CHAINLIT_AUTH_SECRET="replace-with-a-long-random-string"
export CHAINLIT_AUTH_USERNAME="admin"
export CHAINLIT_AUTH_PASSWORD="change-me"With DATABASE_URL, CHAINLIT_AUTH_SECRET, and either CHAINLIT_AUTH_USERS or the legacy username/password pair set:
- users can sign in through Chainlit's native auth screen
- the history sidebar can list and reopen prior chats
- resumed chats default the LangGraph thread ID to the persisted Chainlit thread ID for that conversation
If you leave auth disabled, Chainlit can still persist thread records in Postgres, but the native history bar will stay hidden.
Install dependencies, then either pull an Ollama model or point deepagent.toml at an OpenAI-compatible server such as LM Studio:
uv sync
ollama pull gpt-oss:20bThe dependency manifests require DeepAgents ==0.7.17; uv sync
installs the matching locked release.
PDF downloads are rendered with WeasyPrint. uv sync installs the Python package,
but WeasyPrint also needs native rendering libraries. On macOS, install them with:
brew install weasyprintOn Linux, install the Pango packages listed in the WeasyPrint installation guide for your distribution before starting the app.
Response PDFs include images referenced with public HTTP or HTTPS Markdown image URLs. Each export downloads at most 20 unique images, with a 10 MiB per-image and 25 MiB aggregate limit, and processes at most 50 million raster pixels. Images that are unavailable, invalid, oversized, or hosted on private network addresses are replaced with a labeled placeholder so the rest of the PDF can still be downloaded.
If you are using LM Studio or another OpenAI-compatible server instead of Ollama, skip ollama pull, load a model in that server, and set [model].provider = "openai_compatible" with the server's base_url.
If you are using Claude through Anthropic, set [model].provider = "anthropic" and provide ANTHROPIC_API_KEY or DEEPAGENT_MODEL_API_KEY.
For Snowflake Cortex, use the dedicated snowflake_cortex provider and a Snowflake PAT as shown in Snowflake Cortex.
For Amazon Bedrock, set [model].provider = "bedrock" with a Bedrock model or inference-profile ID and configure AWS credentials and AWS_REGION as shown in Amazon Bedrock. For Claude on Bedrock through the Anthropic Messages API, use provider = "anthropic_bedrock".
If you enable workspace-docs RAG with Ollama embeddings, also pull an embedding model such as:
ollama pull nomic-embed-textThis repo includes a portable deepagent.toml with:
- Ollama at
http://127.0.0.1:11434withgpt-oss:20b - a higher LangGraph recursion limit for longer tool-heavy Deep Agent runs
- recursive file deletion disabled unless
[agent].delete_tool_enabled = true - command execution disabled unless
[agent].execute_tool_enabled = true - RAG, reflection, Langfuse, MCP servers, and subagents disabled until configured
- the repo-local
skills/source for the main agent
See deepagent.toml.example and the sections below for provider and optional integration examples.
Start the Chainlit app:
uv run chainlit run main.py -wStart the FastAPI server:
uv run chainagents-api --host 127.0.0.1 --port 8000Install the locked development environment, then run the same gates used by CI:
uv sync --locked
uv run ruff check chainagents *.py scripts/*.py tests
uv run mypy
uv run pytest
bash scripts/verify-installed-wheel.shThe wheel check builds the real distribution, installs its locked runtime dependencies in a fresh virtual environment, and runs import, CLI, API, and configuration smoke checks from a temporary user working directory.
The API uses the same deepagent.toml and environment settings as the Chainlit
and CLI entrypoints. It serves one trusted owner: tokenless access requires a
loopback peer and safe local Host/Origin, while remote binds require
CHAINAGENTS_API_TOKEN and bearer Authorization headers. See
API access, browser deployment, and request limits.
Useful endpoints include:
curl http://127.0.0.1:8000/health
curl http://127.0.0.1:8000/api/status
THREAD_ID="api-$(uuidgen)"
curl -X POST http://127.0.0.1:8000/api/agent/invoke \
-H "Content-Type: application/json" \
-d "{\"prompt\":\"Summarize this repository\",\"thread_id\":\"$THREAD_ID\"}"
curl -N -X POST http://127.0.0.1:8000/api/agent/stream \
-H "Content-Type: application/json" \
-d "{\"prompt\":\"Summarize this repository\",\"thread_id\":\"$THREAD_ID\"}"JSON requests to /api/agent/invoke and /api/agent/stream accept an optional
command field. When present, prompt is passed to that configured native
command as its argument text. The multipart /api/agent/stream/multipart
endpoint accepts the same optional command form field alongside prompt,
thread_id, and uploaded files.
When an MCP server is unavailable, /api/agent/invoke includes a warnings
array while /api/agent/stream emits an mcp_status warning event before the
agent response. Healthy tools remain available.
The /api/status response sources its starters and their optional command
values from the active deepagent.toml. A client launching a configured starter
should send the starter message as prompt and its command separately; the
runtime then applies the configured command template to that prompt.
Run the same underlying agent from a terminal without the Chainlit UI:
uv run chainagents --prompt "Summarize this repository" --thread-id cliStart the full-screen terminal UI:
uv run chainagents --tuiThe TUI defaults to thread ID tui, keeps the prompt box at the bottom, shows
the conversation in the main pane with Markdown-formatted assistant responses,
and splits reasoning and tool activity in the right sidebar. The prompt editor
supports multiple lines: press Shift+Enter to insert a newline and Enter to send
the complete prompt. Type / to show configured slash commands, and press Tab
to complete the first matching command. Stdio MCP server diagnostics are
written to .files/tui-stderr.log in TUI mode so they do not corrupt the
full-screen interface.
Useful CLI examples:
uv run chainagents --status --no-rag
uv run chainagents --configure
uv run chainagents --tui --reasoning high
uv run chainagents --list-commands
uv run chainagents --command summarize --prompt "Summarize the config entrypoints"
uv run chainagents --stdin --model gpt-oss:20b --reasoning high < prompt.txt
uv run chainagents --rebuild-rag
uv run chainagents --upload-rag notes.md --prompt "Use my uploaded notes"
uv run chainagents --photo scene.jpg --prompt "Describe this photo"Run uv run chainagents --help for all runtime flags, including model provider,
base URL, endpoint URL, API key, temperature, persistence, MCP session scope,
async subagent URL, RAG controls, photo attachments, streaming, reasoning traces,
tool traces, and JSON output.
Core Python code lives under the chainagents/ package. The root-level Python
files are compatibility wrappers and entrypoints so existing imports and commands
continue to work.
chainagents/
runtime/ Core DeepAgents runtime, model setup, config parsing,
MCP/tool loading, persistence backends, and Langfuse.
interfaces/
chainlit/ Chainlit callbacks, UI bridge, auth, persistence,
uploads, async task notifications, and chat settings.
cli/ Terminal CLI parser, status output, command execution,
upload handling, and event rendering.
tui/ Full-screen Textual terminal UI.
api/ FastAPI application, request schemas, and streaming API.
turns/ Shared TurnRunner: one agent turn (commands, uploads,
streaming, generated files) used by every interface.
events/ Shared LangGraph stream normalization used by all
interfaces.
commands/ Native slash-command parsing and dispatch helpers.
rag/ Workspace documentation RAG config, index, uploads,
and search tool.
exports/ Markdown and PDF response export helpers.
langgraph/ Agent Server graph exports.
util/ Shared utility helpers.
Runtime assets stay at the repository root because they are user/configuration content rather than importable Python package code:
deepagent.tomlanddeepagent.toml.example: model, agent, MCP, RAG, Chainlit, Langfuse, and subagent configuration.skills/: Deep Agents skill sources referenced from TOML asskills.prompts/: prompt files referenced by configured subagents.public/and.chainlit/: Chainlit static assets and native Chainlit config.tests/: regression tests for runtime, interfaces, RAG, exports, and events.
Compatibility wrappers such as main.py, deepagent_runtime.py,
chainlit_bridge.py, chainagents_cli.py, chainagents_api.py,
rag_runtime.py, and response_exports.py import the moved package modules.
Prefer new code under chainagents/, but keep the wrappers until external users
no longer rely on the old import paths.
Deprecated: every root-level wrapper except main.py and langgraph_app.py
(which stay silent for chainlit run main.py -w and langgraph.json) now
emits a DeprecationWarning on first import and will be removed in a future
release. Import from the package path instead, e.g.
chainagents.runtime.core instead of deepagent_runtime.
You can keep the model defaults in deepagent.toml:
[model]
provider = "ollama"
base_url = "http://127.0.0.1:11434"
temperature = 0
max_tokens = 4096
repeat_penalty = 1.1
name = "gpt-oss:20b"
models = ["gpt-oss:20b", "gemma4:27b"]
reasoning_effort = "medium"
# Disable streaming only for requests that include tools, which can help
# model servers that emit malformed streamed tool-call chunks.
disable_streaming_for_tool_calls = falseFor LM Studio or another OpenAI-compatible server:
[model]
provider = "openai_compatible"
base_url = "http://127.0.0.1:1234/v1"
temperature = 0
name = "your-loaded-model-id"
reasoning_effort = "medium"
# api_key = "optional"For OpenAI-compatible servers with a non-standard full chat-completions endpoint:
[model]
provider = "openai_compatible"
endpoint_url = "https://api.example.test/openai/deployments/local/chat/completions?api-version=2026-01-01"
name = "your-loaded-model-id"
# api_key = "optional"Snowflake Cortex uses the canonical provider value snowflake_cortex (no aliases).
It requires a key. Set SNOWFLAKE_PAT, DEEPAGENT_MODEL_API_KEY, or [model].api_key,
or pass --api-key for a one-off CLI run. Credential precedence is --api-key,
SNOWFLAKE_PAT, DEEPAGENT_MODEL_API_KEY, then [model].api_key.
Use either the Chat Completions base URL or the complete Chat Completions endpoint:
[model]
provider = "snowflake_cortex"
base_url = "https://<account-identifier>.snowflakecomputing.com/api/v2/cortex/v1"
name = "claude-sonnet-4-5"
max_tokens = 4096
# api_key = "" # optional only when SNOWFLAKE_PAT or DEEPAGENT_MODEL_API_KEY is set[model]
provider = "snowflake_cortex"
endpoint_url = "https://<account-identifier>.snowflakecomputing.com/api/v2/cortex/v1/chat/completions"
name = "claude-sonnet-4-5"For the CLI, the same settings can be supplied without editing TOML:
export SNOWFLAKE_PAT="your-snowflake-pat"
uv run chainagents --provider snowflake_cortex \
--base-url "https://<account-identifier>.snowflakecomputing.com/api/v2/cortex/v1" \
--model claude-sonnet-4-5 --prompt "Summarize this repository"The auto RAG embedding provider is not valid for a Cortex chat model. If RAG is
enabled, set [rag.embedding].provider to ollama or openai_compatible, together
with an appropriate embedding model and base_url (and api_key when needed).
Tool-call IDs are normalized only for Snowflake Cortex outbound Chat Completions requests;
other OpenAI-compatible providers keep their original tool-call IDs.
For Claude through Anthropic:
[model]
provider = "anthropic"
temperature = 0
name = "claude-sonnet-4-6"
models = ["claude-sonnet-4-6", "claude-opus-4-8", "claude-haiku-4-5-20251001"]
reasoning_effort = "medium"
thinking = "auto"
# api_key = "optional-if-ANTHROPIC_API_KEY-or-DEEPAGENT_MODEL_API_KEY-is-set"
# base_url = "https://api.anthropic.com"
# endpoint_url = "https://claude-proxy.example/proxy/v1/messages"For models hosted on Amazon Bedrock (Claude, Amazon Nova, Llama, Mistral, gpt-oss, and others) through the Bedrock Converse API:
[model]
provider = "bedrock"
temperature = 0
name = "us.anthropic.claude-sonnet-5"
models = ["us.anthropic.claude-sonnet-5", "amazon.nova-pro-v1:0"]
reasoning_effort = "medium"
thinking = "auto"
# endpoint_url = "https://vpce-0123.bedrock-runtime.us-east-1.vpce.amazonaws.com"export AWS_REGION="us-east-1"
export AWS_PROFILE="my-profile" # or AWS_ACCESS_KEY_ID / AWS_SECRET_ACCESS_KEY, SSO, or an instance rolenameis a Bedrock model ID (such asamazon.nova-pro-v1:0), a cross-region inference-profile ID (such asus.anthropic.…), or a foundation-model/system inference-profile ARN. Application inference profiles and provisioned models are not supported, as an ID or an ARN, because they hide the underlying model; use that model or cross-region inference-profile ID instead.- Credentials and region come from the standard AWS chain;
api_keyandDEEPAGENT_MODEL_API_KEYare not used. Bedrock API keys work through boto3's ownAWS_BEARER_TOKEN_BEDROCKvariable. base_urlorendpoint_urloptionally overrides the Bedrock runtime endpoint, for example a VPC interface endpoint; leave both unset to use the regional default.provider = "aws_bedrock"andprovider = "amazon_bedrock"are accepted as aliases.reasoning_effortis forwarded for models whose langchain-aws profile declares configurable reasoning (for example Claude Opus/Sonnet 5, Nova 2, and gpt-oss) and ignored for others. For Claude this enables adaptive thinking, and the samplingtemperatureis dropped because Claude rejects it while thinking. Setthinking = "disabled"to turn reasoning off; for Claude Opus 5 and Sonnet 5, which think by default, this sends an explicit disabled-thinking setting. Claude Opus 5.5, Sonnet 5.5 and Fable 5 cannot run without thinking, sothinking = "disabled"is rejected for them. gpt-oss always reasons, so it rejectsthinking = "disabled"too; usereasoning_effort = "low".- Amazon Nova rejects a temperature of exactly 0, so
temperature = 0is sent as 0.00001. At high reasoning effort Nova accepts no sampling settings, so the temperature is dropped. - Reasoning blocks streamed by Bedrock are shown as thinking, separate from the answer.
- Leave
disable_streamingunset to keep langchain-aws's per-model default (for example, models that cannot stream tool calls fall back to non-streaming tool requests); setdisable_streaming = trueor"tool_calling"to force non-streaming requests. - Workspace-docs RAG cannot infer Bedrock embeddings. If
[rag].enabled = true, set[rag.embedding].providertoollamaoropenai_compatible, together with an appropriate embeddingmodelandbase_url(andapi_keywhen needed).
provider = "bedrock" uses the Converse API, which works for every Bedrock model. To run Claude on Bedrock through the Anthropic Messages API instead, use provider = "anthropic_bedrock". It uses the Anthropic SDK's Bedrock client (langchain-aws ChatAnthropicBedrock), so Claude gets the same request handling as provider = "anthropic": effort, adaptive thinking, and Anthropic content blocks.
[model]
provider = "anthropic_bedrock"
name = "us.anthropic.claude-sonnet-4-6"
reasoning_effort = "medium"
thinking = "auto"
# endpoint_url = "https://vpce-0123.bedrock-runtime.us-east-1.vpce.amazonaws.com"- Credentials, region and
endpoint_urlwork the same way as forprovider = "bedrock"; no API key is used. nameis a Bedrock Claude model ID or cross-region inference-profile ID.reasoning_effortis only sent to Claude models that support effort; it is skipped for models such as Claude 3.x and Haiku 4.5.thinking = "disabled"explicitly turns thinking off on Claude Opus 5 and Sonnet 5, which think by default. Claude Opus 5.5, Sonnet 5.5 and Fable 5 cannot run without thinking, so that setting is rejected for them.- Only Anthropic Claude model IDs are accepted; use
provider = "bedrock"for other Bedrock models. Foundation-model and system inference-profile ARNs work, but application inference profiles and provisioned models hide the underlying model, so use the Claude model or inference-profile ID instead. bedrock_anthropicandclaude_bedrockare accepted as aliases.
Named model profiles let the main agent, Chainlit mode picker, and sync subagents use different provider settings from the same config file:
[model]
provider = "openai_compatible"
base_url = "http://127.0.0.1:1234/v1"
name = "local-default"
models = ["local-default"]
modalities = ["text"]
[model.profiles.fast-local]
name = "local-fast"
temperature = 0.1
reasoning_effort = "low"
modalities = ["text", "image"]
[model.profiles.claude-reviewer]
provider = "anthropic"
name = "claude-sonnet-4-6"
thinking = "auto"
# api_key = "optional-if-ANTHROPIC_API_KEY-or-DEEPAGENT_MODEL_API_KEY-is-set"
[agent]
model = "fast-local"
[[subagents]]
name = "reviewer"
description = "Reviews proposed changes."
system_prompt = "Review for bugs, regressions, and missing tests."
model = "claude-reviewer"Notes:
providerselectsChatOllama,ChatOpenAI,ChatAnthropic,ChatBedrockConverse(provider = "bedrock"), orChatAnthropicBedrock(provider = "anthropic_bedrock").provider = "claude"is accepted as an alias forprovider = "anthropic".- Preferred shared fields are
base_url,name,temperature,max_tokens, andreasoning_effort. max_tokensis an optional positive output-token limit. It maps tomax_completion_tokensfor Snowflake Cortex and OpenAI-compatible providers,max_tokensfor Anthropic and Bedrock, andnum_predictfor Ollama.- If the model reaches this limit while producing a tool call, ChainAgents discards the incomplete call and tells the model to shorten or split it once. A second truncated tool call ends that run with a clear message; increasing
max_tokensmay help. repeat_penaltyis optional and currently applies toprovider = "ollama"; when omitted, Ollama defaults are used.disable_streaming = "tool_calling"ordisable_streaming_for_tool_calls = truebypasses model streaming only when tools are attached to the request; use this for providers that have trouble streaming tool-call chunks.disable_streaming = truedisables model streaming for all requests.endpoint_urlis an override for full non-standard model endpoint URLs. OpenAI-compatible paths ending in/chat/completionsor/responsesare normalized to the client base URL and query parameters are forwarded as OpenAI client default query parameters. Anthropic paths ending in/v1/messagesare normalized to the Claude client base URL and query parameters are forwarded as Anthropic client default query parameters.modelsis an optional list of model IDs surfaced in Chainlit settings and modes so users can switch models per session or per message.modalitiesdeclares accepted input types for a model or profile. It defaults to["text"]; add"image"only for models that accept image content.[model.profiles.<name>]defines a named profile. Profiles inherit omitted fields from[model]when they keep the same provider; profiles that switch toopenai_compatiblemust providebase_urlorendpoint_url, profiles that switch toanthropicdefault tohttps://api.anthropic.comunlessbase_urlorendpoint_urlis set, and profiles that switch tobedrockuse the AWS regional endpoint unlessbase_urlorendpoint_urlis set.- Profile names are surfaced in Chainlit settings and modes alongside
[model].models. When a selected value matches a profile name, the full profile is used; otherwise the value is treated as a raw model name using the inherited/default provider settings. [agent].modeloptionally sets the main/supervisor agent's default profile or raw model name. CLI and environment model overrides still take precedence.api_keyis optional forprovider = "openai_compatible"; when omitted, the runtime sends a placeholder token that local servers like LM Studio accept.- Anthropic requires an API key from
ANTHROPIC_API_KEY,DEEPAGENT_MODEL_API_KEY, orapi_key; when multiple are set,ANTHROPIC_API_KEYtakes precedence over the generic key. - When switching from another provider to Anthropic through environment or CLI overrides, provide Anthropic credentials through
ANTHROPIC_API_KEY,DEEPAGENT_MODEL_API_KEY, or--api-key; the runtime will not reuse anapi_keyfrom another provider's TOML config. - Legacy Ollama
endpointandportare still accepted whenprovider = "ollama"or omitted. reasoning_effortsets the default Chainlit reasoning level for new chats. Ollama uses that level directly, Anthropic maps it to Claudeeffort, Bedrock forwards it asreasoning_effortfor models that support it, and OpenAI-compatible servers may ignore it.thinkingcontrols Anthropic adaptive thinking:autoenables it only for known supported Claude models,adaptivealways sendsthinking = {"type": "adaptive"}, anddisablednever sends a thinking parameter.DEEPAGENT_MODEL_PROVIDER,DEEPAGENT_MODEL_BASE_URL,DEEPAGENT_MODEL_ENDPOINT_URL,DEEPAGENT_MODEL_NAME,DEEPAGENT_MODEL_API_KEY,DEEPAGENT_MODEL_REASONING,DEEPAGENT_MODEL_DISABLE_STREAMING, andDEEPAGENT_MODEL_DISABLE_STREAMING_FOR_TOOL_CALLSoverride the TOML defaults when set.OLLAMA_BASE_URL,OLLAMA_MODEL, andOLLAMA_REASONINGstill work as Ollama-only compatibility aliases.
Langfuse tracing is disabled by default. To enable it, set your Langfuse credentials in the environment and turn on the TOML option:
export LANGFUSE_PUBLIC_KEY="pk-lf-..."
export LANGFUSE_SECRET_KEY="sk-lf-..."
export LANGFUSE_BASE_URL="https://cloud.langfuse.com"[langfuse]
enabled = trueWhen enabled, ChainAgents attaches Langfuse's LangChain callback handler to Chainlit, CLI, TUI, and API agent runs. The LangGraph thread ID is also passed as the Langfuse session ID.
Set a LangSmith API key in the environment, then enable the integration in
deepagent.toml:
export LANGSMITH_API_KEY="lsv2_..."
# For a non-default region, also set LANGSMITH_ENDPOINT without a trailing slash.
# For an API key linked to multiple workspaces, set LANGSMITH_WORKSPACE_ID.[langsmith]
enabled = true
project = "chainagents"
background_trace_mode = "linked" # or "separate"The project setting takes precedence over LANGSMITH_PROJECT; when neither
is set, ChainAgents uses chainagents. In linked mode, a local background
subagent run belongs to its parent trace when a LangSmith parent is available.
If the parent cannot be captured, it starts its own trace. In separate mode,
each task starts its own trace and records parent run and trace IDs as metadata.
Both modes identify runs by the conversation session ID and background task ID,
including nested and batch tasks. Search for background_task_id in LangSmith
to find a task, or filter by session_id to see a conversation's tasks. The
background_trace_link metadata says linked, separate, or
parent_unavailable; the last value marks a linked-mode fallback to a root
trace. Each run also records the agent name and path, and nested runs record
their parent task ID. Chainlit reasoning and tool step visibility settings do
not change tracing. Langfuse can remain enabled at the same time.
LangSmith records graph execution. The local task manager remains the source for final cancellation and cleanup status, which can differ from a graph run that already finished successfully. An enabled integration owns a client and flushes buffered traces when the runtime shuts down. Exported graphs flush at application teardown and remain usable in a later lifespan.
The [agent] table configures main-agent runtime behavior:
[agent]
state = "stateful"
delete_tool_enabled = false
execute_tool_enabled = false
recursion_limit = 200
memory_namespace = "filesystem"
memory_files = ["/memories/AGENTS.md"]
skills = ["skills"]
mcp_servers = ["repo"]
# custom_instruction = "Always ask clarifying questions before editing files."
# custom_instruction_file = "prompts/ui_prompts.md"
[agent.reflection]
enabled = true
memory_file = "/memories/AGENTS.md"
max_lesson_chars = 700
tool_failure_mode = "unrecovered"Notes:
state = "stateful"passes the configured LangGraph store and checkpointer to DeepAgents so thread IDs can continue conversation state.state = "stateless"omits those state handles and does not expose/memories/when building the agent graph.delete_tool_enabled = falsepreserves the pre-0.7 filesystem surface. Set it totrueonly when the main agent and local synchronous subagents should receive DeepAgents 0.7's recursivedeletetool. Remote async graphs have their own configuration.execute_tool_enabled = falsekeeps command execution out of the tool surface. Set it totrueonly when the main agent and local synchronous subagents should receive DeepAgents 0.7'sexecutetool. The default ChainAgents backend is not execution-capable, so opting in exposes the tool for a compatible sandbox backend but does not grant host-shell access by itself. Remote async graphs have their own configuration.recursion_limitis the LangGraph step limit for one agent run.- The built-in default is
100; this repo'sdeepagent.tomlsets it to200. DEEPAGENT_RECURSION_LIMIToverrides this value when set.- Increase it for long tool-heavy runs that hit
GraphRecursionError; lower it if you want runaway loops to stop sooner. memory_namespaceis the shared agent-scopedStoreBackendnamespace for/memories/. ChainAgents passes a concrete backend instance and the explicit namespace tuple(memory_namespace,), as required by DeepAgents 0.7. The default isfilesystemto preserve existing memory data from earlier configs. Use only letters, numbers, hyphens, underscores, dots,@,+, colons, and tildes.memory_fileslists/memories/files DeepAgents loads into the startup memory prompt. The default is["/memories/AGENTS.md"]; set it to[]to keep the memory route without startup memory loading.custom_instructionappends an inline instruction to the main/supervisor agent system prompt.custom_instruction_fileloads that appended instruction from a UTF-8 text file. Relative paths are resolved from the activedeepagent.toml; use eithercustom_instructionorcustom_instruction_file, not both. This repo usesprompts/ui_prompts.mdto encourage active Chainlit generated UI panels and next-step action buttons.[agent.reflection]is opt-in. When enabled for stateful agents, ChainAgents proposes a compact lesson formemory_fileafter correction phrases such as "that was wrong" or after unrecovered tool failures. Chainlit asks with Save/Dismiss before writing through the agent; CLI, TUI, and API expose the proposal without mutating memory.
DeepAgents 0.7 no longer installs TodoListMiddleware by default. ChainAgents
adds langchain.agents.middleware.TodoListMiddleware explicitly to every local
main, synchronous, and nested agent stack. This preserves the write_todos
tool, the todos state channel, the planning prompt, and Chainlit task-list
rendering. Separately deployed async graphs must restore todo middleware in
their own runtime if they rely on the same behavior.
ChainAgents uses concrete BackendProtocol instances throughout. Backend
integrations should use the current ls(path), glob(pattern, path=None),
grep(pattern, path=None, glob=None, max_count=None), and
read(file_path, offset=0, limit=2000) -> ReadResult contracts. Consume
ReadResult.file_data and its metadata fields rather than parsing rendered
read_file text.
For rendered tool output, empty ls and glob results are the string
No files found, not []. read_file line numbers no longer use a fixed-width
cat -n gutter, so callers must not parse text by fixed character columns.
ChainAgents does not parse any of these rendered filesystem outputs.
The app can build a local-first RAG index over repo documentation and expose it to the main agent as the search_workspace_knowledge tool.
Example config:
[rag]
enabled = true
persist_directory = ".rag"
include_globs = ["README.md", "chainlit.md", "prompts/**/*.md", "skills/**/*.md"]
exclude_globs = ["AGENTS.md", "AGENT.md"]
chunk_size = 1200
chunk_overlap = 200
top_k = 4
[rag.embedding]
provider = "auto"Notes:
- The default corpus is docs-only:
README.md,chainlit.md,prompts/**/*.md, andskills/**/*.md. AGENTS.mdandAGENT.mdstay out of RAG becauseAGENTS.mdis loaded directly into the main agent prompt when present.- The persisted local index lives under
.rag/and is safe to delete and rebuild. - With
provider = "auto", the embedding backend follows the active chat-model provider. - For Ollama, the default embedding model is
nomic-embed-text. - For OpenAI-compatible embeddings, set
[rag.embedding].modelexplicitly. - For Anthropic or Snowflake Cortex chat models, set
[rag.embedding].providerexplicitly toollamaoropenai_compatible, with an appropriate model and base URL;autois not valid for either provider. - On startup, the UI reports whether RAG is ready and how many files/chunks were indexed.
- The startup message includes a
Rebuild Knowledge Indexaction so you can refresh the index after documentation changes. - The startup message also includes
Upload File For RAG, which lets you add text-based files to the current chat thread's knowledge index. - Composer file attachments are enabled for text-based uploads; attached files are automatically ingested into the current thread's RAG store before the model responds.
- Uploaded files are thread-scoped and persist under
.rag/uploads/, so they do not leak into other chat threads.
This repo also includes an app-specific chainlit.toml for UI behavior that the bridge owns:
[steps]
auto_collapse_delay_seconds = 3Notes:
chainlit.tomlis separate from Chainlit's native.chainlit/config.toml.[steps].auto_collapse_delay_secondscontrols how long completed reasoning and tool steps stay expanded before auto-collapsing.- If
chainlit.tomlis missing or invalid, the app falls back to3seconds.
The runtime now supports Deep Agents skill sources through deepagent.toml.
- Create a skill source directory in the repo, for example:
skills/
├── repo-docs/
│ └── SKILL.md
└── reviewer/
└── SKILL.md
- Add the source directory to
deepagent.toml:
[agent]
skills = ["skills"]Notes:
- Relative paths in
deepagent.tomlare resolved from the config file location. - Relative skill paths are automatically mapped into the Deep Agents virtual filesystem as
/workspace/.... - Each skill source directory should contain one or more skill folders, and each skill folder must contain
SKILL.md. - You can also use explicit virtual paths such as
"/workspace/skills/"if you prefer. - Every loaded skill is also exposed as a Chainlit slash command using the skill
name, for examplereviewerbecomes/reviewer. - Running a skill-backed slash command forces the main agent to read that skill's
SKILL.mdand apply it for that request. - Skills loaded through
[agent].skillsand sync[[subagents]].skillsare both considered for slash commands, but explicit[chainlit].commandstake precedence on name collisions.
Minimal SKILL.md example:
---
name: reviewer
description: Use this skill when reviewing code changes for bugs and missing tests.
---
# reviewer
When asked to review code:
1. Read the relevant files first.
2. Focus on bugs, regressions, and missing tests.
3. Return concise findings with file references.Custom subagents are also loaded from deepagent.toml, and each subagent can have its own skills and mcp_servers.
Example:
[mcp]
tool_name_prefix = true
stateful = true
[mcp.servers.repo]
transport = "stdio"
command = "npx"
args = ["-y", "@modelcontextprotocol/server-filesystem@2025.8.21", "."]
cwd = "."
[agent]
state = "stateful"
recursion_limit = 200
skills = ["skills"]
mcp_servers = ["repo"]
summarization_trigger_tokens = 6000
summarization_keep_tokens = 2400
[[subagents]]
name = "repo-researcher"
description = "Researches the codebase and produces concise implementation guidance."
system_prompt_file = "prompts/repo-researcher.md"
skills = ["skills/research"]
mcp_servers = ["repo"]
nested_subagents = ["reviewer"]
[[subagents.subagents]]
name = "repo-planner"
description = "Turns repository findings into an implementation plan."
system_prompt = "Use repository context to produce a concise implementation plan."
skills = ["skills/research"]
mcp_servers = ["repo"]
[[subagents]]
name = "reviewer"
description = "Reviews proposed changes for bugs and regressions."
system_prompt = """
You are a strict code reviewer.
Focus on correctness, regressions, and missing tests.
Keep findings concise and actionable.
"""
skills = ["skills/reviewer"]
mcp_servers = ["repo"]
# model = "gpt-oss:20b"Supported subagent fields:
name: requireddescription: requiredsystem_promptorsystem_prompt_file: one is requiredskills: optional list of skill source paths for that subagentmcp_servers: optional list of MCP server names to attach to that subagentmodel: optional profile name or raw model name. Profile names can switch provider settings and tool-schema handling for that sync subagent. Raw model names inherit the parent/default provider settings.background: optional boolean, defaulting tofalse. When global local background execution is enabled,trueallows the subagent's direct parent to launch it withspawn_background_taskor include it inrun_subagent_batch. Foregroundtaskdelegation is unaffected.messaging: optional boolean, defaulting tofalse. When[agent.messaging].enabled = true, this grants that synchronous subagent invocation the local message tools. The flag is not inherited by children. Remote Agent Protocol subagents do not support local messaging.nested_subagents: optional list of top-level sync subagent names exposed as children of this subagent[[subagents.subagents]]: optional inline private sync child subagents under a parent subagent
Nested subagents let a synchronous subagent delegate work to its own synchronous child agents. They are useful when a coordinator needs focused helpers but the main agent should not necessarily see every helper directly.
ChainAgents supports two nesting patterns:
- use
[[subagents.subagents]]for a private inline child available only to its parent - use
nested_subagents = ["name"]to reuse a top-level synchronous subagent as a child while keeping it available to the main agent
Place [[subagents.subagents]] immediately after its parent
[[subagents]] entry:
[[subagents]]
name = "research-manager"
description = "Coordinates repository research and planning."
system_prompt = "Delegate focused work and synthesize the results."
[[subagents.subagents]]
name = "repo-planner"
description = "Turns repository findings into an implementation plan."
system_prompt = "Produce a concise, actionable implementation plan."In this configuration:
repo-planneris available toresearch-managerrepo-planneris not exposed directly to the main agent- the child can use the normal synchronous subagent fields, including
skills,mcp_servers, andmodel
Use an inline child when it is an implementation detail of one parent.
Define the child as a normal top-level [[subagents]] entry and reference its
name from the parent:
[[subagents]]
name = "research-manager"
description = "Coordinates repository research and review."
system_prompt = "Delegate research and ask the reviewer to check the result."
nested_subagents = ["reviewer"]
[[subagents]]
name = "reviewer"
description = "Reviews proposed changes for bugs and regressions."
system_prompt = "Return concise findings with actionable file references."In this configuration:
reviewerremains available directly to the main agentresearch-managercan also delegate toreviewer- more than one parent can reference the same top-level subagent
- each referenced name must exactly match a top-level synchronous subagent
Use a reference when a role should be shared by the main agent or multiple
parents. A parent can expose several shared children, for example
nested_subagents = ["planner", "reviewer"].
Inline children accept the same fields as other synchronous subagents:
name, description, system_prompt or system_prompt_file, skills,
mcp_servers, model, background, and their own nested children. A referenced child uses
the configuration from its top-level [[subagents]] entry wherever it is
reused. When model is omitted, the child continues with its parent/default
model configuration.
Nested subagents are synchronous only. Async Agent Protocol subagents must
remain top-level [[async_subagents]] entries.
Configuration loading rejects:
- a
nested_subagentsname that does not match a top-level synchronous subagent - duplicate direct child names, including an inline child and a referenced child with the same name
- reference cycles such as
manager -> reviewer -> manager - a nested child that defines
graph_id, because that represents an async subagent
As a rule of thumb, use an inline child for a parent-private specialist, a
referenced child for a shared synchronous role, and [[async_subagents]] for
remote or background Agent Protocol work.
Main [agent] additions:
state: optional agent state mode. Usestatefulfor checkpointed conversation state, orstatelessto build the DeepAgents graph without a LangGraph store, checkpointer, or/memories/route. Defaults tostateful.recursion_limit: optional positive integer LangGraph step limit for a single agent run. Defaults to100unless overridden byDEEPAGENT_RECURSION_LIMIT.memory_namespace: optional non-empty namespace for agent-scoped/memories/storage. Defaults tofilesystem; allowed characters are letters, numbers,-,_,.,@,+,:, and~.memory_files: optional list of absolute/memories/file paths loaded into the DeepAgents startup memory prompt. Defaults to["/memories/AGENTS.md"]; use[]to disable startup memory loading.delete_tool_enabled: optional boolean controlling DeepAgents 0.7's recursivedeletetool for the main agent and local synchronous subagents. Defaults tofalse.execute_tool_enabled: optional boolean controlling DeepAgents 0.7'sexecutetool for the main agent and local synchronous subagents. Defaults tofalse.[agent.background_subagents]: global opt-in and limits for process-local background execution.enableddefaults tofalse; eligible synchronous subagents must also setbackground = true.stream_activity = trueexposes live reasoning and tool activity as nested Chainlit steps while leaving other interfaces completion-only.batch_result_formatselectsrun_subagent_batchresults fromjson,markdown, ormarkdown_filesand defaults tomarkdown. The three positive integer limits bound running work per conversation, running work across the process, and retained task records per conversation. When the retained limit is reached, the oldest finished tasks are forgotten to make room; unfinished tasks, parents of retained tasks, and tasks whose cleanup has not completed are kept.[agent.messaging]: opt-in process-local mailboxes. The main agent and eachmessaging = truesynchronous subagent receivelist_agent_recipients,send_agent_message,get_agent_message, andwait_for_agent_messages. A message is inserted before the recipient's next model step; an idle agent is not woken. Addresses identify individual invocations, so uselist_agent_recipientswhen same-name agents run concurrently. Finished subagents cannot receive messages. Mailboxes are cleared when the conversation closes or the process exits.[agent.user_input]: opt-in nonblocking user prompts. A conversation still runs one main turn at a time. In Chainlit, busy input offers Steer active turn and Queue next turn; Stop pauses queued turns until Resume. In the TUI, use/steer text,/queue text,/stop, and/resume(or Ctrl+R). The interactive CLI prompts for steer or queue during a run and also accepts those commands. HTTP clients can submit throughPOST /api/agent/inputwithmodeset toturn,steer, orqueue, then query/api/agent/turns/{thread_id}/{turn_id}or its/eventsNDJSON stream. Stop and resume use the corresponding/api/agent/turns/{thread_id}/stopand/resumeendpoints. Input IDs make repeated API submissions idempotent while their turns are retained. Steering is text-only; queued turns retain attachments and run settings. If an active turn ends before it reads a steering note, the note runs as a visible follow-up ahead of queued turns.max_queued_turnsalso reserves room for pending steering follow-ups;max_completed_turnsbounds retained results and idempotency records.[agent.clarification]: opt-in clarifying questions. Whenenabled = trueon a stateful runtime, the main agent gets anask_usertool and is told to call it once, before delegating, when a request is ambiguous in ways that would change what subagents do. The tool pauses the run with a LangGraph interrupt; your next message on the same thread is the answer and resumes the paused run instead of starting a new turn. A bare option number such as2picks that suggested option. While a question is pending, only plain text answers it: slash commands, configured response actions, and replies with attachments are refused with a reminder, and turns queued earlier through[agent.user_input]wait until the answer has run. When the model emits other tool calls next toask_user, they are dropped so no subagent starts before the answer. Chainlit shows the question with one button per option; the CLI and TUI print numbered options; the HTTP API reportsstatus: "awaiting_input"with the pendingclarificationson/api/agent/invokeand on the stream'sdoneevent, and aclarification_requestedstream event carries the question. Subagents never get the tool, and it is disabled whenagent.state = "stateless"because resuming needs a checkpointer. With the in-memory checkpointer a pending question is lost on restart.model: optional profile name or raw model name for the main/supervisor agent. CLI and environment model overrides take precedence.[agent.reflection]: optional correction-learning workflow.enabled = truerequiresstate = "stateful"and amemory_fileunder/memories/;max_lesson_charslimits proposal size;tool_failure_mode = "unrecovered"only proposes lessons for failed tool calls that do not produce a later final response.AGENTS.md: optional repo-root file that is automatically appended to the main/supervisor agent system prompt when present. It is not applied to separately configured async graph prompts.custom_instruction: optional string appended to the main/supervisor agent system prompt. This setting does not get applied to separately configured prompts such as theasync_researchergraph prompt.custom_instruction_file: optional UTF-8 text file loaded as the main-agent custom instruction. Relative paths resolve from the activedeepagent.toml. This is mutually exclusive withcustom_instruction.- ChainAgents explicitly restores
TodoListMiddlewarefor the main agent and local sync subagents because DeepAgents 0.7 no longer includes it by default. - DeepAgents still provides its own summarization middleware in the main agent and sync subagents.
summarization_trigger_tokens: optional positive integer token threshold for DeepAgents' built-in summarization middleware.summarization_keep_tokens: optional positive integer token budget to keep after DeepAgents summarizes conversation history.- Legacy
summarization_middleware_enabledentries are still parsed for compatibility, but ChainAgents no longer injects a second summarization middleware.
Local background execution lets the main agent or a nested synchronous agent start one of its configured children and continue immediately. Enable it with:
[agent.background_subagents]
enabled = true
stream_activity = true
batch_result_format = "markdown"
max_running_per_session = 4
max_running_total = 16
max_tasks_per_session = 100Then opt in each synchronous subagent that its direct parent may launch in the background:
[[subagents]]
name = "research-manager"
description = "Coordinates repository research and planning."
system_prompt = "Delegate focused work and synthesize the results."
background = true
[[subagents.subagents]]
name = "repo-planner"
description = "Turns repository findings into an implementation plan."
system_prompt = "Produce a concise, actionable implementation plan."
background = trueAn unmarked subagent remains available through the blocking task tool but is
rejected by spawn_background_task and run_subagent_batch. Marking a parent
does not implicitly mark its children; each background launch target opts in
independently.
The agent receives five tools:
spawn_background_task(description, subagent_type)starts an allowed direct child and returns a task ID immediatelyrun_subagent_batch(tasks)starts every independent task concurrently, waits for all of them, and returns their terminal reports in the globally configured format and in input orderlist_background_tasks()lists tasks visible to the calling agentget_background_task(task_id, wait_seconds=0)returns current state or waits up to 60 secondscancel_background_task(task_id)cancels the task and all descendants
For example, with the research-manager and private repo-planner nesting
shown above, the main agent can call:
spawn_background_task("Investigate the failing API tests", "research-manager")
When several tasks are independent, one tool call can fan out to separate subagent conversations:
run_subagent_batch(tasks=[
{"subagent_type": "research-manager", "description": "Trace the API failures."},
{"subagent_type": "research-manager", "description": "Check the related tests."}
])
The tasks input is unchanged across all result modes. Set
batch_result_format once under [agent.background_subagents]; it is not a
per-call argument. Entries stay in request order even when children finish in
another order, and a failed child does not discard successful sibling reports.
The default, batch_result_format = "markdown", returns one newline-delimited
Markdown document. Each section includes the agent name, task ID, terminal
status, original request, and either the report or error. Empty and cancelled
reports are marked explicitly. The line-oriented format also gives oversized
batches useful head/tail previews and lets the agent page through a complete
offloaded result with read_file.
batch_result_format = "json" restores the original structured contract. It
returns every terminal snapshot, including session, task-tree, and timing
metadata:
{
"results": [
{
"task_id": "bg-123",
"session_id": "thread-1",
"agent_name": "research-manager",
"description": "Trace the API failures.",
"agent_path": ["research-manager"],
"parent_task_id": null,
"status": "success",
"result": "The failure starts in ...",
"error": null,
"created_at": 1750000000.0,
"completed_at": 1750000001.5
}
]
}batch_result_format = "markdown_files" writes one standalone Markdown file
for every terminal result—including errors, cancellations, and empty reports—
then returns an ordered manifest:
{
"files": [
{
"task_id": "bg-123",
"agent_name": "research-manager",
"status": "success",
"path": "/workspace/.files/outputs/subagent-batches/batch-call-unique/01-research-manager-bg-123.md"
}
]
}Each file contains its agent name, task ID, status, original request, and report
or error. Batch directories are unique, filenames are sanitized and numbered in
request order, and descriptions never enter filenames. Files are written
through the generated-output backend under .files/outputs/subagent-batches,
so the existing UI can discover and download them. They persist across
conversation teardown and process restarts until removed as workspace outputs;
they are not temporary large-result offloads. The batch returns only after every
file succeeds; a write failure rolls back completed files, and a cleanup failure
is reported together with the original error.
The manager validates capacity for the whole batch before launch, so a batch that exceeds a configured limit starts no children. Cancelling the waiting call cancels its unfinished children and their descendants. In file mode, cancellation during output writes waits for the active write and removes every file completed by that interrupted batch.
Batch delegation is provider-independent. It is useful when a supervisor model, including a Snowflake Cortex model, can emit only one tool call per assistant turn: that single call starts separate child graph runs, and each child uses its own model request stream.
The running research-manager can independently launch its private child with
spawn_background_task("Identify the smallest fix", "repo-planner"). Either
agent can continue its current response, call list_background_tasks() later,
retrieve a terminal result with get_background_task("bg-..."), or stop its
visible subtree with cancel_background_task("bg-...").
The main agent can inspect the whole conversation task tree. Nested agents can inspect only their own task subtree and can spawn only their configured direct children. Background runs receive an isolated user message and checkpoint thread while retaining their configured model, skills, MCP tools, workspace, and shared memory access. Their state and streamed tokens are not merged into the parent response. Concurrent children can therefore observe the same workspace and memory resources; prompts should assign non-overlapping writes or otherwise coordinate shared updates.
By default, Chainlit and the interactive CLI/TUI post one status-only notice
when a task finishes; successful output remains available through
get_background_task instead of being copied into the notice. With
stream_activity = true, Chainlit additionally renders each background task as
a parent step with nested reasoning and tool-call steps, then closes that tree
before posting the same single terminal notice. Background reasoning and tool
steps follow [chainlit].reasoning_steps_enabled and
[chainlit].tool_steps_enabled, including the current chat settings switches;
the parent step and terminal notice remain visible when either is disabled.
The setting does not expose live background activity through the CLI, TUI, or
HTTP API. A one-shot CLI invocation prints the main response first, then waits
for its remaining background work. One-shot text output prints terminal results because
the process is about to exit; JSON output includes a background_tasks array.
The HTTP API exposes conversation-scoped list, get, cancel, and close operations
under /api/background-tasks.
curl "http://127.0.0.1:8000/api/background-tasks?thread_id=$THREAD_ID"
curl "http://127.0.0.1:8000/api/background-tasks/bg-123?thread_id=$THREAD_ID&wait_seconds=10"
curl -X DELETE "http://127.0.0.1:8000/api/background-tasks/bg-123?thread_id=$THREAD_ID"
curl -X DELETE "http://127.0.0.1:8000/api/background-tasks?thread_id=$THREAD_ID"Tasks are retained until the conversation closes (or until the oldest finished
tasks are evicted to stay within max_tasks_per_session) and are cancelled before its
offloaded large tool results and MCP resources are released. Cleanup is scoped
by thread ID, so closing one conversation does not remove another conversation's
artifacts; workspace files, generated downloads, memories, and uploads are not
part of this cleanup. Tasks and offload ownership are stored only in the current
process and do not survive restarts. Agent Server deployments therefore require
session affinity when multiple workers are used. The custom Agent Server app
exposes DELETE /background-tasks/sessions/{thread_id} for explicit cleanup and
closes all remaining managers and tracked offloads during server shutdown.
You can configure slash-style commands that run from the Chainlit composer before the model call. Chainlit also auto-generates slash commands for loaded skills.
Place this config in deepagent.toml or whatever file DEEPAGENT_CONFIG points to.
Do not put it in the app UI file chainlit.toml or Chainlit's native .chainlit/config.toml.
Example:
[chainlit]
# Set false to hide model selection in chat settings and Modes.
model_mode_enabled = true
# Set false to disable per-message reasoning overrides from the Modes picker.
reasoning_mode_enabled = true
# Set false to hide streamed reasoning step panels and reasoning task entries.
reasoning_steps_enabled = true
# Set false to hide streamed tool step panels and tool task entries.
tool_steps_enabled = true
# Set false to hide the initial startup status message ("Workspace agent ready...").
startup_status_enabled = true
# Set false to hide generated Chainlit CustomElement panels and remove the render tool.
generative_ui_enabled = true
# Set false to keep legacy non-chronological streaming order in Chainlit.
chronological_ui_enabled = true
commands = [
{ name = "ask-researcher", description = "Delegate to repo-researcher.", target = "subagent", value = "repo-researcher", template = "{input}" },
{ name = "repo-readme", description = "Run an MCP tool directly.", target = "mcp_tool", value = "repo_read_file", mcp_server = "repo", template = "{\"path\":\"README.md\"}" },
{ name = "summarize", description = "Apply a prompt template.", target = "prompt", value = "Summarize the input", template = "Summarize this:\n{input}" }
]
starters = [
{ label = "Explain this repo", message = "Explain the architecture of this repository and identify the most important files.", command = "ask-researcher", icon = "book-open" },
{ label = "Review current changes", message = "Review the current working tree changes for bugs, regressions, and missing tests." }
]
response_actions = [
{ name = "summarize", label = "Summarize", icon = "list", prompt = "Summarize this response concisely:\n\n{response}" },
{ name = "explain", label = "Explain", icon = "book-open", prompt = "Explain this response in more detail.\n\nOriginal request:\n{prompt}\n\nResponse:\n{response}" }
]target modes:
prompt: rewrites the user prompt before sending it to the agent.subagent: rewrites the user prompt to direct the runtime to delegate via the configured subagent.mcp_tool: invokes the configured MCP tool directly and returns tool output in chat.
Notes:
- The
[chainlit]table for native commands belongs indeepagent.toml, alongside[model],[agent],[mcp],[[subagents]], and[[async_subagents]]. [chainlit].model_mode_enabled = falsehides the Model selector in chat settings and the Model mode group, and ignores per-message model overrides from UI modes.[chainlit].reasoning_mode_enabled = falsehides the Reasoning mode group and ignores per-message reasoning overrides from UI modes.[chainlit].reasoning_steps_enabled = falsehides streamed reasoningcl.Steppanels and reasoning task-list entries while preserving model reasoning settings.[chainlit].tool_steps_enabled = falsehides streamed toolcl.Steppanels and tool task-list entries while preserving tool execution.[chainlit].startup_status_enabled = falsedisables the initial startup status message that summarizes runtime configuration.[chainlit].generative_ui_enabled = falsehides generated Chainlit CustomElement panels and removes therender_chainlit_uitool from the main agent.[chainlit].chronological_ui_enabled = falsedisables chronological UI ordering so response tokens stream immediately and reasoning steps are not force-rolled at tool boundaries.- Command
nameis invoked as/<name>and must be unique. templateis optional and may include{input}.- For
mcp_tool, user command arguments must be valid JSON, e.g./repo-readme {"path":"README.md"}. - Each discovered skill also becomes
/<skill-name>automatically. For example, a skill withname: revieweris available as/reviewer. - Skill-backed commands always force the main agent to read the selected
SKILL.mdfirst and use it for that turn. - If a configured
[chainlit].commandsentry and a skill share the same slash name, the configured command wins. startersdefine starter prompts shown by Chainlit before the first message in a thread.- Starter
labelandmessageare required. Startercommandandiconare optional. response_actionsappear after Markdown and PDF beneath completed Chainlit replies. Each action needs a uniquename, alabel, and aprompt;iconandtooltipare optional. Omit the list or set it to[]to hide custom actions.- A response-action prompt can use
{response}for the clicked reply and{prompt}for the request that produced it. Other braces stay literal. The agent receives the expanded prompt in the current conversation and displays its reply normally, but Chainlit does not show a user message or prompt text in reasoning panels for the action. The prompt remains part of agent state and persisted response context. - Saved Chainlit chats restore response actions when response context is available. Restart the app after editing
deepagent.tomlto reload action definitions.
Async subagents are loaded from deepagent.toml as background Agent Protocol jobs. They are useful for long-running or remote work where the main agent should return a task ID immediately and let you check, update, cancel, or list tasks later.
Example:
[[async_subagents]]
name = "remote-researcher"
description = "Runs longer research jobs in the background on an Agent Protocol server."
graph_id = "researcher"
# Omit url for ASGI transport in a co-deployed LangGraph setup.
# Set url for HTTP transport to a remote Agent Protocol server.
# url = "https://researcher-deployment.langsmith.dev"
# headers = { Authorization = "Bearer ${RESEARCHER_TOKEN}" }Supported async subagent fields:
name: requireddescription: requiredgraph_id: required graph or assistant ID on the Agent Protocol serverurl: optional remote Agent Protocol server URL; omit for ASGI transport in co-deployed LangGraph setupsheaders: optional request headers for remote/self-hosted Agent Protocol servers
For compatibility with DeepAgents' native discriminator, a [[subagents]] entry with a graph_id is also treated as an async subagent. Async subagents cannot define sync-only fields such as system_prompt, skills, mcp_servers, or model; those capabilities are configured on the remote graph.
This repo includes a LangGraph co-deployment entrypoint for ASGI transport:
- langgraph.json registers
supervisorandasync-researcher - langgraph_app.py exports both graphs
- omit
urlindeepagent.tomlwhen running through LangGraph Agent Server
Run the co-deployed Agent Protocol server with enough worker capacity for the supervisor plus background tasks:
uv run --with "langgraph-cli[inmem]" langgraph dev --n-jobs-per-worker 10ASGI transport is only available in this LangGraph server path. If you launch the UI with chainlit run main.py -w, use HTTP transport instead by setting url = "http://127.0.0.1:2024" on the async subagent.
Chainlit also starts a background notifier for launched async tasks. It polls the Agent Protocol server and posts a message when a task reaches success, error, cancelled, interrupted, or timeout. If deepagent.toml omits url for ASGI co-deployment, Chainlit defaults to http://127.0.0.1:2024, the usual langgraph dev URL. Override it with:
export CHAINLIT_ASYNC_SUBAGENT_URL="http://127.0.0.1:2024"Optional:
export CHAINLIT_ASYNC_TASK_POLL_SECONDS="5"MCP servers are defined once in deepagent.toml and then attached by name to the main agent or any subagent.
Example:
[mcp]
tool_name_prefix = true
stateful = true
[mcp.servers.repo]
transport = "stdio"
command = "npx"
args = ["-y", "@modelcontextprotocol/server-filesystem@2025.8.21", "."]
cwd = "."
[mcp.servers.docs]
transport = "http"
url = "http://localhost:8000/mcp"
[mcp.servers.github]
transport = "sse"
url = "http://localhost:8080/sse"
headers = { Authorization = "Bearer ${GITHUB_TOKEN}" }
[agent]
mcp_servers = ["repo"]
[[subagents]]
name = "repo-researcher"
description = "Researches the repo and docs."
system_prompt = "Use the repo and docs MCP servers to answer questions."
mcp_servers = ["repo", "docs"]
[[subagents]]
name = "release-assistant"
description = "Works with repository metadata and hosted services."
system_prompt = "Use the GitHub MCP server when release metadata is needed."
mcp_servers = ["github"]Supported MCP config fields:
- top-level
[mcp] tool_name_prefix = true|false - top-level
[mcp] stateful = true|false [mcp.servers.<name>] transport- for
stdio:command,args, optionalcwd, optionalenv - for
http,streamable_http,streamable-http:url, optionalheaders - for
sse:url, optionalheaders - for
websocket:url
Notes:
mcp_serverson[agent]attaches those MCP tools to the main agent.mcp_serverson[[subagents]]attaches those MCP tools only to that subagent.mcp_serverson[[subagents.subagents]]attaches those MCP tools only to that nested child subagent.- Skills and MCP servers are independent. You can use neither, either, or both on any subagent.
- If one MCP server cannot load tools, the agent continues with tools from healthy servers and retries the failed server on the next run. The affected run shows an MCP warning in Chainlit, CLI/TUI, and API output. Tool invocation errors are returned to the model as recoverable tool errors.
- Relative
cwdvalues are resolved from the location ofdeepagent.toml. tool_name_prefix = trueis recommended when multiple MCP servers expose overlapping tool names.stateful = truekeeps MCP sessions open per conversation scope while the app process is running. Chainlit uses the effective thread ID, so reopening a saved chat reuses its MCP tools and transport. Turns on the same thread run in order when stateful MCP is enabled.- After the last Chainlit session leaves a conversation, its resources remain available for 10 minutes. At most four idle conversations are retained; the oldest idle scope is closed when a fifth becomes idle. Active conversations are never evicted by this limit.
- MCP servers are discovered concurrently, and each
(scope, server)discovery is shared by callers already waiting for it. Failed servers remain retryable on a later run. stateful = falserecreates the MCP session for every tool call.
Current scope of this config support:
- it exposes
ls,read_file,write_file,edit_file,glob, andgrepby default, pluswrite_todos, subagent tools, config-driven skills, and MCP tools - it exposes recursive
deleteonly when[agent].delete_tool_enabled = true - it exposes
executeonly when[agent].execute_tool_enabled = true; the configured backend must also implement sandbox execution - it supports config-driven sync subagents and async Agent Protocol subagents
- it does not yet provide a config-driven registry for custom Python tools per subagent beyond MCP
- if you need custom Python tools, define them alongside the generated UI tool in chainagents/runtime/commands.py, then add them in
build_main_toolsin graph.py; both the static LangGraph graph and the live runtime agent (lifecycle.py) assemble their tools through that one function
See deepagent.toml.example for a portable baseline and the sections above for optional MCP and subagent examples.
/workspace/maps to this repo on disk./memories/is available in stateful mode under the configured agent-scoped namespace and durable across LangGraph threads only whenDATABASE_URLis configured.- any other absolute path is treated as ephemeral scratch space by the deep agent backend.
- Native Chainlit history is available when
DATABASE_URL,CHAINLIT_AUTH_SECRET, and eitherCHAINLIT_AUTH_USERSor the legacy username/password pair are configured. - If
DATABASE_URLis set but authentication is not configured, Chainlit still persists thread records, but they are not browseable from the UI. - When
DATABASE_URLis unset, thread IDs only persist while the process stays alive. - When
DATABASE_URLis set, durable state is available through LangGraph thread IDs. You can reuse a thread ID from the chat settings panel to continue the same checkpointed thread. - When
[agent].state = "stateless", thread IDs still identify requests and MCP/RAG scopes, but the agent graph does not checkpoint conversation state, receive a LangGraph store, or expose/memories/. - MCP stateful sessions are process-local. They survive tool calls and saved-chat navigation in the same Chainlit thread until idle eviction, but not an app restart.
- Switching between saved Chainlit chats keeps active turns and local background subagents running. Returning during the same app process restores recent terminal background-task notices that finished while the chat was away. Idle conversation resources still follow the configured retention window.
- On startup, the UI shows how many skill sources, MCP servers, custom subagents, and async subagents were loaded from
deepagent.toml.
The public runtime API is chainagents.runtime. runtime/core.py explicitly
re-exports the same objects for compatibility. The root-level deepagent_runtime
module remains an alias of that facade but is deprecated and will be removed in a
future release; import from chainagents.runtime instead.
Modules in chainagents/runtime/ |
Responsibility |
|---|---|
constants.py, types.py |
Shared defaults, workspace root, and configuration records |
model_config.py |
Provider settings, endpoint normalization, and model profiles |
extension_config.py, config.py |
Extension parsing and TOML/environment configuration |
providers.py, models.py |
Provider adapters and configured model construction |
backends.py |
Workspace paths and filesystem/storage backend routing |
middleware.py |
Tool resilience and summarization middleware |
commands.py |
Generated UI tools and skill command discovery |
artifacts.py |
Session-scoped storage for offloaded large tool results |
graph.py |
Shared agent assembly (main tools, subagent specs, agent kwargs) used by both the static graph and the live runtime |
tracing.py |
Langfuse callbacks and LangGraph run configuration |
mcp_sessions.py |
MCP session pool, stateful transport ownership, and tool discovery caching |
rag_ops.py |
RAG index status, rebuild, and thread-scoped upload operations |
background_tasks/ |
Process-local background execution of configured synchronous subagents |
lifecycle.py, reflection.py |
Agent/MCP/persistence lifecycle and correction reflection |
Implementation modules import their lower-level owners directly; they never import
core.py or the package facade. Cross-module function calls resolve through the
owning module so tests can patch that owner (for example,
chainagents.runtime.models.build_model). Facade globals are not patch seams.
Warning filters are installed by the runtime package before SDK imports.
This hobby, personal project is made available under the MIT License and is built with open source Python libraries. See LICENSE for the full text.
