Skip to content

About

This project runs a local-first LangChain DeepAgent behind a Chainlit UI and dedicated CLI.

Topics

Resources

Stars

1 star

Watchers

0 watching

Forks

Latest commit

 

History

455 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

CI Python License Chainlit LangChain

This hobby, personal project powers a highly configurable, local-first LangChain DeepAgent. We also provide interface to the agent though Chainlit, CLI and TUI.

Implemented features include:

  • Sub-agents — delegate tasks to agents with isolated context windows
  • Filesystem — read, write, edit, or search over pluggable local, sandboxed, or remote backends
  • Context management — summarize long threads and offload tool outputs to disk
  • Persistent memory — pluggable state and store backends for cross-session recall
  • Skills — reusable behaviors the agent can load on demand
  • Tools — bring your own functions or any MCP server

Highlights

  • ChatOllama, ChatOpenAI, or ChatAnthropic with configurable model backends
  • native Chainlit streaming for reasoning, tool calls, and final response
  • optional Langfuse tracing through the LangChain callback handler
  • Chainlit image uploads sent to vision-capable models as photo attachments for OCR or image analysis
  • Chainlit OCR/image uploads accept PNG, JPEG, WEBP, and GIF files
  • config-driven synchronous and async DeepAgents subagents
  • DeepAgents ==0.7.17 with explicit todo planning and safe filesystem defaults
  • per-response download buttons for Markdown and PDF exports
  • configurable Chainlit response actions that ask the agent without showing a user prompt
  • Postgres-backed LangGraph checkpoints and durable /memories/ when DATABASE_URL is set
  • repo files mounted for the agent under /workspace/
  • Chainlit Modes support for per-message reasoning selection (Low, Medium, High)

Environment

Set these variables before starting the app if you want environment-based overrides:

export DATABASE_URL="postgresql://USER:PASSWORD@HOST:5432/DBNAME?sslmode=disable"
export DEEPAGENT_MODEL_PROVIDER="ollama"
export DEEPAGENT_MODEL_BASE_URL="http://127.0.0.1:11434"
# export DEEPAGENT_MODEL_ENDPOINT_URL="https://api.example.test/custom/v1/messages"
# export DEEPAGENT_MODEL_DISABLE_STREAMING_FOR_TOOL_CALLS="true"
export DEEPAGENT_MODEL_NAME="gpt-oss:20b"
export DEEPAGENT_MODEL_REASONING="medium"
export DEEPAGENT_RECURSION_LIMIT="200"
# export DEEPAGENT_MODEL_API_KEY="optional-for-secured-openai-compatible-servers"
# export ANTHROPIC_API_KEY="required-for-provider-anthropic-unless-DEEPAGENT_MODEL_API_KEY-is-set"
# export SNOWFLAKE_PAT="required-for-provider-snowflake_cortex-unless-a-CLI-or-generic-key-is-set"
# export AWS_REGION="us-east-1"  # provider = "bedrock" uses the standard AWS credential chain
export DEEPAGENT_CONFIG="deepagent.toml"
export CHAINLIT_AUTH_SECRET="replace-with-a-long-random-string"
export CHAINLIT_AUTH_USERS='{"admin":"change-me","alice":"alice-password"}'
# export LANGFUSE_PUBLIC_KEY="pk-lf-..."
# export LANGFUSE_SECRET_KEY="sk-lf-..."
# export LANGFUSE_BASE_URL="https://cloud.langfuse.com"

DATABASE_URL is optional now:

  • when set, LangGraph checkpoints and /memories/ are persisted in Postgres
  • when unset, the app falls back to in-memory persistence for the current process only
  • if [agent].state = "stateless", LangGraph checkpoint and store handles are not opened or passed to the agent graph even when DATABASE_URL is set

DEEPAGENT_CONFIG is optional:

  • defaults to deepagent.toml in the project root
  • if the file is missing, the app falls back to built-in model defaults and runs without extra skills, MCP servers, or custom subagents

DEEPAGENT_MODEL_* variables are optional:

  • they override the matching [model] values in deepagent.toml
  • DEEPAGENT_MODEL_API_KEY is used for secured OpenAI-compatible servers and can also supply the Anthropic API key when ANTHROPIC_API_KEY is unset
  • ANTHROPIC_API_KEY is read first when provider = "anthropic" or provider = "claude", so stale generic keys do not override the Claude credential
  • SNOWFLAKE_PAT is read first when provider = "snowflake_cortex"; Cortex still requires a key, resolved in this order: --api-key, SNOWFLAKE_PAT, DEEPAGENT_MODEL_API_KEY, then [model].api_key
  • provider = "bedrock" and provider = "anthropic_bedrock" ignore API-key variables and uses the standard AWS credential chain (AWS_ACCESS_KEY_ID/AWS_SECRET_ACCESS_KEY, AWS_PROFILE, SSO, or an instance/task role); set the region with AWS_REGION or AWS_DEFAULT_REGION
  • when switching to Anthropic with DEEPAGENT_MODEL_PROVIDER, unset stale DEEPAGENT_MODEL_BASE_URL; use DEEPAGENT_MODEL_ENDPOINT_URL with the /v1/messages path for env-based Anthropic proxy switches, or pass --base-url explicitly from the CLI
  • DEEPAGENT_MODEL_DISABLE_STREAMING accepts true, false, or tool_calling; DEEPAGENT_MODEL_DISABLE_STREAMING_FOR_TOOL_CALLS=true is a convenience alias for tool_calling
  • OLLAMA_BASE_URL, OLLAMA_MODEL, and OLLAMA_REASONING remain supported as Ollama-only compatibility aliases

DEEPAGENT_RECURSION_LIMIT is optional:

  • it overrides [agent].recursion_limit in deepagent.toml
  • it controls the maximum LangGraph steps for a single agent run
  • raise it when long tool-heavy Deep Agent runs hit GraphRecursionError

CHAINLIT_AUTH_SECRET and Chainlit user credentials are optional:

  • when CHAINLIT_AUTH_SECRET and CHAINLIT_AUTH_USERS are set, the app enables Chainlit password authentication for each configured user
  • CHAINLIT_AUTH_USERS must be a JSON object mapping usernames to passwords, e.g. {"admin":"change-me","alice":"alice-password"}
  • the legacy CHAINLIT_AUTH_USERNAME and CHAINLIT_AUTH_PASSWORD pair still works for a single user when CHAINLIT_AUTH_USERS is unset
  • together with DATABASE_URL, that unlocks the native Chainlit history bar and chat resume UI
  • when auth credentials are unset, the app stays unauthenticated and the history bar remains unavailable

Optional: Install Postgres

You only need Postgres if you want durable LangGraph checkpoints and /memories/. If DATABASE_URL is unset, the app runs fully in memory for the current process.

This repo includes a Compose file for a local Postgres instance:

docker compose up -d postgres

Point the app at that database:

export DATABASE_URL="postgresql://chainagents:chainagents@127.0.0.1:5432/chainagents?sslmode=disable"

Optional verification:

docker compose exec postgres psql -U chainagents -d chainagents -c "select 1;"

Notes:

  • The Compose file lives at compose.yaml and creates a persistent postgres-data volume.
  • If you already have Postgres installed locally, create an empty database and set DATABASE_URL to that instance instead.
  • No separate migration step is required for this app. On startup it calls the LangGraph Postgres store/checkpointer setup() routines and creates any missing Chainlit persistence tables ("User", "Thread", "Step", "Feedback", and "Element") automatically.
  • If you manage the Chainlit schema externally, set CHAINLIT_SCHEMA_BOOTSTRAP=false before launching the app to skip the automatic Chainlit table bootstrap.

Optional: Enable Native Chainlit History

Chainlit only shows its built-in history sidebar when both persistence and authentication are enabled.

This app includes a simple password-based auth callback driven by environment variables:

export CHAINLIT_AUTH_SECRET="replace-with-a-long-random-string"
export CHAINLIT_AUTH_USERS='{"admin":"change-me","alice":"alice-password"}'

For compatibility, a single user can still be configured with:

export CHAINLIT_AUTH_SECRET="replace-with-a-long-random-string"
export CHAINLIT_AUTH_USERNAME="admin"
export CHAINLIT_AUTH_PASSWORD="change-me"

With DATABASE_URL, CHAINLIT_AUTH_SECRET, and either CHAINLIT_AUTH_USERS or the legacy username/password pair set:

  • users can sign in through Chainlit's native auth screen
  • the history sidebar can list and reopen prior chats
  • resumed chats default the LangGraph thread ID to the persisted Chainlit thread ID for that conversation

If you leave auth disabled, Chainlit can still persist thread records in Postgres, but the native history bar will stay hidden.

Setup

Install dependencies, then either pull an Ollama model or point deepagent.toml at an OpenAI-compatible server such as LM Studio:

uv sync
ollama pull gpt-oss:20b

The dependency manifests require DeepAgents ==0.7.17; uv sync installs the matching locked release.

PDF downloads are rendered with WeasyPrint. uv sync installs the Python package, but WeasyPrint also needs native rendering libraries. On macOS, install them with:

brew install weasyprint

On Linux, install the Pango packages listed in the WeasyPrint installation guide for your distribution before starting the app.

Response PDFs include images referenced with public HTTP or HTTPS Markdown image URLs. Each export downloads at most 20 unique images, with a 10 MiB per-image and 25 MiB aggregate limit, and processes at most 50 million raster pixels. Images that are unavailable, invalid, oversized, or hosted on private network addresses are replaced with a labeled placeholder so the rest of the PDF can still be downloaded.

If you are using LM Studio or another OpenAI-compatible server instead of Ollama, skip ollama pull, load a model in that server, and set [model].provider = "openai_compatible" with the server's base_url. If you are using Claude through Anthropic, set [model].provider = "anthropic" and provide ANTHROPIC_API_KEY or DEEPAGENT_MODEL_API_KEY. For Snowflake Cortex, use the dedicated snowflake_cortex provider and a Snowflake PAT as shown in Snowflake Cortex. For Amazon Bedrock, set [model].provider = "bedrock" with a Bedrock model or inference-profile ID and configure AWS credentials and AWS_REGION as shown in Amazon Bedrock. For Claude on Bedrock through the Anthropic Messages API, use provider = "anthropic_bedrock".

If you enable workspace-docs RAG with Ollama embeddings, also pull an embedding model such as:

ollama pull nomic-embed-text

This repo includes a portable deepagent.toml with:

  • Ollama at http://127.0.0.1:11434 with gpt-oss:20b
  • a higher LangGraph recursion limit for longer tool-heavy Deep Agent runs
  • recursive file deletion disabled unless [agent].delete_tool_enabled = true
  • command execution disabled unless [agent].execute_tool_enabled = true
  • RAG, reflection, Langfuse, MCP servers, and subagents disabled until configured
  • the repo-local skills/ source for the main agent

See deepagent.toml.example and the sections below for provider and optional integration examples.

Run

Start the Chainlit app:

uv run chainlit run main.py -w

Start the FastAPI server:

uv run chainagents-api --host 127.0.0.1 --port 8000

Development validation

Install the locked development environment, then run the same gates used by CI:

uv sync --locked
uv run ruff check chainagents *.py scripts/*.py tests
uv run mypy
uv run pytest
bash scripts/verify-installed-wheel.sh

The wheel check builds the real distribution, installs its locked runtime dependencies in a fresh virtual environment, and runs import, CLI, API, and configuration smoke checks from a temporary user working directory.

The API uses the same deepagent.toml and environment settings as the Chainlit and CLI entrypoints. It serves one trusted owner: tokenless access requires a loopback peer and safe local Host/Origin, while remote binds require CHAINAGENTS_API_TOKEN and bearer Authorization headers. See API access, browser deployment, and request limits. Useful endpoints include:

curl http://127.0.0.1:8000/health
curl http://127.0.0.1:8000/api/status
THREAD_ID="api-$(uuidgen)"
curl -X POST http://127.0.0.1:8000/api/agent/invoke \
  -H "Content-Type: application/json" \
  -d "{\"prompt\":\"Summarize this repository\",\"thread_id\":\"$THREAD_ID\"}"
curl -N -X POST http://127.0.0.1:8000/api/agent/stream \
  -H "Content-Type: application/json" \
  -d "{\"prompt\":\"Summarize this repository\",\"thread_id\":\"$THREAD_ID\"}"

JSON requests to /api/agent/invoke and /api/agent/stream accept an optional command field. When present, prompt is passed to that configured native command as its argument text. The multipart /api/agent/stream/multipart endpoint accepts the same optional command form field alongside prompt, thread_id, and uploaded files. When an MCP server is unavailable, /api/agent/invoke includes a warnings array while /api/agent/stream emits an mcp_status warning event before the agent response. Healthy tools remain available.

The /api/status response sources its starters and their optional command values from the active deepagent.toml. A client launching a configured starter should send the starter message as prompt and its command separately; the runtime then applies the configured command template to that prompt.

Run the same underlying agent from a terminal without the Chainlit UI:

uv run chainagents --prompt "Summarize this repository" --thread-id cli

Start the full-screen terminal UI:

uv run chainagents --tui

The TUI defaults to thread ID tui, keeps the prompt box at the bottom, shows the conversation in the main pane with Markdown-formatted assistant responses, and splits reasoning and tool activity in the right sidebar. The prompt editor supports multiple lines: press Shift+Enter to insert a newline and Enter to send the complete prompt. Type / to show configured slash commands, and press Tab to complete the first matching command. Stdio MCP server diagnostics are written to .files/tui-stderr.log in TUI mode so they do not corrupt the full-screen interface.

Useful CLI examples:

uv run chainagents --status --no-rag
uv run chainagents --configure
uv run chainagents --tui --reasoning high
uv run chainagents --list-commands
uv run chainagents --command summarize --prompt "Summarize the config entrypoints"
uv run chainagents --stdin --model gpt-oss:20b --reasoning high < prompt.txt
uv run chainagents --rebuild-rag
uv run chainagents --upload-rag notes.md --prompt "Use my uploaded notes"
uv run chainagents --photo scene.jpg --prompt "Describe this photo"

Run uv run chainagents --help for all runtime flags, including model provider, base URL, endpoint URL, API key, temperature, persistence, MCP session scope, async subagent URL, RAG controls, photo attachments, streaming, reasoning traces, tool traces, and JSON output.

Project Structure

Core Python code lives under the chainagents/ package. The root-level Python files are compatibility wrappers and entrypoints so existing imports and commands continue to work.

chainagents/
  runtime/              Core DeepAgents runtime, model setup, config parsing,
                        MCP/tool loading, persistence backends, and Langfuse.
  interfaces/
    chainlit/           Chainlit callbacks, UI bridge, auth, persistence,
                        uploads, async task notifications, and chat settings.
    cli/                Terminal CLI parser, status output, command execution,
                        upload handling, and event rendering.
    tui/                Full-screen Textual terminal UI.
    api/                FastAPI application, request schemas, and streaming API.
  turns/                Shared TurnRunner: one agent turn (commands, uploads,
                        streaming, generated files) used by every interface.
  events/               Shared LangGraph stream normalization used by all
                        interfaces.
  commands/             Native slash-command parsing and dispatch helpers.
  rag/                  Workspace documentation RAG config, index, uploads,
                        and search tool.
  exports/              Markdown and PDF response export helpers.
  langgraph/            Agent Server graph exports.
  util/                 Shared utility helpers.

Runtime assets stay at the repository root because they are user/configuration content rather than importable Python package code:

  • deepagent.toml and deepagent.toml.example: model, agent, MCP, RAG, Chainlit, Langfuse, and subagent configuration.
  • skills/: Deep Agents skill sources referenced from TOML as skills.
  • prompts/: prompt files referenced by configured subagents.
  • public/ and .chainlit/: Chainlit static assets and native Chainlit config.
  • tests/: regression tests for runtime, interfaces, RAG, exports, and events.

Compatibility wrappers such as main.py, deepagent_runtime.py, chainlit_bridge.py, chainagents_cli.py, chainagents_api.py, rag_runtime.py, and response_exports.py import the moved package modules. Prefer new code under chainagents/, but keep the wrappers until external users no longer rely on the old import paths.

Deprecated: every root-level wrapper except main.py and langgraph_app.py (which stay silent for chainlit run main.py -w and langgraph.json) now emits a DeprecationWarning on first import and will be removed in a future release. Import from the package path instead, e.g. chainagents.runtime.core instead of deepagent_runtime.

Model Config

You can keep the model defaults in deepagent.toml:

[model]
provider = "ollama"
base_url = "http://127.0.0.1:11434"
temperature = 0
max_tokens = 4096
repeat_penalty = 1.1
name = "gpt-oss:20b"
models = ["gpt-oss:20b", "gemma4:27b"]
reasoning_effort = "medium"
# Disable streaming only for requests that include tools, which can help
# model servers that emit malformed streamed tool-call chunks.
disable_streaming_for_tool_calls = false

For LM Studio or another OpenAI-compatible server:

[model]
provider = "openai_compatible"
base_url = "http://127.0.0.1:1234/v1"
temperature = 0
name = "your-loaded-model-id"
reasoning_effort = "medium"
# api_key = "optional"

For OpenAI-compatible servers with a non-standard full chat-completions endpoint:

[model]
provider = "openai_compatible"
endpoint_url = "https://api.example.test/openai/deployments/local/chat/completions?api-version=2026-01-01"
name = "your-loaded-model-id"
# api_key = "optional"

Snowflake Cortex

Snowflake Cortex uses the canonical provider value snowflake_cortex (no aliases). It requires a key. Set SNOWFLAKE_PAT, DEEPAGENT_MODEL_API_KEY, or [model].api_key, or pass --api-key for a one-off CLI run. Credential precedence is --api-key, SNOWFLAKE_PAT, DEEPAGENT_MODEL_API_KEY, then [model].api_key.

Use either the Chat Completions base URL or the complete Chat Completions endpoint:

[model]
provider = "snowflake_cortex"
base_url = "https://<account-identifier>.snowflakecomputing.com/api/v2/cortex/v1"
name = "claude-sonnet-4-5"
max_tokens = 4096
# api_key = ""  # optional only when SNOWFLAKE_PAT or DEEPAGENT_MODEL_API_KEY is set
[model]
provider = "snowflake_cortex"
endpoint_url = "https://<account-identifier>.snowflakecomputing.com/api/v2/cortex/v1/chat/completions"
name = "claude-sonnet-4-5"

For the CLI, the same settings can be supplied without editing TOML:

export SNOWFLAKE_PAT="your-snowflake-pat"
uv run chainagents --provider snowflake_cortex \
  --base-url "https://<account-identifier>.snowflakecomputing.com/api/v2/cortex/v1" \
  --model claude-sonnet-4-5 --prompt "Summarize this repository"

The auto RAG embedding provider is not valid for a Cortex chat model. If RAG is enabled, set [rag.embedding].provider to ollama or openai_compatible, together with an appropriate embedding model and base_url (and api_key when needed). Tool-call IDs are normalized only for Snowflake Cortex outbound Chat Completions requests; other OpenAI-compatible providers keep their original tool-call IDs.

For Claude through Anthropic:

[model]
provider = "anthropic"
temperature = 0
name = "claude-sonnet-4-6"
models = ["claude-sonnet-4-6", "claude-opus-4-8", "claude-haiku-4-5-20251001"]
reasoning_effort = "medium"
thinking = "auto"
# api_key = "optional-if-ANTHROPIC_API_KEY-or-DEEPAGENT_MODEL_API_KEY-is-set"
# base_url = "https://api.anthropic.com"
# endpoint_url = "https://claude-proxy.example/proxy/v1/messages"

Amazon Bedrock

For models hosted on Amazon Bedrock (Claude, Amazon Nova, Llama, Mistral, gpt-oss, and others) through the Bedrock Converse API:

[model]
provider = "bedrock"
temperature = 0
name = "us.anthropic.claude-sonnet-5"
models = ["us.anthropic.claude-sonnet-5", "amazon.nova-pro-v1:0"]
reasoning_effort = "medium"
thinking = "auto"
# endpoint_url = "https://vpce-0123.bedrock-runtime.us-east-1.vpce.amazonaws.com"
export AWS_REGION="us-east-1"
export AWS_PROFILE="my-profile"  # or AWS_ACCESS_KEY_ID / AWS_SECRET_ACCESS_KEY, SSO, or an instance role
  • name is a Bedrock model ID (such as amazon.nova-pro-v1:0), a cross-region inference-profile ID (such as us.anthropic.…), or a foundation-model/system inference-profile ARN. Application inference profiles and provisioned models are not supported, as an ID or an ARN, because they hide the underlying model; use that model or cross-region inference-profile ID instead.
  • Credentials and region come from the standard AWS chain; api_key and DEEPAGENT_MODEL_API_KEY are not used. Bedrock API keys work through boto3's own AWS_BEARER_TOKEN_BEDROCK variable.
  • base_url or endpoint_url optionally overrides the Bedrock runtime endpoint, for example a VPC interface endpoint; leave both unset to use the regional default.
  • provider = "aws_bedrock" and provider = "amazon_bedrock" are accepted as aliases.
  • reasoning_effort is forwarded for models whose langchain-aws profile declares configurable reasoning (for example Claude Opus/Sonnet 5, Nova 2, and gpt-oss) and ignored for others. For Claude this enables adaptive thinking, and the sampling temperature is dropped because Claude rejects it while thinking. Set thinking = "disabled" to turn reasoning off; for Claude Opus 5 and Sonnet 5, which think by default, this sends an explicit disabled-thinking setting. Claude Opus 5.5, Sonnet 5.5 and Fable 5 cannot run without thinking, so thinking = "disabled" is rejected for them. gpt-oss always reasons, so it rejects thinking = "disabled" too; use reasoning_effort = "low".
  • Amazon Nova rejects a temperature of exactly 0, so temperature = 0 is sent as 0.00001. At high reasoning effort Nova accepts no sampling settings, so the temperature is dropped.
  • Reasoning blocks streamed by Bedrock are shown as thinking, separate from the answer.
  • Leave disable_streaming unset to keep langchain-aws's per-model default (for example, models that cannot stream tool calls fall back to non-streaming tool requests); set disable_streaming = true or "tool_calling" to force non-streaming requests.
  • Workspace-docs RAG cannot infer Bedrock embeddings. If [rag].enabled = true, set [rag.embedding].provider to ollama or openai_compatible, together with an appropriate embedding model and base_url (and api_key when needed).

Claude on Amazon Bedrock (Anthropic Messages API)

provider = "bedrock" uses the Converse API, which works for every Bedrock model. To run Claude on Bedrock through the Anthropic Messages API instead, use provider = "anthropic_bedrock". It uses the Anthropic SDK's Bedrock client (langchain-aws ChatAnthropicBedrock), so Claude gets the same request handling as provider = "anthropic": effort, adaptive thinking, and Anthropic content blocks.

[model]
provider = "anthropic_bedrock"
name = "us.anthropic.claude-sonnet-4-6"
reasoning_effort = "medium"
thinking = "auto"
# endpoint_url = "https://vpce-0123.bedrock-runtime.us-east-1.vpce.amazonaws.com"
  • Credentials, region and endpoint_url work the same way as for provider = "bedrock"; no API key is used.
  • name is a Bedrock Claude model ID or cross-region inference-profile ID.
  • reasoning_effort is only sent to Claude models that support effort; it is skipped for models such as Claude 3.x and Haiku 4.5.
  • thinking = "disabled" explicitly turns thinking off on Claude Opus 5 and Sonnet 5, which think by default. Claude Opus 5.5, Sonnet 5.5 and Fable 5 cannot run without thinking, so that setting is rejected for them.
  • Only Anthropic Claude model IDs are accepted; use provider = "bedrock" for other Bedrock models. Foundation-model and system inference-profile ARNs work, but application inference profiles and provisioned models hide the underlying model, so use the Claude model or inference-profile ID instead.
  • bedrock_anthropic and claude_bedrock are accepted as aliases.

Named model profiles let the main agent, Chainlit mode picker, and sync subagents use different provider settings from the same config file:

[model]
provider = "openai_compatible"
base_url = "http://127.0.0.1:1234/v1"
name = "local-default"
models = ["local-default"]
modalities = ["text"]

[model.profiles.fast-local]
name = "local-fast"
temperature = 0.1
reasoning_effort = "low"
modalities = ["text", "image"]

[model.profiles.claude-reviewer]
provider = "anthropic"
name = "claude-sonnet-4-6"
thinking = "auto"
# api_key = "optional-if-ANTHROPIC_API_KEY-or-DEEPAGENT_MODEL_API_KEY-is-set"

[agent]
model = "fast-local"

[[subagents]]
name = "reviewer"
description = "Reviews proposed changes."
system_prompt = "Review for bugs, regressions, and missing tests."
model = "claude-reviewer"

Notes:

  • provider selects ChatOllama, ChatOpenAI, ChatAnthropic, ChatBedrockConverse (provider = "bedrock"), or ChatAnthropicBedrock (provider = "anthropic_bedrock").
  • provider = "claude" is accepted as an alias for provider = "anthropic".
  • Preferred shared fields are base_url, name, temperature, max_tokens, and reasoning_effort.
  • max_tokens is an optional positive output-token limit. It maps to max_completion_tokens for Snowflake Cortex and OpenAI-compatible providers, max_tokens for Anthropic and Bedrock, and num_predict for Ollama.
  • If the model reaches this limit while producing a tool call, ChainAgents discards the incomplete call and tells the model to shorten or split it once. A second truncated tool call ends that run with a clear message; increasing max_tokens may help.
  • repeat_penalty is optional and currently applies to provider = "ollama"; when omitted, Ollama defaults are used.
  • disable_streaming = "tool_calling" or disable_streaming_for_tool_calls = true bypasses model streaming only when tools are attached to the request; use this for providers that have trouble streaming tool-call chunks. disable_streaming = true disables model streaming for all requests.
  • endpoint_url is an override for full non-standard model endpoint URLs. OpenAI-compatible paths ending in /chat/completions or /responses are normalized to the client base URL and query parameters are forwarded as OpenAI client default query parameters. Anthropic paths ending in /v1/messages are normalized to the Claude client base URL and query parameters are forwarded as Anthropic client default query parameters.
  • models is an optional list of model IDs surfaced in Chainlit settings and modes so users can switch models per session or per message.
  • modalities declares accepted input types for a model or profile. It defaults to ["text"]; add "image" only for models that accept image content.
  • [model.profiles.<name>] defines a named profile. Profiles inherit omitted fields from [model] when they keep the same provider; profiles that switch to openai_compatible must provide base_url or endpoint_url, profiles that switch to anthropic default to https://api.anthropic.com unless base_url or endpoint_url is set, and profiles that switch to bedrock use the AWS regional endpoint unless base_url or endpoint_url is set.
  • Profile names are surfaced in Chainlit settings and modes alongside [model].models. When a selected value matches a profile name, the full profile is used; otherwise the value is treated as a raw model name using the inherited/default provider settings.
  • [agent].model optionally sets the main/supervisor agent's default profile or raw model name. CLI and environment model overrides still take precedence.
  • api_key is optional for provider = "openai_compatible"; when omitted, the runtime sends a placeholder token that local servers like LM Studio accept.
  • Anthropic requires an API key from ANTHROPIC_API_KEY, DEEPAGENT_MODEL_API_KEY, or api_key; when multiple are set, ANTHROPIC_API_KEY takes precedence over the generic key.
  • When switching from another provider to Anthropic through environment or CLI overrides, provide Anthropic credentials through ANTHROPIC_API_KEY, DEEPAGENT_MODEL_API_KEY, or --api-key; the runtime will not reuse an api_key from another provider's TOML config.
  • Legacy Ollama endpoint and port are still accepted when provider = "ollama" or omitted.
  • reasoning_effort sets the default Chainlit reasoning level for new chats. Ollama uses that level directly, Anthropic maps it to Claude effort, Bedrock forwards it as reasoning_effort for models that support it, and OpenAI-compatible servers may ignore it.
  • thinking controls Anthropic adaptive thinking: auto enables it only for known supported Claude models, adaptive always sends thinking = {"type": "adaptive"}, and disabled never sends a thinking parameter.
  • DEEPAGENT_MODEL_PROVIDER, DEEPAGENT_MODEL_BASE_URL, DEEPAGENT_MODEL_ENDPOINT_URL, DEEPAGENT_MODEL_NAME, DEEPAGENT_MODEL_API_KEY, DEEPAGENT_MODEL_REASONING, DEEPAGENT_MODEL_DISABLE_STREAMING, and DEEPAGENT_MODEL_DISABLE_STREAMING_FOR_TOOL_CALLS override the TOML defaults when set.
  • OLLAMA_BASE_URL, OLLAMA_MODEL, and OLLAMA_REASONING still work as Ollama-only compatibility aliases.

Optional: Enable Langfuse Tracing

Langfuse tracing is disabled by default. To enable it, set your Langfuse credentials in the environment and turn on the TOML option:

export LANGFUSE_PUBLIC_KEY="pk-lf-..."
export LANGFUSE_SECRET_KEY="sk-lf-..."
export LANGFUSE_BASE_URL="https://cloud.langfuse.com"
[langfuse]
enabled = true

When enabled, ChainAgents attaches Langfuse's LangChain callback handler to Chainlit, CLI, TUI, and API agent runs. The LangGraph thread ID is also passed as the Langfuse session ID.

Optional: Enable LangSmith Tracing

Set a LangSmith API key in the environment, then enable the integration in deepagent.toml:

export LANGSMITH_API_KEY="lsv2_..."
# For a non-default region, also set LANGSMITH_ENDPOINT without a trailing slash.
# For an API key linked to multiple workspaces, set LANGSMITH_WORKSPACE_ID.
[langsmith]
enabled = true
project = "chainagents"
background_trace_mode = "linked" # or "separate"

The project setting takes precedence over LANGSMITH_PROJECT; when neither is set, ChainAgents uses chainagents. In linked mode, a local background subagent run belongs to its parent trace when a LangSmith parent is available. If the parent cannot be captured, it starts its own trace. In separate mode, each task starts its own trace and records parent run and trace IDs as metadata. Both modes identify runs by the conversation session ID and background task ID, including nested and batch tasks. Search for background_task_id in LangSmith to find a task, or filter by session_id to see a conversation's tasks. The background_trace_link metadata says linked, separate, or parent_unavailable; the last value marks a linked-mode fallback to a root trace. Each run also records the agent name and path, and nested runs record their parent task ID. Chainlit reasoning and tool step visibility settings do not change tracing. Langfuse can remain enabled at the same time.

LangSmith records graph execution. The local task manager remains the source for final cancellation and cleanup status, which can differ from a graph run that already finished successfully. An enabled integration owns a client and flushes buffered traces when the runtime shuts down. Exported graphs flush at application teardown and remain usable in a later lifespan.

Agent Runtime Config

The [agent] table configures main-agent runtime behavior:

[agent]
state = "stateful"
delete_tool_enabled = false
execute_tool_enabled = false
recursion_limit = 200
memory_namespace = "filesystem"
memory_files = ["/memories/AGENTS.md"]
skills = ["skills"]
mcp_servers = ["repo"]
# custom_instruction = "Always ask clarifying questions before editing files."
# custom_instruction_file = "prompts/ui_prompts.md"

[agent.reflection]
enabled = true
memory_file = "/memories/AGENTS.md"
max_lesson_chars = 700
tool_failure_mode = "unrecovered"

Notes:

  • state = "stateful" passes the configured LangGraph store and checkpointer to DeepAgents so thread IDs can continue conversation state. state = "stateless" omits those state handles and does not expose /memories/ when building the agent graph.
  • delete_tool_enabled = false preserves the pre-0.7 filesystem surface. Set it to true only when the main agent and local synchronous subagents should receive DeepAgents 0.7's recursive delete tool. Remote async graphs have their own configuration.
  • execute_tool_enabled = false keeps command execution out of the tool surface. Set it to true only when the main agent and local synchronous subagents should receive DeepAgents 0.7's execute tool. The default ChainAgents backend is not execution-capable, so opting in exposes the tool for a compatible sandbox backend but does not grant host-shell access by itself. Remote async graphs have their own configuration.
  • recursion_limit is the LangGraph step limit for one agent run.
  • The built-in default is 100; this repo's deepagent.toml sets it to 200.
  • DEEPAGENT_RECURSION_LIMIT overrides this value when set.
  • Increase it for long tool-heavy runs that hit GraphRecursionError; lower it if you want runaway loops to stop sooner.
  • memory_namespace is the shared agent-scoped StoreBackend namespace for /memories/. ChainAgents passes a concrete backend instance and the explicit namespace tuple (memory_namespace,), as required by DeepAgents 0.7. The default is filesystem to preserve existing memory data from earlier configs. Use only letters, numbers, hyphens, underscores, dots, @, +, colons, and tildes.
  • memory_files lists /memories/ files DeepAgents loads into the startup memory prompt. The default is ["/memories/AGENTS.md"]; set it to [] to keep the memory route without startup memory loading.
  • custom_instruction appends an inline instruction to the main/supervisor agent system prompt.
  • custom_instruction_file loads that appended instruction from a UTF-8 text file. Relative paths are resolved from the active deepagent.toml; use either custom_instruction or custom_instruction_file, not both. This repo uses prompts/ui_prompts.md to encourage active Chainlit generated UI panels and next-step action buttons.
  • [agent.reflection] is opt-in. When enabled for stateful agents, ChainAgents proposes a compact lesson for memory_file after correction phrases such as "that was wrong" or after unrecovered tool failures. Chainlit asks with Save/Dismiss before writing through the agent; CLI, TUI, and API expose the proposal without mutating memory.

DeepAgents 0.7 Compatibility

DeepAgents 0.7 no longer installs TodoListMiddleware by default. ChainAgents adds langchain.agents.middleware.TodoListMiddleware explicitly to every local main, synchronous, and nested agent stack. This preserves the write_todos tool, the todos state channel, the planning prompt, and Chainlit task-list rendering. Separately deployed async graphs must restore todo middleware in their own runtime if they rely on the same behavior.

ChainAgents uses concrete BackendProtocol instances throughout. Backend integrations should use the current ls(path), glob(pattern, path=None), grep(pattern, path=None, glob=None, max_count=None), and read(file_path, offset=0, limit=2000) -> ReadResult contracts. Consume ReadResult.file_data and its metadata fields rather than parsing rendered read_file text.

For rendered tool output, empty ls and glob results are the string No files found, not []. read_file line numbers no longer use a fixed-width cat -n gutter, so callers must not parse text by fixed character columns. ChainAgents does not parse any of these rendered filesystem outputs.

Optional: Enable Workspace Docs RAG

The app can build a local-first RAG index over repo documentation and expose it to the main agent as the search_workspace_knowledge tool.

Example config:

[rag]
enabled = true
persist_directory = ".rag"
include_globs = ["README.md", "chainlit.md", "prompts/**/*.md", "skills/**/*.md"]
exclude_globs = ["AGENTS.md", "AGENT.md"]
chunk_size = 1200
chunk_overlap = 200
top_k = 4

[rag.embedding]
provider = "auto"

Notes:

  • The default corpus is docs-only: README.md, chainlit.md, prompts/**/*.md, and skills/**/*.md.
  • AGENTS.md and AGENT.md stay out of RAG because AGENTS.md is loaded directly into the main agent prompt when present.
  • The persisted local index lives under .rag/ and is safe to delete and rebuild.
  • With provider = "auto", the embedding backend follows the active chat-model provider.
  • For Ollama, the default embedding model is nomic-embed-text.
  • For OpenAI-compatible embeddings, set [rag.embedding].model explicitly.
  • For Anthropic or Snowflake Cortex chat models, set [rag.embedding].provider explicitly to ollama or openai_compatible, with an appropriate model and base URL; auto is not valid for either provider.
  • On startup, the UI reports whether RAG is ready and how many files/chunks were indexed.
  • The startup message includes a Rebuild Knowledge Index action so you can refresh the index after documentation changes.
  • The startup message also includes Upload File For RAG, which lets you add text-based files to the current chat thread's knowledge index.
  • Composer file attachments are enabled for text-based uploads; attached files are automatically ingested into the current thread's RAG store before the model responds.
  • Uploaded files are thread-scoped and persist under .rag/uploads/, so they do not leak into other chat threads.

Chainlit App Config

This repo also includes an app-specific chainlit.toml for UI behavior that the bridge owns:

[steps]
auto_collapse_delay_seconds = 3

Notes:

  • chainlit.toml is separate from Chainlit's native .chainlit/config.toml.
  • [steps].auto_collapse_delay_seconds controls how long completed reasoning and tool steps stay expanded before auto-collapsing.
  • If chainlit.toml is missing or invalid, the app falls back to 3 seconds.

Add Skills

The runtime now supports Deep Agents skill sources through deepagent.toml.

  1. Create a skill source directory in the repo, for example:
skills/
├── repo-docs/
│   └── SKILL.md
└── reviewer/
    └── SKILL.md
  1. Add the source directory to deepagent.toml:
[agent]
skills = ["skills"]

Notes:

  • Relative paths in deepagent.toml are resolved from the config file location.
  • Relative skill paths are automatically mapped into the Deep Agents virtual filesystem as /workspace/....
  • Each skill source directory should contain one or more skill folders, and each skill folder must contain SKILL.md.
  • You can also use explicit virtual paths such as "/workspace/skills/" if you prefer.
  • Every loaded skill is also exposed as a Chainlit slash command using the skill name, for example reviewer becomes /reviewer.
  • Running a skill-backed slash command forces the main agent to read that skill's SKILL.md and apply it for that request.
  • Skills loaded through [agent].skills and sync [[subagents]].skills are both considered for slash commands, but explicit [chainlit].commands take precedence on name collisions.

Minimal SKILL.md example:

---
name: reviewer
description: Use this skill when reviewing code changes for bugs and missing tests.
---

# reviewer

When asked to review code:
1. Read the relevant files first.
2. Focus on bugs, regressions, and missing tests.
3. Return concise findings with file references.

Add Subagents

Custom subagents are also loaded from deepagent.toml, and each subagent can have its own skills and mcp_servers.

Example:

[mcp]
tool_name_prefix = true
stateful = true

[mcp.servers.repo]
transport = "stdio"
command = "npx"
args = ["-y", "@modelcontextprotocol/server-filesystem@2025.8.21", "."]
cwd = "."

[agent]
state = "stateful"
recursion_limit = 200
skills = ["skills"]
mcp_servers = ["repo"]
summarization_trigger_tokens = 6000
summarization_keep_tokens = 2400

[[subagents]]
name = "repo-researcher"
description = "Researches the codebase and produces concise implementation guidance."
system_prompt_file = "prompts/repo-researcher.md"
skills = ["skills/research"]
mcp_servers = ["repo"]
nested_subagents = ["reviewer"]

[[subagents.subagents]]
name = "repo-planner"
description = "Turns repository findings into an implementation plan."
system_prompt = "Use repository context to produce a concise implementation plan."
skills = ["skills/research"]
mcp_servers = ["repo"]

[[subagents]]
name = "reviewer"
description = "Reviews proposed changes for bugs and regressions."
system_prompt = """
You are a strict code reviewer.
Focus on correctness, regressions, and missing tests.
Keep findings concise and actionable.
"""
skills = ["skills/reviewer"]
mcp_servers = ["repo"]
# model = "gpt-oss:20b"

Supported subagent fields:

  • name: required
  • description: required
  • system_prompt or system_prompt_file: one is required
  • skills: optional list of skill source paths for that subagent
  • mcp_servers: optional list of MCP server names to attach to that subagent
  • model: optional profile name or raw model name. Profile names can switch provider settings and tool-schema handling for that sync subagent. Raw model names inherit the parent/default provider settings.
  • background: optional boolean, defaulting to false. When global local background execution is enabled, true allows the subagent's direct parent to launch it with spawn_background_task or include it in run_subagent_batch. Foreground task delegation is unaffected.
  • messaging: optional boolean, defaulting to false. When [agent.messaging].enabled = true, this grants that synchronous subagent invocation the local message tools. The flag is not inherited by children. Remote Agent Protocol subagents do not support local messaging.
  • nested_subagents: optional list of top-level sync subagent names exposed as children of this subagent
  • [[subagents.subagents]]: optional inline private sync child subagents under a parent subagent

Nested Subagents

Nested subagents let a synchronous subagent delegate work to its own synchronous child agents. They are useful when a coordinator needs focused helpers but the main agent should not necessarily see every helper directly.

ChainAgents supports two nesting patterns:

  • use [[subagents.subagents]] for a private inline child available only to its parent
  • use nested_subagents = ["name"] to reuse a top-level synchronous subagent as a child while keeping it available to the main agent

Add a Private Inline Child

Place [[subagents.subagents]] immediately after its parent [[subagents]] entry:

[[subagents]]
name = "research-manager"
description = "Coordinates repository research and planning."
system_prompt = "Delegate focused work and synthesize the results."

[[subagents.subagents]]
name = "repo-planner"
description = "Turns repository findings into an implementation plan."
system_prompt = "Produce a concise, actionable implementation plan."

In this configuration:

  • repo-planner is available to research-manager
  • repo-planner is not exposed directly to the main agent
  • the child can use the normal synchronous subagent fields, including skills, mcp_servers, and model

Use an inline child when it is an implementation detail of one parent.

Reuse a Top-Level Subagent as a Child

Define the child as a normal top-level [[subagents]] entry and reference its name from the parent:

[[subagents]]
name = "research-manager"
description = "Coordinates repository research and review."
system_prompt = "Delegate research and ask the reviewer to check the result."
nested_subagents = ["reviewer"]

[[subagents]]
name = "reviewer"
description = "Reviews proposed changes for bugs and regressions."
system_prompt = "Return concise findings with actionable file references."

In this configuration:

  • reviewer remains available directly to the main agent
  • research-manager can also delegate to reviewer
  • more than one parent can reference the same top-level subagent
  • each referenced name must exactly match a top-level synchronous subagent

Use a reference when a role should be shared by the main agent or multiple parents. A parent can expose several shared children, for example nested_subagents = ["planner", "reviewer"].

Configuration and Inheritance

Inline children accept the same fields as other synchronous subagents: name, description, system_prompt or system_prompt_file, skills, mcp_servers, model, background, and their own nested children. A referenced child uses the configuration from its top-level [[subagents]] entry wherever it is reused. When model is omitted, the child continues with its parent/default model configuration.

Limitations and Validation

Nested subagents are synchronous only. Async Agent Protocol subagents must remain top-level [[async_subagents]] entries.

Configuration loading rejects:

  • a nested_subagents name that does not match a top-level synchronous subagent
  • duplicate direct child names, including an inline child and a referenced child with the same name
  • reference cycles such as manager -> reviewer -> manager
  • a nested child that defines graph_id, because that represents an async subagent

As a rule of thumb, use an inline child for a parent-private specialist, a referenced child for a shared synchronous role, and [[async_subagents]] for remote or background Agent Protocol work.

Main [agent] additions:

  • state: optional agent state mode. Use stateful for checkpointed conversation state, or stateless to build the DeepAgents graph without a LangGraph store, checkpointer, or /memories/ route. Defaults to stateful.
  • recursion_limit: optional positive integer LangGraph step limit for a single agent run. Defaults to 100 unless overridden by DEEPAGENT_RECURSION_LIMIT.
  • memory_namespace: optional non-empty namespace for agent-scoped /memories/ storage. Defaults to filesystem; allowed characters are letters, numbers, -, _, ., @, +, :, and ~.
  • memory_files: optional list of absolute /memories/ file paths loaded into the DeepAgents startup memory prompt. Defaults to ["/memories/AGENTS.md"]; use [] to disable startup memory loading.
  • delete_tool_enabled: optional boolean controlling DeepAgents 0.7's recursive delete tool for the main agent and local synchronous subagents. Defaults to false.
  • execute_tool_enabled: optional boolean controlling DeepAgents 0.7's execute tool for the main agent and local synchronous subagents. Defaults to false.
  • [agent.background_subagents]: global opt-in and limits for process-local background execution. enabled defaults to false; eligible synchronous subagents must also set background = true. stream_activity = true exposes live reasoning and tool activity as nested Chainlit steps while leaving other interfaces completion-only. batch_result_format selects run_subagent_batch results from json, markdown, or markdown_files and defaults to markdown. The three positive integer limits bound running work per conversation, running work across the process, and retained task records per conversation. When the retained limit is reached, the oldest finished tasks are forgotten to make room; unfinished tasks, parents of retained tasks, and tasks whose cleanup has not completed are kept.
  • [agent.messaging]: opt-in process-local mailboxes. The main agent and each messaging = true synchronous subagent receive list_agent_recipients, send_agent_message, get_agent_message, and wait_for_agent_messages. A message is inserted before the recipient's next model step; an idle agent is not woken. Addresses identify individual invocations, so use list_agent_recipients when same-name agents run concurrently. Finished subagents cannot receive messages. Mailboxes are cleared when the conversation closes or the process exits.
  • [agent.user_input]: opt-in nonblocking user prompts. A conversation still runs one main turn at a time. In Chainlit, busy input offers Steer active turn and Queue next turn; Stop pauses queued turns until Resume. In the TUI, use /steer text, /queue text, /stop, and /resume (or Ctrl+R). The interactive CLI prompts for steer or queue during a run and also accepts those commands. HTTP clients can submit through POST /api/agent/input with mode set to turn, steer, or queue, then query /api/agent/turns/{thread_id}/{turn_id} or its /events NDJSON stream. Stop and resume use the corresponding /api/agent/turns/{thread_id}/stop and /resume endpoints. Input IDs make repeated API submissions idempotent while their turns are retained. Steering is text-only; queued turns retain attachments and run settings. If an active turn ends before it reads a steering note, the note runs as a visible follow-up ahead of queued turns. max_queued_turns also reserves room for pending steering follow-ups; max_completed_turns bounds retained results and idempotency records.
  • [agent.clarification]: opt-in clarifying questions. When enabled = true on a stateful runtime, the main agent gets an ask_user tool and is told to call it once, before delegating, when a request is ambiguous in ways that would change what subagents do. The tool pauses the run with a LangGraph interrupt; your next message on the same thread is the answer and resumes the paused run instead of starting a new turn. A bare option number such as 2 picks that suggested option. While a question is pending, only plain text answers it: slash commands, configured response actions, and replies with attachments are refused with a reminder, and turns queued earlier through [agent.user_input] wait until the answer has run. When the model emits other tool calls next to ask_user, they are dropped so no subagent starts before the answer. Chainlit shows the question with one button per option; the CLI and TUI print numbered options; the HTTP API reports status: "awaiting_input" with the pending clarifications on /api/agent/invoke and on the stream's done event, and a clarification_requested stream event carries the question. Subagents never get the tool, and it is disabled when agent.state = "stateless" because resuming needs a checkpointer. With the in-memory checkpointer a pending question is lost on restart.
  • model: optional profile name or raw model name for the main/supervisor agent. CLI and environment model overrides take precedence.
  • [agent.reflection]: optional correction-learning workflow. enabled = true requires state = "stateful" and a memory_file under /memories/; max_lesson_chars limits proposal size; tool_failure_mode = "unrecovered" only proposes lessons for failed tool calls that do not produce a later final response.
  • AGENTS.md: optional repo-root file that is automatically appended to the main/supervisor agent system prompt when present. It is not applied to separately configured async graph prompts.
  • custom_instruction: optional string appended to the main/supervisor agent system prompt. This setting does not get applied to separately configured prompts such as the async_researcher graph prompt.
  • custom_instruction_file: optional UTF-8 text file loaded as the main-agent custom instruction. Relative paths resolve from the active deepagent.toml. This is mutually exclusive with custom_instruction.
  • ChainAgents explicitly restores TodoListMiddleware for the main agent and local sync subagents because DeepAgents 0.7 no longer includes it by default.
  • DeepAgents still provides its own summarization middleware in the main agent and sync subagents.
  • summarization_trigger_tokens: optional positive integer token threshold for DeepAgents' built-in summarization middleware.
  • summarization_keep_tokens: optional positive integer token budget to keep after DeepAgents summarizes conversation history.
  • Legacy summarization_middleware_enabled entries are still parsed for compatibility, but ChainAgents no longer injects a second summarization middleware.

Local Background Subagents

Local background execution lets the main agent or a nested synchronous agent start one of its configured children and continue immediately. Enable it with:

[agent.background_subagents]
enabled = true
stream_activity = true
batch_result_format = "markdown"
max_running_per_session = 4
max_running_total = 16
max_tasks_per_session = 100

Then opt in each synchronous subagent that its direct parent may launch in the background:

[[subagents]]
name = "research-manager"
description = "Coordinates repository research and planning."
system_prompt = "Delegate focused work and synthesize the results."
background = true

[[subagents.subagents]]
name = "repo-planner"
description = "Turns repository findings into an implementation plan."
system_prompt = "Produce a concise, actionable implementation plan."
background = true

An unmarked subagent remains available through the blocking task tool but is rejected by spawn_background_task and run_subagent_batch. Marking a parent does not implicitly mark its children; each background launch target opts in independently.

The agent receives five tools:

  • spawn_background_task(description, subagent_type) starts an allowed direct child and returns a task ID immediately
  • run_subagent_batch(tasks) starts every independent task concurrently, waits for all of them, and returns their terminal reports in the globally configured format and in input order
  • list_background_tasks() lists tasks visible to the calling agent
  • get_background_task(task_id, wait_seconds=0) returns current state or waits up to 60 seconds
  • cancel_background_task(task_id) cancels the task and all descendants

For example, with the research-manager and private repo-planner nesting shown above, the main agent can call:

spawn_background_task("Investigate the failing API tests", "research-manager")

When several tasks are independent, one tool call can fan out to separate subagent conversations:

run_subagent_batch(tasks=[
  {"subagent_type": "research-manager", "description": "Trace the API failures."},
  {"subagent_type": "research-manager", "description": "Check the related tests."}
])

The tasks input is unchanged across all result modes. Set batch_result_format once under [agent.background_subagents]; it is not a per-call argument. Entries stay in request order even when children finish in another order, and a failed child does not discard successful sibling reports.

The default, batch_result_format = "markdown", returns one newline-delimited Markdown document. Each section includes the agent name, task ID, terminal status, original request, and either the report or error. Empty and cancelled reports are marked explicitly. The line-oriented format also gives oversized batches useful head/tail previews and lets the agent page through a complete offloaded result with read_file.

batch_result_format = "json" restores the original structured contract. It returns every terminal snapshot, including session, task-tree, and timing metadata:

{
  "results": [
    {
      "task_id": "bg-123",
      "session_id": "thread-1",
      "agent_name": "research-manager",
      "description": "Trace the API failures.",
      "agent_path": ["research-manager"],
      "parent_task_id": null,
      "status": "success",
      "result": "The failure starts in ...",
      "error": null,
      "created_at": 1750000000.0,
      "completed_at": 1750000001.5
    }
  ]
}

batch_result_format = "markdown_files" writes one standalone Markdown file for every terminal result—including errors, cancellations, and empty reports— then returns an ordered manifest:

{
  "files": [
    {
      "task_id": "bg-123",
      "agent_name": "research-manager",
      "status": "success",
      "path": "/workspace/.files/outputs/subagent-batches/batch-call-unique/01-research-manager-bg-123.md"
    }
  ]
}

Each file contains its agent name, task ID, status, original request, and report or error. Batch directories are unique, filenames are sanitized and numbered in request order, and descriptions never enter filenames. Files are written through the generated-output backend under .files/outputs/subagent-batches, so the existing UI can discover and download them. They persist across conversation teardown and process restarts until removed as workspace outputs; they are not temporary large-result offloads. The batch returns only after every file succeeds; a write failure rolls back completed files, and a cleanup failure is reported together with the original error.

The manager validates capacity for the whole batch before launch, so a batch that exceeds a configured limit starts no children. Cancelling the waiting call cancels its unfinished children and their descendants. In file mode, cancellation during output writes waits for the active write and removes every file completed by that interrupted batch.

Batch delegation is provider-independent. It is useful when a supervisor model, including a Snowflake Cortex model, can emit only one tool call per assistant turn: that single call starts separate child graph runs, and each child uses its own model request stream.

The running research-manager can independently launch its private child with spawn_background_task("Identify the smallest fix", "repo-planner"). Either agent can continue its current response, call list_background_tasks() later, retrieve a terminal result with get_background_task("bg-..."), or stop its visible subtree with cancel_background_task("bg-...").

The main agent can inspect the whole conversation task tree. Nested agents can inspect only their own task subtree and can spawn only their configured direct children. Background runs receive an isolated user message and checkpoint thread while retaining their configured model, skills, MCP tools, workspace, and shared memory access. Their state and streamed tokens are not merged into the parent response. Concurrent children can therefore observe the same workspace and memory resources; prompts should assign non-overlapping writes or otherwise coordinate shared updates.

By default, Chainlit and the interactive CLI/TUI post one status-only notice when a task finishes; successful output remains available through get_background_task instead of being copied into the notice. With stream_activity = true, Chainlit additionally renders each background task as a parent step with nested reasoning and tool-call steps, then closes that tree before posting the same single terminal notice. Background reasoning and tool steps follow [chainlit].reasoning_steps_enabled and [chainlit].tool_steps_enabled, including the current chat settings switches; the parent step and terminal notice remain visible when either is disabled. The setting does not expose live background activity through the CLI, TUI, or HTTP API. A one-shot CLI invocation prints the main response first, then waits for its remaining background work. One-shot text output prints terminal results because the process is about to exit; JSON output includes a background_tasks array. The HTTP API exposes conversation-scoped list, get, cancel, and close operations under /api/background-tasks.

curl "http://127.0.0.1:8000/api/background-tasks?thread_id=$THREAD_ID"
curl "http://127.0.0.1:8000/api/background-tasks/bg-123?thread_id=$THREAD_ID&wait_seconds=10"
curl -X DELETE "http://127.0.0.1:8000/api/background-tasks/bg-123?thread_id=$THREAD_ID"
curl -X DELETE "http://127.0.0.1:8000/api/background-tasks?thread_id=$THREAD_ID"

Tasks are retained until the conversation closes (or until the oldest finished tasks are evicted to stay within max_tasks_per_session) and are cancelled before its offloaded large tool results and MCP resources are released. Cleanup is scoped by thread ID, so closing one conversation does not remove another conversation's artifacts; workspace files, generated downloads, memories, and uploads are not part of this cleanup. Tasks and offload ownership are stored only in the current process and do not survive restarts. Agent Server deployments therefore require session affinity when multiple workers are used. The custom Agent Server app exposes DELETE /background-tasks/sessions/{thread_id} for explicit cleanup and closes all remaining managers and tracked offloads during server shutdown.

Chainlit Native Commands

You can configure slash-style commands that run from the Chainlit composer before the model call. Chainlit also auto-generates slash commands for loaded skills.

Place this config in deepagent.toml or whatever file DEEPAGENT_CONFIG points to. Do not put it in the app UI file chainlit.toml or Chainlit's native .chainlit/config.toml.

Example:

[chainlit]
# Set false to hide model selection in chat settings and Modes.
model_mode_enabled = true
# Set false to disable per-message reasoning overrides from the Modes picker.
reasoning_mode_enabled = true
# Set false to hide streamed reasoning step panels and reasoning task entries.
reasoning_steps_enabled = true
# Set false to hide streamed tool step panels and tool task entries.
tool_steps_enabled = true
# Set false to hide the initial startup status message ("Workspace agent ready...").
startup_status_enabled = true
# Set false to hide generated Chainlit CustomElement panels and remove the render tool.
generative_ui_enabled = true
# Set false to keep legacy non-chronological streaming order in Chainlit.
chronological_ui_enabled = true
commands = [
  { name = "ask-researcher", description = "Delegate to repo-researcher.", target = "subagent", value = "repo-researcher", template = "{input}" },
  { name = "repo-readme", description = "Run an MCP tool directly.", target = "mcp_tool", value = "repo_read_file", mcp_server = "repo", template = "{\"path\":\"README.md\"}" },
  { name = "summarize", description = "Apply a prompt template.", target = "prompt", value = "Summarize the input", template = "Summarize this:\n{input}" }
]
starters = [
  { label = "Explain this repo", message = "Explain the architecture of this repository and identify the most important files.", command = "ask-researcher", icon = "book-open" },
  { label = "Review current changes", message = "Review the current working tree changes for bugs, regressions, and missing tests." }
]
response_actions = [
  { name = "summarize", label = "Summarize", icon = "list", prompt = "Summarize this response concisely:\n\n{response}" },
  { name = "explain", label = "Explain", icon = "book-open", prompt = "Explain this response in more detail.\n\nOriginal request:\n{prompt}\n\nResponse:\n{response}" }
]

target modes:

  • prompt: rewrites the user prompt before sending it to the agent.
  • subagent: rewrites the user prompt to direct the runtime to delegate via the configured subagent.
  • mcp_tool: invokes the configured MCP tool directly and returns tool output in chat.

Notes:

  • The [chainlit] table for native commands belongs in deepagent.toml, alongside [model], [agent], [mcp], [[subagents]], and [[async_subagents]].
  • [chainlit].model_mode_enabled = false hides the Model selector in chat settings and the Model mode group, and ignores per-message model overrides from UI modes.
  • [chainlit].reasoning_mode_enabled = false hides the Reasoning mode group and ignores per-message reasoning overrides from UI modes.
  • [chainlit].reasoning_steps_enabled = false hides streamed reasoning cl.Step panels and reasoning task-list entries while preserving model reasoning settings.
  • [chainlit].tool_steps_enabled = false hides streamed tool cl.Step panels and tool task-list entries while preserving tool execution.
  • [chainlit].startup_status_enabled = false disables the initial startup status message that summarizes runtime configuration.
  • [chainlit].generative_ui_enabled = false hides generated Chainlit CustomElement panels and removes the render_chainlit_ui tool from the main agent.
  • [chainlit].chronological_ui_enabled = false disables chronological UI ordering so response tokens stream immediately and reasoning steps are not force-rolled at tool boundaries.
  • Command name is invoked as /<name> and must be unique.
  • template is optional and may include {input}.
  • For mcp_tool, user command arguments must be valid JSON, e.g. /repo-readme {"path":"README.md"}.
  • Each discovered skill also becomes /<skill-name> automatically. For example, a skill with name: reviewer is available as /reviewer.
  • Skill-backed commands always force the main agent to read the selected SKILL.md first and use it for that turn.
  • If a configured [chainlit].commands entry and a skill share the same slash name, the configured command wins.
  • starters define starter prompts shown by Chainlit before the first message in a thread.
  • Starter label and message are required. Starter command and icon are optional.
  • response_actions appear after Markdown and PDF beneath completed Chainlit replies. Each action needs a unique name, a label, and a prompt; icon and tooltip are optional. Omit the list or set it to [] to hide custom actions.
  • A response-action prompt can use {response} for the clicked reply and {prompt} for the request that produced it. Other braces stay literal. The agent receives the expanded prompt in the current conversation and displays its reply normally, but Chainlit does not show a user message or prompt text in reasoning panels for the action. The prompt remains part of agent state and persisted response context.
  • Saved Chainlit chats restore response actions when response context is available. Restart the app after editing deepagent.toml to reload action definitions.

Add Async Subagents

Async subagents are loaded from deepagent.toml as background Agent Protocol jobs. They are useful for long-running or remote work where the main agent should return a task ID immediately and let you check, update, cancel, or list tasks later.

Example:

[[async_subagents]]
name = "remote-researcher"
description = "Runs longer research jobs in the background on an Agent Protocol server."
graph_id = "researcher"
# Omit url for ASGI transport in a co-deployed LangGraph setup.
# Set url for HTTP transport to a remote Agent Protocol server.
# url = "https://researcher-deployment.langsmith.dev"
# headers = { Authorization = "Bearer ${RESEARCHER_TOKEN}" }

Supported async subagent fields:

  • name: required
  • description: required
  • graph_id: required graph or assistant ID on the Agent Protocol server
  • url: optional remote Agent Protocol server URL; omit for ASGI transport in co-deployed LangGraph setups
  • headers: optional request headers for remote/self-hosted Agent Protocol servers

For compatibility with DeepAgents' native discriminator, a [[subagents]] entry with a graph_id is also treated as an async subagent. Async subagents cannot define sync-only fields such as system_prompt, skills, mcp_servers, or model; those capabilities are configured on the remote graph.

This repo includes a LangGraph co-deployment entrypoint for ASGI transport:

  • langgraph.json registers supervisor and async-researcher
  • langgraph_app.py exports both graphs
  • omit url in deepagent.toml when running through LangGraph Agent Server

Run the co-deployed Agent Protocol server with enough worker capacity for the supervisor plus background tasks:

uv run --with "langgraph-cli[inmem]" langgraph dev --n-jobs-per-worker 10

ASGI transport is only available in this LangGraph server path. If you launch the UI with chainlit run main.py -w, use HTTP transport instead by setting url = "http://127.0.0.1:2024" on the async subagent.

Chainlit also starts a background notifier for launched async tasks. It polls the Agent Protocol server and posts a message when a task reaches success, error, cancelled, interrupted, or timeout. If deepagent.toml omits url for ASGI co-deployment, Chainlit defaults to http://127.0.0.1:2024, the usual langgraph dev URL. Override it with:

export CHAINLIT_ASYNC_SUBAGENT_URL="http://127.0.0.1:2024"

Optional:

export CHAINLIT_ASYNC_TASK_POLL_SECONDS="5"

Add MCP Servers

MCP servers are defined once in deepagent.toml and then attached by name to the main agent or any subagent.

Example:

[mcp]
tool_name_prefix = true
stateful = true

[mcp.servers.repo]
transport = "stdio"
command = "npx"
args = ["-y", "@modelcontextprotocol/server-filesystem@2025.8.21", "."]
cwd = "."

[mcp.servers.docs]
transport = "http"
url = "http://localhost:8000/mcp"

[mcp.servers.github]
transport = "sse"
url = "http://localhost:8080/sse"
headers = { Authorization = "Bearer ${GITHUB_TOKEN}" }

[agent]
mcp_servers = ["repo"]

[[subagents]]
name = "repo-researcher"
description = "Researches the repo and docs."
system_prompt = "Use the repo and docs MCP servers to answer questions."
mcp_servers = ["repo", "docs"]

[[subagents]]
name = "release-assistant"
description = "Works with repository metadata and hosted services."
system_prompt = "Use the GitHub MCP server when release metadata is needed."
mcp_servers = ["github"]

Supported MCP config fields:

  • top-level [mcp] tool_name_prefix = true|false
  • top-level [mcp] stateful = true|false
  • [mcp.servers.<name>] transport
  • for stdio: command, args, optional cwd, optional env
  • for http, streamable_http, streamable-http: url, optional headers
  • for sse: url, optional headers
  • for websocket: url

Notes:

  • mcp_servers on [agent] attaches those MCP tools to the main agent.
  • mcp_servers on [[subagents]] attaches those MCP tools only to that subagent.
  • mcp_servers on [[subagents.subagents]] attaches those MCP tools only to that nested child subagent.
  • Skills and MCP servers are independent. You can use neither, either, or both on any subagent.
  • If one MCP server cannot load tools, the agent continues with tools from healthy servers and retries the failed server on the next run. The affected run shows an MCP warning in Chainlit, CLI/TUI, and API output. Tool invocation errors are returned to the model as recoverable tool errors.
  • Relative cwd values are resolved from the location of deepagent.toml.
  • tool_name_prefix = true is recommended when multiple MCP servers expose overlapping tool names.
  • stateful = true keeps MCP sessions open per conversation scope while the app process is running. Chainlit uses the effective thread ID, so reopening a saved chat reuses its MCP tools and transport. Turns on the same thread run in order when stateful MCP is enabled.
  • After the last Chainlit session leaves a conversation, its resources remain available for 10 minutes. At most four idle conversations are retained; the oldest idle scope is closed when a fifth becomes idle. Active conversations are never evicted by this limit.
  • MCP servers are discovered concurrently, and each (scope, server) discovery is shared by callers already waiting for it. Failed servers remain retryable on a later run.
  • stateful = false recreates the MCP session for every tool call.

Current scope of this config support:

  • it exposes ls, read_file, write_file, edit_file, glob, and grep by default, plus write_todos, subagent tools, config-driven skills, and MCP tools
  • it exposes recursive delete only when [agent].delete_tool_enabled = true
  • it exposes execute only when [agent].execute_tool_enabled = true; the configured backend must also implement sandbox execution
  • it supports config-driven sync subagents and async Agent Protocol subagents
  • it does not yet provide a config-driven registry for custom Python tools per subagent beyond MCP
  • if you need custom Python tools, define them alongside the generated UI tool in chainagents/runtime/commands.py, then add them in build_main_tools in graph.py; both the static LangGraph graph and the live runtime agent (lifecycle.py) assemble their tools through that one function

See deepagent.toml.example for a portable baseline and the sections above for optional MCP and subagent examples.

Workspace Contract

  • /workspace/ maps to this repo on disk.
  • /memories/ is available in stateful mode under the configured agent-scoped namespace and durable across LangGraph threads only when DATABASE_URL is configured.
  • any other absolute path is treated as ephemeral scratch space by the deep agent backend.

Notes

  • Native Chainlit history is available when DATABASE_URL, CHAINLIT_AUTH_SECRET, and either CHAINLIT_AUTH_USERS or the legacy username/password pair are configured.
  • If DATABASE_URL is set but authentication is not configured, Chainlit still persists thread records, but they are not browseable from the UI.
  • When DATABASE_URL is unset, thread IDs only persist while the process stays alive.
  • When DATABASE_URL is set, durable state is available through LangGraph thread IDs. You can reuse a thread ID from the chat settings panel to continue the same checkpointed thread.
  • When [agent].state = "stateless", thread IDs still identify requests and MCP/RAG scopes, but the agent graph does not checkpoint conversation state, receive a LangGraph store, or expose /memories/.
  • MCP stateful sessions are process-local. They survive tool calls and saved-chat navigation in the same Chainlit thread until idle eviction, but not an app restart.
  • Switching between saved Chainlit chats keeps active turns and local background subagents running. Returning during the same app process restores recent terminal background-task notices that finished while the chat was away. Idle conversation resources still follow the configured retention window.
  • On startup, the UI shows how many skill sources, MCP servers, custom subagents, and async subagents were loaded from deepagent.toml.

Runtime module structure

The public runtime API is chainagents.runtime. runtime/core.py explicitly re-exports the same objects for compatibility. The root-level deepagent_runtime module remains an alias of that facade but is deprecated and will be removed in a future release; import from chainagents.runtime instead.

Modules in chainagents/runtime/ Responsibility
constants.py, types.py Shared defaults, workspace root, and configuration records
model_config.py Provider settings, endpoint normalization, and model profiles
extension_config.py, config.py Extension parsing and TOML/environment configuration
providers.py, models.py Provider adapters and configured model construction
backends.py Workspace paths and filesystem/storage backend routing
middleware.py Tool resilience and summarization middleware
commands.py Generated UI tools and skill command discovery
artifacts.py Session-scoped storage for offloaded large tool results
graph.py Shared agent assembly (main tools, subagent specs, agent kwargs) used by both the static graph and the live runtime
tracing.py Langfuse callbacks and LangGraph run configuration
mcp_sessions.py MCP session pool, stateful transport ownership, and tool discovery caching
rag_ops.py RAG index status, rebuild, and thread-scoped upload operations
background_tasks/ Process-local background execution of configured synchronous subagents
lifecycle.py, reflection.py Agent/MCP/persistence lifecycle and correction reflection

Implementation modules import their lower-level owners directly; they never import core.py or the package facade. Cross-module function calls resolve through the owning module so tests can patch that owner (for example, chainagents.runtime.models.build_model). Facade globals are not patch seams. Warning filters are installed by the runtime package before SDK imports.

License

This hobby, personal project is made available under the MIT License and is built with open source Python libraries. See LICENSE for the full text.

About

This project runs a local-first LangChain DeepAgent behind a Chainlit UI and dedicated CLI.

Topics

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Used by

Contributors

Languages