Ask AI
Ask AI Start conversation ↵

I'm an AI assistant with Mewbo's codebase and documentation in context.

Ask me anything about Mewbo.

EXAMPLE QUESTIONS

Every configuration key

Mewbo is configured via configs/app.json. This reference is auto-generated from the JSON Schema at configs/app.schema.json.

Copy configs/app.example.json to configs/app.json to get started. See Get Started for the full setup walkthrough.

Runtime

Top-level key: runtime

Runtime environment settings.

Key Type Default Description
envmode string dev Free-text label for this deployment (e.g. dev, staging, prod). It becomes the Langfuse tracing environment, so traces from one deployment can be filtered and compared without staging traffic polluting production aggregates. Lowercased and punctuation-stripped on the way out, since Langfuse rejects other shapes. The trace release is the running Mewbo version and is not set here.
log_level string DEBUG Logging verbosity. One of DEBUG, INFO, WARNING, ERROR, CRITICAL.
log_style string "" Override for the CLI's terminal log format; prefer cli_log_style. Only the literal value 'dark' has any effect (dims the log line style for a dark background); anything else uses the plain format. Takes priority over cli_log_style when set; leave empty to use that instead.
cli_log_style string dark CLI terminal log format: only the literal value 'dark' has any effect (dims the log line style for a dark background); any other value uses the plain format. Overridden by runtime.log_style when that's set.
preflight_enabled boolean false Run connectivity checks for LLM, Langfuse, and Home Assistant on startup.
developer_mode boolean false Enable developer mode. Unlocks graph-only (no-LLM) repository indexing so contributors can inspect AST graph construction without documentation generation. Read by the API, console, and CLI.
cache_dir string "" Directory for tool caches. Defaults to $MEWBO_HOME/cache. ⚠️
session_dir string "" Directory for session transcripts. Defaults to $MEWBO_HOME/sessions. ⚠️
config_dir string "" Root configuration directory. Defaults to $MEWBO_HOME. ⚠️
result_export_dir string "" Directory for large tool result exports. Empty to disable.
projects_home string "" Directory for virtual project folders. Defaults to $MEWBO_HOME/projects. ⚠️

Storage

Top-level key: storage

Session storage backend configuration.

Key Type Default Description
driver string json Storage driver: 'json' (filesystem) or 'mongodb'. Overridden by the MEWBO_STORAGE_DRIVER environment variable.
mongodb MongoDB MongoDB connection settings (used when driver is 'mongodb').
mongodb.uri string mongodb://localhost:27017 MongoDB connection URI (includes host, port, credentials). Overridden by the MEWBO_MONGODB_URI environment variable.
mongodb.database string mewbo MongoDB database name for session storage. Overridden by the MEWBO_MONGODB_DATABASE environment variable.

Language Model

Top-level key: llm

LLM provider connection and model selection.

Key Type Default Description
api_base string "" Optional base URL override. Leave empty for direct provider access (LiteLLM routes automatically from the model prefix). Set only when using a proxy (e.g. LiteLLM, Bifrost).
api_key string API key for the LLM provider (e.g. Anthropic, OpenAI) or proxy master key. ⚠️
default_model string gpt-5.2 Model ID using 'provider/model' syntax. LiteLLM auto-routes to the right API endpoint. When using a proxy, adjust the prefix to match its routing.
action_plan_model string "" Model ID the orchestrator uses to generate a session's initial plan and, when no explicit model is passed in, as that session's default model. Falls back to default_model when empty.
tool_model string "" Model ID used by individual tools. Falls back to default_model when empty.
title_model string "" Model ID for session-title generation. Falls back to default_model when empty.
compact_models list[string] Priority-ordered list of models for context compaction. On failure, the next model in the list is tried. The keyword "default" resolves to the running agent's model. Example: ["anthropic/claude-haiku-4-5-20251001", "default"]
fallback_models list[string] Flat ordered list of fallback model IDs. Prefer 'fallback' (typed, explicit opt-in). A non-empty value here is still honored, and is treated as fallback enabled.
fallback Fallback Opt-in cross-model fallback policy (see FallbackConfig).
fallback.enabled boolean false Enable automatic fallback to other models when the primary is exhausted or hits a hopeless-here error. Off by default.
fallback.models list[string] Ordered fallback model IDs, tried after the primary is exhausted.
fallback.self_steering boolean false Allow the agent to steer its own model routing at runtime via the model_control tool (switch down the declared fallback ladder, or up when allow_upgrade is set). Off by default; the automatic fallback ladder still operates regardless. Every deliberate switch is bounded by max_switches and the existing retry budget / circuit breaker.
fallback.max_switches integer 2 Maximum deliberate model switches the model_control tool may perform in one run. 0 disables switching while leaving status/list readable.
fallback.allow_upgrade boolean false Permit model_control switches UP the declared ladder (toward the primary). Off by default, so self-steering is a one-way ratchet downward — the direction that heals a failing primary without re-provoking it.
proxy_model_prefix string openai LiteLLM provider prefix prepended to model names when api_base is set. LiteLLM strips this prefix before forwarding the model name to the proxy, so the proxy receives the model ID it advertises in /v1/models. Leave as 'openai' for LiteLLM proxy, Bifrost, and OpenRouter. Only relevant when api_base is configured.
request_timeout number 900.0 Seconds handed to the LiteLLM client as its HTTP timeout. Because the loop STREAMS, this behaves as the maximum gap BETWEEN response chunks, not as a ceiling on total call duration — a long generation that keeps emitting chunks never trips it. The total-duration ceiling is agent.llm_call_timeout. Raise this only when a provider is slow to send its FIRST chunk.
reasoning_effort string "" Reasoning effort hint for supported models. One of low, medium, high, none, or empty.
reasoning_effort_models list[string] Additional model IDs (or 'prefix*' patterns) that should receive the reasoning_effort parameter, on top of the built-in match for gpt-5, o3, Claude, and Gemini models. Use this to opt in a model the built-in detection doesn't recognize yet.
structured_patch_models list[string] Model IDs (or glob prefixes ending in '*') that prefer the structured_patch edit tool over search_replace_block. Runtime override layer ON TOP of the controllable model→tool-variant map (mewbo_core/prompts/model_variants.yaml), which now holds the built-in defaults (GPT-5/o3/o4/Codex/GPT-4); only set this to override or extend without editing that file.

Context

Top-level key: context

Context window selection and event filtering.

Key Type Default Description
recent_event_limit integer 8 Maximum number of recent events injected into the context window.
selection_threshold number 0.8 Relevance score threshold (0.0-1.0) for the context selector to keep an event.
selection_enabled boolean true Enable LLM-based context event selection. When false, all recent events are used.
context_selector_model string "" Model ID for context selection. Falls back to llm.default_model when empty.

Speech

Top-level key: speech

Which gateway models handle speech, and which gateway serves them.

The three connection fields are all optional and all empty by default, because the common deployment has one gateway: speech falls back to llm.api_base/llm.api_key whenever these are blank, so an install that never writes a speech block still works. They exist for the deployment that genuinely splits the two, which is a real shape — a self-hosted synthesis backend beside a hosted chat provider — and refusing to represent it would only push the operator into running one gateway they do not want.

Key Type Default Description
api_base string "" Base URL of the gateway that serves the speech models. Leave it empty to use the same gateway as the language models.

Set this only when speech is served somewhere other than the endpoint under Language Model, such as a synthesis service running beside a hosted chat provider.
api_key string Key for the speech gateway. Leave it empty to reuse the language model key.

Only needed alongside a separate speech endpoint above. Setting one here without the other is almost always a mistake, because the key is then sent to the language model gateway that already had one. ⚠️
timeout number 90.0 Seconds to wait for the speech gateway before giving up.

Synthesis is not instant and scales with the length of the text: a sentence takes about half a second and a paragraph about four, so a short timeout cuts off long answers. Raising it has a cost too, because a request that is going to fail holds a server slot for the whole wait.
tts Text to speech Reading an answer aloud.
tts.model string supertonic-3 Model that turns text into audio. Leave it empty to turn read aloud off.

This is a model id your LLM gateway advertises, and speech models are a separate family from the chat models the answer itself runs on. A chat model named here is refused by the gateway at the moment someone presses play, not when this page is saved.
tts.voice string nova Voice the reader speaks in. Type the name your gateway knows it by.

This was once a fixed list of eleven, which was wrong for a self-hosted gateway: a backend can carry its own trained voice style, and a name absent from that list was refused here before the gateway ever saw it. The gateway decides what a voice is. A name it does not know fails the request when someone presses play, the same way an unknown model does.
tts.response_format string wav Audio format the gateway returns. WAV is the safe default and FLAC is the same audio in a smaller file.

Only these two are accepted. Asking for MP3, Opus, AAC or raw PCM fails the request, so they are not offered here.
stt Speech to text Turning a recording into text.
stt.model string nova-3 Model that transcribes a recording. Leave it empty to turn dictation off.

This is a model id your LLM gateway advertises. Transcription models are separate from both chat models and the text-to-speech model above, so the id here will not appear in the model picker used for answers.

Token Budget

Top-level key: token_budget

Token budget and auto-compaction thresholds.

Key Type Default Description
default_context_window integer 128000 Default context window size in tokens used when the model is not listed in model_context_windows.
auto_compact_threshold number 0.8 Fraction of the context window (0.0-1.0) that triggers automatic conversation compaction.
model_context_windows dict[str, integer] Override only: per-model context window in tokens. Keys are model names (with or without provider prefix). The authoritative source is LiteLLM's model catalogue; populate this only to cap below the model's real max, or for models LiteLLM doesn't know yet — which includes every model renamed by a proxy, since the catalogue is keyed on the real name. A model the catalogue cannot resolve falls back to default_context_window and logs a one-time warning naming it; that number then drives compaction and every utilisation reading, so an unlisted model silently budgets against a guess.

Compaction

Top-level key: compaction

Summarization prompt selection for conversation compaction.

caveman_mode enables a rule-augmented "caveman" prompt that instructs the summarizer LLM to drop articles, filler, pleasantries, and hedging while preserving code, paths, URLs, and error strings verbatim. Reduces output tokens in the compaction summary without changing the <analysis>/<summary> response structure downstream parsers expect.

Key Type Default Description
caveman_mode boolean false Enable caveman-style terse summarization prompt. Drops articles, filler, pleasantries, and hedging in the compacted summary while preserving code, file paths, URLs, and error strings verbatim. Reduces compaction output tokens without changing the response structure downstream parsers expect.

Langfuse

Top-level key: langfuse

Langfuse LLM observability integration.

Key Type Default Description
enabled boolean false Enable Langfuse tracing for all LLM calls.
host string "" Langfuse server URL.
project_id string "" Langfuse project ID for constructing dashboard URLs.
public_key string Langfuse project public key. ⚠️
secret_key string Langfuse project secret key. ⚠️

Home Assistant

Top-level key: home_assistant

Home Assistant smart-home integration.

Key Type Default Description
enabled boolean false Enable the Home Assistant tool for smart-home control.
url string "" Home Assistant API base URL.
token string Long-lived access token for Home Assistant authentication. ⚠️

Permissions

Top-level key: permissions

Tool execution permission policy.

Key Type Default Description
policy_path string "" Path to a JSON or TOML permission policy file. Empty uses built-in defaults.
approval_mode string ask Default approval mode: 'ask' prompts the user, 'allow' auto-approves, 'deny' blocks.

Safety Plane

Top-level key: safety

Master switch for the operator-owned tool-call gate and session observer.

OFF by default. While off, .mewbo/policy/ and .mewbo/monitor/ are never read, no rule is built and no event is emitted — a deployment that never sets this section pays nothing. There is deliberately no per-project override and no other knob here: discovery path, rule precedence and the built-in self-protection rule are fixed by the plane itself, not config-tunable, because a knob that could redirect discovery or reorder rules would be a knob that could weaken the guardrail it configures.

Key Type Default Description
enabled boolean false Master switch for the policy gate and session-budget observer read from .mewbo/policy/ and .mewbo/monitor/. OFF by default: no evaluation, no tokens, no latency until an operator turns this on.

CLI

Top-level key: cli

Terminal CLI display and interaction settings.

Key Type Default Description
disable_textual boolean false Disable the Textual TUI and fall back to plain Rich output.
remote CLI Remote Opt-in remote session sync + product tools (CLI-only).
remote.base_url string "" Base URL of the remote Mewbo deployment, a reverse-proxy root that serves the REST API under /api and the Mewbo MCP server under /mcp. Empty (default) ⇒ the CLI is fully local.
remote.token string API token presented to the remote deployment: sent as X-API-Key to the REST API for transcript sync and as a Bearer token to the Mewbo MCP server (which forwards it to the REST API). ${ENV_VAR} references are expanded by the CLI at use time. ⚠️

API Server

Top-level key: api

REST API authentication.

Key Type Default Description
master_token string msk-strong-password Bearer token required for all REST API requests. Change from the default before deploying. The MEWBO_MASTER_API_TOKEN environment variable OVERRIDES whatever is set here: when it is present, the API and the MCP server both use it and this value is never consulted — which is the case in every containerised deployment. To keep the token out of this file and say so plainly, set this to ${MEWBO_MASTER_API_TOKEN}. Any value may name an environment variable that way, and a name that is not set in the environment is refused at startup rather than read as empty. ⚠️
allow_external_cwd boolean false Allow callers to anchor sessions in arbitrary host paths via the cwd field on POST /api/sessions and POST /api/sessions/{id}/query. A directory belonging to a configured project, managed project or worktree, or registered repository checkout is always accepted, as is a re-send of the session's own bound directory; this flag governs host paths a caller names that the server does not already own.
max_concurrent_streams integer 6 How many Server-Sent Events streams the server keeps open at once. A stream holds one of the server's request threads for as long as it stays open rather than for the work it does, so without a bound enough of them starve every other endpoint and the server stops answering at all. Past this many, a new stream is refused with a retryable 503 and the rest of the API keeps serving. Keep it below the worker's thread count so ordinary requests always have headroom; 0 removes the bound.
apps_token_secret string Signing secret for Mewbo Apps render tokens — the short-lived, app-scoped read tokens the served app frontend presents on the read-only data/system endpoints. Set this to sign (and rotate) app tokens independently of the master token; when left empty it falls back to the master token, logging one startup warning. ⚠️
apps_exec_binaries list[string] Command-line programs a Mewbo App's code pipeline may run. A pipeline must still declare the ones it needs, so this is a ceiling the deployment sets rather than a grant: a program absent here cannot be reached however the pipeline is written. Widening it is an operator decision — pipeline code runs in-process, so a program added here runs with the API server's own file and network access, and a program that can be steered into running other programs effectively grants a shell.
apps_max_concurrent_pipelines integer 4 How many Mewbo App pipelines may execute at once. A pipeline runs synchronously and holds one of the server's request threads for its whole duration, so without a bound enough concurrent invocations starve every other endpoint. Past this many, an invocation is refused with a retryable 429 and the rest of the API keeps serving. 0 removes the bound.
auth Authentication Identity & access management (opt-in; off by default). Configures authenticators, roles, and sessions. Documented in docs/authentication.md.
auth.enabled boolean false Master switch for identity & access management. When off (the default), every request resolves to the built-in full-power identity and the server behaves exactly as it did before IAM. Turn on only after configuring at least one authenticator.
auth.authenticators list[Authenticator] Ordered list of identity sources (local API keys, OIDC, trusted reverse-proxy headers, LDAP, SAML). Each entry is an object whose kind field selects the authenticator type, plus that type's own settings. Validated in full at server startup; an invalid entry stops the server from booting. Each authenticator type's own settings are documented in docs/authentication.md.
auth.role_mappings object null Rules mapping identity-provider group names to Mewbo roles at login: an object with an ordered rules list and a default_role applied when no rule matches. The rule format is documented in docs/authentication.md.
auth.team_mappings object null Rules mapping identity-provider group names to team slugs at login: an object with an ordered rules list. The rule format is documented in docs/authentication.md.
auth.bootstrap object null Cold-start admin rule: grants the admin role to the first users matching an identity-provider group or an explicit subject allowlist, so an administrator exists before any role has been assigned. Documented in docs/authentication.md.
auth.session Session Browser session/cookie settings for federated logins.
auth.session.cookie_name string mewbo_session Name of the browser session cookie issued after a federated login.
auth.session.ttl_seconds integer 28800 Lifetime of a browser session, in seconds.
auth.session.secret string Signing secret for the browser session cookie. Required once any non-API-key authenticator is configured. Write-only: never returned by the config API. ⚠️
auth.avatars object null Avatar-resolution policy: whether to fall back to Gravatar for users without a profile picture, and the default image style. Documented in docs/authentication.md.
auth.scim SCIM SCIM 2.0 provisioning settings.
auth.scim.enabled boolean false Whether the SCIM 2.0 provisioning endpoint is served.
auth.scim.secret string Bearer secret an identity provider presents to the SCIM endpoint. Write-only: never returned by the config API. ⚠️
auth.audit object null Auth audit-trail settings: an object with an enabled flag; on by default once IAM is enabled. The events recorded are listed in docs/authentication.md.

Agent

Top-level key: agent

Sub-agent hypervisor settings.

Key Type Default Description
max_depth integer 5 Maximum nesting depth for sub-agent delegation (1 = no sub-agents).
max_concurrent integer 20 Maximum number of sub-agents allowed to run concurrently.
default_sub_model string "" Default LLM model for sub-agents. Falls back to the root agent's model when empty.
allowed_models list[string] Allowlist of model names sub-agents may use. Empty means all models are allowed.
model_tiers dict[str, string] Coarse model-cost tier -> concrete model id map (e.g. {'economy': 'anthropic/claude-haiku-4-5', 'frontier': 'anthropic/claude-opus-4-6'}), resolved by a spawned agent's DelegationContract.model_tier. An explicit spawn model arg or an agent_type's configured model always wins; a declared tier with no map entry is a silent no-op fallthrough.
session_step_budget integer 0 Ceiling on total tool-execution steps across every agent in a session (root + all sub-agents combined), enforced by the hypervisor: a warning is injected as the budget nears, and the run hard-stops at exhaustion. 0 = unlimited.
attestation_enabled boolean true Record a best-effort provenance hash chain over every spawn/terminal transition in a session's agent tree — bounded scalars + a contract snapshot + a summary FINGERPRINT only, never task text or raw summary content. Additive and never fatal: a failed or absent chain simply records no attestation. Default ON; this is the kill switch.
default_workspace_mode string full_access Root filesystem-containment tier every session starts at, narrowed per sub-agent by spawn_agent's workspace_mode. One of 'read_only' (reads confined to the workspace, no writes), 'workspace_write' (reads + writes confined to the workspace), or 'full_access' (no path restriction). The root stays 'full_access' because containment attenuates privilege ACROSS A SPAWN: the root agent acts directly for the operator, while a sub-agent defaults to 'workspace_write' and can only ever narrow further. Set this to confine the root session itself as well.
workspace_enforcement boolean true Master switch for workspace_mode filesystem containment. While on, an agent whose tier is narrower than 'full_access' resolves every path against its own workspace root plus the Mewbo-owned scratch roots, and the workspace root it was handed is authoritative — a wider root supplied in tool arguments is ignored rather than honoured. Turn it off to restore the older behaviour, where a path resolves against the union of every configured project root and an agent's tier governs nothing.
shell_sandbox boolean true Confine shell subprocesses using the kernel's Landlock LSM, so a command cannot read another task's files no matter which binary it runs. This is a DENY-list, not an allowlist: every configured project other than the session's own active one is denied, together with anything listed in agent.shell_denied_paths. Everything else — the interpreter, system libraries, the CLI toolbox and HOME — stays reachable, so nothing has to be enumerated to keep the shell working. Normal filesystem permissions still apply on top of this; Landlock only removes access, never grants it. On a kernel without Landlock this logs once and changes nothing.
shell_denied_paths list[string] Extra absolute directories the shell tool may never read or write, on top of the other-projects denial shell_sandbox already applies — for example the harness's own source and config directory on a deployment where those sit outside every configured project. A path that does not exist is ignored. This only affects commands run through the shell tool; every other tool is already argument-validated.
harness_self_deny boolean true Hide Mewbo's own source trees and the directory holding app.json from every session, so a command cannot read the file that holds the API keys. The directories are derived from where Mewbo is installed rather than configured, and a directory containing the Python runtime is never denied — that would stop every command instead of confining it. Turn this off to develop Mewbo itself, where its packages are the work rather than internals to hide; the session then reaches them like any other project. It denies unconditionally by default, including when the session's own project contains the installation, because the alternative was to guess when to make an exception: the guess is invisible on a deployment, where an unnoticed exception exposes the keys, while an unwanted denial is immediate and local on a workstation, where it is fixed by turning this off.
path_scope_to_active_project boolean true Scope the path-taking tools — file read, file edit, directory listing and LSP — to the session's active project, so a path argument resolves only under that project plus the extra directories its allowed_paths re-admits. While off, a path resolves against the union of every configured project root and the API host's own working directory, which lets one session read another project's files by naming them. This is the argument-validated half of the boundary agent.shell_sandbox enforces in the kernel for shell subprocesses, and both halves read the same active project. It is a separate switch because shell_sandbox also decides whether a kernel mechanism is applied at all, so a deployment that turns that one off — an older kernel, or a ruleset that broke a command — should not silently lose this check too. The Mewbo-owned scratch roots stay reachable either way, and a call with no active project, such as a direct library caller, is unaffected.
server_sandbox boolean false Extend the shell sandbox to the MCP and language servers Mewbo configures but does not spawn itself. Their command line is prefixed with a small launcher that applies the same Landlock ruleset to itself and is then replaced by the real server, which inherits the confinement. Each server is scoped to the workspace it serves — a language server to the project root it resolves, an MCP server to its own configured working directory — and denied every other configured project, so an MCP server that names no working directory reaches none of them. Servers reached over HTTP spawn no process here and are unaffected. This needs agent.shell_sandbox as well, because it applies the same kernel mechanism: a deployment that turned that one off must not have it reappear underneath its language servers. Off by default because a denied path surfaces inside a server as a missing file rather than as a refusal, and which servers a deployment runs is not knowable in advance.
stall_threshold_s number 120.0 Seconds of no tool-execution progress before the watchdog flags an agent as stalled and injects an NL warning into its message queue.
stall_check_interval_s number 30.0 Seconds between watchdog stall-detection sweeps.
write_progress_signal_step_threshold integer 25 Consecutive non-write tool-execution steps before the write-progress signal fires telemetry for a write-capable agent. 0 disables.
write_progress_signal_event_interval integer 10 Steps between repeat write-progress signal events once the threshold is crossed.
write_progress_signal_max_events integer 2 Maximum write-progress signal events emitted before it goes quiet.
write_progress_signal_reminder_enabled boolean false When the write-progress signal fires, also inject a criterion-blind objective-restatement reminder (states the task goal only — never the signal or its criteria). Default off: telemetry alone is the observe-only default.
verification_enabled boolean false Master switch for verifier-gated completion. OFF by default (staged): while off, a spawn's verification spec is carried but never run, so every natural completion is accepted unchanged. Flip on to gate a write-capable agent's claimed completion behind a ground-truth command check before its text is accepted.
verification_max_retries integer 2 Maximum times a failed completion verifier re-drives the agent before its text is accepted, honestly flagged verification_failed. Clamped to [0, 10] and bounded ALSO by the step/wall budget, whichever is tighter. 0 = one check, no retry.
verification_timeout_s number 60.0 Ceiling in seconds on a single verifier subprocess; a spawn's per-spec timeout_s is clamped down to this at run. Clamped to [1, 600].
llm_call_timeout number 300.0 Ceiling in seconds for a single model.ainvoke() call. Covers extended-thinking models (raised from 60s because bare timeouts were the largest single failure class). On timeout, the call is retried up to llm_call_retries times before cascading to fallback models. Deployments where one call legitimately runs long (slow local inference, a saturated proxy) can lift the ceiling without editing this file via MEWBO_AGENT_LLM_CALL_TIMEOUT.
llm_first_token_timeout number 0.0 Ceiling in seconds on the wait for the FIRST chunk of a streamed model call. 0 (the default) defers to llm_call_timeout, so one bound covers the whole call. A large prompt on an extended-thinking model can legitimately take minutes to emit its first token, so a tighter value here manufactures failures unless a deployment has measured its own first-token latency.
llm_stream_idle_timeout number 240.0 Maximum seconds between chunks once a streamed model call has started producing. This is the bound llm_call_timeout cannot express: a provider that returns 200 and then goes silent is otherwise unbounded until the total ceiling, and a total ceiling loose enough for a long healthy generation is far too loose for a dead one. 0 disables it, leaving llm_call_timeout as the only bound.
llm_call_liveness_s number 600.0 Seconds one outstanding model call may go silent before the loop emits an llm_call_stalled event. This DETECTS a wedged call; it does not recover one — asyncio's attempt cap can only cancel a coroutine that reaches a cancellation point, and a provider read wedged below the event loop never does. Keep it above llm_call_timeout, or a normally timing-out call reports as stalled.
llm_call_retries integer 2 Maximum attempts for the primary model before cascading to fallback models (default 2 = one try + one retry). Each fallback model gets retry.fallback_retries attempts. A rescue model that wins is pinned for the rest of the run. Backoff/budget/circuit-breaker live under agent.retry.
retry Retry Automatic LLM-call retry / fallback resilience knobs.
retry.backoff_base number 1.0 Base seconds for full-jitter backoff: random(0, min(cap, base*2^(n-1))).
retry.backoff_cap number 60.0 Maximum backoff delay in seconds.
retry.retry_after_cap number 60.0 Upper bound in seconds applied to a server's Retry-After header before the loop sleeps on it; caps how long one misbehaving response can stall a run.
retry.turn_deadline number 1500.0 Wall-clock seconds budget for one logical LLM call across all retries and fallbacks. Checked before each attempt AND before advancing to the next model, so a spent budget stops the chain without claiming a model it never called. Must exceed llm_call_timeout x llm_call_retries plus one more attempt, or no fallback model is ever reachable. 0 disables.
retry.fallback_retries integer 1 Attempts per fallback model after the primary is exhausted.
retry.circuit_breaker_threshold integer 3 Consecutive per-model failures before that model is cooled down and skipped (when an alternative exists). 0 disables this heuristic. It does NOT disable the cooldown a provider declares for itself: a model whose quota the provider reports as exhausted is still skipped until that quota resets, since that is a stated fact rather than an inference this threshold tunes.
retry.circuit_breaker_cooldown number 30.0 Seconds a model is skipped after tripping the circuit breaker. A provider-declared reset time overrides this for that model.
retry.budget_capacity number 24.0 Token-bucket retry budget for transient failures (timeouts, 5xx, connection errors). Retries stop once the bucket drops to half capacity, so a sustained outage fails fast instead of storming.
retry.rate_limit_budget_capacity number 4.0 Separate, much smaller token-bucket budget for rate-limit (429) retries. Retrying a server error bets that the server recovers in seconds and usually pays; retrying a throttle bets against a rate window the provider controls, and each attempt re-sends the whole prompt. Once this bucket is spent the run moves to the next model in the ladder rather than failing, because that model draws on a different quota pool. Retries stop at half capacity, as above.
retry.doom_loop_threshold integer 3 Halt cleanly when the model repeats the same tool + identical input this many times in a row (no progress). 0 disables.
default_denied_tools list[string] Tool IDs denied to all sub-agents by default (e.g. spawn_agent).
edit_tool string "" File editing mechanism override: 'search_replace_block' (Aider-style SEARCH/REPLACE blocks) or 'structured_patch' (per-file exact string replacement). Leave empty (default) to auto-select based on the active model via llm.structured_patch_models.
plan_mode_shell_allowlist list[string] Shell command prefixes allowed during plan mode. Each entry matches a command at a word boundary (e.g. 'git log' matches 'git log --oneline' but not 'git logger'). Commands containing pipes, redirects, variable expansion, command substitution, or chaining (|, >, <, &, ;, $, backtick) are always rejected. Set to an empty list to disable shell in plan mode entirely.
web_ide Web IDE null Optional 'Open in Web IDE' feature config (code-server containers).
web_ide.enabled boolean false Turn on the 'Open in Web IDE' feature (per-session code-server containers via Docker). Also requires a MongoDB-backed session store; toggling this needs an API process restart to take effect since the /api/ide routes are registered at startup.
web_ide.image string codercom/code-server:latest Docker image used to launch each session's code-server container.
web_ide.default_lifetime_hours integer 1 Hours a new Web IDE container stays up before it self-terminates, unless the session extends it first.
web_ide.max_lifetime_hours integer 8 Hard ceiling on a session's total Web IDE lifetime across all extensions; a request to extend past this is rejected.
web_ide.cpus number 1.0 CPU core limit for each Web IDE container, e.g. 1.0 = one core (maps to Docker's --cpus / nano_cpus).
web_ide.memory string 1g Memory limit for each Web IDE container, in Docker's --memory syntax: digits followed by m or g, e.g. '1g' or '512m'.
web_ide.pids_limit integer 512 Maximum number of processes/threads allowed inside a Web IDE container; bounds a runaway process from exhausting the host.
web_ide.network string mewbo-ide Docker network each Web IDE container joins. Must be the same network the ide-proxy is attached to, or the proxy can't reach the container.
web_ide.proxy_url string http://127.0.0.1:5126 Base URL the API uses to reach the ide-proxy for readiness probes. The default suits a host-networked API, where the proxy is published on loopback 127.0.0.1:5126. A bridge-networked API (e.g. one joined to extra Docker networks via a compose override) cannot reach that loopback: attach it to the network above and point this at the proxy's in-network name, e.g. http://mewbo-ide-proxy:8080.
web_ide.state_dir string /tmp/mewbo-ide Host directory where each Web IDE container's expiry-deadline file is written; the container's internal watchdog reads it to self-terminate on schedule.
web_ide.broker_url string "" Base URL of the IDE broker service, e.g. http://127.0.0.1:5128. When this is set and the MEWBO_IDE_BROKER_TOKEN environment variable is present, the API delegates every container operation to the broker and needs no Docker access of its own — the broker holds the socket and builds each container spec from its own configuration. Leave empty to drive Docker directly from the API process, which requires giving that process the socket. The MEWBO_IDE_BROKER_URL environment variable overrides this value; the shared secret is read from the environment only, never from this file.
lsp LSP Language Server Protocol integration settings.
lsp.enabled boolean true Master switch for the native LSP tool (hover/diagnostics/go-to-definition). When off, or when the pygls dependency isn't installed, the tool is never registered and the agent works from grep/read alone.
lsp.servers dict[str, object] Override or extend built-in server definitions. Set {"pyright": {"disabled": true}} to disable a built-in, or add custom servers with command/extensions/root_markers.
tool_search Tool Search Deferred tool loading via on-demand schema fetching.
tool_search.mode string on 'on' (the default) always defers MCP tools and any spec with metadata.deferred=True; the model loads schemas on demand via tool_search, so no user MCP tool occupies the context window until it is actually needed. 'off' keeps every tool's schema in the initial bind. 'auto' defers only when the number of deferrable tools exceeds auto_threshold. Note that 'on' costs a zero-MCP session nothing: deferral only engages when the deferrable set is non-empty, so the two modes bind an identical list there. The range where they differ is 1..auto_threshold tools, where 'auto' spends ~240 tokens per tool every turn.
tool_search.auto_threshold integer 25 In 'auto' mode, defer tool schemas only when more than this many deferrable tools (MCP + metadata.deferred specs) are registered. Ignored when mode is 'off' or 'on'.
Deprecated fields
Key Type Default Description
max_iters integer 30 Deprecated. The tool-use loop now runs until natural completion (model returns text without tool calls). This field is retained for API backward compatibility but is not enforced.

Wiki

Top-level key: wiki

Operator-facing knobs for the wiki subsystem.

Key Type Default Description
default_model string "" Model the wiki picker pre-selects for indexing (the wizard). Overrides llm.default_model for the wizard. Empty string means: fall back to llm.default_model.
default_qa_model string "" Model the Q&A composer pre-selects. Typically smaller/faster than default_model because Q&A is a tight read-only loop where latency matters more than depth. Empty string means: fall back to default_model, then to llm.default_model.
default_qa_fast_model string "" Model the Q&A composer's fast mode pre-selects. Fast mode holds the retrieval surface itself with no probe fan-out, so a smaller/faster model than default_qa_model often suffices. Empty string means: fall back to default_qa_model, then default_model, then llm.default_model.
qa_fast_step_budget integer 15 Tool-step ceiling for a fast-mode Q&A run — the root itself retrieves with no probe fan-out, so it needs far fewer steps than deep mode's session-wide budget. Budgets are config-tunable, never hardcoded.
default_depth string "" Indexing depth the wizard pre-selects. Empty string means: use the wizard's own default (comprehensive).
default_language string "" Language code the wizard pre-selects (e.g. en, es). Empty string means: use the wizard's own default.
embedding Embedding Embedding settings for the wiki indexer.
embedding.enabled boolean true When false, wiki_build_graph skips embedding generation and retrieval falls back to BM25 + graph traversal only.
embedding.model string openai/text-embedding-3-small Embedding model ID routed through the LLM proxy. Must support the OpenAI /v1/embeddings shape (LiteLLM normalises Gemini and others to this shape). Pin a fast model here to speed up indexing, since embedding is per-node and runs synchronously.
embedding.batch_size integer 64 Number of graph nodes embedded per API call during indexing. Embedding throughput per connection is roughly flat regardless of batch size — a larger batch just takes proportionally longer per call — so use concurrency to speed up indexing, and use this knob only to stay under the provider's request payload limit.
embedding.concurrency integer 4 Maximum number of embedding requests issued to the provider concurrently during indexing. Embedding is I/O-bound — the indexer spends its time waiting on the network, not on CPU — so this, not batch_size, is what determines indexing speed. Keep it conservative: issuing too many requests at once trades throughput for HTTP 429 responses and retries. Raise it only after confirming headroom against your provider's rate limits.
embedding.requests_per_minute integer null Optional ceiling on embedding requests issued per minute. When set, the indexer paces itself against this budget — queuing work rather than firing it — instead of relying on concurrency alone. Leave unset to rely on concurrency plus automatic backoff on 429 responses.
embedding.tokens_per_minute integer null Optional ceiling on embedding tokens processed per minute, estimated from input text length. Paced the same way as requests_per_minute. Leave unset to rely on concurrency plus automatic backoff on 429 responses.
embedding.max_retries integer 5 Number of times a rate-limited (HTTP 429) embedding request is retried before the indexing job fails. A retry honours the provider's Retry-After header when present, and falls back to exponential backoff with jitter otherwise; sustained rate-limiting also reduces concurrency for the remainder of the run so a rate-limited pass degrades to slower rather than failing outright.
memory Memory Multiplex memory-layer knobs (atomic insights over the graph).
memory.enabled boolean true Master switch for the memory layer (gates wiki_submit_insight).
memory.model string "" Chat model for condense + LLM dedup on the human/REST/MCP path. Empty → falls back to default_qa_model, then default_model.
memory.max_insight_chars integer 200 Hard cap on a memory note's length. Lowering this shortens notes; it cannot be raised above 200, which is the length the stored note model itself declares — a larger value would let a note past this check and then fail validation as it was written.
memory.max_anchors integer 8 Max code anchors per note.
memory.dedup_k integer 5 kNN candidate window for fuzzy + LLM dedup tiers.
memory.dedup_cosine number 0.6 Cosine floor for the LLM dedup tier.
memory.fuzzy_jaccard number 0.85 Jaccard floor for the fuzzy dedup tier.
memory.fusion_w_ppr number 0.1 Weight applied to a code node's score when it's surfaced only by following a memory note's anchor rather than direct text/code search. Raise it to rank memory-anchored context higher relative to direct hits; lower it toward 0 to favor direct hits.
memory.hub_degree integer 50 Degree above which an anchor is hub-damped.
memory.expansion_hops integer 1 Structural hops to expand from an anchor.
refresh Refresh On-demand incremental-refresh thresholds.
refresh.default_mode string auto Default re-index strategy when none is requested.
refresh.closure_max_depth integer 4 Reverse-dependency closure depth cap.
refresh.drift_keep number 0.9 Cosine ≥ this keeps a memory anchor (no LLM).
refresh.drift_invalidate number 0.75 Cosine < this invalidates a memory anchor.
refresh.page_keep number 0.05 Doc staleness < this → keep.
refresh.page_edit number 0.35 Doc staleness < this → edit.
refresh.page_regen number 0.7 Doc staleness ≥ this → regenerate + review.
refresh.new_page_min integer 5 Uncovered public symbols to propose a new page.
phase_timeouts Phase timeouts How long each long-running indexing phase may run before it is called wedged.
phase_timeouts.clone_s number 1800.0 Seconds wiki_clone_repo waits for the checkout. Sized from the credential chain's own worst case rather than from a clone's duration: the chain tries up to five candidates and each git attempt is capped at 300s, so a repository whose every stored credential has been revoked legitimately spends 1500s before it reaches the anonymous attempt that succeeds. Raise it for very large repositories on slow links.
phase_timeouts.scan_s number 900.0 Seconds wiki_scan_tree waits for the tree walk. The scan reads and hashes every file it keeps, so its cost tracks total bytes on disk rather than file count, and it touches no network. Raise it for a very large monorepo or a slow filesystem.
phase_timeouts.graph_build_s number 3600.0 Seconds wiki_build_graph waits for the tree-sitter parse, the graph write and the embedding pass. The longest of the phases and the one with the widest spread: it scales with parseable source size, and its embedding leg waits on the LLM proxy. Raise it for a large repository or a slow embedding model.

SCG

Top-level key: scg

Operator-facing knobs for the Source Capability Graph (agentic search).

Key Type Default Description
enabled boolean false Master switch for the SCG feature (source mapping and orchestrated agentic-search runs).
traversal Traversal Traversal defaults (the per-run search tier).
traversal.default_tier string auto Default search tier, one budget knob over decomposition depth and probe fan-out. Overridable per run.
traversal.tier_models Tier models Which LLM each search tier runs on (fast/auto/deep).
traversal.tier_models.fast string openai/gpt-oss-120b Model for fast tier runs (the tier still sets the low-latency budget).
traversal.tier_models.auto string openai/gpt-oss-120b Model for auto tier runs (the tier still sets the balanced budget).
traversal.tier_models.deep string openai/gpt-oss-120b Model for deep tier runs (the tier still sets the exhaustive budget).

Hooks

Top-level key: hooks

External shell hooks fired during the session lifecycle.

Command hooks run unsandboxed shell commands with the API/CLI process's own privileges (see HookEntry.command's docstring) — a caller who can PATCH this section can execute arbitrary code on the host. x-protected puts the whole section in the same never-read-never-written-via-API tier as the other host-level settings in this file: settable only by editing the config file directly, never over the network regardless of credential (see ConfigSchemaView in apps/mewbo_api).

Key Type Default Description
pre_tool_use list[Hook] Hooks executed before each tool invocation.
post_tool_use list[Hook] Hooks executed after each tool invocation.
on_session_start list[Hook] Hooks executed when a new session begins.
on_session_end list[Hook] Hooks executed when a session ends.
on_event list[Hook] Hooks executed (fire-and-forget) for every event appended to a session transcript. The matcher fnmatches the event type.

Plugins

Top-level key: plugins

Plugin system configuration.

Key Type Default Description
enabled boolean true Turn the whole plugin system on, including Mewbo's own built-in suites such as widget_builder.

While this is off, no plugin contributes anything to the agent, whether it is built in or installed from a marketplace: no agent definitions, no skills, no hooks, no MCP tools. The enabled_plugins and marketplaces settings below are ignored entirely until you turn it back on.
enabled_plugins list[string] Plugin names to enable. Empty = all installed plugins. Format: 'plugin-name' or 'plugin-name@marketplace'.
marketplaces list[string] Marketplace catalogs holding a marketplace.json plugin index, on any git host. Each entry is a full git URL (https/ssh/git, or scp-style git@host:owner/repo), a 'host/owner/repo' shorthand, or a bare 'owner/repo' (cloned from marketplace_default_host).
marketplace_default_host string github.com Default git host for bare 'owner/repo' marketplace entries. Full URLs and 'host/owner/repo' entries ignore this.
install_path string "" Override install path for Mewbo-managed plugins. Defaults to $MEWBO_HOME/plugins/ (via resolve_mewbo_home).

Triggers

Top-level key: triggers

Reverse-invocation trigger subsystem.

The durable peer of the sub-agent hypervisor: a background watcher that fires time / cron / CI / forge-PR / webhook triggers and re-invokes the sessions that armed them. OFF by default (enabled=False), so the feature is un-enableable until an operator turns it on and a stock deployment pays nothing. The lower half of this section is the admission policy the schedule_trigger tool + the arm route enforce (mirrors mewbo_core.triggers.policy.TriggerPolicy field-for-field; to_policy builds one).

Key Type Default Description
enabled boolean false Turn the trigger watcher on. Nothing fires until you do.

A trigger is how Mewbo starts a session later, on its own, with nobody watching: at a set time, on a repeating schedule, or when a CI run finishes, a pull request changes, or a webhook calls in. The watcher is the background loop that notices those moments and wakes the session that asked to be woken. While it is off, the trigger routes still work, so a session can arm a trigger and you can list, pause, or cancel it, but no trigger ever fires. Armed triggers simply wait until you turn the watcher on.
tick_interval_seconds number 5.0 How often the watcher wakes up to look at the schedule, in seconds.

On each pass it expires the triggers whose deadline has gone by and fires the time and cron triggers that have come due. A shorter interval wakes a session closer to the moment it asked for; a longer one costs the server less. This is also the cadence at which the forge poll below gets a chance to run.
poll_interval_seconds number 60.0 How often the watcher asks the forge about CI runs and pull requests, in seconds.

Time and cron triggers can be judged from the clock alone, but ci.workflow and forge.pr triggers cannot: the watcher has to call the forge's REST API to see what changed. Those calls are rate-limited and cost a round trip each, so they run on this deliberately coarser cadence rather than on every pass. Raise it if you are bumping into API limits; lower it if you want CI results picked up sooner.
max_consecutive_failures integer 5 How many errors in a row one trigger may hit before it is given up on.

When a fire or a forge poll raises, the watcher records the error on the trigger and leaves it armed, so a passing outage never throws away a schedule. Once a trigger has failed this many times back to back without a single success in between, the watcher stops retrying it and moves it to failed. Any success resets the count to zero.
max_armed_per_session integer 20 The most triggers one session may have armed at the same time.

Triggers are armed by the agent from inside a session, so this ceiling is what keeps a single session from filling the schedule with wakes. An attempt to arm one past the limit is refused, and the agent is told why. Cancelling a trigger, or letting one finish, frees the slot again.
max_fires_cap integer 100 The ceiling on how many times any single trigger may fire.

A repeating trigger, a cron schedule for instance, can name its own max_fires limit when it is armed. This is the ceiling on that request: an attempt to arm a trigger asking for more is refused. A trigger that reaches its own limit completes and stops firing.
default_expiry_days number 7.0 How long an armed trigger lives when it names no expiry of its own, in days.

Every trigger expires eventually, so that a wake nobody remembers arming cannot linger forever. When the agent arms one without setting an expiry date, this many days from the moment of arming is stamped on it. Once that moment passes, the watcher expires the trigger instead of firing it.
cron_min_interval_seconds integer 60 The shortest gap allowed between two fires of a cron trigger, in seconds.

A cron expression can be written to fire far more often than a session is worth waking, so this is the floor. When a cron trigger is armed, the gap between its first two fires is measured, and a schedule tighter than this is rejected there and then rather than being throttled later.
webhook_payload_max_bytes integer 200000 How much of an incoming webhook body the woken session gets to see, in bytes.

A webhook can carry a large payload, and all of it becomes context the session has to read. A body bigger than this is truncated rather than rejected: the call still fires the trigger, the session receives the first part of the body, and it is told the payload was cut short. When a signature is configured, it is checked against the whole body before any truncation happens.

Channels

Top-level key: channels

Channel adapters, keyed by name (nextcloud-talk, email).

Structure varies by entry. See the schema source for details.

Projects

Top-level key: projects

Directories you've already created, registered here by hand under a short name; sessions reference them by that name (e.g. "project": "<key>"). Distinct from Mewbo-managed ("virtual") projects, which the API creates and owns itself and which sessions reference as managed:<project_id>: entries here are never created, modified, or deleted by Mewbo, only pointed at.

Structure varies by entry. See the schema source for details.