Inference Gateway CLI
The Inference Gateway CLI (infer) is a powerful Go-based command-line tool providing comprehensive access to the Inference Gateway with interactive chat, autonomous agents, Computer Use tools, and development workflows.
Versioning: the CLI is pre-1.0 and breaking changes are expected until it stabilises. The commands on this page install
@latest, so there is no version number to keep in sync here - see the releases page for whatlatestcurrently resolves to.
Key Features
- Zero-Configuration Setup - Add API keys and start chatting
- Autonomous Agent Mode - Delegate complex tasks with iterative execution
- Computer Use Tools - GUI automation with screenshot, mouse, and keyboard control
- Screen Recording - Record the screen, a window, or a region to MP4 from chat (Learn more)
- Rich Tool Integration - File operations, code search, web access, GitHub via the
ghCLI - Smart Safety System - Configurable approval workflow with diff visualization
- Beautiful TUI - Scrollable interface with syntax highlighting and multiple themes
- Web Terminal - Browser-based interface with tabbed sessions
- Remote Messaging Channels - Control the agent from Telegram and other platforms (Learn more)
- Agent Skills - Reusable, model-readable instruction folders loaded on demand, portable across vendors (Learn more)
- Cost Tracking - Real-time token usage and cost calculation
Installation
npm / npx (Recommended)
Run the CLI without installing anything (requires Node.js >= 18). The matching native binary is downloaded and cached on first use:
npx @inference-gateway/cli@latest --help
npx @inference-gateway/cli@latest chatOr install it globally:
npm install -g @inference-gateway/cli
infer --helpNot recommended for production - prefer the install script or building from source. Prebuilt binaries cover Linux and macOS on amd64/arm64 (on Windows, use WSL).
Install Script (Recommended)
# Latest version
curl -fsSL https://raw.githubusercontent.com/inference-gateway/cli/main/install.sh | bash
# Specific version
curl -fsSL https://raw.githubusercontent.com/inference-gateway/cli/main/install.sh | bash -s -- --version v0.97.0
# Custom directory
curl -fsSL https://raw.githubusercontent.com/inference-gateway/cli/main/install.sh | bash -s -- --install-dir $HOME/.local/binGo Install
go install github.com/inference-gateway/cli@latestManual Download
Download binaries from the GitHub releases page. Binaries are signed with Cosign for verification.
Build from Source
git clone https://github.com/inference-gateway/cli.git
cd cli
CGO_ENABLED=0 go build -tags purego -o infer ./cmd/inferThe build is fully cgo-free on macOS, Linux, and Windows - no C toolchain and no macOS SDK are required, and the same host cross-compiles every target. The purego build tag selects the pure-Go display and input backend and is used on all three platforms.
Shell Completions
The CLI ships an infer completion subcommand (provided by fang) that generates completion scripts for bash, zsh, fish, and powershell. Enabling completions adds tab-completion for subcommands, flags, and many flag values.
# Zsh (current session)
source <(infer completion zsh)
# Zsh (persistent) - write to a directory on $fpath
infer completion zsh > "${fpath[1]}/_infer"
# Bash (current session)
source <(infer completion bash)
# Bash (persistent)
infer completion bash > /etc/bash_completion.d/infer
# Fish
infer completion fish > ~/.config/fish/completions/infer.fish
# PowerShell
infer completion powershell | Out-String | Invoke-ExpressionRun infer completion --help to list the supported shells. After writing a persistent completion file, start a new shell (or re-source your shell rc) for it to take effect. If completions do not appear, see Shell Completions Not Working.
Quick Start

# Initialize configuration
infer init
# Generate AGENTS.md documentation for AI agents (recommended for new projects)
infer chat
> /init
# Check gateway status
infer status
# Start interactive chat
infer chat
# Launch web terminal
infer chat --web
# Headless mode
infer headless "Analyze this codebase and suggest improvements"
# Get help (styled output)
infer --help
# Show version
infer --version
# Enable shell completions for the current shell (zsh example)
source <(infer completion zsh)Generating AGENTS.md
For new projects, use the /init shortcut to automatically generate an AGENTS.md file. This file provides structured documentation that helps AI agents understand your project:
infer chat
> /initThe agent will:
- Analyze your project structure with the Tree tool
- Examine configuration files, build systems, and documentation
- Generate comprehensive
AGENTS.mdincluding:- Project overview and technologies
- Architecture and structure
- Development environment setup
- Key commands (build, test, lint, run)
- Testing instructions
- Project conventions and coding standards
- Important files and configurations
This documentation helps other AI agents (and developers) quickly understand how to work with your project.
Help and Version Output
The CLI's help, error, and version output are rendered with fang, so every command produces styled, colorized output. The samples below are shown as plain text; in a real terminal the headings, flags, and errors are colorized.
Help
infer --help (and --help on any subcommand) prints a styled usage page grouped into usage, commands, and flags. Note that -v is --verbose; version is the long-form --version flag.
infer
A powerful command-line interface for managing and interacting with
the Inference Gateway.
USAGE
infer [command] [--flags]
COMMANDS
init Initialize user configuration
status Check gateway health and resource usage
chat Interactive chat session (TUI)
headless Headless task execution
config Configuration management
tools Run and inspect agent tools directly
completion Generate the autocompletion script for the specified shell
version Show version information
FLAGS
-h --help Show help
-v --verbose Verbose output
--version Print version informationErrors
Unknown commands and flags exit non-zero with a styled error message and no noisy usage dump (fang sets cobra's SilenceErrors/SilenceUsage and renders the error itself):
$ infer badcmd
Error: unknown command "badcmd" for "infer"
$ echo $?
1Version
infer --version prints the version, styled by fang:
$ infer --version
infer version vX.Y.ZThe standalone version subcommand is kept for backwards compatibility and prints the same information:
infer versionThe manual
--versionboolean flag was replaced by fang's built-in version handling (fang.WithVersion). Bothinfer --versionand theinfer versionsubcommand remain supported.
Core Commands
| Command | Description | Key Features |
|---|---|---|
infer init | Seed the userspace baseline | Creates ~/.infer/ defaults - writes nothing to the project |
infer status | Check gateway health | Shows resource usage and connectivity |
infer chat | Interactive chat TUI | Streaming, scrolling, tool expansion, mode switching |
infer chat --web | Web-based terminal | Browser interface, tabbed sessions, remote access |
infer headless <task> | Autonomous task execution | Background operation, task planning, validation |
infer config <cmd> | Configuration management | Generic get/set for any config key |
infer tools <cmd> | Run agent tools directly | Execute a tool or validate a bash command |
infer stats | Summarize local telemetry | Token usage, tool outcomes, and cost across sessions |
infer traces | View a session's trace span tree | Offline span-tree viewer, --list and JSON output |
Chat Interface Features
Navigation:
- Shift + Arrow Down/Up: Scroll chat history
- Ctrl+R: Toggle tool result expansion
- Shift+Tab: Cycle agent modes (Standard -> Plan -> Auto-Accept -> Auto+Judge)
- Ctrl+K: Toggle model thinking blocks
Capabilities:
- Real-time streaming with syntax highlighting
- Mouse wheel and keyboard scrolling
- Model switching during conversation
- Tool result inspection
- Cost tracking in status bar
- Collapsible thinking blocks
- GitHub issue references - type
#to insert and expand#Ntokens (see below)
GitHub Issue References (#)
Type # in the chat input to open a dropdown of the current repository's open issues - each entry shows the issue number, title, and state. The list is resolved through the gh CLI from the repo's git remote, newest first. Selecting an issue inserts a highlighted #N token into your message.
On submit, every #N token is expanded inline into that issue's title, body, and most recent comments (up to the latest 20) before the message is sent to the model - so the agent works from full issue context without a redundant gh issue view lookup.
infer chat
> Summarize #123 and propose a fix
# "#123" expands into the issue title, body, and recent comments before sendingPrerequisite: the gh CLI must be installed and authenticated, and the working directory must be a git repository with a remote. The feature gracefully no-ops when gh is missing, the directory is not a git repo, the repo has no remote, or authentication has expired - the dropdown simply shows nothing.
Resuming a session (--session-id)
infer chat --session-id <id> loads a persisted conversation before the TUI starts, letting you pick up where you left off. The session must have been persisted by a storage backend (storage.enabled: true).
# Find session IDs from saved conversations
infer conversations list
# Resume a specific session
infer chat --session-id abc-123-defSession ID resolution:
- A literal UUID is used as-is.
- Any non-UUID value is treated as a session group key and resolved through the session rollover chain to the latest session in that group.
Fallback: If the session cannot be loaded (e.g. the ID does not exist or storage is disabled), the CLI prints a visible notice and starts a new session under the requested ID — the same semantics as infer headless --session-id.
Ignored in non-interactive modes: The flag is ignored (with a printed notice) in --web mode and when input is piped (non-interactive).
Status indicator row
Below the chat input, a row of status indicators shows the current agent state. You can interact with these indicators using the keyboard:
| Key | Action |
|---|---|
Down arrow | Focus the indicator row from the chat input |
Left / Right arrow | Cycle between indicators |
Enter | Open the matching view for the selected indicator |
Esc / Up arrow | Return focus to the chat input |
The selected indicator is highlighted as an accent-colored pill.
Indicator labels:
| Indicator | Label format | Description |
|---|---|---|
| Tools | Tools: N (mode) | N is the number of tools available in the current agent mode. mode is the active mode name (Standard, Plan, Auto-Accept, or Auto+Judge - the latter shown as AUTO+JUDGE - <model>). Opens the /tools view. |
| A2A | A2A: X/Y | X is the number of connected A2A agents, Y is the total number of configured agents. Opens the /agents view. When liveness probes are enabled, X counts down as agents fail and counts back up when they recover - the indicator stays live for the session lifetime. |
| Theme | Theme | Opens the theme selector to change the TUI color scheme. |
| Reconnect | Reconnecting... / Reconnecting (N/M) | Shown in red when the stream has stalled and the CLI is reconnecting. N is the current attempt, M is client.retry.max_attempts. Input is blocked until the stream recovers or all attempts are exhausted. |
Switching models (/model)
/model is the unified model command:
/model <name>- permanently switch the active model for the rest of the session./model <name> <prompt...>- run a single message with<name>, then restore the session model afterward. Handy for sending one hard question to a stronger model without changing your default./model(no argument) - open the model picker.
infer chat
> /model deepseek/deepseek-v4-flash # switch the session model
> /model anthropic/claude-opus-4-8 Explain this stack trace # one-off, then restoreModel picker labels
When you open the model picker (/model with no argument), each model may show a suffix label indicating its capabilities:
| Label | Meaning |
|---|---|
vision | The model accepts image input natively - pasted or @-referenced images are seen directly, no ImageDecode needed. |
audio | The model takes audio input (speech-to-text, multimodal chat). |
video | The model takes video input. |
image-gen | The model generates images (e.g. DALL-E, GPT-Image) rather than text. |
view-only | The model cannot serve /chat/completions and therefore cannot be selected - see below. |
Labels are derived from the gateway-reported modalities (/v1/models?include=modalities), not from name patterns. A model gets the vision label when its input modalities include both text and image; it gets the image-gen label when its output modalities include image without text. Labels are appended to the metadata suffix after the context window and price:
ollama_cloud/glm-5.3-flash (1M, vision, video)
groq/whisper-large-v3 (?, $0.04/$0.00 per MTok, audio, view-only)Gateway version requirement: Modalities-driven labels and the chat-capability filter require gateway v0.47+. Against an older gateway, models report no modalities, so no labels are shown and no model is chat-capable. Upgrade the running gateway to v0.47+ to see your catalog.
View-only models
Non-chat models - speech-to-text (e.g. groq/whisper-*), text-to-speech (e.g. openai/tts-*, groq/playai-tts*), image generation and video - are listed in the picker with a view-only marker, after the chat-capable rows. They are visible so you can see what the gateway offers, but pressing Enter on one is rejected with a notice (<model> does not support chat and cannot be selected) and the picker stays open. Only chat-capable models can be selected, and /model <name>, model validation and autocomplete are unaffected.
A model the gateway reports no modalities for ("modalities": null) is treated as not chat-capable, which is how most speech and embedding models arrive.
Filter tabs
The picker has two tab rows that are ANDed together, so Free + Vision lists only free vision models:
| Row | Keys | Tabs |
|---|---|---|
| Pricing | 1-4 | [1] All, [2] Free, [3] Pay-as-you-go, [4] Subscription |
| Capability | 5-8 | [5] Any, [6] Vision, [7] Audio, [8] Video |
Vision matches image input. Audio and Video match models that work with that modality on either side, so both speech-to-text and text-to-speech models appear under Audio. See Model Categories for how the pricing tabs are derived.
The footer help line lists the full keymap:
↑↓ navigate · Enter select · / search · esc clear · 1-4 pricing · 5-8 capability · Ctrl+C cancelDigits typed while the search input is active (/) go into the query, not the tabs.
Diff viewer and git staging
When the agent proposes file changes (or you open a diff), the diff viewer supports patch-level staging - select individual lines, split hunks, and stage or unstage everything at once. All keys are configurable in .infer/keybindings.yaml (category diff_viewer); the defaults:
| Key | Action |
|---|---|
space / v | Start or clear a line-range selection within the current hunk |
a / u / enter | Apply (stage/unstage) the selected lines - or the whole hunk if none selected |
s | Split the current hunk into smaller, independently stageable blocks |
] / [ | Jump to the next / previous hunk |
A | Stage all changes (git add -A, including untracked files and deletions) |
U | Unstage all changes (git reset -q HEAD) |
Select a range with space/v, navigate, then apply to stage just those lines - or split a mixed hunk with s and stage each block separately. The footer hint reflects whether a selection is active, and hides discard when a staged file is selected (discard only applies to unstaged changes).
Agent Modes
Toggle between modes anytime during chat using Shift+Tab.
| Mode | Tools | Approval | Best For |
|---|---|---|---|
| Standard (Default) | All configured | Required for Write/Edit/Delete/Bash | General development, collaborative coding |
| Plan (Read-Only) | Read, Grep, Tree only | None | Code reviews, architecture analysis, planning |
| Auto-Accept (YOLO) | All configured | None - immediate execution | Trusted environments, rapid prototyping, automation |
| Auto+Judge | All configured | An LLM judge answers every gate | Unattended CI/headless runs that still want a gate |
Standard Mode
Full tool access with safety controls and approval prompts for sensitive operations.
infer chat
> "Refactor the authentication module to use environment variables"
# Agent analyzes code, proposes changes, requests approval before modifyingPlan Mode
Analysis and planning without execution. Safe exploration of unfamiliar codebases.
infer chat
# Press Shift+Tab to switch to Plan Mode
> "How should I implement user authentication with JWT tokens?"
# Agent explores code structure and provides detailed planWhile planning, the agent can pause to ask you up to four multiple-choice clarifying questions with the AskUserQuestion tool, then fold your answers into the plan it submits for approval. The same tool is available in the other interactive modes, where the answers come back mid-task instead.
How mode instructions are delivered
Plan-mode instructions are not a system prompt. The system prompt at message[0] is byte-stable for the whole session - including across Shift+Tab mode switches - so the provider/local prompt (KV) cache keeps its prefix hits. The per-mode instructions ride along in the built-in mode-change-reminder (on_mode_change trigger) that fires when the mode changes, substituted into its {guidance} placeholder.
prompts.agent.mode_adjustment_plan and prompts.agent.mode_adjustment_auto (env: INFER_PROMPTS_AGENT_MODE_ADJUSTMENT_PLAN / _AUTO) let you override that guidance. They ship empty - the built-in texts live in the mode-change reminder's guidance map. Precedence per mode, highest first:
| Priority | Source |
|---|---|
| 1 (Highest) | guidance.<mode> in your reminders.yaml |
| 2 | prompts.agent.mode_adjustment_<mode> in prompts.yaml |
| 3 (Lowest) | Built-in default guidance |
Deprecated names.
prompts.agent.system_prompt_plan/system_prompt_auto(envINFER_PROMPTS_AGENT_SYSTEM_PROMPT_PLAN/_AUTO) still load into the new fields with a deprecation warning; when both are set the new names win.Trade-off: since the mode-change reminder is the sole carrier of mode-specific instructions, setting
enabled: falseinreminders.yaml(orINFER_REMINDERS_ENABLED=false) means those instructions are never delivered. Tool restrictions still apply - only the guidance text is lost.
Approving a plan
When the plan is ready, the agent calls RequestPlanApproval; the chat TUI renders the saved plan in a dedicated panel and shows a status line - use the arrow keys to select an option and Enter to confirm. Three options are offered:
| Option | Key | Resulting mode | Effect |
|---|---|---|---|
| Accept (default) | Enter / y | Auto-Accept | Executes the plan with no per-action approval prompts. |
| Approve Each Step | s | Standard | Executes the plan but prompts for approval on each Write/Edit/Delete/Bash action. |
| Reject | n | - | Ends the session; reply with feedback and the agent re-iterates the plan. |
The default Accept enables Auto-Accept mode - unrestricted execution with no per-action approval. Pick Approve Each Step to accept the plan but keep the Standard-mode approval gate on every action. The plan file stays on disk whichever you choose; rejecting a plan does not delete it.
Auto-Accept Mode
Zero approval prompts for maximum speed. Use with caution in version-controlled environments.
infer chat
# Press Shift+Tab twice to switch to Auto-Accept Mode
> "Run the test suite, fix all failing tests, and commit the changes"
# Agent executes everything immediatelyImportant for Auto-Accept: Ensure clean git working tree and backups.
Because the per-action approval gate is off in this mode, the agent runs under a dedicated destructive-action policy delivered by the mode-change reminder (see How mode instructions are delivered; override with prompts.agent.mode_adjustment_auto). It is told to stop and confirm before irreversible operations - deletes, git push --force, git reset --hard, dropping databases, rm -rf, publishing or releasing - to prefer the reversible path when no user is reachable, and never to print or publish a secret value. With reminders disabled this policy text is not delivered.
Auto+Judge Mode
Autonomous like Auto-Accept, but the calls that would prompt a human are decided by an LLM judge instead of being waved through. Press Shift+Tab once more from Auto-Accept, or select the mode explicitly in headless runs:
infer headless --mode auto-with-judge "fix issue #42"The judge is configured in its own judge.yaml file, and the same gate is available in any mode via tools.safety.approval_behaviour: judge. Both are off by default. See Judge Mode for the verdict contract, on_error semantics, and observability.
Headless Agent Stream Output
infer headless <task> runs the agent non-interactively and writes a newline-delimited JSON (JSONL) stream to stdout. Each line is one JSON object with a type discriminator, intended for programmatic consumers such as the infer-action GitHub Action. The stream is additive: new type values may be introduced over time and consumers should ignore any type they do not recognize.
infer headless "Refactor the authentication module"Secure by default. A headless run executes in standard mode, so off-list or mutating actions are not auto-run - they are blocked (when no approver is reachable) or sent for IPC approval (under a channel manager). See Headless secure-by-default to opt into more autonomy.
Output format (--format)
The output format is controlled by the --format flag (renamed from the legacy --output-format). It accepts four values:
json(default) - newline-delimited JSON (JSONL), one compact object per line, suitable for programmatic consumptionjson-pretty- same per-turn stream asjsonwith each object indented across multiple lines for human readingag-ui- spec-compliant AG UI framed output withTEXT_MESSAGE_START/CONTENT/ENDframing and a fresh message id per turn. Atoken_usagecustom event follows every LLM step. The terminalRUN_FINISHEDevent carries the session totals in itsresulttext- human-readable plain text output
# Default JSON output
infer headless "Refactor the authentication module"
# Human-readable pretty-printed JSON
infer headless "Analyze the codebase" --format json-prettyMachine consumers should keep using
jsonfor performance;json-prettyis for debugging and human inspection.
Stream event types
Every line in the json (or json-pretty) stream is a JSON object with a type discriminator. The stream is additive - new types may be introduced and consumers should ignore any they do not recognize.
| Type | When emitted | Description |
|---|---|---|
info | Once at start | Run metadata (model, session id) |
assistant | Per turn | Assistant response with content, reasoning_content, tool_calls, token_usage |
tool | Per tool call | Tool execution result. content holds the bare marshaled result (no legacy Result of tool call: / Tool execution failed: envelope). Structured tool_execution metadata (tool_name, success, error, rejected, duration) and the failed flag ride alongside. |
approval_request | When approval is needed | Approval request metadata |
judge_verdict | Per judge decision | Emitted under judge mode or approval_behaviour: judge - tool, decision, reason, model, turn. Mirrored as a custom event in ag-ui. |
computer_use_paused | On a pause control message | Computer use was paused; carries the session request_id - see pause and resume control |
computer_use_resumed | On a resume control message | Computer use was resumed; carries the session request_id |
session_stats | Once at end | Token usage and cost summary (detailed below) |
agent_error | Before stream starts | Machine-readable error when the run fails before any turn (gateway unavailable, unknown model). For ag-ui format this is emitted as RUN_ERROR. |
One assistant line per LLM turn
Each completed LLM turn emits exactly one assistant line, in turn order - an empty turn (no content, no reasoning, no tool calls) emits none. That holds for a text-only turn the agent continued past because a post_stream continuation nudge fired (todo continuation, truncation continuation, empty-response continuation): the nudged turn keeps its own assistant line instead of being merged into the next turn's line. Consumers that walk the stream by turn can therefore treat assistant lines as the sequence of turns that produced output, and read the final answer from the content of the last assistant line.
The hidden system reminder a continuation nudge appends is not emitted to the stream. It exists only in the model-facing message history, so it never shows up as an assistant or tool line. Reminder text visible in a trace or under logging.debug is log output, not a stream line.
In the ag-ui format the same contract holds one level up: one TEXT_MESSAGE_START / CONTENT / END triple with a fresh message id per LLM turn, nudge-continued turns included.
Exit codes:
| Code | Meaning |
|---|---|
| 0 | Task completed successfully |
| 1 | Task failed |
| 2 | Max turns exhausted |
Headless behavior:
- Subagent mode: Headless runs always spawn subagents in
headlessmode, never tmux/interactive, regardless oftools.agent.mode. - Background tasks: The run waits for in-flight background tasks (background shells, A2A tasks) before completing, bounded by
a2a.task.agent_mode_max_wait_seconds. - Session rollover: Long conversations automatically roll over into a new session id (session compact). A non-UUID
--session-idis treated as a session-group key that follows these rollovers - useful for channel-based workflows.
Pause and resume control (IPC)
A host UI (the desktop app, or any process that owns the infer headless subprocess) can pause and resume an in-flight computer-use run by writing a control line to the agent's stdin, the same IPC pattern as approval_response:
{ "type": "computer_use_control", "action": "pause" }
{ "type": "computer_use_control", "action": "resume" }- Works with the
json,json-pretty, andag-uiformats, with or without--require-approval, and regardless of thecomputer_use.approvallevel. - Pause cancels the in-flight request and emits
computer_use_paused. - Resume restarts the run over the same conversation with a hidden
Please continue from where you left off.message, and emitscomputer_use_resumed. The resume also clears the error the cancelled run carried, so a paused run that is resumed still finishes cleanly. - A malformed line, an unknown
type, or an unrecognizedactionis ignored (logged as a warning), so the stdin stream can carryapproval_responseandcomputer_use_controlmessages interleaved.
Events out. In the json / json-pretty stream both events are plain JSON lines carrying the session request_id:
{ "type": "computer_use_paused", "request_id": "b7a1c3d2-..." }
{ "type": "computer_use_resumed", "request_id": "b7a1c3d2-..." }In the ag-ui stream they arrive as CUSTOM events with the same names:
{ "type": "CUSTOM", "name": "computer_use_paused", "value": { "request_id": "b7a1c3d2-..." } }
{ "type": "CUSTOM", "name": "computer_use_resumed", "value": { "request_id": "b7a1c3d2-..." } }Example. Pause the run, then resume it a few seconds later:
{
sleep 5
echo '{"type":"computer_use_control","action":"pause"}'
sleep 5
echo '{"type":"computer_use_control","action":"resume"}'
} | infer headless --format json "Open the settings window and enable dark mode"{"type":"info","model":"...","session_id":"b7a1c3d2-..."}
{"type":"assistant","content":"Taking a screenshot to find the settings window..."}
{"type":"computer_use_paused","request_id":"b7a1c3d2-..."}
{"type":"computer_use_resumed","request_id":"b7a1c3d2-..."}
{"type":"assistant","content":"Continuing - clicking the appearance tab..."}
{"type":"session_stats","message":"Session complete","...":"..."}Session stats summary line
When a session completes, the CLI emits a single session_stats line summarizing token usage and computed dollar cost for the run. This lets consumers report real run cost without re-implementing the per-model pricing table.
{
"type": "session_stats",
"message": "Session complete",
"timestamp": "2026-05-29T17:48:55+02:00",
"model": "deepseek/deepseek-v4-flash",
"prompt_tokens": 21000,
"completion_tokens": 1260,
"total_tokens": 22260,
"requests": 7,
"cost": { "input": 0.0021, "output": 0.0008, "total": 0.0029, "currency": "USD" }
}Fields:
| Field | Type | Description |
|---|---|---|
type | string | Always session_stats for this line. |
message | string | Human-readable status, currently Session complete. |
timestamp | string | RFC 3339 timestamp at which the line was emitted. |
model | string | Model used for the run. A single model is attributed per run. |
prompt_tokens | number | Sum of input tokens across all requests in the run. |
completion_tokens | number | Sum of output tokens across all requests in the run. |
total_tokens | number | prompt_tokens + completion_tokens. |
requests | number | Number of LLM requests (turns that reported usage) in the run. |
cost | object | Computed dollar cost for the run - see Cost object. |
Cost object
| Field | Type | Description |
|---|---|---|
input | number | Cost attributed to prompt_tokens using the configured pricing table. |
output | number | Cost attributed to completion_tokens using the configured pricing table. |
total | number | input + output. |
currency | string | ISO 4217 currency code from pricing.currency. Defaults to USD. |
When pricing.enabled: false (or pricing data is unavailable for the model), input, output, and total are all 0 while currency is still populated. The cost object is always present, giving consumers a stable schema.
Behavior notes:
- The line is additive - it does not replace any existing stream output.
- It is emitted once per run, at session completion (including on early errors).
- It is always emitted in
agentmode - there is no flag to enable or disable it. - Cost is attributed to a single model per run.
- Consumers should ignore unknown
typevalues to remain forward-compatible.
AG-UI RUN_FINISHED result
In the ag-ui format the same totals ride on the terminal RUN_FINISHED event instead of a session_stats line. result carries the per-session totals, cumulative across the session (not just the current run):
{
"type": "RUN_FINISHED",
"threadId": "b7a1c3d2-...",
"runId": "9f4e1a0b-...",
"result": {
"inputTokens": 21000,
"outputTokens": 1260,
"cacheReadTokens": 18400,
"totalToolCalls": 7,
"cost": 0.0029,
"lastInputTokens": 12800,
"contextWindow": 128000
}
}| Key | Type | Description |
|---|---|---|
inputTokens | number | Total input tokens across the session. |
outputTokens | number | Total output tokens across the session. |
cacheReadTokens | number | Tokens served from the prompt cache. |
totalToolCalls | number | Tool calls issued across the session. |
cost | number | Total session cost, in the configured currency. |
lastInputTokens | number | Input tokens of the most recent request. |
contextWindow | number | Model context window in tokens, omitted when the model's window is unknown. |
The event has no result when the run made no model request, so consumers must treat it as optional.
The same stats object is also streamed as a CUSTOM event named token_usage after every LLM step, before tool execution continues, so consumers can track cumulative usage while the run progresses. Its value keys match the RUN_FINISHED result above, including contextWindow being omitted when the model's window is unknown, and the event is not emitted until the session has made at least one model request:
{
"type": "CUSTOM",
"name": "token_usage",
"value": {
"inputTokens": 21000,
"outputTokens": 1260,
"cacheReadTokens": 18400,
"totalToolCalls": 7,
"cost": 0.0029,
"lastInputTokens": 12800,
"contextWindow": 128000
}
}For AG-UI consumers (the desktop app sidecar, or any process hosting infer headless --format ag-ui): the per-step token_usage events stream the cumulative stats above, so live indicators stay current during the run instead of jumping once at RUN_FINISHED. Drive the context-percentage indicator from lastInputTokens / contextWindow - lastInputTokens is the live occupancy of the window, where inputTokens is a session-wide sum and will overshoot it. Hide the indicator when contextWindow is absent. Drive the cost indicator from cost, which is already the computed dollar total and needs no per-model pricing table on the client.
Writing the result to a file (--result-file)
infer headless accepts a --result-file <path> flag that atomically writes the final assistant message and the run outcome as JSON to <path> on exit. The Agent tool uses it to harvest the result of a detached (tmux pane) subagent, but it is useful on its own whenever a script needs the final answer as a file rather than by parsing the stdout stream.
infer headless "Summarize the open PRs" --result-file /tmp/result.jsonResuming a session (--session-id)
infer headless --session-id <id> loads a persisted conversation before the agent starts, letting you resume a previous session. The session must have been persisted by a storage backend (storage.enabled: true).
# Find session IDs from saved conversations
infer conversations list
# Resume a specific session
infer headless "Continue the refactoring" --session-id abc-123-defSession ID resolution:
- A literal UUID is used as-is.
- Any non-UUID value is treated as a session group key and resolved through the session rollover chain to the latest session in that group.
Fallback: If the session cannot be loaded (e.g. the ID does not exist or storage is disabled), the CLI prints a visible notice and starts a new session under the requested ID.
The same flag is also available on infer chat --session-id for resuming chat sessions interactively.
Computer Use
GUI automation and visual understanding capabilities for interacting with applications and desktop environments.
Display Server Support
Automatic display server detection - no configuration needed:
| Platform | Supported Servers | Notes |
|---|---|---|
| macOS | Quartz (native), X11 (XQuartz) | Quartz automatically detected and used |
| Linux | X11, Wayland | Auto-detection handles both protocols |
| Windows | Native | No configuration required |
Display server type is automatically detected at runtime. No manual configuration required.
Computer Use Tools
Computer use exposes two tools: Computer, an action-based desktop tool, and GetLatestFrame, registered whenever a frame source exists (screenshot streaming, or a directory source under vision.sources).
Computer takes an action:
| Action | Description |
|---|---|
accessibility | Preferred first observation. Returns compact {role,label,state,bbox} elements from the accessibility tree. Read-only. |
press | Presses the first element whose label matches exactly, via its accessibility action - the cursor never moves. |
screenshot | Captures the screen, or a native-resolution region. Use when the accessibility tree is empty or insufficient. |
cursor | Reports the current pointer position. |
move, click, double_click, triple_click, scroll | Pointer control. |
type, key | Type text or send key combinations (Ctrl+C, Cmd+V). |
Targets. accessibility and press accept an optional target: frontmost (default), dock, menubar, pid:<N>, app:<name>, or a bare application name. Application identifiers are pid:<N> on every platform - the macOS bundle ID form is gone, as are the older GetFocusedApp and ActivateApp tools.
Coordinate space. Accessibility bounding boxes use the same frame coordinate space as screenshots and pointer actions, so an element's center can be passed straight to a click.
Pressing beats clicking. press is the reliable way to activate dock items, buttons, and menu titles: it drives the element's own accessibility action instead of aiming the pointer at a coordinate, and it takes no screenshot.
macOS Accessibility permission. The AX tools return elements only after infer is granted permission in System Settings > Privacy & Security > Accessibility. Without it - or on a helper crash, timeout, or unavailable tree - the tool returns screenshot fallback guidance to the agent rather than failing the run. The macOS bridge is pure Go (PureGo calling CoreFoundation, CoreGraphics, and AXUIElement) running in a short-lived helper process; there is no cgo, Swift, or Objective-C.
Linux and Windows. AT-SPI and UIA providers are not implemented yet: accessibility and press report unsupported; use screenshot there, while every other Computer action works normally.
Clipboard. Clipboard support is text-only; image clipboard was removed with the cgo-free rewrite.
Screen Recording
Two tools record the screen to an MP4 file (H.264, yuv420p - it plays in browsers and QuickTime):
RecordStart- begins a recording of the whole primary screen, a single window, or a region, and returns the output path and the captured rectangle. Recording continues in the background.RecordStop- finalizes the recording and returns the path, duration and file size.
Typical uses are an unattended audit trail of a computer-use run, and recording a tutorial or a bug reproduction while the agent drives the desktop.
Disabled by default. While
computer_use.recording.enabledisfalse, neither tool is registered, so they cost zero prompt tokens. Recording captures whatever is on your screen - that is why it is opt-in.
Recording lives under recording in .infer/computer_use.yaml (project) or ~/.infer/computer_use.yaml (user). It works on its own - computer_use.enabled is not required.
# .infer/computer_use.yaml
recording:
enabled: true # register the RecordStart/RecordStop tools (default: false)
max_duration: 120 # seconds; the recording stops and finalizes itself at this cap
output_dir: '' # empty = ~/.infer/tmp/recordings
framerate: 24 # frames per secondEvery key has an INFER_COMPUTER_USE_RECORDING_-prefixed environment variable that takes precedence over the YAML value.
| Config key | Environment variable | Type | Default | Notes |
|---|---|---|---|---|
computer_use.recording.enabled | INFER_COMPUTER_USE_RECORDING_ENABLED | bool | false | Feature flag - both tools are absent from the LLM payload when false |
computer_use.recording.max_duration | INFER_COMPUTER_USE_RECORDING_MAX_DURATION | int | 120 | Seconds; the recording finalizes itself at the cap. Must be positive |
computer_use.recording.output_dir | INFER_COMPUTER_USE_RECORDING_OUTPUT_DIR | string | ~/.infer/tmp/recordings | Where MP4s are written, as <timestamp>.mp4; created on first use |
computer_use.recording.framerate | INFER_COMPUTER_USE_RECORDING_FRAMERATE | int | 24 | Capture frame rate. Must be positive |
Requirements. The recorder shells out to ffmpeg with libx264 and the platform's screen grabber, found on PATH - for example brew install ffmpeg (macOS) or apt install ffmpeg (Debian/Ubuntu). A minimal ffmpeg build without the encoder or the grabber cannot record.
| Platform | Grabber | Extra requirements |
|---|---|---|
| macOS | avfoundation | Your terminal app needs Screen Recording permission, plus Accessibility for window mode (System Settings > Privacy & Security) |
| Linux | x11grab | An X11 session. Wayland is not supported yet |
| Windows | gdigrab | None |
Capture modes. RecordStart takes an optional mode:
screen(default) - the entire primary display.window- a single window, selected bywindow:frontmost(default),app:<name>,pid:<number>, or a bare application name. This is the same target syntax as theComputertool.region- a rectangle, given asregionwithx,y,widthandheightin the frame coordinate space (the same space asComputerscreenshots and accessibility bounding boxes).
{ "mode": "screen" }
{ "mode": "window", "window": "app:Safari" }
{ "mode": "region", "region": { "x": 0, "y": 80, "width": 1280, "height": 720 } }window mode captures the window's bounds at the moment the recording starts: anything drawn over that area is recorded too, and moving the window afterwards does not move the capture.
Behaviour. RecordStart returns the file path and the captured rectangle and leaves ffmpeg running in the background. RecordStop takes no arguments and returns the path, duration and size.
- One recording at a time, machine wide. While ffmpeg runs, the recording holds
~/.infer/run/screen-recording.lock. ARecordStartfrom any otherinferprocess - another chat, a Desktop app session, a channel run - fails with an error naming the owning pid, and the active recording is left alone.RecordStopwith nothing recording is an error too. - Self-finalizing. A recording stops and finalizes itself at
max_duration(120s by default);RecordStopstill returns that file. One that stops on its own - the cap, an ffmpeg exit, or a stop from/tasks- queues a note asking the agent to collect it withRecordStop. - Clean exit. On normal exit, Ctrl+C or SIGTERM the CLI finalizes an active recording and leaves no ffmpeg process behind. Call
RecordStopin the same session; channel and scheduled runs each start a fresh process. - Visibility. The chat status bar shows a
RECbadge while a recording runs, and the recording is listed in/tasksunder Screen Recordings as a background job of kindrecording. infer tools executerefuses both tools.
Approval. RecordStart always requires approval, except in auto-accept (auto) mode:
- In chat you get the usual approval prompt.
- In headless mode it follows
approval_behaviour: sent over IPC when the run has--require-approval, decided by the LLM judge inauto-with-judgemode, otherwise blocked.
RecordStop follows computer_use.approval: under never (the default) and destructive it runs without a prompt - it counts as an observation - and only always asks.
Headless and AG-UI hosts. A running recording is a background job, so infer headless waits for it (up to a2a.task.agent_mode_max_wait_seconds, 300s by default) instead of exiting and cutting it short. How the stop request reaches the recording depends on the caller:
- Desktop app and other stdin hosts send a
user_messageline on stdin, which reachesRecordStopin the same process. - Channel, scheduled and heartbeat runs do not forward follow-up messages into a running process, so there a recording runs until
max_durationand the agent collects the file afterwards.
With --format ag-ui (the format the desktop app consumes):
| Event | Payload |
|---|---|
CUSTOM event screen_recording - on RecordStart, RecordStop, or the max_duration cap | { "active": true } / { "active": false } |
CUSTOM event background_tasks | The recording appears among jobs with kind recording |
approval_request | RecordStart when the run has --require-approval; without an approver it is blocked |
A recording still running when the run ends is finalized after the terminal event, so no screen_recording event with active: false follows it - clear the indicator on RUN_FINISHED or RUN_ERROR.
Screenshot Tool Features
Streaming Mode:
- Maintains circular buffer of recent screenshots
- Configurable buffer size (default: 60)
- Configurable capture interval (default: 3 seconds)
- Efficient memory management
- Fast access to recent captures
Image Optimization:
- Automatic resolution scaling to fit the target box (default: 1024x768), preserving the aspect ratio so coordinates map back to the screen
- JPEG compression with configurable quality (default: 85%)
- Reduces bandwidth and storage requirements
Region Selection:
- Full screen capture
- Custom region coordinates (x, y, width, height)
- Multiple monitor support
Visualizing Agent Activity
The desktop app is the visualization layer for computer use: it shows the live screen monitor, the on-screen action overlay, and the approval prompts driven by computer_use.approval. Run the CLI from the desktop app to watch a computer-use run as it happens.
The CLI's own macOS floating window was removed. The computer_use.floating_window config section and the INFER_COMPUTER_USE_FLOATING_WINDOW_* environment variables no longer exist - existing config files that still carry those keys keep loading, the keys are simply ignored.
Computer Use Configuration
Computer use has its own file, .infer/computer_use.yaml, seeded by infer init. It is off by default; missing keys fall back to the in-code defaults shown here.
# .infer/computer_use.yaml
---
enabled: true
approval: never # never | destructive | always
screenshot:
enabled: true
target_width: 1024
target_height: 768
format: jpeg
quality: 85
streaming_enabled: true # also registers the GetLatestFrame tool
capture_interval: 3
buffer_size: 60
temp_dir: ''
rate_limit:
enabled: true
max_actions_per_minute: 60
window_seconds: 60Computer Use Approval
computer_use.approval decides which computer-use actions go through the approval gate before they run. It applies to both interactive chat and infer headless.
| Value | Behavior |
|---|---|
never | Default. Computer-use actions bypass the approval gate and run immediately. |
destructive | The observations accessibility, screenshot, and cursor bypass approval; press and the input actions require it. |
always | Every computer-use action requires approval. |
# .infer/computer_use.yaml
approval: destructiveexport INFER_COMPUTER_USE_APPROVAL=always- Env override:
INFER_COMPUTER_USE_APPROVALtakes precedence over the YAML value. - Fail closed: an unknown value is rejected at config load with a validation error, and if one ever reaches the approval policy it is treated as
alwaysrather than silently bypassing the gate. - Under
destructiveoralways, a headless run needs an approver reachable over IPC - otherwise the gated calls are blocked, same as any other approval-requiring tool. See Headless secure-by-default.
Safety and Rate Limiting
Rate Limiting:
- Default: 60 actions per minute
- Prevents runaway automation
- Configurable threshold
Safety Controls:
- Approval prompts in Standard Mode
- Auto-approve in YOLO mode
- Activity logging for audit trails
- Command execution monitoring
Best Practices:
- Use Standard Mode for initial exploration
- Enable logging for debugging
- Set appropriate rate limits
- Monitor activity logs
- Test in safe environments first
Example Use Cases
infer chat
> "Take a screenshot and analyze the error dialog"
> "Click the Submit button in the center of the screen"
> "Type 'Hello World' and press Enter"
> "Switch to the Terminal app and run ls command"
> "Find the Save button and click it"Tools & Capabilities
When tools are enabled, LLMs have access to a comprehensive suite across multiple categories.
Tool Categories
| Category | Tools | Description |
|---|---|---|
| File System | Read, Write, Edit, MultiEdit, Delete, Tree, Grep | File operations and search with safety controls |
| Command Execution | Bash, BashOutput, KillShell, ListShells, Wait | Allow-listed shell execution (including gh for GitHub), background shell control, and blocking wait for conditions |
| Web | WebSearch, WebFetch | Internet research and content fetching |
| Workflow | TodoWrite, Schedule, RequestPlanApproval, AskUserQuestion, RequestApproval, Memory | Task tracking, cron jobs, plan-mode approval, clarifying questions, judge-rejection escalation, and persistent cross-session memory |
| A2A Integration | A2A_QueryAgent, A2A_SubmitTask, A2A_QueryTask | Delegate to external specialized agents - see A2A |
| Local Subagents | Agent | Fan out short-lived local subagents in parallel - see Local Subagents |
| Computer Use | Computer, GetLatestFrame, RecordStart, RecordStop | Accessibility-tree reads, presses, screenshots, and pointer/keyboard control - see the Computer Use section above; screen recording is opt-in via computer_use.recording.enabled, see Screen Recording |
| Image | ImageGeneration, ImageEdit, ImageVariation | Generate, edit, and vary images using the configured image model - independent of the chat session model |
| Audio | TextToSpeech, TextToMusic, TextToSFX | Local speech synthesis and voice cloning - opt-in via text_to_speech.enabled, see Text-to-Speech; music composition and sound-effect generation through the gateway - opt-in via text_to_music.enabled and text_to_sfx.enabled |
| Video | TextToVideo, CreateAvatar | Prompt and lip-synced avatar renders through the gateway, plus avatar-library creation - opt-in via text_to_video.enabled and text_to_video.create_avatar, see Text-to-Video and Avatars |
| MCP | MCP_<server>_<tool> | Dynamically registered tools from MCP servers - see MCP |
File System Tools
Read
Read a file from the local filesystem with an optional line range. Handles text files and PDFs.
- Parameters:
file_path(required, absolute or relative),limit(default 2000 lines),offset(default 1) - Approval: not required (read-only)
- Notes: lines longer than 2000 characters are truncated; output is returned in
cat -nformat
Write
Write content to a file on disk. Overwrites the existing file at the given path.
- Parameters:
file_path(required, absolute),content(required) - Approval: required by default
- Notes: if the file exists, the Read tool must have been used first; respects configured path exclusions (
.git/,*.env,.infer/)
Edit
Perform an exact string replacement in a single file.
- Parameters:
file_path(required),old_string(required - must match exactly and be unique unlessreplace_allis set),new_string(required - must differ fromold_string),replace_all(defaultfalse) - Approval: required by default
- Notes: the file must have been Read at least once in the conversation; indentation must be preserved exactly
MultiEdit
Apply a sequence of edits to a single file atomically - either all succeed or none are applied.
- Parameters:
file_path(required),edits(required array; each item hasold_string,new_string, optionalreplace_all) - Approval: required by default
- Notes: edits are applied in order, each operating on the result of the previous one - plan them so earlier edits don't invalidate later matches
Delete
Delete a file or directory. Wildcards are supported when enabled.
- Parameters:
path(required - supports patterns like*.txtortemp/*),recursive(defaultfalse),force(defaultfalse),format(textorjson) - Approval: required by default
- Notes: restricted to the current working directory for safety
Tree
Display a directory tree, similar to the Unix tree command.
- Parameters:
path(default.),max_depth(1-10, default 3),max_files(1-1000, default 100),respect_gitignore(defaulttrue),show_hidden(defaultfalse),format(textorjson) - Approval: not required
- Notes: uses the system
treebinary when available, otherwise falls back to a built-in implementation
Grep
Powerful regex search across files. Uses ripgrep when available, otherwise a built-in Go implementation.
- Parameters:
pattern(required regex),path(default cwd),glob(e.g.*.ts,**/*.tsx),type(e.g.go,py,rust),output_mode(content|files_with_matches|count, defaultfiles_with_matches),-i,-n,-A,-B,-C,multiline,head_limit - Approval: not required
- Backend: configurable via
tools.grep.backend(auto|ripgrep|go) - Notes: respects
.gitignore; auto-excludes.git,node_modules,.infer,vendor,dist,build,target
Command Execution
Bash
Execute a bash command that matches the active mode's allowed-list. Matching is default-deny: a command auto-runs only when it matches the allowed-list for the current agent mode. Anything unmatched falls through to an approval prompt in chat, or is rejected with an actionable hint in headless mode. There is no separate deny list.
- Parameters:
command(required),format(textorjson) - Approval: configurable via
tools.bash.require_approval
Per-mode allowed-list
The allowed-list is configured per agent mode under tools.bash.mode.<mode>.allow. The effective list for a mode is mode.all.allow (the every-mode baseline) unioned with that mode's own entries:
tools:
bash:
enabled: true
require_approval: false
mode:
all: # baseline applied in every mode
allow:
- ls( .*)?
- pwd( .*)?
- git status( .*)?
- git diff( .*)?
plan: # read-only analysis - usually adds nothing
allow: []
standard: # default interactive mode
allow:
- npm (install|test|run).*
auto: # Auto-Accept / YOLO mode
allow:
- .* # unrestricted sentinel- Default-deny. Out of the box only
mode.autoships the.*sentinel;mode.planandmode.standardadd nothing on top ofmode.all, so they reduce to the read-only baseline. GitHub writes (gh issue/pr create|edit|comment),git push, andgit commitare not in the defaults - they fall through to approval until you add them. - Full-command matching. Each entry matches the whole command, so a bare token like
ghallows onlygh- nevergh issue list. Opt into arguments explicitly with a pattern (gh issue.*,npm (install|test|run).*); the default entries use a( .*)?suffix to allow trailing arguments. - The
.*sentinel means unrestricted: any single command runs and the clean-command guard below is skipped. It is the default formode.auto(chat's Auto-Accept mode, toggled with Shift+Tab) and is an explicit opt-in - never a headless default.
Clean-command guard
For every mode except the .* sentinel, each command passes a clean-command guard before the allowed-list is consulted. The guard rejects, regardless of the list:
- Command substitution -
$(...), backticks,<(...),>(...). - Multi-command chains and pipelines - a top-level
|,&&,||,;,&, or newline. Operators inside quotes don't count, sojq '.a | .b'stays a single command. (This closes the oldecho x | xargs rmprefix hole.) - File-write redirections -
>and>>. Benign stream redirects (2>&1,>/dev/null) are stripped first and remain allowed. - Dangerous
findactions --exec,-delete, and the like. A barefindfor read-only discovery is fine. - Environment-variable leaks - a printing or publishing command (
echo,printf,gh issue|pr create|comment|edit) may not expand$VAR. Soecho $AWS_SECRET_ACCESS_KEYis blocked, whilels $DIRstays allowed. A single-quoted or escaped$is treated literally.
A rejected command returns an actionable hint naming what tripped the guard; the model is told to stop and ask, or use an allowed alternative, rather than retry the same call.
Append-only override (CI)
The mode.all baseline takes an append-only override so CI can add a few safe commands without rewriting config or shipping .*:
# Comma- or newline-separated; the env var wins over the flag
export INFER_TOOLS_BASH_ALLOW_APPEND="git commit,git push"
# Flag form
infer headless "Release the changelog" --tools-bash-allow-append "git commit,git push"The extra commands merge onto mode.all.allow, so they auto-run in every mode. There is no replace override - the old tools.bash.whitelist.commands key, the INFER_TOOLS_BASH_WHITELIST_COMMANDS[_APPEND] env vars, and the --tools-bash-whitelist-commands* flags were removed.
BashOutput, KillShell, ListShells
Background-shell management. These tools are only registered when tools.bash.background_shells.enabled: true.
- BashOutput -
bash_id(required),filter(optional regex). Returns only new output since the last read. - KillShell -
shell_id(required). Sends SIGTERM, then SIGKILL after 5 seconds if the shell doesn't exit. - ListShells - no parameters. Lists all running and recently completed background shells with their IDs, state, and elapsed time.
Wait
Block inside a single tool execution until a condition is met (shells exit, file event, or check command succeeds), then return once with the outcome. Waiting costs zero chat completions - no LLM round-trip per iteration.
- Approval: not required (passive utility tool - no side effects)
- Enabled by default: yes (gated by
tools.wait.enabled) - Common parameters:
timeout_seconds(required, number) - maximum time to wait in seconds, bounded by the config ceiling (tools.wait.max_timeout_seconds, default 600)
Return value: A structured result with condition (the condition type), reason (outcome: condition_met, timeout, cancelled, check_failed, no_shells, error, not_allowed), elapsed_seconds (time spent waiting), and condition-specific details (exit codes, last output, shell states) - included even on failure so the model can see why the wait ended.
Cancellation: Pressing Esc in chat or session cancel interrupts the wait immediately (reason: cancelled).
Condition: Shells (condition=shells)
Block until the given background shell ID(s) exit. When shell_ids is omitted, waits for all currently running background shells.
- Parameters:
shell_ids(optional, array of strings) - specific shell IDs to wait for. Omit to wait for all pending background shells. - Returns: Exit codes and tail output (last 4096 bytes) for each shell.
Use cases:
- Wait for a long-running build or test to finish
- Wait for multiple parallel background tasks to complete
- Coordinate sequential steps that depend on background work
Condition: File (condition=file)
Block until a file path is created, modified, or removed (uses fsnotify for efficient inotify/FSEvents-based watching).
- Parameters:
path(required, string) - file path to watch.event(optional, string, enum:create,modify,remove,any) - file event to wait for. Default:any. - Behavior: For
create: checks if the file already exists first (returns immediately if so). Watches the parent directory for the target filename. Uses OS-native file system notifications (no polling).
Use cases:
- Wait for a download to complete
- Wait for a log file to be created
- Wait for a lock file to be removed
- Wait for a build artifact to appear
Condition: Command (condition=command)
Re-run a check command server-side at a fixed interval until it exits 0. The check command goes through the same bash allow-list as the Bash tool.
- Parameters:
command(required, string) - check command to re-run until it exits 0.pending_exit_codes(optional, array of numbers) - non-zero exit codes that mean "still pending, keep polling". Any other non-zero exit ends the wait immediately with reasoncheck_failed. - Behavior: First run happens immediately (no initial delay). Subsequent runs at
command_poll_interval_msinterval (default 2s). Each run has a 30-second per-execution timeout. Not available on Windows (requires bash). - Exit code classification:
- Exit 0:
condition_met- success - Exit in
pending_exit_codes: keep polling - Exit non-zero (not in pending):
check_failed- ends immediately - No pending_exit_codes specified: keep polling on any non-zero exit
- Exit 0:
Use cases:
- Wait for a service to become healthy:
curl -sf localhost:8080/health - Wait for CI to complete:
gh pr checkswithpending_exit_codes=[8] - Wait for a file to contain specific content:
grep -q 'ready' /var/log/app.log - Wait for a port to open:
nc -z localhost 3000
Configuration
# .infer/config.yaml
tools:
wait:
enabled: true # Enable/disable the Wait tool
max_timeout_seconds: 600 # Maximum allowed timeout (ceiling)
command_poll_interval_ms: 2000 # Poll interval for command condition (ms)Examples
# Wait for background build to finish
Wait(condition=shells, shell_ids=["build-1"], timeout_seconds=300)
# Wait for a file to be created
Wait(condition=file, path="/tmp/result.json", event="create", timeout_seconds=60)
# Wait for service health
Wait(condition=command, command="curl -sf http://localhost:8080/health", timeout_seconds=120)
# Wait for CI checks with pending exit codes
Wait(condition=command, command="gh pr checks 792 --repo inference-gateway/cli", pending_exit_codes=[8], timeout_seconds=600)
# Wait for all background shells
Wait(condition=shells, timeout_seconds=300)Web Tools
WebSearch
Search the web via DuckDuckGo or Google.
- Parameters:
query(required),engine(duckduckgo|google, defaults to the configured engine),limit(1-50, defaults to configuredmax_results),format(textorjson)
tools:
web_search:
enabled: true
default_engine: duckduckgo
max_results: 10
engines: [duckduckgo, google]
timeout: 10WebFetch
Fetch content from an allowed URL. Optionally save the response to disk.
- Parameters:
url(required),format(textorjson),download(defaultfalse- whentrue, saves under.infer/artifacts/<session-id>/) - Notes: only allowed domains can be fetched; responses are cached (default 15-minute TTL)
tools:
web_fetch:
enabled: true
allowed_domains:
- golang.org
- github.com
- agents.md
safety:
max_size: 8192
timeout: 30
cache:
enabled: true
ttl: 3600Image Tools
ImageGeneration, ImageEdit, and ImageVariation all write their PNG to the session artifacts directory - .infer/artifacts/<session-id>/image-<timestamp>.png - so the images produced by a conversation stay grouped together. See Artifacts directory.
ImageEdit Tool
Edit an existing image and save the result as a PNG under .infer/artifacts/<session-id>/. The chat model calls the tool when the user asks to edit an image; the tool reads the input image from a local file path and sends a plain one-off request to /v1/images/edits using the configured image model - no system prompt, no tools, independent of the model selected for the chat session.
Parameters:
image(required): Local file path of the image to editprompt(required): Text description of the desired editmask(optional): Local file path to a PNG whose fully transparent areas (alpha = 0) mark the editable region; all other pixels are preserved exactly. Must be a PNG with the same dimensions as the input image. Non-PNG paths are rejected by tool validation. Omit to let the model localize the change from the prompt alone - useful for models that repaint areas the user did not ask to change, and required for dall-e-2 edits without transparency.quality(optional):auto(default),low,medium,high, orstandardsize(optional):1024x1024(default),1536x1024, or1024x1536
Configuration:
tools:
image_edit:
enabled: true
model: openai/gpt-image-2
require_approval: falseExample with mask:
{
"image": "photo.png",
"prompt": "replace the sky with a sunset",
"mask": "sky-mask.png"
}ImageVariation Tool
Create a variation of an existing image and save the result as a PNG under .infer/artifacts/<session-id>/. The chat model calls the tool when the user asks for a variation; the tool reads the input image from a local file path and sends a plain one-off request to /v1/images/edits using the configured image model - no system prompt, no tools, independent of the model selected for the chat session.
Parameters:
image(required): Local file path of the image to base the variation onsize(optional):1024x1024(default),1536x1024, or1024x1536
Configuration:
tools:
image_variation:
enabled: true
model: openai/gpt-image-2
require_approval: falseVision-capable models and ImageDecode
Vision-capable models (those with the vision label in the model picker) see pasted or @-referenced images natively - the image data is sent directly to the model, which can inspect it without a separate tool call. For these models, the ImageDecode tool is hidden from the advertised tool list to avoid steering the model toward a text annotation when it can see the image directly.
The tool remains executable if called - a model that invokes it from conversation history or right after a model switch gets a working tool, not an error. Text-only models keep the tool and the "use ImageDecode to inspect it" note unchanged.
Audio Tools
TextToSpeech Tool
Synthesize speech from text with a local TTS engine and save it as a WAV file. The chat model calls the tool when the user asks to say something aloud or to clone a voice; synthesis shells out to llama.cpp's llama-tts binary running Qwen3-TTS GGUF models, fully offline. Disabled by default - while text_to_speech.enabled is false, the tool definition is not sent to the LLM at all.
Parameters:
text(required): The text to speakvoice_sample(optional): Bare file name of a WAV of the target speaker (~10-30s of clean speech) to clone, resolved against the working directory first and then the voice samples library at~/.infer/models/tts/samples/; a name found in neither fails with an error listing both paths triedoutput_path(optional): Destination WAV; defaults to a timestamped file undertext_to_speech.output_dir(~/.infer/tmp/tts/)
Configuration:
text_to_speech:
enabled: trueSee Text-to-Speech for prerequisites, the full configuration reference, model presets, and voice-cloning guidance.
TextToMusic Tool
Compose a music clip from a text prompt and save it as an MP3 file. The chat model calls the tool when the user asks for background music, a loop, or a jingle. Generation goes through the gateway's Music API (POST /v1/audio/music) using the configured provider/model - the CLI holds no provider key, the gateway does, so requests show up in gateway logs and traces. Disabled by default - while text_to_music.enabled is false, the tool definition is not sent to the LLM at all.
Parameters:
prompt(required): Description of the music - genre, mood, instruments, temposeconds(optional): Clip length in seconds; omitted lets the provider pick a length that fits the promptinstrumental(optional):trueto guarantee the clip has no vocalsoutput_path(optional): Bare file name (no directories, no absolute paths) for the generated MP3, written insidetext_to_music.output_dir; defaults to a timestampedmusic-*.mp3
The clip is always MP3 - the format the gateway's music providers serve (ElevenLabs offers mp3, opus, and pcm, not wav). To place it elsewhere, compose first and copy the returned file.
Configuration:
text_to_music:
enabled: true
# model: elevenlabs/music_v2_5 # gateway provider/model id
# output_dir: ~/.infer/tmp/music
require_approval: false # optional; unset = no approval, like the image toolsEvery key also has an INFER_TEXT_TO_MUSIC_-prefixed environment variable that takes precedence over the config file:
| Config key | Environment variable | Type | Default | Notes |
|---|---|---|---|---|
text_to_music.enabled | INFER_TEXT_TO_MUSIC_ENABLED | bool | false | Feature flag - must be true for the TextToMusic tool to reach the LLM |
text_to_music.model | INFER_TEXT_TO_MUSIC_MODEL | string | elevenlabs/music_v2_5 | Gateway provider/model id; a bare model name fails validation once the feature is enabled |
text_to_music.output_dir | INFER_TEXT_TO_MUSIC_OUTPUT_DIR | string | ~/.infer/tmp/music | Where generated MP3s are written |
text_to_music.require_approval | INFER_TEXT_TO_MUSIC_REQUIRE_APPROVAL | bool | unset (no approval) | Tri-state: unset keeps the tool's own default, an explicit value pins the policy either way |
Gateway requirements: the music endpoint is part of the gateway's Audio API, so the gateway must run with AUDIO_ENABLED=true and hold credentials for the provider behind model (an ElevenLabs API key for the default). The CLI-managed local gateway is started with AUDIO_ENABLED=true automatically when text_to_music.enabled is on, and an already-running instance without the Audio API is restarted. Point the CLI at an externally managed gateway and you set AUDIO_ENABLED=true and the provider key there yourself.
A gateway without the endpoint, or a provider that rejects the request, fails the tool call with a one-line error naming the configured model (music generation with elevenlabs/music_v2_5 failed: ...). The agent run still completes and no partial file is left behind.
TextToSFX Tool
Generate a short sound effect or ambience clip from a text prompt and save it as a WAV file. The chat model calls the tool when the user asks for a whoosh, a click, a riser or room tone - non-speech audio that TextToSpeech (it would read the words aloud) and TextToMusic (it composes songs) cannot cover. Generation goes through the gateway's SFX API (POST /v1/audio/sfx) using the configured provider/model - the CLI holds no provider key, the gateway does, so requests show up in gateway logs and traces. Disabled by default - while text_to_sfx.enabled is false, the tool definition is not sent to the LLM at all.
Parameters:
prompt(required): Description of the sound - the event or atmosphere and its character (a whoosh, a click, a riser, room tone, distant thunder)seconds(optional): Clip length in seconds,0.5-30; omitted lets the provider pick a length that fits the promptloop(optional):truefor a clip that loops seamlessly - useful for ambience bedsoutput_path(optional): Bare file name (no directories, no absolute paths) for the generated WAV, written insidetext_to_sfx.output_dir; defaults to a timestamped file
The same output_path rule as TextToMusic applies: the clip always lands in the configured output directory. To place it elsewhere, generate first and copy the returned file.
Configuration:
text_to_sfx:
enabled: true
# model: elevenlabs/eleven_text_to_sound_v2 # gateway provider/model id
# output_dir: ~/.infer/tmp/sfx
require_approval: false # optional; unset = no approval, like the image toolsEvery key also has an INFER_TEXT_TO_SFX_-prefixed environment variable that takes precedence over the config file:
| Config key | Environment variable | Type | Default | Notes |
|---|---|---|---|---|
text_to_sfx.enabled | INFER_TEXT_TO_SFX_ENABLED | bool | false | Feature flag - must be true for the TextToSFX tool to reach the LLM |
text_to_sfx.model | INFER_TEXT_TO_SFX_MODEL | string | elevenlabs/eleven_text_to_sound_v2 | Gateway provider/model id; a bare model name fails validation once the feature is enabled |
text_to_sfx.output_dir | INFER_TEXT_TO_SFX_OUTPUT_DIR | string | ~/.infer/tmp/sfx | Where generated WAVs are written |
text_to_sfx.require_approval | INFER_TEXT_TO_SFX_REQUIRE_APPROVAL | bool | unset (no approval) | Tri-state: unset keeps the tool's own default, an explicit value pins the policy either way |
Gateway requirements: the SFX endpoint is part of the gateway's Audio API (gateway v0.54.0 or newer), so the gateway must run with AUDIO_ENABLED=true and hold credentials for the provider behind model (an ElevenLabs API key for the default). The CLI-managed local gateway is started with AUDIO_ENABLED=true automatically when text_to_sfx.enabled is on, and an already-running instance without the Audio API is restarted. Point the CLI at an externally managed gateway and you set AUDIO_ENABLED=true and the provider key there yourself.
A gateway without the endpoint, or a provider that rejects the request, fails the tool call with a one-line error naming the configured model (sfx generation with elevenlabs/eleven_text_to_sound_v2 failed: ...). The agent run still completes and no partial file is left behind.
Video Tools
TextToVideo Tool
Render a short video clip and save it as an MP4. The chat model calls the tool when the user asks for a video clip or for an avatar to say something. Rendering goes through the gateway's Videos API (POST /v1/videos, polled with GET /v1/videos/{id} and downloaded from GET /v1/videos/{id}/content) using the configured provider/model - the CLI holds no provider key, the gateway does. Disabled by default - while text_to_video.enabled is false, the tool definition is not sent to the LLM at all, because avatar renders send the user's face and voice to a third-party provider.
Two modes:
- Prompt render - a text prompt, optionally with a portrait used as the first frame, rendered with
text_to_video.model. - Avatar render (lip-sync) - a portrait plus a
.wavor.mp3clip, rendered withtext_to_video.avatar_modelso the face speaks the audio.
Parameters:
prompt(required for a prompt render): Description of the clip - subject, action, camera, moodavatar(optional): Name of an avatar folder in the avatar library at~/.infer/avatars/, or a local image file, used as the portrait. Withaudioit is the portrait to lip-sync (first image in sort order); withoutaudioa library avatar is sent as reference images instead, while a bare image file stays the first frameaudio(optional): Local path to a.wavor.mp3clip to lip-sync; providing it selects the avatar render mode and requires a portraitoutput_path(optional): Bare file name (no directories, no absolute paths) for the generated MP4, written insidetext_to_video.output_dir; defaults to a timestamped file
Configuration:
text_to_video:
enabled: true
# model: elevenlabs/veo-3.1-fast-generate-001 # prompt renders
# avatar_model: elevenlabs/creatify-aurora # lip-synced avatar renders
# size: '' # "widthxheight" passthrough
# output_dir: ~/.infer/tmp/video
# timeout: 900
# poll_interval: 5
# create_avatar: false # register the CreateAvatar tool as well
require_approval: false # optional; unset = no approval, like the image toolsGateway requirements: the gateway must run with VIDEOS_ENABLED=true (gateway v0.54.0 or newer) and hold credentials for the provider behind the configured models. The CLI-managed local gateway is started with VIDEOS_ENABLED=true automatically while text_to_video.enabled is on; an externally managed gateway is your responsibility.
The portrait and the audio clip travel in one request, so together they must fit the gateway's 10 MiB request body limit, and creatify-aurora renders 480p or 720p keeping the portrait's aspect ratio.
See Text-to-Video and Avatars for the full configuration reference, the INFER_TEXT_TO_VIDEO_* environment variables, the avatar library layout, and infer avatars.
CreateAvatar Tool
Build an avatar folder at ~/.infer/avatars/<name>/ from a photo - the agent-side twin of infer avatars create. It stores the photo as the primary image and generates extra views of the same face through the gateway's image edit API using tools.image_edit.model, so a portrait the agent just generated, a selfie sent over a channel, or a photo in the project can become an avatar without leaving the chat.
Opt-in twice. The tool is registered only when both text_to_video.enabled and text_to_video.create_avatar (INFER_TEXT_TO_VIDEO_CREATE_AVATAR) are true; create_avatar defaults to false. It requires approval by default - unlike the other media tools - and an explicit text_to_video.require_approval overrides that either way.
Parameters:
name(required): Avatar name, which becomes the folder name under~/.infer/avatars/photo(required): Bare file name of the source photo, looked up in the working directory and then in the session artifacts directory; absolute paths and..are rejectedangles(optional): Which extra views to generate - defaults to both three-quarter angles, left/right profiles are available, and[]stores the photo only, with no image-edit callquality(optional): Image quality passed to the image edit API; defaults tohighsize(optional): Size of the generated views; defaults to1024x1536
It never overwrites an existing avatar - a name already in the library fails the call - and there is no delete counterpart, so the agent cannot remove an avatar. A failed view generation removes the half-built folder.
Privacy: generating views sends the photo to the image-edit provider (OpenAI by default). Before anything is stored or uploaded, a JPEG is turned upright per its EXIF orientation and re-encoded without its metadata, so camera and GPS tags never reach the library or a provider. PNG and WebP pass through unchanged.
GitHub Operations
There is no built-in GitHub tool. The agent performs all GitHub work - issues, pull requests, releases, repository metadata, and the raw API - through the gh CLI run via the Bash tool.
# Issues and pull requests
gh issue view 123
gh issue list --state open
gh pr create --title "fix: handle nil channel" --body "Closes #123"
gh pr diff 456
# Raw API (read-only / GET)
gh api repos/inference-gateway/cli/issues
gh api user --jq .loginRequirements: gh must be installed and authenticated. It uses the standard gh credential chain - run gh auth login, or set GITHUB_TOKEN (or GH_TOKEN). No separate token configuration exists anymore.
Default gh allowed-list
GitHub operations run through Bash, so they obey the Bash allowed-list. The default mode.all baseline auto-approves common read-only gh commands only:
| Auto-approved by default | Examples |
|---|---|
| Read-only reads | gh issue list, gh pr view 5, gh pr diff, gh repo view, gh release view v1 |
| Auth status | gh auth status |
| Search | gh search issues kind:bug, gh search code "func main" |
| Read-only project boards | gh project list, gh project view 3, gh project item-list 3 |
GitHub writes and destructive operations are deliberately left off the defaults. They are not auto-approved - they fall through to the standard approval prompt (in chat) or are blocked (in headless mode) until you add them to an allowed-list:
- Issue / PR writes -
gh issue create|edit|comment,gh pr create. - Project writes -
gh project item-add|item-edit. - Destructive -
gh pr merge,gh pr close,gh issue delete,gh repo delete,gh release create,gh run cancel,gh auth login. - Raw
gh api- any call. The previous GET-wildcard auto-approval was dropped; a raw-API need is now opt-in per repo.
Behavior change. Earlier defaults auto-approved
gh issue/prwrites and read-onlygh api. They now require approval - add the specific commands you trust to an allowed-list, or use the append override.
The shipped mode.all baseline:
tools:
bash:
mode:
all:
allow:
- gh (issue|pr|repo|release|run|workflow) (list|view|status|diff|checks)( .*)?
- gh auth status( .*)?
- gh search (issues|code|prs|repos|commits)( .*)?
- gh project (list|view|item-list|field-list)( .*)?Migration: the built-in GitHub tool was removed
Breaking change. The built-in
Githubtool was removed in favor of theghCLI. Thetools.githubconfig block and theinfer config tools githubcommands no longer exist. Existing configs that still contain atools.githubsection are ignored - unknown keys are dropped, so they do not error and need no manual cleanup. Replace any scripted use of the old tool with the matchingghcommand (for examplegh issue view,gh pr create,gh api).
Workflow Tools
TodoWrite
Create and update a structured task list for the current session. Use for complex multi-step work to track progress and surface intent to the user.
- Parameters:
todos(required array; each item hascontent,status∈pending|in_progress|completed, and optionalid) - Approval: not required
- Best practice: keep at most one task in
in_progressat a time; mark itemscompletedimmediately on finishing
Schedule
Create, list, get, update, or delete cron jobs that fire on a schedule. Jobs are persisted as YAML under ~/.infer/schedules/ and executed by the infer daemon (which reconciles its cron entries against storage every 2 seconds, so new or hand-edited jobs are picked up without a restart). When created from a channel session (e.g. Telegram), output is delivered back to that channel; otherwise the job is record-only - run history is persisted and viewable through the configured storage backend.
- Parameters:
operation(required:create|list|get|update|delete),job_id(required for get/update/delete),cron_expression(5-field crontab or@every <duration>),prompt,run_once(defaultfalse- whentrue, the job is deleted after firing once),name,description,model(optional model override) - Approval: required by default
- Notes: each fire creates a brand-new agent session - no context is carried between runs. Every fire persists a
RunRecord(session_id,job_id,status,error,started_at,finished_at) through the configured storage backend; the newest 200 records are retained.channelandrecipient_idon a job are optional delivery targets derived from the session, never passed by the LLM.
"0 8 * * *" every day at 08:00
"*/15 * * * *" every 15 minutes
"0 9 * * 1-5" weekdays at 09:00
"@every 1h" every hourJobs can also run in the cloud instead of the local daemon: set
scheduler.backend: githubto materialize each job as a GitHub Actions scheduled workflow. See Scheduling.
AskUserQuestion
Pause and ask the user 1-4 multiple-choice clarifying questions as an interactive, keyboard-driven form. The agent reaches for this whenever a task is ambiguous - in Plan Mode it resolves the ambiguity before it calls RequestPlanApproval, so your answers feed straight back into the plan it then proposes; in the other interactive modes the answers come back mid-task. It is read-only with no approval gate.
- Parameters:
questions(required array, 1-4 items). Each question has:header(required) - short chip label shown above the question, <= 12 charactersquestion(required) - the full question textoptions(required array, 2-4 items) - each option is{ label, description }multiSelect(optional, defaultfalse) - allow more than one answer to be selected
- Approval: not required (read-only)
- Availability: every interactive mode - Plan, Standard, Auto-Accept and auto-with-judge - when
tools.ask_user_question.enabledistrue. Headless runs have no interactive form and get the plain-text degradation below instead.
The form always appends an "Other" free-text choice to every question, so the user can answer outside the offered options. Suffix a label with (Recommended) to preselect that option when the question opens.
{
"questions": [
{
"header": "Datastore",
"question": "Which datastore should the new service use?",
"multiSelect": false,
"options": [
{
"label": "PostgreSQL (Recommended)",
"description": "Relational, strong consistency, already used by the gateway."
},
{ "label": "MongoDB", "description": "Document store with a flexible schema." },
{ "label": "Redis", "description": "In-memory, best for ephemeral or cache data." }
]
}
]
}Keyboard controls:
| Key | Action |
|---|---|
Up / Down | Move between options. For single-select questions the radio selection follows the cursor. |
Space | Toggle the highlighted option (multi-select questions). |
Enter | Confirm the current question and advance - or submit on the last question. |
Esc / Ctrl+C | Cancel the whole prompt. |
Headless graceful-degrade. When no interactive user is reachable to answer - a CI run, a heartbeat, or a scheduled job - the tool does not block. It returns a "proceed with assumptions" result so the agent keeps moving and picks a reasonable default instead of hanging.
RequestPlanApproval
Submit a completed plan for user approval. Available only in Plan Mode.
- Parameters:
plan(required - the complete, detailed plan text) - Behavior: pauses execution and offers three choices (see Approving a plan) - Accept (
Enter/y) switches to Auto-Accept mode and executes with no per-action approval, Approve Each Step (s) executes in Standard mode with approval on each action, and Reject (n) ends the session so you can reply with feedback.
RequestApproval
Ask the user to override an LLM judge rejection so the rejected tool call can run once. The agent reaches for it after a judge rejection (in auto-with-judge mode or under tools.safety.approval_behaviour: judge) - the rejection result itself hints that the path exists.
- Parameters:
tool(required - name of the rejected tool),arguments(required object - the exact arguments of the rejected call;{}for a call without arguments),what(required - what permission is needed, one sentence),why(required - why the action serves the user's request) - Behavior: shows the rejected call in the regular approval box with the judge's reason and the agent's justification. Approve runs that exact call once with the judge bypassed; Reject or dismissal denies and the turn continues with the decision in context.
- Eligibility: only calls the judge actually rejected can be escalated, and each one only once - anything else returns
not_rejectedoralready_escalated. The tool is advertised in every mode (a mode switch never invalidates the prompt cache), but outside judge mode nothing is ever rejected, so it has nothing to escalate. - Headless graceful-degrade: with no interactive approver (CI, headless, channels without an approval form) it returns a distinguishable "no approver reachable" result instead of blocking.
See Escalating a rejection for the full flow.
Local Subagents (Agent tool)
The Agent tool lets the main agent - in chat or headless mode - spawn one or more local subagents that run work in parallel and fold their results back into the main conversation. A subagent is just an infer headless subprocess with its own isolated session, so it is cheap, isolated, and session-persisted. The tool is enabled by default and gated by the tools.agent.* config block.
This is the lightweight, local complement to the A2A tools (A2A_SubmitTask / A2A_QueryTask / A2A_QueryAgent), which target external A2A servers:
| Reach for... | When |
|---|---|
| Agent (local subagents) | Short-lived helpers for the task at hand - parallel exploration, fan-out edits, scoped research - with no server to run. Each is a local infer headless subprocess. |
A2A tools (A2A_SubmitTask, ...) | Delegating to external, long-running, specialized A2A servers (calendar, docs, ...) discovered over the network. See A2A. |
Tool parameters
The model calls the tool with either a batch of tasks or a single description:
tasks- an array of subagent tasks run in parallel, each with:description(required) - the task for that subagentlabel(optional) - short label shown in progress output / tmux panesmodel(optional) - per-subagent model overridesystem_prompt(optional) - gives that subagent a specialized role/personaagent(optional) - name of a Markdown subagent preset that supplies the system prompt, model, and tool allowlist
description(optional) - shorthand for a single-task call (an alternative totasks)system_prompt(optional) - system prompt for the single-descriptionformagent(optional) - preset name for the single-descriptionform
Each subagent runs in its own isolated session id of the form subagent-<parentSession>-<uuid>. Parallel fan-out is capped by max_parallel (default 4) concurrent subagents per call.
Markdown subagent presets
Instead of spelling out a system prompt on every call, a subagent can be defined once as a Markdown file with YAML frontmatter and delegated to by name through the agent parameter. The format is the same one Claude Code (.claude/agents/*.md) and Gemini CLI (.gemini/agents/*.md) use, so an existing agent file from either tool loads unchanged once copied in. Every preset shows up as a local row in the /agents view.
---
name: code-reviewer
description: Reviews a diff for correctness bugs. Use after making code changes.
model: deepseek/deepseek-v4-pro
tools: Read, Grep, Tree
---
You are a senior reviewer. Read the changed files and report findings as a
numbered list ordered by severity.The Markdown body is the subagent's system prompt. Frontmatter keys:
| Key | Required | Meaning |
|---|---|---|
name | yes | Identifier passed as the agent argument. Lowercase letters, digits, -, _; up to 64 characters |
description | yes | Shown to the main agent so it knows when to delegate |
model | no | provider/model, or inherit to use the parent turn's model |
tools | no | Tools the subagent may use - a YAML list or a comma-separated string. Omitted means inherit all |
disallowedTools | no | Tools removed from the resolved list. Same format as tools |
Any other key (color, temperature, max_turns, mcpServers, permissionMode, ...) is accepted and ignored, so files written for other orchestrators load as-is. These are presets for the Agent tool, entirely separate from .infer/agents.yaml, the A2A agent registry.
Locations and precedence. Definitions are looked up in this order, first match wins on a name collision:
- Project:
.infer/agents/<name>.md- commit it to share the agent with the repository - User-global:
~/.infer/agents/<name>.md- stays personal
A project preset therefore overrides a personal one of the same name, exactly like skills.
Tool allowlist. The resolved allowlist is enforced inside the spawned subagent in one place: a disallowed tool is neither offered to the model nor executable by naming it. Unknown tool names (for example Claude's Glob) log a warning and are dropped, disallowedTools entries are always honored, and a restriction that resolves to no known tool skips the file rather than silently falling back to all tools. The preset's capability follows from the allowlist - read-only when every allowed tool is read-only, otherwise read-write (mutations still go through approval). Listing Agent itself lets a preset spawn subagents, still bounded by max_depth.
Model resolution. The file's model wins over a per-task model argument. A value without a provider prefix - a Claude alias such as sonnet, or a bare model ID - logs a warning and falls back to inherit. When the file sets no model, the normal order applies: per-task model, then tools.agent.model, then the parent turn's model.
Files are scanned once per session. A file with broken frontmatter, a missing or invalid name/description, or an unusable tools list is skipped with a warning naming the file and the reason - an invalid preset never fails startup.
v1 limits.
.claude/agents/and.gemini/agents/are not read in place (copy or symlink the files into.infer/agents/), there is no hot reload, and per-agentmcpServers,temperature,max_turns,permissionMode, and tool wildcards such asmcp_*are not honored. MCP tools register after session start, so MCP tool names in atoolslist are dropped as unknown.
Result modes: async and wait-all
- Wait-all (
wait: true) - the shipped default. The call blocks until every spawned subagent reaches a terminal state, then returns the aggregated results in one tool result. - Async (
wait: false) - the call returns immediately with the subagent ids; when each subagent finishes, its final result is injected back into the main conversation (mirroringA2A_SubmitTasknotify behavior). In chat, running/completed status is surfaced in the sticky progress area.
Execution surfaces: headless and interactive (tmux)
The mode controls where subagents run. Either way the result aggregates back into the main context exactly the same - interactive is "headless plus a tmux pane attached to the live process":
headless- subagents run in the background; results aggregate back into the main context.interactive(the shipped default) - each subagent runs in a live tmux pane/window you can watch while it works.
tmux is an optional runtime dependency, required only for interactive mode (headless needs nothing extra). Interactive mode must be run from inside tmux ($TMUX set). When you are not inside tmux (or tmux is not installed), the interactive.fallback setting decides what happens:
fallback: headless(default) - warn and run headless.fallback: error- fail the call instead.
Agent tool configuration
The new tools.agent.* block, with its shipped defaults (regenerated by infer init):
tools:
agent:
enabled: true
require_approval: true # spawning work that can edit files is a mutating action
mode: interactive # headless | interactive (default when a call omits it)
wait: true # block and return aggregated results by default
max_parallel: 4 # cap on concurrent subagents per call
max_depth: 1 # recursion guard; a subagent is itself an `infer headless`
model: '' # default subagent model (inherits parent if blank)
inherit_mock: true # when gateway.mock is on, spawn subagents against the embedded mock too
interactive:
multiplexer: tmux # tmux only
layout: vertical # vertical | horizontal | window
fallback: headless # headless | error (when not inside tmux)Every key has an INFER_TOOLS_AGENT_* environment-variable override, consistent with the rest of the config:
| Setting | Environment variable |
|---|---|
enabled | INFER_TOOLS_AGENT_ENABLED |
require_approval | INFER_TOOLS_AGENT_REQUIRE_APPROVAL |
mode | INFER_TOOLS_AGENT_MODE |
wait | INFER_TOOLS_AGENT_WAIT |
max_parallel | INFER_TOOLS_AGENT_MAX_PARALLEL |
max_depth | INFER_TOOLS_AGENT_MAX_DEPTH |
model | INFER_TOOLS_AGENT_MODEL |
inherit_mock | INFER_TOOLS_AGENT_INHERIT_MOCK |
interactive.multiplexer | INFER_TOOLS_AGENT_INTERACTIVE_MULTIPLEXER |
interactive.layout | INFER_TOOLS_AGENT_INTERACTIVE_LAYOUT |
interactive.fallback | INFER_TOOLS_AGENT_INTERACTIVE_FALLBACK |
# Toggle the tool, or switch the default execution surface to headless
infer config set tools.agent.enabled true
infer config set tools.agent.mode headlessMock inheritance (inherit_mock, default true). When the parent CLI runs against the embedded mock gateway (gateway.mock: true), spawned subagents inherit mock mode: the CLI passes INFER_GATEWAY_MOCK=true to each subagent - both the headless environment and the interactive tmux-pane command - so they exercise the same mock instead of talking to the real gateway. Set tools.agent.inherit_mock: false (or INFER_TOOLS_AGENT_INHERIT_MOCK=false) to opt out and have subagents always target the configured gateway.
Approval and security
- Subagents run in standard bash mode (the restricted allowed-list), exactly like every other headless run - an off-list or mutating action is blocked in CI/heartbeat (no approver reachable) or sent for IPC approval under a channel (for example Telegram). See Headless secure-by-default.
- The Agent tool is in the approval policy and requires approval by default (
require_approval: true), with a per-tool override - consistent withA2A_SubmitTask. Spawning work that can edit files is treated as a mutating action. - A depth guard (
max_depth, default1) prevents subagent fork-bombs: a subagent cannot itself spawn further subagents at the default cap.
Tracing
When telemetry is enabled, a headless subagent inherits the caller's trace context (TRACEPARENT), so its own session span nests under the caller's execute_tool Agent span - the whole fan-out reads as one cross-process trace in infer traces:
session (standard, success) 152ms
|-- execute_tool Agent 89ms
| `-- session (readonly, success) 52ms
| `-- chat openai/gpt-4o 13msInteractive (tmux-pane) subagents are not stitched into the caller's trace; use mode: headless when you need the subagent's spans in the same trace. See Trace context propagation to subprocesses for the full contract.
v1 scope. Subagents do not nest (depth capped at 1), a subagent's tool-approval prompt is not routed back to the main chat TUI, only tmux is supported (no screen/zellij), and there is no CLI command to list or create Markdown presets (the
/agentsview lists them read-only).
Security Features
- Command allow-listing: Default-deny, per-mode allowed-list for the Bash tool
- Approval Prompts: Safety confirmations for Write/Edit/Delete/Bash
- Path Protection: Sensitive directories automatically excluded (
.git/,*.env,.infer/) - Sandbox Controls: Restrict tool operations to allowed directories
- Domain allow-listing: Control web fetch access
- Diff Preview: Colored, syntax-aware diff before file modifications
Tool Configuration
Tool settings are read and written with the generic config commands - there are no per-setting subcommands.
# Enable/disable all tool execution for LLMs
infer config set tools.enabled true
infer config set tools.enabled false
# Enable/disable an individual tool (for example bash)
infer config set tools.bash.enabled true
# Require approval before any tool runs
infer config set tools.safety.require_approval true
# Require approval for a specific tool only (for example bash)
infer config set tools.bash.require_approval true
# Sandbox directories - comma-separated; the whole list is replaced
infer config set tools.sandbox.directories ".,/protected/path"
# Inspect the resulting tools config
infer config get toolsRunning Tools Directly
Run any enabled tool outside a chat session, or check whether a bash command would pass the allowed-list, with the top-level infer tools command.
# Execute a tool by name with JSON arguments (tool names are case-insensitive)
infer tools execute Read '{"file_path":"README.md"}'
infer tools execute grep '{"pattern":"func main","path":"."}'
# Validate whether a bash command is allowed (without running it)
infer tools validate "git status"infer tools execute <tool> [json-args] resolves tool names case-insensitively in the CLI - the agent itself still uses the exact PascalCase names. infer tools validate <command> reports whether a bash command would be permitted by the configured allowed-list, without executing it.
infer tools executeandinfer tools validatemoved fromconfig tools exec/config tools validateto the top-levelinfer toolscommand.
Configuration
Two-layer configuration system with precedence from highest to lowest:
Configuration Precedence
| Priority | Source | Example |
|---|---|---|
| 1 (Highest) | Environment Variables | INFER_GATEWAY_URL, INFER_AGENT_MODEL |
| 2 | Command Line Flags | --model, --debug |
| 3 | Project Config | .infer/config.yaml |
| 4 | User Config | ~/.infer/config.yaml |
| 5 (Lowest) | Built-in Defaults | Internal defaults |
Configuration Files
infer init seeds the userspace baseline in ~/.infer/ (config.yaml, mcp.yaml, prompts.yaml, agents.yaml, shortcuts/, skills/ and so on) and writes nothing into the project. Project-level overrides are created on demand with infer config set --project <key> <value>, or with --project on infer mcp / infer agents. Configuration is split across purpose-specific YAML files rather than one giant file:
| File | Scope | Purpose | Where it is documented |
|---|---|---|---|
config.yaml | Project/user | Main config - agent, tools, storage, pricing, and everything config get/set touches. | Configuration |
prompts.yaml | Project/user | System prompts (prompts.agent.system_prompt), per-mode adjustments, and tool descriptions (prompts.tools.<Tool>.description) - edited, not set. | Configuration Commands |
mcp.yaml | Project/user | MCP server definitions and connection settings. | MCP Integration |
keybindings.yaml | Project/user | Keybindings for the TUI and diff viewer (category diff_viewer). | Diff viewer and git staging |
hooks.yaml | Project/user | User-defined shell commands run at agent-loop hook points (feature-flagged off by default). | Command Hooks |
reminders.yaml | Project/user | System reminders injected into the conversation on a schedule. | System Reminders |
judge.yaml | Project/user | LLM judge that decides approval-requiring tool calls (model, timeout, prompts, on_error). | Judge Mode |
memory.yaml | Project/user | Persistent, cross-session agent memory - fact-files plus the MEMORY.md index. | Persistent Memory |
shortcuts/*.yaml | Project | Custom slash shortcuts - simple commands, subcommands, and AI-powered snippets. | Custom Shortcuts |
skills/ | Project/user | Agent Skills folders (name/SKILL.md) discovered and injected on demand. | Agent Skills |
schedules/ | User | Persisted cron jobs created by the Schedule tool, run by the daemon. | Schedule |
artifacts/ | Project/user | Agent deliverables, grouped per session. | Artifacts directory |
logs/ | User | CLI and gateway log files (~/.infer/logs, overridable via logging.dir). | Key Configuration Areas |
bin/ | User | Downloaded binaries - the gateway server, plus optional helpers like ffmpeg. | Key Configuration Areas |
insights/ | User | Saved infer insights reports. Written with secrets redacted, and kept by /reset. | Insights Shortcut |
avatars/ | User | Avatar portrait folders (avatars/<name>/*.png) used by TextToVideo lip-sync renders. Kept by /reset. | Text-to-Video |
auth.yaml | User | Fallback provider API keys, used when a key is not in the environment or the project .env. | Provider API keys |
tmp/ | User | Userspace scratch - generated speech (tmp/tts), retained recordings (tmp/voice), screen recordings (tmp/recordings), channel media (tmp/media). Wiped by /reset. | Userspace tmp tree |
No migration.
logs/andbin/are userspace-only: they live under~/.infer/and are shared by every project. Older versions wrote them into the project's.infer/directory; those directories are orphaned by design and safe to delete. The.infer/.gitignoreseeded byinfer initno longer listsbin/orlogs/*.log.
Provider API keys
Provider API keys (ANTHROPIC_API_KEY, OPENAI_API_KEY, GROQ_API_KEY, ...) are resolved per key, first hit wins:
| Priority | Source | Notes |
|---|---|---|
| 1 (Highest) | System environment | Exported in your shell, CI secrets, and so on. |
| 2 | Project .env | .env in the project directory. |
| 3 (Lowest) | ~/.infer/auth.yaml | Userspace fallback, shared by every project. |
Because resolution is per key, sources mix: OPENAI_API_KEY can come from the environment while ANTHROPIC_API_KEY comes from auth.yaml in the same run. The fallback applies when the CLI starts the gateway in both container and binary modes, and to A2A agent containers.
auth.yaml is a flat YAML map of environment-variable names to values:
ANTHROPIC_API_KEY: sk-ant-...
OPENAI_API_KEY: sk-...
DEEPSEEK_API_KEY: sk-...Create it once and keep it private:
mkdir -p ~/.infer
$EDITOR ~/.infer/auth.yaml
chmod 600 ~/.infer/auth.yaml- Permissions:
0600is recommended. Broader permissions still work but log a warning. - Sandboxed:
~/.infer/auth.yamlis on the protected paths list - agent tools cannot read or edit it. - Graceful degradation: a missing or unreadable file changes nothing; a malformed file is ignored with a logged warning. Key resolution never fails because of
auth.yaml. - Legacy
auth.json: the old JSON file is still read as a fallback whenauth.yamlis absent, so existing credentials keep working. Move your keys intoauth.yaml- JSON is valid YAML, so the contents can be pasted as-is.
Artifacts directory
Files the agent produces for you land in .infer/artifacts/<session-id>/, under the project config dir when there is one and the userspace config dir (~/.infer/artifacts/) otherwise. Each conversation gets its own subdirectory, so a session's output stays grouped and is easy to find, keep, or delete as a unit.
What lands there:
- Images from the ImageGeneration, ImageEdit, and ImageVariation tools -
image-<timestamp>.png. - WebFetch downloads and binary fetches.
- Auto-downloaded A2A task artifacts.
.infer/artifacts is not the same as .infer/tmp:
| Directory | Contents | Lifetime |
|---|---|---|
.infer/artifacts | Intended agent deliverables, grouped by session | Yours to keep - nothing prunes it for you |
.infer/tmp | Internal scratch only - chunk staging, skill downloads, screenshots, pastes | Disposable working data |
Session IDs are sanitized before use, so a directory can never escape the artifacts root. The tool sandbox carves the artifacts directory out as writable, so tools can save there even when it sits outside the sandbox directory list.
Userspace tmp tree
Disposable runtime output that is not tied to a single project lives under ~/.infer/tmp/:
| Directory | Contents | Config default of |
|---|---|---|
~/.infer/tmp/tts | Generated speech WAVs | text_to_speech.output_dir |
~/.infer/tmp/music | Generated music MP3s | text_to_music.output_dir |
~/.infer/tmp/sfx | Generated sound-effect WAVs | text_to_sfx.output_dir |
~/.infer/tmp/video | Rendered MP4 clips | text_to_video.output_dir |
~/.infer/tmp/voice | Retained inbound voice recordings | speech_to_text.recordings_dir |
~/.infer/tmp/media | Retained inbound Telegram media | channels.telegram.media.dir |
The whole ~/.infer/tmp tree is agent-readable and writable by design - retained recordings and media are assets the agent consumes, and generated speech is output it can reference. The rest of ~/.infer/ stays on the protected paths list. /reset empties the tree through the tmp parent; the owning subsystems recreate the subdirectories on next use. Directories you explicitly point outside ~/.infer are left alone.
Existing installs. Older releases placed these three directories directly under
~/.infer/(tts/,voice/,media/). Nothing migrates automatically - they hold only disposable output, so delete them, ormvtheir contents under~/.infer/tmp/to keep the retained files. Explicitoutput_dir/recordings_dir/media.diroverrides are unaffected.
Key Configuration Areas
Gateway Settings:
- Gateway URL and API key
- Timeout and retry configuration (see Client retry and stream reconnection below)
- OCI image for auto-running gateway
- Model filtering:
gateway.include_models(allowlist) andgateway.exclude_models(blocklist). Both default to[].gateway.exclude_modelsis opt-in - the model picker already hides non-chat-capable models via the gateway-reported modalities, so use it only to hide specific models the gateway reports (large or costly chat models, for example); there is no shipped default blocklist. Set exclusions are passed to the gateway as theDISALLOWED_MODELSenvironment variable.
The downloaded gateway binary lands at ~/.infer/bin/inference-gateway - one shared copy per machine, reused by every project (the staleness check still re-downloads when the pinned version changes).
Logging Configuration:
Logging settings control log output, file location, and automatic log archiving:
logging:
debug: false
dir: '' # Override log directory (defaults to ~/.infer/logs)
stdout: false # Also write logs to stdout/stderr in addition to the log file
archive:
enabled: true # Automatically archive oversized log files (default: true)
max_size_mb: 1024 # Threshold in MB; files exceeding this are gzip-compressed and truncated (default: 1024 = 1 GB)- logging.debug: Enable debug logging for verbose output
- logging.dir: Override the log directory. Defaults to
~/.infer/logs(INFER_LOGGING_DIR) - CLI and gateway logs are machine-scoped and never written to the project directory. - logging.stdout: Also write logs to stdout/stderr in addition to the log file (default:
false) - logging.archive.enabled: Enable automatic log archiving (default:
true). When enabled, log files exceeding the size threshold are gzip-compressed to a timestamped.gzarchive and the original file is truncated so logging continues at the same path. The check runs at process startup. Set viaINFER_LOGGING_ARCHIVE_ENABLED. - logging.archive.max_size_mb: Maximum log file size in MB before archiving is triggered (default:
1024, i.e. 1 GB). A value of0or less disables archiving. Set viaINFER_LOGGING_ARCHIVE_MAX_SIZE_MB.
Agent Configuration:
- Default model for operations
- System prompt (
prompts.agent.system_prompt) and per-mode adjustments (mode_adjustment_plan/mode_adjustment_auto, delivered by the mode-change reminder) - System reminders interval
- Max turns and tokens
- Parallel tool execution (default: 5 concurrent)
The built-in system prompt includes a Current date: line (date-only, no time) so provider-side prompt-prefix caching stays effective across turns. When the agent needs the current time, it runs the date command via Bash.
Tool Settings:
- Enable/disable individual tools
- Approval requirements per tool (whether) and delivery via
tools.safety.approval_behaviour(how) - Per-mode bash allowed-lists (
tools.bash.mode.<mode>.allow) - Sandbox directories
- Protected paths
Storage Backends:
- JSONL (default) - append-only files for portable, inspectable history
- PostgreSQL - shared database for teams
- Redis - high-performance caching
- SQLite - local file storage
- Cloudflare D1 - external SQLite-compatible store over HTTP (for ephemeral CI runners)
- In-memory - temporary sessions
Conversation Features:
- Automatic history with search
- AI-generated titles
- Token optimization and compaction
- Export/import capabilities
Client Retry and Stream Reconnection
The CLI automatically reconnects when a stream stalls or drops, using the client.retry.* settings for exponential backoff and attempt limits. A stall is detected when no progress is made - no response while connecting, no SSE chunk while streaming - for longer than client.stall_threshold_sec.
# .infer/config.yaml
client:
stall_threshold_sec: 30 # seconds without progress before reconnecting (0 disables)
retry:
enabled: true
max_attempts: 5 # maximum number of retry attempts (default: 5)
retryable_status_codes: [408, 429, 500, 502, 503, 504] # transient errors only
initial_backoff_sec: 1
max_backoff_sec: 60
backoff_multiplier: 2Settings:
- client.stall_threshold_sec: Seconds without progress - no response while connecting, no SSE chunk while streaming - before the CLI drops the connection and reconnects (default:
30,0disables). Reconnects reuseclient.retry.*for attempts and exponential backoff. Keep this above your provider's worst first-token latency, since a stall retry restarts the response from scratch. Set viaINFER_CLIENT_STALL_THRESHOLD_SEC. - client.retry.enabled: Enable automatic retries for failed requests (default:
true). - client.retry.max_attempts: Maximum number of retry attempts (default:
5). Set viaINFER_CLIENT_RETRY_MAX_ATTEMPTS. - client.retry.retryable_status_codes: HTTP status codes that trigger retries (default:
[408, 429, 500, 502, 503, 504]); non-transient errors such as 401 fail fast with their real message. Set viaINFER_CLIENT_RETRY_RETRYABLE_STATUS_CODES. - client.retry.initial_backoff_sec: Initial delay between retries in seconds (default:
1). - client.retry.max_backoff_sec: Maximum delay between retries in seconds (default:
60). - client.retry.backoff_multiplier: Backoff multiplier for exponential delay (default:
2).
Reconnecting indicator: In the chat TUI, the status bar shows a red Reconnecting... indicator when a stall is detected, then Reconnecting (N/M) per attempt (where N is the current attempt and M is max_attempts). Input is blocked until the stream recovers or all attempts are exhausted. See Status indicator row.
System Reminders
System reminders are YAML-configured prompts injected into the conversation on a schedule or in response to tool outcomes. They are defined in reminders.yaml (project or user scope) and can also be supplied inline or from an arbitrary path.
Reminders YAML Schema
# .infer/reminders.yaml (or ~/.infer/reminders.yaml)
enabled: true
reminders:
- name: memory-consult
hook: pre_tool
trigger: always
text: 'Consult the Memory tool before making changes.'
- name: fail-nudge
hook: post_tool
trigger: on_failure
text: 'A failed call means the change did not happen. Re-try or ask the user.'| Field | Type | Description |
|---|---|---|
enabled | boolean | Master switch. Default true. Can also be toggled via INFER_REMINDERS_ENABLED. |
name | string | Unique identifier for the reminder. Used for deduplication and logging. |
hook | string | When the reminder fires. One of pre_tool (before each tool call) or post_tool (after each tool call completes). |
trigger | string | Condition under which the reminder fires. See Trigger catalog below. |
text | string | The reminder text injected into the conversation. Supports os.ExpandEnv environment variable interpolation ($VAR or ${VAR}). |
guidance | map | on_mode_change only. Maps a mode key (standard, plan, auto) to the text substituted for the {guidance} placeholder in text. Omitted keys keep their built-in defaults. |
Trigger catalog
| Trigger | Hook requirement | Description |
|---|---|---|
always | Any | Fire on every hook invocation. |
on_failure | post_tool | Fire only when the tool call that just ran failed (returned an error). Requires hook: post_tool; validation rejects other hooks. |
on_mode_change | Any | Fire when the agent mode changes (Shift+Tab). Used by the built-in mode-change-reminder, which is the sole carrier of mode-specific instructions - see How mode instructions are delivered. |
Configuration sources and precedence
Reminders are resolved with the following precedence (highest first):
| Priority | Source | Description |
|---|---|---|
| 1 (Highest) | INFER_REMINDERS_CONFIG | Inline YAML string. When set, it replaces all file-loaded reminders. |
| 2 | --reminders-file PATH | Load reminders from an arbitrary file path. Available on infer headless and infer chat. |
| 3 | Project config | ./.infer/reminders.yaml |
| 4 | User config | ~/.infer/reminders.yaml |
| 5 (Lowest) | Built-in defaults | The CLI ships a built-in memory-consult reminder that nudges the agent to consult the Memory tool. |
INFER_REMINDERS_ENABLED toggles the master switch on top of all sources — set it to false to disable all reminders regardless of the resolved config.
Disabling reminders drops mode instructions. The mode-change reminder is the only carrier of per-mode instructions, so with
enabled: false(orINFER_REMINDERS_ENABLED=false) the Plan and Auto-Accept guidance - including the destructive-action policy - is never delivered. Mode tool restrictions are unaffected.
INFER_REMINDERS_CONFIG
Set this environment variable to supply the full reminders YAML inline, without writing a reminders.yaml file. This is especially useful for CI/CD and embedded consumers (e.g. infer-action) that cannot write to ~/.infer/.
export INFER_REMINDERS_CONFIG='enabled: true
reminders:
- name: fail-nudge
hook: post_tool
trigger: on_failure
text: "A failed call means the change did not happen"
'When INFER_REMINDERS_CONFIG is set, it replaces the file-loaded config entirely — it is not merged. To disable it and fall back to file-based config, unset the variable.
--reminders-file
The --reminders-file PATH flag is available on infer headless and infer chat. It loads reminders from an arbitrary YAML file path, bypassing the default file resolution.
infer headless "Refactor the module" --reminders-file /path/to/custom-reminders.yaml
infer chat --reminders-file ./ci-reminders.yamlExample: CI with inline reminders
INFER_REMINDERS_CONFIG='enabled: true
reminders:
- name: fail-nudge
hook: post_tool
trigger: on_failure
text: "A failed call means the change did not happen"
' INFER_REMINDERS_ENABLED=true infer headless "..."Essential Environment Variables
export INFER_GATEWAY_URL="http://localhost:8080"
export INFER_GATEWAY_API_KEY="your-api-key"
export INFER_AGENT_MODEL="deepseek/deepseek-v4-flash"
export INFER_LOGGING_DEBUG="true"
export INFER_LOGGING_ARCHIVE_ENABLED="true"
export INFER_LOGGING_ARCHIVE_MAX_SIZE_MB="1024"
export INFER_CLIENT_STALL_THRESHOLD_SEC="30" # seconds without stream progress before reconnecting (0 disables)
export GITHUB_TOKEN="your-github-token" # used by the gh CLI credential chain for GitHub operations
# Append a few commands onto the bash allowed-list baseline (comma- or newline-separated)
export INFER_TOOLS_BASH_ALLOW_APPEND="git commit,git push"
# How a needed approval is delivered: prompt | ipc | judge | block
export INFER_TOOLS_SAFETY_APPROVAL_BEHAVIOUR="prompt"
# Inline reminders YAML (replaces file-loaded reminders)
export INFER_REMINDERS_CONFIG='enabled: true
reminders:
- name: fail-nudge
hook: post_tool
trigger: on_failure
text: "A failed call means the change did not happen"
'
# Master switch for reminders (default: true)
export INFER_REMINDERS_ENABLED=true
# Per-mode adjustment instructions, delivered by the mode-change reminder
# (supersede the deprecated INFER_PROMPTS_AGENT_SYSTEM_PROMPT_PLAN / _AUTO)
export INFER_PROMPTS_AGENT_MODE_ADJUSTMENT_PLAN="Investigate first; do not propose edits."
export INFER_PROMPTS_AGENT_MODE_ADJUSTMENT_AUTO="Confirm before any irreversible operation."Configuration Commands
Configuration uses a generic key/value interface. infer config get reads the effective value of any key; infer config set writes one to the userspace baseline (~/.infer/config.yaml) by default, or to the project config (./.infer/config.yaml) when you pass --project. Keys are dotted paths into the config (for example agent.model, tools.bash.enabled).
# Initialize configuration
infer config init
# Print the whole effective config (defaults + ~/.infer + .infer + INFER_* env)
infer config get
# Print a single key
infer config get agent.model
# Print as JSON instead of YAML
infer config get --format json
# Set a value - parsed to the field's type (bool, integer, number, or string)
infer config set agent.model deepseek/deepseek-v4-flash
infer config set agent.max_turns 50
infer config set agent.verbose_tools true
# List-valued keys take a comma-separated value (the whole list is replaced)
infer config set tools.sandbox.directories ".,/work/project"
infer config set tools.web_fetch.allowed_domains "golang.org,github.com"
# config set writes the userspace baseline (~/.infer/config.yaml) by default
infer config set agent.model deepseek/deepseek-v4-flash
# Target the project config (./.infer/config.yaml) instead - it overrides the baseline key-by-key
infer config set agent.model deepseek/deepseek-v4-flash --project
# Recreate config.yaml from defaults
infer config init --overwriteSystem prompts are not set via
config set- they live inprompts.yaml(for exampleprompts.agent.system_prompt, and the per-modeprompts.agent.mode_adjustment_plan/mode_adjustment_auto) and are edited there. The same file holds the tool descriptions sent to the model underprompts.tools.<Tool>.description- for exampleprompts.tools.RequestApproval.description, which ships with a built-in default explaining when a judge rejection may be escalated. Any key you leave out keeps its built-in text.
Command Mapping
The per-setting subcommands were removed in favor of config get/config set and the top-level infer tools command:
| Old command | New command |
|---|---|
config agent set-model X | config set agent.model X |
config agent set-max-turns N | config set agent.max_turns N |
config agent verbose-tools enable | config set agent.verbose_tools true |
config agent skills enable | config set agent.skills.enabled true |
config tools enable | config set tools.enabled true |
config tools bash enable | config set tools.bash.enabled true |
config tools safety enable | config set tools.safety.require_approval true |
config tools safety set bash enabled | config set tools.bash.require_approval true |
config tools sandbox add DIR | config set tools.sandbox.directories ".,DIR" |
config tools grep set-backend rg | config set tools.grep.backend ripgrep |
config tools web-fetch add-domain D | config set tools.web_fetch.allowed_domains "D" |
config export set-model X | config set export.summary_model X |
config show | config get |
config tools exec <tool> | tools execute <tool> |
config tools validate <cmd> | tools validate <cmd> |
See the full configuration reference for detailed options.
Telemetry
The CLI records OpenTelemetry signals (metrics, traces, and logs) for every session. Telemetry is always on when the CLI runs - there is no opt-out switch. Data is written to local files under ~/.infer/telemetry/ and can optionally be exported to an OTLP/HTTP collector.
Local files
All three signals write per-session JSONL files to ~/.infer/telemetry/:
| Signal | File pattern | Description |
|---|---|---|
| Metrics | ~/.infer/telemetry/\<session-id\>.jsonl | Token usage, tool outcomes, session duration, and cost (delta temporality) |
| Traces | ~/.infer/telemetry/\<session-id\>-traces.jsonl | One root span per session, child spans for each LLM turn and tool call |
| Logs | ~/.infer/telemetry/\<session-id\>-logs.jsonl | Structured log entries emitted during the session |
The \<session-id\> is the same UUID that appears in the CLI's session output and conversation storage. Local files use the OTLP/semconv JSON format as-is - no custom encoding.
Metrics
The CLI records the following metrics using the OpenTelemetry GenAI semantic conventions:
- gen_ai.client.token.usage - input and output token counts per LLM request
- gen_ai.execute_tool.duration - tool execution duration in seconds
- infer.agent.runs - total agent session count
- infer.agent.run.duration - session duration in seconds
- infer.client.cost - computed dollar cost per session
Metrics use delta temporality (required by the gateway ingest, and what makes the local files trivially summable by infer stats).
Traces
The CLI emits one trace per session with the following span hierarchy:
- session (root span) - carries
infer.execution.mode,infer.agent.mode,infer.run.outcome, andgen_ai.conversation.id - chat <model> (CLIENT kind) - one per LLM request, with
gen_ai.request.model,gen_ai.provider.name,gen_ai.usage.input_tokens,gen_ai.usage.output_tokens - execute_tool <name> (INTERNAL kind) - one per tool call, with
gen_ai.tool.name,gen_ai.tool.type,infer.tool.outcome
No prompt or response content is recorded in spans. Failed spans carry error.type and Error status per the OpenTelemetry recording-errors conventions.
Spans emitted by subprocesses - Bash tool commands, skill scripts, and headless subagents - are ingested and nested under their originating execute_tool span, so a single trace can span multiple processes. See Trace context propagation to subprocesses.
Logs
The CLI emits structured log entries as OTel log records. Each entry includes a timestamp, severity level, message, and contextual fields such as session_id, model, and request_id.
OTLP/HTTP export
All three signals can be exported to an OpenTelemetry Collector or any OTLP/HTTP-compatible backend by setting the standard OpenTelemetry environment variables:
# Base endpoint (all signals append their own path: /v1/metrics, /v1/traces, /v1/logs)
export OTEL_EXPORTER_OTLP_ENDPOINT=http://otel-collector:4318
# Optional: per-signal endpoint overrides (takes precedence over the base)
export OTEL_EXPORTER_OTLP_METRICS_ENDPOINT=http://otel-collector:4318/v1/metrics
export OTEL_EXPORTER_OTLP_TRACES_ENDPOINT=http://otel-collector:4318/v1/traces
export OTEL_EXPORTER_OTLP_LOGS_ENDPOINT=http://otel-collector:4318/v1/logs
# Headers (e.g. for authentication)
export OTEL_EXPORTER_OTLP_HEADERS="Authorization=Bearer%20my-token"
# Service identity (merged into the resource attributes on every signal)
export OTEL_SERVICE_NAME=infer-cli
export OTEL_RESOURCE_ATTRIBUTES="deployment.environment=production,actor=ci-bot"When OTEL_EXPORTER_OTLP_ENDPOINT is set, the CLI exports all three signals to that endpoint. Per-signal env vars (OTEL_EXPORTER_OTLP_METRICS_ENDPOINT, OTEL_EXPORTER_OTLP_TRACES_ENDPOINT, OTEL_EXPORTER_OTLP_LOGS_ENDPOINT) override the base for that signal. Headers and timeouts follow the standard OTel exporter env-var spec.
Local file export is always active regardless of OTLP configuration - remote export is additive, not a replacement.
Local OTLP receiver (telemetry.receiver_address)
By default, the CLI runs an ephemeral loopback OTLP/HTTP receiver (127.0.0.1:0) that accepts spans from subprocesses (Bash tool commands, headless subagents) and persists them into the per-session trace file. This receiver is loopback-only and auto-allocated, so it is invisible to external processes.
When you need an external OpenTelemetry Collector (or any OTLP producer) to feed spans back into the CLI's local trace store - for example to see a remote A2A agent's spans in infer traces - set telemetry.receiver_address to a fixed address:
# .infer/config.yaml
telemetry:
receiver_address: 0.0.0.0:4318# Environment variable
export INFER_TELEMETRY_RECEIVER_ADDRESS=0.0.0.0:4318When set, the receiver binds to that address at startup (instead of a random loopback port) and accepts OTLP/HTTP POST /v1/traces (protobuf) from any reachable producer. Received spans are filtered to the session's trace ID and appended to the local trace file, where they appear in infer traces nested under the originating execute_tool span.
Defensive limits: 4 MiB per request body, 5000 spans per session. The receiver always responds 200 - receiver failures never fail the tool.
Example: OpenTelemetry Collector
To receive all three signals from the CLI, configure an OpenTelemetry Collector with OTLP/HTTP receivers:
receivers:
otlp:
protocols:
http:
endpoint: 0.0.0.0:4318
exporters:
debug:
verbosity: detailed
otlp/jaeger:
endpoint: jaeger-collector:4317
tls:
insecure: true
service:
pipelines:
metrics:
receivers: [otlp]
exporters: [debug]
traces:
receivers: [otlp]
processors: [batch]
exporters: [otlp/jaeger, debug]
logs:
receivers: [otlp]
exporters: [debug]
processors:
batch:
send_batch_size: 1024
timeout: 5sThen run the CLI with the endpoint set:
export OTEL_EXPORTER_OTLP_ENDPOINT=http://otel-collector:4318
infer headless "Analyze the codebase"Trace context propagation to subprocesses
When telemetry is enabled (telemetry.enabled: true, the default), the CLI propagates its active trace context into the processes it spawns, so spans emitted by tools and subagents stitch into the same trace as the CLI session. Propagation is built entirely on open standards - W3C Trace Context and W3C Baggage on the way out, plain OTLP spans on the way back - so any OpenTelemetry-instrumented tool participates with zero custom code. When telemetry is disabled, none of these variables are set.
Environment variables set for child processes
The Bash tool (and therefore skill scripts) and headless subagents receive:
| Variable | Standard | Contents |
|---|---|---|
TRACEPARENT / TRACESTATE | W3C Trace Context | The current trace, parented on the active execute_tool span, so child spans nest under it. |
BAGGAGE | W3C Baggage | infer.session.id and infer.tool.call.id, for correlating child spans back to the call. |
Where child spans are sent (OTLP sink selection)
The CLI selects the span sink automatically so children report to the same place the CLI does:
- An OTLP endpoint is configured (
telemetry.otlp.endpoint, or the standardOTEL_EXPORTER_OTLP_ENDPOINT) - the endpoint and its headers are passed through to the child environment asOTEL_EXPORTER_OTLP_ENDPOINT/OTEL_EXPORTER_OTLP_HEADERS, so child spans land in the same remote backend and stitch into the trace there. - No endpoint configured (the default) - the CLI runs an ephemeral localhost OTLP/HTTP receiver for the session and points children at it. Received spans are persisted into the local per-session trace store and appear in
infer traces//traces, nested under the tool span that spawned them.
Either way you get one coherent trace - local by default, in your backend when you have one.
Example: any OTel-instrumented command
Because propagation is standard OTLP, an off-the-shelf tool such as otel-cli emits a correctly parented child span with no wiring. Run it inside a Bash tool call:
otel-cli exec --name "go test" -- go test ./...The resulting go test span shows up as a child of execute_tool Bash in infer traces:
session (standard, success) 12.4s
`-- execute_tool Bash 9.8s
`-- go test 9.6sThe same holds for a program instrumented with an OpenTelemetry SDK: as long as it reads the standard OTEL_EXPORTER_OTLP_ENDPOINT and TRACEPARENT variables - which every conformant SDK does - its spans nest under the execute_tool span automatically, with no CLI-specific integration.
Subagents: one cross-process trace
infer headless honors an inherited TRACEPARENT, so a headless subagent's own session span nests under the caller's execute_tool Agent span - a single trace that spans both processes:
session (standard, success) 152ms
|-- execute_tool Agent 89ms
| `-- session (readonly, success) 52ms
| `-- chat openai/gpt-4o 13msRemote A2A agents
Outbound A2A requests carry traceparent, tracestate, and baggage HTTP headers, so a remote agent can continue the trace. This is context-out only: remote spans stitch into your trace through a shared OTLP collector that both sides export to - they are not written back into the local per-session trace store. Point both the CLI and the remote agent at the same telemetry.otlp.endpoint to see the full cross-service trace.
A2A traces example
The examples/a2a-traces directory in the CLI repository demonstrates end-to-end distributed telemetry between the CLI and a remote A2A agent (mock-agent), with an OpenTelemetry Collector in the middle. It exercises both telemetry models:
- Push (OTLP): the CLI exports its traces and metrics to the collector; the mock-agent exports its traces to the collector.
- Pull (Prometheus): the collector scrapes the mock-agent's
:9090/metricsendpoint.
The CLI injects W3C trace context into every outgoing A2A request, so the mock-agent's spans share the CLI's trace ID and parent under the CLI's execute_tool span. The collector fans all traces back to the CLI's local OTLP receiver (via INFER_TELEMETRY_RECEIVER_ADDRESS: 0.0.0.0:4318), so infer traces shows the full distributed trace, including the mock-agent's spans:
session (standard, success) 152ms
|-- chat deepseek/deepseek-v4-flash 5ms
|-- execute_tool A2A_SubmitTask 89ms
| `-- a2a.request [mock-agent] 52ms
`-- chat deepseek/deepseek-v4-flash 12msIngested foreign spans are labeled name [service] from the producer's service.name resource attribute - for example a2a.request [mock-agent].
To run the example:
cd examples/a2a-traces
cp .env.example .env # set at least one provider API key
docker compose up -d --build
docker compose attach cli
# In the CLI, ask: "Ask the mock-agent to summarize the current project."
# Detach with ctrl-p ctrl-q, then:
docker compose exec cli infer tracesSee the example README for full configuration notes and troubleshooting.
Limitations
- Interactive (tmux-pane) subagents are not stitched into the caller's trace - only headless subagents inherit the context. Use
mode: headlesswhen you need a subagent's spans in the same trace. - Detached background shells that outlive the CLI lose late spans - once the CLI (and its ephemeral receiver) exits, spans emitted afterward have nowhere local to land. Configure a persistent
telemetry.otlp.endpointfor long-running detached work.
Viewing telemetry data
Local files - the JSONL files under
~/.infer/telemetry/are plain text and can be inspected with any JSON tool:bash# List telemetry files ls ~/.infer/telemetry/ # View metrics for a session cat ~/.infer/telemetry/\<session-id\>.jsonl | jq . # View traces for a session cat ~/.infer/telemetry/\<session-id\>-traces.jsonl | jq .infer stats- aggregates metrics from local files into a summary of tool outcomes, token usage, and sessions. Trace and log files are excluded from the aggregate. The per-tool Avg column renders microseconds (for example432us) when the mean duration is below 1ms and milliseconds otherwise, so fast tools no longer collapse to0ms. Ininfer stats --format jsonthe matchingavg_msfield can be fractional (for example0.432) rather than an integer. The same applies toinfer insightsand the/statsshortcut.infer traces- renders the span tree of a session from its local trace file. See Viewing traces.infer insights- analyzes past sessions for repeatable workflows and recurring tool failures. The verbatim error and log samples it collects are redacted before the model call - PEM private-key blocks, GitHub token shapes, and the values of provider-secret env vars (plain and JSON-escaped) are replaced with[redacted], in the digest and in the saved report alike. See Insights Shortcut.Remote backend - when OTLP export is configured, data appears in your collector's configured backend (Jaeger for traces, your metrics store, your log aggregator).
Viewing traces
infer traces renders the span tree of a session - the session root span, its LLM turns, and the tool calls nested under each turn - with a duration next to every span. It reads the local per-session trace file (~/.infer/telemetry/\<session-id\>-traces.jsonl) directly, so it works fully offline with no OTLP collector required.
# Render the most recent session's span tree
infer traces
# Render a specific session by ID
infer traces abc-123-def
# List the sessions that have trace files
infer traces --list
# Emit the tree as structured JSON
infer traces --format jsonWith no session ID, infer traces picks the most recent session. The default output is an indented span tree:
session (standard, success) 38.8s
|-- chat ollama_cloud/deepseek-v4-flash 3.4s
|-- execute_tool Read 162us
`-- chat ollama_cloud/deepseek-v4-flash 8.1sSpans that finished with an error status (or carry an error.type attribute) are flagged inline with an [error: ...] marker showing the error type - for example [error: context_deadline_exceeded].
Spans ingested from subprocesses appear in the same tree, nested under the execute_tool span that spawned them - Bash tool commands (including otel-cli and OpenTelemetry-SDK-instrumented programs) and headless subagents. See Trace context propagation to subprocesses for how the context is passed and where child spans are collected.
Spans ingested from external producers (for example a remote A2A agent or an OpenTelemetry Collector) are labeled with the producer's service.name resource attribute: name [service]. For example, a span named a2a.request from a service called mock-agent renders as a2a.request [mock-agent] in the tree. This makes it easy to identify which service produced each span in a distributed trace.
--format json emits the same tree as structured output - each node carries the span name, start time, duration in fractional milliseconds, attributes, and an array of child spans - ready to pipe into jq or a custom viewer:
{
"name": "session",
"start_time": "2026-07-14T09:12:03.145Z",
"duration_ms": 38800.0,
"attributes": {
"infer.execution.mode": "standard",
"infer.run.outcome": "success"
},
"children": [
{
"name": "chat ollama_cloud/deepseek-v4-flash",
"start_time": "2026-07-14T09:12:03.150Z",
"duration_ms": 3400.0,
"attributes": { "gen_ai.request.model": "deepseek-v4-flash" },
"children": []
}
]
}The same view is available inside a chat session with the /traces shortcut.
Shortcuts
The CLI provides built-in commands, YAML shortcuts, and Agent Skills - all reachable by typing / in chat.
Autocomplete kinds
Typing / opens the autocomplete list, and every row carries a kind label telling you what it is and where it comes from (the same labels the desktop composer uses):
| Label | What it is | Where it lives |
|---|---|---|
command | Built-in shortcut compiled into the CLI - views and wizards it implements itself | The binary - see Built-in Commands |
shortcut | YAML shortcut, including the ones infer init seeds | .infer/shortcuts/*.yaml - see YAML Shortcuts |
skill | Installed Agent Skill - /<name> activates it for the turn | ./.infer/skills/ or ~/.infer/skills/ |
remote skill | Catalog skill not installed yet - /<name> downloads it first (confirmed in chat) | The skills catalog |
A command always wins over a YAML shortcut of the same name, so a stale shortcut file can never shadow a built-in view. See Shortcut precedence.
Built-in Commands
These show as command in autocomplete:
| Shortcut | Description | Example |
|---|---|---|
/model [name] [msg] | Switch the active model, or run one message with another model | /model deepseek/deepseek-v4-flash |
/init | Fill the input with the AGENTS.md project-analysis prompt | /init |
/install-opentask | Ask the agent to install or update the infer-action workflow and open a PR | /install-opentask |
/agents | View every configured agent - local Markdown presets and remote A2A agents - with their state | /agents |
/tasks | View background work (A2A tasks, shells, subagents) with live status and captured output | /tasks |
/tools | View a filterable list of tools available in the current agent mode, including MCP tools | /tools |
/stats | Summarize the session's token usage, tool outcomes, and cost (mirrors infer stats) | /stats |
/traces [id] | Render a session's trace span tree offline (mirrors infer traces) | /traces, /traces abc-123-def |
/voice [seconds] | Record the mic and transcribe to the input field (requires speech-to-text) | /voice, /voice 8 |
YAML Shortcuts
infer init seeds these into ~/.infer/shortcuts/, so they show as shortcut in autocomplete. They are ordinary custom shortcuts - editable, removable, and replaceable with your own:
| Shortcut | Description | Example |
|---|---|---|
/git <cmd> | Git operations | /git status, /git commit, /git push |
/scm <cmd> | GitHub operations | /scm issues, /scm issue 123 |
/mcp <cmd> | Manage MCP servers | /mcp list, /mcp add |
/skills <cmd> | Manage Agent Skills | /skills list, /skills install <url> |
/shells | List running and recent background shells | /shells |
/export | Export the current conversation to Markdown | /export |
/env | Generate a .env.example with the provider API keys | /env |
/insights [since] | Analyze past sessions for repeatable workflows and recurring tool failures (secrets redacted before the model call and before the report is saved) | /insights, /insights 7d |
/reset [arg] | Wipe all local runtime state on this machine and start a fresh session | /reset, /reset confirm |
Git Shortcuts
# Execute git commands
/git status
/git branch
# AI-generated commit message
/git commit
# Push to remote
/git push origin mainSCM (GitHub) Shortcuts
# List GitHub issues
# Runs: gh issue list --json ... --limit 20
/scm issues
# View issue details
# Runs: gh issue view 123 --json ...
/scm issue 123The seeded scm.yaml defines only these two subcommands. For pull-request work, ask the agent directly - gh pr create is a write, so it falls through to approval - or add your own subcommand to the shortcut file.
Telemetry Shortcuts
/stats and /traces surface the CLI's local telemetry from inside a chat session, mirroring the infer stats and infer traces commands.
/statssummarizes the session's token usage, tool outcomes, and cost. Its markdown and vertical views use the same duration formatting asinfer stats: the Avg column shows microseconds (for example432us) below 1ms and milliseconds otherwise./traces [session-id]renders the span tree of a session - the session root, its LLM turns, and the tool calls under each turn - with per-span durations. With no argument it shows the most recent session.
session (standard, success) 38.8s
|-- chat ollama_cloud/deepseek-v4-flash 3.4s
|-- execute_tool Read 162us
`-- chat ollama_cloud/deepseek-v4-flash 8.1sBoth read the local files under ~/.infer/telemetry/, so they work fully offline. See Viewing traces for the infer traces command, its --list and --format json flags, and how error spans are marked.
Insights Shortcut
/insights [since] (and the equivalent infer insights) analyzes past sessions for repeatable workflows worth turning into a skill and for recurring tool failures. The optional since argument limits the window (for example /insights 7d). It reads local sessions and telemetry, builds a digest, sends that digest to the configured model, and writes the report to ~/.infer/insights/.
Secrets are redacted locally before any network call. The digest is masked before it reaches the analysis model, and the saved report is masked before it is written to disk. Redaction covers:
- PEM private-key blocks (always on).
- GitHub token shapes:
ghp_,gho_,ghu_,ghs_,ghr_, andgithub_pat_. - The values of provider-secret environment variables set in the environment, both plain and JSON-escaped (so a key quoted inside a captured JSON payload is masked too).
Each match is replaced with [redacted]. Because the masking runs before the model call and before the file write, pointing insights at collected sessions does not ship verbatim API keys or private keys to the provider, and the report you keep or share carries the same guarantee. Redaction targets known secret shapes and configured provider-secret values - it is not a substitute for keeping unrelated sensitive data out of session transcripts.
If conversation storage is disabled or no model is configured, the analysis is skipped with a notice.
Reset Shortcut
/reset wipes the CLI's local runtime state on the whole machine - not just the current project - and starts a fresh session. It is a two-step shortcut: a bare /reset only previews, and /reset confirm performs the wipe.
# Preview what would be deleted (deletes nothing)
/reset
# Analyze the sessions first, then preview
/reset insights
# Actually delete everything listed in the preview
/reset confirmBoth steps end with a disk-space total: the preview ends with Total reclaimable space: <size>, and /reset confirm ends with Total reclaimed space: <size> after the wipe. The confirm total counts only targets that were actually removed, so anything the wipe failed to delete is excluded. Sizes are formatted like docker system prune - decimal units and 4 significant digits (for example 1.653GB, 4.096kB). The same totals appear when running infer reset and infer reset confirm directly.
/reset confirm only deletes after a preview was shown in the same session (within the last 5 minutes). A cold /reset confirm prints the preview instead, so a tab-completed confirm cannot wipe anything by accident.
Deleted, for every project under ~/.infer/projects/: conversations, plans, scratch dirs, artifacts, history, backups, exports, logs, telemetry, schedules, pid/lock files, and the userspace tmp tree (generated speech, retained recordings, channel media). With the SQLite backend the conversation database and its WAL sidecars go too.
Preserved: configuration (config.yaml, custom shortcuts, skills, projects.yaml), the avatar library under ~/.infer/avatars/, and saved insights reports under ~/.infer/insights/ (written redacted). Directories you pointed outside ~/.infer (for example a text_to_speech.output_dir of /data/tts) are left alone.
Remote conversation stores are skipped. If storage.type is postgres, redis, or d1, /reset clears local state only and prints a notice that the remote store was left untouched - it is not an error.
/reset insights runs the /insights analysis (repeatable workflows worth turning into a skill, recurring tool failures) before the preview, so you can capture what past sessions were worth learning from before deleting them. The report is written to ~/.infer/insights/ - with secrets redacted, as for any insights run - and survives the reset. If conversation storage is disabled or no model is configured, the analysis is skipped with a notice and the preview is still shown.
Voice Shortcut
The /voice shortcut records audio from your microphone, transcribes it locally with whisper.cpp, and places the text into the input field - ready to review and send. It is disabled by default and only appears when speech_to_text.enabled is true.
# Record until you go quiet (or the max cap), then transcribe
/voice
# Record for at most 8 seconds
/voice 8Recording stops automatically a couple of seconds after you stop speaking (speech_to_text.silence_timeout), at the max_recording_seconds cap, or at the per-call override. See Speech-to-Text for prerequisites, configuration, and model selection.
GitHub Action Setup
/install-opentask [owner/repo] [extra context...] asks the agent to install or update the infer-action GitHub Action workflow in a repository and open a pull request with it. It is not a wizard - the shortcut submits a task into the conversation, so the run streams like any other turn and every write goes through normal tool approval.
For a full reference of
infer-actioninputs, outputs, and workflow recipes (PR review, scheduled summaries, release notes), see the GitHub Action documentation.
With no argument it targets the repository of the current checkout; pass owner/repo to target another one. Anything after the repository is handed to the agent as workflow configuration, for example a model or a timeout to use.
infer chat
> /install-opentask
> /install-opentask my-org/my-service model: deepseek/deepseek-v4-flashWhat the agent does:
- Reads the
opentaskcatalog skill, which carries the canonicalinfer-actionworkflow and its example workflows - it works from the skill rather than fetching docs over the network. - Adds a git worktree under
/tmpwhen the target is the current checkout, or shallow-clones the target repository, so your working branch is never touched. - Checks out the fixed branch
infer/install-github-action, on top of the remote branch if one already exists. - Creates or updates
.github/workflows/tasks.yml, following both the skill's canonical workflow and the repository's existing CI conventions. - Summarizes the diff, then commits, pushes, and opens or updates the pull request - asking for your approval on the push and on the PR.
Re-running the shortcut updates that same branch and pull request instead of opening a duplicate.
Prerequisites: the GitHub CLI installed and authenticated (gh auth login), with write access to the target repository.
Repository secrets: the generated workflow expects these in the target repository, or in its organization:
INFER_APP_ID- GitHub App ID, when authenticating as a GitHub App instead of with the defaultGITHUB_TOKENINFER_APP_PRIVATE_KEY- the App's private key (.pemcontents)- Provider API keys (
ANTHROPIC_API_KEY,OPENAI_API_KEY,DEEPSEEK_API_KEY, and so on)
For organization repositories the CLI checks whether the INFER_APP_* org secrets already exist, so an App registered once can be reused across repositories. See Secrets and least-privilege for why an App token beats GITHUB_TOKEN.
Usage in issues: once the workflow is merged, mention @infer in any issue or issue comment to activate the agent:
@infer Please analyze this bug and suggest a fixFor more information on the infer-action GitHub Action, see the GitHub Action documentation or the upstream repository.
Custom Shortcuts
Create YAML files in .infer/shortcuts/ directory. Shortcuts support three types:
1. Simple Commands
Execute a single command:
# .infer/shortcuts/simple.yaml
shortcuts:
- name: hello
description: 'Say hello'
command: echo
args:
- 'Hello from Inference Gateway!'2. Shortcuts with Subcommands
Group related commands under a parent shortcut:
# .infer/shortcuts/dev.yaml
shortcuts:
- name: dev
description: 'Development operations'
command: bash
subcommands:
- name: test
description: 'Run all tests'
command: bash
args:
- -c
- 'go test ./...'
- name: build
description: 'Build the project'
command: bash
args:
- -c
- 'go build -o app .'A subcommand that declares its own command: runs that command with its args verbatim - /dev test runs bash -c 'go test ./...'. A subcommand without its own command: reuses the parent's command and gets its name (then its args) appended to the parent's args, so spell out the full invocation in every subcommand.
Usage: /dev test, /dev build
3. AI-Powered Snippets
Use LLM to generate dynamic content based on command output. The snippet.prompt can reference JSON fields from command output using {fieldName} placeholders, and snippet.template uses {llm} for the AI-generated response:
# .infer/shortcuts/ai-commit.yaml
shortcuts:
- name: ai-commit
description: 'AI-generated commit message'
command: bash
args:
- -c
- |
diff=$(git diff --cached)
jq -n --arg diff "$diff" '{"diff": $diff}'
snippet:
prompt: "Generate commit message for:\n{diff}"
template: '!git commit -m "{llm}"'The command must output JSON. Fields are accessible in the prompt template via {fieldName} syntax. The LLM response is accessible via {llm} in the template.
Name collisions with built-ins
A built-in command always wins over a custom YAML shortcut of the same name - the custom one is skipped with a warning naming the file, so a stale shortcut file can never shadow a built-in view such as /agents. Rename the custom shortcut to reach it again.
Advanced Features
Cost Tracking
Real-time token usage and cost calculation displayed in the status bar.
Features:
- Per-model pricing calculation
- Cumulative session costs
- Input and output token tracking
- Status bar indicator
- Custom pricing support
View Costs:
# Costs displayed in status bar during chat
infer chat
# Status bar shows model and current cost
# Inspect a saved conversation's entries (per-entry metadata, including model)
infer conversations show <session-id>
# Same, as one JSON object per line for piping into jq
infer conversations show <session-id> --format json | jq .Pricing Configuration
Pricing lives under the pricing key in .infer/config.yaml. The custom_prices map overrides or adds entries to the built-in per-model pricing table, keyed by model name.
# .infer/config.yaml
pricing:
enabled: true
currency: 'USD'
custom_prices:
'ollama_cloud/deepseek-v4-pro':
input_price_per_mtoken: 0.0
output_price_per_mtoken: 0.0
requires_pro: true| Field | Type | Description |
|---|---|---|
input_price_per_mtoken | number | Cost per 1M prompt (input) tokens, in currency. |
output_price_per_mtoken | number | Cost per 1M completion (output) tokens, in currency. |
requires_pro | boolean | Marks the model as gated behind a paid Pro subscription. Defaults to false. |
Subscription-gated models are billed at zero session cost regardless of the rates on the entry. Use non-zero rates with requires_pro: true to keep an informational rate in the picker label while a flat-fee subscription covers the actual billing:
pricing:
enabled: true
custom_prices:
ollama_cloud/glm-5.3-flash:
input_price_per_mtoken: 0.15
output_price_per_mtoken: 0.50
requires_pro: true # informational rate only, billed as flat-fee subscriptionOverride caveat: a
custom_pricesentry fully replaces the default for that model - it is not merged field by field. Omittingrequires_proin a custom override therefore resets it tofalse, even when the model is flagged Pro by default. Setrequires_pro: trueexplicitly when overriding the pricing of a Pro model.
Model Categories (Free / Pay-as-you-go / Subscription)
The model picker's pricing tab row - [1] All, [2] Free, [3] Pay-as-you-go, [4] Subscription - groups models into three disjoint categories:
| Category | Meaning |
|---|---|
| Free | No per-token cost and not subscription-gated. |
| Pay-as-you-go | Billed per token. |
| Subscription | Gated behind a paid subscription rather than per-token billing. |
Subscription is an axis orthogonal to price: an Ollama Cloud model is not metered per token but is not free, so it is labelled pro subscription rather than free. A subscription model may still carry a per-token rate - the gateway keeps the provider's pay-as-you-go rate for reference - in which case the picker shows the rate with a subscription suffix. The marker appears both in the picker rows and in /model autocomplete descriptions:
ollama_cloud/deepseek-v4-pro (1M, pro subscription)
ollama_cloud/glm-5.3-flash (1M, $0.15/$0.50 per MTok, subscription)
deepseek/deepseek-v4-flash (1M, $1.74/$3.48 per MTok)Subscription models cost zero per session. The session cost line and the status-bar cost report $0.00 for any subscription-gated model, even when the picker label shows a rate - the rate is informational and the flat-fee subscription covers usage. This applies whether the model is flagged by the gateway (pricing.subscription: true) or by a custom_prices entry with requires_pro: true, as in the example above.
The classification comes from the gateway's pricing metadata - the pricing.subscription flag reported per model - so it tracks the gateway catalog with no CLI-side list to maintain. A custom_prices entry still wins: set requires_pro: true to gate a model the gateway does not flag, or false to un-gate one (remembering the override caveat that an entry fully replaces the default).
The gateway flag replaced the previous hardcoded list of Pro models. Subscription models with per-token rates bill at zero cost while keeping their published rates.
Model Thinking Visualization
Collapsible thinking blocks for models that support thinking (Claude, o1, etc.).
Features:
- Collapsible blocks with first sentence preview
- Ctrl+K keyboard shortcut to toggle
- Theme-aware styling
- Performance optimization (long thinking blocks collapsed by default)
Usage:
infer chat
# Ask complex question requiring reasoning
> "Design a scalable microservices architecture for e-commerce"
# Model's thinking process displayed in collapsible blocks
# Press Ctrl+K to expand/collapse thinkingConversation Management
Storage Backends:
- JSONL (default): append-only files under
~/.infer/projects/<project-slug>/conversations/ - SQLite: one shared database at
~/.infer/conversations.db - PostgreSQL: Shared team database
- Redis: High-performance caching
- Cloudflare D1: External SQLite over Cloudflare's HTTP query API
- In-memory: Temporary sessions
Features:
- Automatic conversation history
- AI-generated titles (batch: 10 messages)
- Token optimization with compaction
- Backend-agnostic inspection via the storage layer (works the same across
jsonl,sqlite,postgres,redis,d1, andmemory)
Where conversations live:
Conversations are stored in your home directory, grouped per project - nothing conversation-related is written to the project directory:
- jsonl:
~/.infer/projects/<project-slug>/conversations/<id>.jsonl, where<project-slug>is the absolute project path with the separators replaced (/home/alice/repobecomes-home-alice-repo). - sqlite: one shared database at
~/.infer/conversations.db, with each row carrying aprojectcolumn. - postgres, redis, d1: one shared store, grouped by the same
projectfield in the conversation metadata.
Every session records the project it ran in (its absolute working directory), and listings scope to the current project by default: the /conversations TUI picker shows only this project's conversations, and so does infer conversations list - pass --all-projects to list every project's conversations.
An explicit storage.jsonl.path or storage.sqlite.path always overrides these defaults, and such a store only ever lists itself.
No migration. Stores written by older versions under the project's
.infer/conversationsor.infer/conversations.dbare orphaned by design - the CLI no longer reads or writes them. Delete them manually when you no longer need them. The generated.infer/.gitignoreno longer ignoresconversations/conversations.db*.
Subcommands:
list: List saved conversations with metadata (id, title, message/request counts, tokens, cost). Scoped to the current project;--all-projectslists every project's conversations.show <session-id>: Print a single conversation's entries in chronological order (role, timestamp, content, andtool_call_idfor tool results).delete <session-id>: Remove a conversation from the storage backend. Runs non-interactively with no confirmation prompt. Unknown or missing session id exits non-zero with an error from the storage layer.
delete notes:
- Runs non-interactively - no confirmation prompt, so scripts and the desktop app can shell out to it.
- Unknown or missing session id exits non-zero with a clear error from the storage layer (e.g.
conversation not found: <id>). - Session id resolution follows the same rules as
show(see below).
show flags:
--include-hidden: Include entries persisted as hidden - system reminders, plan-approval prompts, drained background-task results, and the synthetic verify message injected byinfer headless. Off by default.--format text|json:text(default) is human-readable;jsonemits one JSON object per line (NDJSON), matching theinfer headlessstdout shape for piping intojqor log scrapers.
Session id resolution:
<session-id> is resolved the same way as infer headless --session-id and infer chat --session-id: a literal UUID is used as-is, while any other value is treated as a session group key and resolved to that group's current session id (registering the group if it is new). This means you can show a conversation by group name such as channel-telegram-12345.
Commands:
# List this project's conversations to find a session id
infer conversations list
# List conversations from every project
infer conversations list --all-projects
# Show a conversation's entries (hidden entries omitted by default)
infer conversations show 12345678-1234-1234-1234-123456789abc
# Show by session group name (for example a channel group key)
infer conversations show channel-telegram-12345
# Include hidden entries such as system reminders
infer conversations show <session-id> --include-hidden
# One JSON object per line for piping into jq
infer conversations show <session-id> --format json | jq .
# Delete a conversation by literal UUID
infer conversations delete 12345678-1234-1234-1234-123456789abc
# Delete by session group key (for example a channel group key)
infer conversations delete channel-telegram-12345Cloudflare D1 backend
Cloudflare D1 is an external, SQLite-compatible store the CLI writes to over D1's HTTP query API. It is built for ephemeral CI runners (for example a headless infer headless run on GitHub Actions): unlike sqlite, jsonl, and memory - which live on the runner's disk and are wiped on recycle - D1 persists off-runner and stays readable by the gateway through its native binding. Unlike postgres and redis, it needs no wire-protocol connection, just HTTPS.
Set storage.type: d1 and configure the storage.d1 block:
storage:
enabled: true
type: d1
d1:
account_id: '<cloudflare-account-id>'
database_id: '<d1-database-id>'
api_token: '<api-token-with-d1-edit>' # inject via INFER_STORAGE_D1_API_TOKEN
base_url: 'https://api.cloudflare.com/client/v4' # optionalEnvironment variables:
| Variable | Description |
|---|---|
INFER_STORAGE_D1_ACCOUNT_ID | Cloudflare account id that owns the D1 database. |
INFER_STORAGE_D1_DATABASE_ID | Target D1 database id. |
INFER_STORAGE_D1_API_TOKEN | API token with D1 edit permission. Secret - inject, never commit. |
INFER_STORAGE_D1_BASE_URL | Optional API base URL. Defaults to https://api.cloudflare.com/client/v4. |
Notes:
- No manual migration. Like
jsonl,redis, andmemory, D1 creates its schema automatically on first connect - there is no separate migration step to run. - Schema parity. The D1 driver runs the SQLite migrations verbatim over HTTP, so the
conversationsandsession_groupstables stay byte-for-byte compatible with the SQLite backend - either side can initialise the database. - UTC timestamps. Timestamps are stored as UTC RFC3339 so
ORDER BY updated_at DESCsorts stably across runners in any timezone and external reads stay unambiguous. - Secret handling.
api_tokenfollows the existing plaintext-config + env-override convention (like the Postgres password) and is never logged - inject it viaINFER_STORAGE_D1_API_TOKEN.
Persistent Memory
The Memory tool gives the agent durable, cross-session memory: facts it learns in one session survive into the next. Each fact is a single Markdown fact-file (with YAML frontmatter) stored under a configurable directory - ~/.infer/memory by default - and catalogued by a MEMORY.md index. That index is injected into context at the start of every session, so the agent always knows what it has recorded; it then reads or writes individual facts on demand. A default system reminder (memory-consult) nudges it to consult and keep memory current. Memory is enabled by default.
Per-project fact organization
Facts are organized by project to keep memory relevant and scoped:
- Project facts live in per-project subdirectories:
<project-slug>/<slug>.md(e.g.inference-gateway-cli/build-commands.md). - Global facts stay at the memory directory root (e.g.
user-preferences.md). - Legacy flat files (pre-existing fact-files at the root) keep working as global facts - no migration needed.
The project is detected automatically:
- Git remote origin is resolved to
org/repo(e.g.inference-gateway/cli). - If no git remote is found, the current working directory basename is used.
- If neither is available, the fact is stored as global.
Fact frontmatter
Each fact-file includes YAML frontmatter that records its metadata:
---
name: build-commands
description: how to build the CLI
metadata:
type: project
project: inference-gateway/cli
session: channel-telegram-12345
---| Field | Description |
|---|---|
name | Short slug identifying the fact (e.g. build-commands). |
description | One-line summary shown in the MEMORY.md index. |
metadata.type | Fact type: user, feedback, project, or reference. |
metadata.project | Human-readable org/repo that owns this fact. Omitted for global facts. |
metadata.session | The session ID that last wrote the fact (e.g. channel-telegram-12345). |
Index filtering
The MEMORY.md index injected at session start is filtered to show only:
- Entries for the current project (facts under the detected project slug).
- Global entries (facts at the memory directory root).
A single summary line at the bottom names other projects that have facts but are not shown. The full unfiltered index (all projects) is always available by calling Memory read with no name.
Not the same as
storage.type: memory. This is the agent's knowledge memory - durable facts on disk under~/.infer/memory. Thememoryconversation storage backend is unrelated: an in-RAM transcript store that is wiped when the process exits.
The Memory tool
Memory is a Workflow tool whose operation parameter selects one of three actions:
| Operation | Parameters | Effect |
|---|---|---|
read | name (optional) | With no name, returns the MEMORY.md index; with a name, that fact-file. |
write | name, description, type, content (all required), project (optional) | Creates or updates a fact-file and its index entry. |
delete | name (required) | Removes a fact-file and its index entry. |
name is a short slug (for example build-commands) for global facts, or project/slug (for example inference-gateway-cli/build-commands) for project facts - exactly as shown in the MEMORY.md index. description is the one-line summary shown in the MEMORY.md index, content is the Markdown fact body, and type is one of user, feedback, project, or reference.
The optional project argument on write controls where the fact is filed:
project value | Behavior |
|---|---|
| (omitted) | Defaults by type: user facts are global; feedback/project/reference go under the detected project. |
global | Forces the fact to be stored at the memory directory root (a global fact). |
org/repo | Files the fact under another project's subdirectory (e.g. inference-gateway/cli). |
Configuration (memory.yaml)
Runtime knobs live in memory.yaml (seeded by infer init; the in-code defaults apply when the file is absent):
# .infer/memory.yaml (or ~/.infer/memory.yaml)
enabled: true
dir: '' # "" => ~/.infer/memory
max_chars: 2000 # cap on the MEMORY.md index injected into context (truncates at line boundary)
max_entry_chars: 2000 # per-fact write cap (0 = default)| Key | Default | Environment variable | Description |
|---|---|---|---|
enabled | true | INFER_MEMORY_ENABLED | Master switch - registers the Memory tool and the index injection. |
dir | ~/.infer/memory | INFER_MEMORY_DIR | Directory holding the fact-files and MEMORY.md. "" = default. |
max_chars | 2000 | INFER_MEMORY_MAX_CHARS | Upper bound on the MEMORY.md index injected at session start. Truncation respects line boundaries. |
max_entry_chars | 2000 | INFER_MEMORY_MAX_ENTRY_CHARS | Per-fact character cap on write content. 0 means use the default. |
The memory directory is local by default. To back it with a git remote - pull on run start, commit and push on change - configure a Sync backend.
Sync backend
By default the memory directory lives on a single machine (backend.type: local, a pure no-op). Point the backend at a git remote to share one memory across machines, CI runners, channels, and scheduled runs: the CLI pulls on run start and commits + pushes on change.
# .infer/memory.yaml (or ~/.infer/memory.yaml)
enabled: true
dir: '' # "" => ~/.infer/memory
max_chars: 4000
backend:
type: local # local (default) | git
git:
repo: 'git@github.com:my-org/agent-memory.git'
branch: main
commit_message: 'chore(memory): sync'
timeout: 60 # seconds per git op
sync:
on_start: pull # pull (default) | off
on_finish: push # push (default) | offtype: local is the default and a pure no-op - existing users see no change.
| Key | Default | Environment variable | Notes |
|---|---|---|---|
backend.type | local | INFER_MEMORY_BACKEND_TYPE | local (no-op) or git. |
backend.git.repo | '' | INFER_MEMORY_BACKEND_GIT_REPO | Remote URL. Required when type: git. |
backend.git.branch | main | INFER_MEMORY_BACKEND_GIT_BRANCH | Branch to track. |
backend.git.commit_message | chore(memory): sync | INFER_MEMORY_BACKEND_GIT_COMMIT_MESSAGE | Deterministic, non-LLM commit message. |
backend.git.timeout | 60 | INFER_MEMORY_BACKEND_GIT_TIMEOUT | Seconds per git op (prevents credential hangs). |
backend.git.sync.on_start | pull | INFER_MEMORY_BACKEND_GIT_SYNC_ON_START | pull or off. |
backend.git.sync.on_finish | push | INFER_MEMORY_BACKEND_GIT_SYNC_ON_FINISH | push or off. |
Validation: when memory is enabled and type: git, repo is required; on_start must be pull or off, and on_finish must be push or off.
How it syncs
- On run start (SyncIn). Clones the repo when the memory dir is missing, otherwise fast-forward / rebase pulls. An
ls-remoteprobe decides clone vs. init-in-place - an empty remote, or a pre-existing local memory dir, is initialized in place instead of cloned. - On change (SyncOut). Commits and pushes only when
git status --porcelainreports changes, through a bounded push -> pull-rebase -> retry loop. A per-hostflockserializes concurrent runs (channels / scheduler / heartbeat) so they do not clobber each other. - Where the push happens. In chat, the
Memorytool pushes on each write / delete - not a post-session hook, which would commit-storm once per message. In headless mode, it pulls on start and pushes once at run finish. Either way it works across channel, scheduler, and heartbeat subprocess runs. - Best-effort, never fatal. A failed clone / pull / push is logged and the run continues - sync never aborts the agent run.
Authentication
Sync uses the ambient git credential chain - ssh-agent, a git credential helper, or GIT_* environment variables. The backend injects no ssh key or env override of its own, so pick whichever your environment already uses:
- SSH (preferred) - a
git@github.com:...remote plus a loaded ssh-agent key. gh auth- the GitHub CLI credential helper forhttps://remotes.- Token in URL - works, but the CLI logs a warning, because credentials embedded in the remote URL persist in
.git/config. Prefer SSH.
The per-op timeout (default 60 seconds) keeps an interactive credential prompt from hanging a run.
Disabling memory
Turn it off in memory.yaml:
# .infer/memory.yaml (or ~/.infer/memory.yaml)
enabled: falseor via the environment, without touching config:
export INFER_MEMORY_ENABLED=falseWhen disabled, the Memory tool is not registered, no MEMORY.md index is injected, and the memory-consult reminder is pruned automatically.
MCP Integration
Connect to Model Context Protocol servers for extended capabilities. MCP provides stateless tool execution for external services like databases, file systems, and APIs.
Setup:
Seed the userspace baseline, which includes ~/.infer/mcp.yaml:
infer initConfigure MCP servers in ~/.infer/mcp.yaml (or add --project to the infer mcp commands below to write a project-level .infer/mcp.yaml instead):
enabled: true
connection_timeout: 30
discovery_timeout: 30
liveness_probe_enabled: true
liveness_probe_interval: 10
servers:
# Auto-start MCP server in container (recommended)
- name: 'demo-server'
enabled: true
run: true
oci: 'mcp-demo-server:latest'
description: 'Demo MCP server'
# Connect to external MCP server - the endpoint is spelled out field by field,
# there is no `url` key (an entry without them resolves to http://localhost/mcp)
- name: 'filesystem'
scheme: 'http'
host: 'localhost'
port: 3000
path: '/sse'
enabled: true
description: 'File system operations'
exclude_tools:
- 'delete_file'infer mcp add <name> <url> splits a URL into those fields for you, so you rarely need to write them by hand.
CLI Commands:
# Add auto-start MCP server
infer mcp add my-server --run --oci=my-mcp:latest
# Add an external MCP server by URL (split into scheme/host/port/path)
infer mcp add filesystem http://localhost:3000/sse
# List MCP servers, or show connection status
infer mcp list
infer mcp status
# Enable or disable a server
infer mcp enable my-server
infer mcp disable my-server
# Start or stop an auto-start (OCI) server
infer mcp start my-server
infer mcp stop my-server
# Update a server definition, or remove it
infer mcp update my-server --oci=my-mcp:2.0.0
infer mcp remove my-serverinfer mcp enable-global / disable-global toggle MCP support as a whole rather than a single server.
Using MCP Tools:
MCP tools appear as MCP_<server>_<tool> in chat. Example:
infer chat
> "Use the MCP_demo-server_get_time tool to get current time"See MCP documentation for detailed integration guide and server development.
Agent Skills
Reusable, model-readable instruction folders that the agent loads on demand. The CLI uses the same on-disk format as Gemini CLI and OpenAI Codex CLI, so a skill authored for any of those tools drops into .infer/skills/ unchanged. Skills are discovered from three locations, in precedence order: project .infer/skills/, the .agents/skills/ open standard (a shared cross-tool convention), then user-global ~/.infer/skills/. Skills are enabled by default - discovered skills are injected into the system prompt out of the box. Only the lightweight metadata (name + description) is added; each SKILL.md body is read on demand. Turn them off with agent.skills.enabled: false, or skip individual skills with disabled_skills.
# .infer/config.yaml
agent:
skills:
enabled: true # default
max_chars: 4000 # cap on the rendered AVAILABLE SKILLS block (0 disables the cap)
disabled_skills: [] # optional list of skill names to skip# Discover, install, and remove skills (also available in chat as /skills ...)
infer skills list
infer skills install acme/internal-comms --user # or a bare name, or a github tree URL
infer skills uninstall internal-commsOnce enabled, invoke a skill explicitly with /<name> (for example /pdf-helper) or by asking the agent to "use the <name> skill"; the CLI deterministically activates it by injecting the skill's metadata and pointing the agent at its SKILL.md. Installed skills under ~/.infer/skills and ./.infer/skills stay readable by the Read tool through a sandbox carve-out, so they load even when the agent runs outside the project directory (for example in CI).
See the full Agent Skills guide for the on-disk layout, the SKILL.md frontmatter contract, install flags, activation triggers, and the sandbox carve-out. To publish a skill in the shared index, see the Skills Catalog.
/tools view
The /tools shortcut opens a read-only, filterable list of the tools available to the agent in the current agent mode. Each row shows the tool name plus a word-wrapped description (up to two lines, with an ellipsis on overflow).
- Filtering: type
/to start filtering; matching terms are underlined in the results. - Status line: shows
N toolsat the bottom. - Live refresh: the list refreshes on every entry, so agent-mode changes and asynchronously registered MCP tools are reflected immediately.
- Mode-aware: Plan mode hides mutating tools (Write, Edit, Delete, Bash), showing only read-only tools.
- MCP tools: dynamically registered MCP tools appear as
MCP_<server>_<tool>once their server is connected.
infer chat
> /tools/agents view
The /agents shortcut opens the Agents view - the single, filterable list of every agent the chat can delegate to, local and remote alike. There is no separate A2A-only view: each row carries a type chip that says where it comes from.
| Chip | Row | State column | Detail line |
|---|---|---|---|
local | A Markdown subagent preset (.infer/agents/*.md) that the Agent tool spawns in-process | read-only or read-write | Tool allowlist, model (or inherit), and the source file |
a2a | A remote A2A agent from agents.yaml | Its readiness state | The agent URL, or the failure/progress detail |
Rows are sorted by name, and typing / filters on the chip as well as the name - local narrows the list to presets, a2a to remote agents.
The a2a rows stay live for the session lifetime when liveness probes are enabled:
- An agent still pulling its image shows pull progress (
<done>/<total> layers). - An agent that was down at startup turns green automatically when it becomes reachable.
- An agent that goes down mid-session shows the failure detail inline.
- A recovered agent shows a "Recovered" status.
The status bar A2A: X/Y indicator reflects the same live state and opens this view - X counts down on failures and counts back up on recovery.
infer chat
> /agents # View local presets and remote A2A agents/tasks view
The /tasks shortcut opens the background-work panel - a live list of everything running or recently finished outside the current chat turn. Rows are grouped into one table per kind: A2A tasks, background shells, and subagents. Each row shows a status (Running, Completed, or Failed) and an Elapsed column that updates live about once per second while any work is still in flight (the ticker stops once everything is terminal).
infer chat
> /tasksSelecting a row opens a detail panel with that task's metadata - ID, Detail (the command for a shell, or the task description for a subagent), Status, Started, and Elapsed. A2A rows additionally show their task history and the agent's Final Result; background-shell and subagent rows render an Output section instead.
Output detail section
Selecting a background shell row shows the shell's captured stdout/stderr:
- While the shell is running, the output streams live, refreshed about once per second along with the Elapsed column.
- Once the shell finishes, the panel shows the full captured output, bounded by the shell's output ring buffer.
Selecting a subagent row shows the subagent's result:
- Headless subagents show their final result message.
- Interactive (tmux-pane) subagents show the last harvested turn.
Rendering is bounded so a chatty shell cannot overflow the panel: the Output section shows at most the last 10KB of output. When the captured output is larger, it is prefixed with (truncated, showing last 10KB) and only the trailing 10KB is displayed.
The Output section brings background shells and subagents to parity with the A2A "Final Result" panel. Rows for jobs that expose no output (an A2A task keeps its own Final Result panel) show no Output section.
A2A Integration
Delegate specialized tasks to Agent-to-Agent compatible agents.
Setup:
# Initialize agents configuration
infer agents init
# Add remote agent
infer agents add calendar-agent http://calendar.example.com
# Add local agent with Docker
infer agents add my-agent http://localhost:8081 --oci ghcr.io/myorg/agent:latest --run
# Add any catalog agent by name only (URL and image derived from its catalog entry)
infer agents add grafana-agent
# Add a built-in agent by name only (browser-agent ships one tag per browser engine)
infer agents add browser-agent --tag lightpanda
# Pin a built-in agent to a released version
infer agents add browser-agent --tag chromium-0.8.0
# List agents
infer agents list
# View agent details
infer agents show calendar-agent
# Update an agent's image tag
infer agents update browser-agent --tag firefoxBare names resolve any agent published in the A2A Registry catalog. The five agents with built-in defaults (browser-agent, mock-agent, google-calendar-agent, documentation-agent, n8n-agent) take precedence and work offline; every other name is looked up in the catalog, with the URL and OCI image derived from its published metadata. A name in neither place needs an explicit URL - see built-in agents vs other catalog agents.
Usage:
infer chat
> "Schedule a meeting tomorrow at 2 PM using the calendar agent"
> /agents # View connected agentsSee A2A documentation for creating custom agents, or use the ADL CLI to scaffold new A2A agents from YAML definitions.
A2A Liveness Probes
When A2A agents are configured, the CLI periodically re-probes them for the lifetime of the session instead of checking them only once at startup. This means an agent that was down at startup turns green automatically when it becomes reachable, and an agent that goes down mid-session is reflected in the A2A: X/Y indicator and the a2a rows of the /agents view.
How probes work:
- External agents (URLs in
a2a.agents) are probed via their agent card (/.well-known/agent-card.json). - Local Docker agents are probed via
GET <url>/health. - Status updates are emitted only on state changes (no noisy per-probe output).
- Probes shut down cleanly when the session ends.
Configuration:
# .infer/config.yaml
a2a:
enabled: true
agents:
- http://localhost:8081
liveness_probe_enabled: true # default: true
liveness_probe_interval: 30 # seconds, default: 30| Setting | Default | Environment variable | Description |
|---|---|---|---|
liveness_probe_enabled | true | INFER_A2A_LIVENESS_PROBE_ENABLED | Enable recurring liveness probes. Set false for one-shot checks |
liveness_probe_interval | 30 | INFER_A2A_LIVENESS_PROBE_INTERVAL | Interval in seconds between probes |
Setting liveness_probe_enabled: false restores the old behavior where agents are checked only once at startup.
Parallel Tool Execution
Execute up to 5 tools concurrently for improved performance.
Configuration:
agent:
max_concurrent_tools: 5 # Default: 5Benefits:
- Faster multi-file operations
- Concurrent web fetches
- Parallel code searches
- Reduced total execution time
Workflows
Bug Investigation and Fix
infer chat
# Shift+Tab to Plan Mode
> "Analyze bug in issue #123 and create fix plan"
# Shift+Tab to Standard Mode
> "Implement the fix according to the plan"
# Test and commit
> "Run test suite to verify"
> "/git commit"Feature Development
infer chat
> "Read CONTRIBUTING.md and understand project structure"
# Shift+Tab to Plan Mode
> "Design implementation for user profile feature with avatar upload"
# Shift+Tab twice to Auto-Accept Mode
> "Implement the user profile feature according to the plan"
# Shift+Tab to Standard Mode
> "Review changes and run all tests"Code Review and Refactoring
infer chat
# Plan Mode for analysis
> "Review authentication module for security issues and code quality"
# Standard Mode for implementation
> "Refactor based on recommendations, prioritize security issues"GitHub Issue Resolution
infer headless "Fix the bug described in GitHub issue #456"
# Agent autonomously:
# 1. Fetches issue details
# 2. Analyzes relevant code
# 3. Implements fix
# 4. Runs tests
# 5. Creates commit referencing issueBest Practices
For Beginners
- Start with Plan Mode for unfamiliar code
- Always work in git repositories
- Review diff visualizations before approving
- Begin with simple tasks
For Power Users
- Use Auto-Accept for trusted, repetitive tasks
- Create custom shortcuts for frequent commands
- Combine with scripts for automation
- Leverage A2A for specialized workflows
Performance Tips
- Be specific with file paths and function names
- Use Grep to narrow down relevant files first
- Break large tasks into smaller subtasks
- Provide context with references
Safety
- Review diffs before approving modifications
- Run tests after significant changes
- Have backups before extensive Auto-Accept usage
- Allow-list only trusted commands
- Add sensitive directories to protected paths
Security
Command allow-listing
The Bash tool is default-deny: a command auto-runs only when it matches the per-mode allowed-list for the active agent mode. The effective list is mode.all.allow (the every-mode baseline) unioned with the active mode's own entries:
tools:
bash:
mode:
all:
allow:
- ls( .*)?
- pwd( .*)?
- tree( .*)?
- git status( .*)?
- git diff( .*)?
- npm (install|test|run).*
auto:
allow:
- .* # unrestricted - Auto-Accept mode onlyRead-only gh operations are in the baseline so the agent can inspect GitHub out of the box; writes (gh issue/pr create|edit|comment) and destructive operations (for example gh pr merge, gh repo delete) are not - they fall through to approval. See Default gh allowed-list for the full list.
Entries match the whole command, and a clean-command guard rejects command substitution, multi-command chains/pipelines, file-write redirects, dangerous find actions, and environment-variable leaks before matching. The only thing that lifts the guard is the .* sentinel (Auto-Accept mode).
Protected Paths
Automatically excluded from tool access:
.git/- Repository data*.env- Environment files.infer/- Configuration directory~/.infer/auth.yaml(and the legacy~/.infer/auth.json) - Fallback provider API keys~/.infer/tmp/is the exception: it is agent-readable and writable, see Userspace tmp tree- Custom paths via sandbox config
Approval Workflow
Tool approval has two independent layers - whether an action needs approval, and how that approval is delivered:
- Whether -
tools.safety.require_approval(with per-tool overrides liketools.bash.require_approval/tools.write.require_approval, and for Bash the per-mode allowed-list). - How -
tools.safety.approval_behaviour, one of:
approval_behaviour | How a needed approval is delivered |
|---|---|
prompt (default) | Prompt in the chat TUI; under a channel manager, deliver over IPC; otherwise block. |
ipc | Deliver over IPC when a broker is attached (e.g. the channel manager); otherwise block. |
judge | One LLM judge call decides. Always reachable - never downgraded to block. |
block | Always reject an approval-requiring action with a reason - never prompt. |
infer config set tools.safety.require_approval true
infer config set tools.safety.approval_behaviour prompt
judgeis the only behavior that requires a resolvable judge model - config validation fails at startup when neitherjudge.model(injudge.yaml) noragent.modelis set. Theauto-with-judgemode is validated separately by the headless runner at start (--mode/INFER_AGENT_MODE).
LLMs request approval before executing Write/Edit/Delete/Bash operations, with a colored, syntax-aware diff preview for file edits.
Headless secure-by-default
infer headless runs in standard mode, so an off-list or mutating action is not auto-run. With no approver reachable (CI, heartbeat) it is blocked with a reason; under a channel manager (--require-approval) it is sent for IPC approval (for example a Telegram confirmation). There is no .* default - full autonomy is an explicit opt-in (a curated allowed-list, the append override, or mode.auto / .*). To keep a gate without a human, run --mode auto-with-judge (or set approval_behaviour: judge) and let an LLM judge answer instead of blocking.
For a CI agent that should edit files and run a curated command set with no interactive approver, use the controlled-autonomy profile - block everything that would need approval, but let the agent write files and run a vetted allowed-list:
tools:
safety:
approval_behaviour: block # reject anything that would otherwise prompt
write:
require_approval: false # ...but let the agent write/edit files freely
bash:
mode:
all:
allow: # curate exactly what may run unattended
- git status( .*)?
- git add( .*)?
- go (build|test)( .*)?Add a couple more commands without touching config via INFER_TOOLS_BASH_ALLOW_APPEND="git commit,git push".
Troubleshooting
Connection Issues
# Check configuration
infer config get
# Verify gateway status
infer status
# Debug mode
infer --debug chatPermission Issues
# Check configuration directory
ls -la ~/.infer/
# Recreate config.yaml from defaults
infer config init --overwrite
# Re-initialize the project
infer initTool Execution Problems
# Inspect tool configuration
infer config get tools
# Check whether a bash command is allowed (without running it)
infer tools validate "git status"
# Enable debug logging
export INFER_LOGGING_DEBUG=true
infer headless "your task"Computer Use Issues
# Verify display server
echo $DISPLAY # Linux/X11
# Check permissions (macOS) - required before the accessibility tree returns elements
# System Settings > Privacy & Security > Accessibility
# Test screenshot
infer chat
> "Take a screenshot and describe what you see"If Computer with action: accessibility returns an empty tree or screenshot fallback guidance on macOS, the Accessibility permission for infer is the first thing to check. On Linux and Windows that fallback is expected - the accessibility providers are not implemented there yet.
Shell Completions Not Working
# Confirm the completion script generates
infer completion zsh | head
# Zsh: the file must live on a directory in $fpath and be named _infer,
# then start a fresh shell
infer completion zsh > "${fpath[1]}/_infer" && exec zsh
# Bash: source the generated file (or place it under a bash-completion dir)
source <(infer completion bash)If completions still do not appear, the shell rc is usually not sourcing the completion file. Verify compinit is called for zsh (or bash-completion is installed for bash), confirm the file path is on $fpath/a bash-completion directory, then start a fresh shell.
Command Reference
| Command | Description |
|---|---|
infer init | Seed the userspace baseline in ~/.infer/ |
infer status | Check gateway health and resource usage |
infer chat | Interactive chat session (TUI) |
infer chat --web | Web-based terminal interface |
infer headless <task> | Autonomous task execution |
infer skills <subcommand> | Manage Agent Skills (list, install, uninstall) |
infer daemon | Start the daemon - scheduler, channel listener, and heartbeat (Channels) |
infer config <subcommand> | Configuration management (init, get, set) |
infer tools <subcommand> | Run agent tools directly (execute, validate) |
infer agents <subcommand> | A2A agent management |
infer conversations <subcommand> | Conversation history management (list, show, delete) |
infer avatars <subcommand> | Avatar library management (list, create, delete) for TextToVideo renders |
infer completion <shell> | Generate a shell completion script (bash, zsh, fish, powershell) |
infer version | Show version information (backwards-compatible subcommand) |
infer --version | Show version information (styled by fang) |
infer --help | Display styled help information |
Support and Resources
- Repository: github.com/inference-gateway/cli
- Issues: GitHub Issues
- Releases: GitHub Releases
- Documentation: Full Configuration Reference
The CLI is actively developed with regular updates and new features. Check the repository for the latest releases and announcements. s.
