GitHub Action (infer-action)
inference-gateway/infer-action is the official GitHub Action wrapper for the infer CLI. It lets you run the Inference Gateway agent from a GitHub Actions workflow so that mentioning a trigger phrase in an issue or comment kicks off an automated, AI-driven response: plan posting, code edits, branch creation, and a pull request - all without leaving GitHub.
Versioning: the examples on this page reference
inference-gateway/infer-action@main, which always tracks the latest action code - copy them as-is and you get the newest release, with no version numbers to maintain. If you prefer reproducible runs, substitute a tag from the releases page.
When to use it
Use infer-action when you want CI-driven inference instead of an interactive terminal session. Typical scenarios:
- Issue triage and automated fixes - mention
@inferin an issue, the agent reads it, makes the change, and opens a PR. - Automated code review - have the agent review pull requests on a schedule or on push.
- Scheduled agents - cron-driven release notes, changelog drafts, dependency upgrades, drift reports.
- Advisory-only workflows - run the agent in comment-only mode (
enable-git-operations: false) to post suggestions without modifying the repo.
For local, interactive use the CLI remains the right tool; infer-action is the headless, event-driven counterpart.
Quick Start
Create .github/workflows/infer.yml:
name: Infer Agent
on:
issues:
types: [opened, edited]
issue_comment:
types: [created]
permissions:
issues: write
contents: write
pull-requests: write
jobs:
infer:
runs-on: ubuntu-24.04
steps:
- uses: actions/checkout@v7.0.1
- uses: inference-gateway/infer-action@main
with:
github-token: ${{ secrets.GITHUB_TOKEN }}
model: anthropic/claude-opus-4-8
anthropic-api-key: ${{ secrets.ANTHROPIC_API_KEY }}Open an issue (or comment on one) containing @infer and the workflow takes over. The agent posts a "cooking" placeholder comment, runs, makes changes if needed, and finishes by either updating the comment with results or opening a pull request.
How it works
- Trigger detection - the action inspects
github.event.issue.title,github.event.issue.body, orgithub.event.comment.bodyfor the configuredtrigger-phrase(default@infer). Comments authored by bot users are ignored to prevent recursion. - Reaction + cooking message - on a hit, the action adds an
:eyes:reaction to the trigger comment and posts a placeholder "I'm cooking..." comment. Stale cooking messages from earlier runs are cleaned up. - CLI install - downloads
inferat the configuredversionand runsinfer init --overwrite. - Git config - sets
git user.name/user.emailto thegithub-actions[bot]identity so any commits the agent makes have a valid author, and (by default) wires the repo's own git hooks path - see Repo git hooks. - Agent run - executes the agent with the selected
modeland provider API keys. The bash allow-list is augmented withghandgitunlessenable-git-operations: false. - PR creation - when the agent produces file changes, it creates
fix/issue-{number}, commits, pushes, and opens a PR titledFix #{number}: ...withResolves #{number}in the body. - Result posting - the final comment summarises completed work and the model used, links the PR if any, and appends a footer with token usage, per-session cost, and the agent's tool-call count and success rate (see Result comment).
Native reminders
The action composes a default set of CLI system reminders and passes them to the CLI via INFER_REMINDERS_CONFIG. The composed default includes:
- Periodic context nudge - a
pre_streamreminder on an interval that keeps the agent focused on the task. - Turns-before-max wrap-up - a
turns_before_maxreminder that fires near the turn limit, prompting the agent to wrap up. - Post-tool failure nudge - a
post_tool/on_failurereminder that fires only after a failed tool call on writable runs, reminding the agent to retry or ask the user. - Memory nudges - when
memory-repois set, the CLI's built-inmemory-consultandmemory-hygienereminders are re-emitted. - Demo recording pointer - when
record-demoistrue, apre_stream/oncereminder namedinfer-action-record-demofires on the first turn, telling the agent to read and follow the CLI's built-indemoskill for the recording procedure.
To override the composed default and supply your own full set of reminders, use the reminders-config input. When set, it replaces the action's default entirely (it is not merged), so any built-in behaviour you want must be re-declared.
- uses: inference-gateway/infer-action@main
with:
reminders-config: |
enabled: true
reminders:
- name: fail-nudge
hook: post_tool
trigger: on_failure
text: "A failed call means the change did not happen"See the CLI reminders documentation for the full YAML schema, trigger catalog (always, interval, turns_before_max, once, on_failure), and the INFER_REMINDERS_CONFIG / --reminders-file / on_failure trigger reference.
System prompt override
The action bundles a default system prompt for each context kind (issue, PR, fork PR, direct). You can override it with one of four inputs:
| Input | Context | Template variables |
|---|---|---|
system-prompt-issue | Issue-driven runs | |
system-prompt-pr | PR-driven runs (non-fork) | , |
system-prompt-pr-fork | Fork PR runs (view-only) | , , , |
system-prompt-direct | Direct-prompt (manual) runs | (none) |
When set, the input replaces the action's bundled default prompt for that context (it is not merged). The resulting prompt (override or bundled) then replaces the CLI's own base system prompt text. The CLI still appends its dynamic context block - skills, memory, tools, sandbox, and bash allow-list information - because the action pins INFER_AGENT_SYSTEM_PROMPT_WITH_DEFAULTS=true (the CLI default, pinned explicitly so a consumer config cannot turn it off).
The bundled defaults carry git-safety instructions (branch-first, commit-per-todo, push, draft PR, finish checklist). If your override omits those instructions, the action emits a ::warning:: in the run log so the lost-work guard is not dropped silently. Prefer custom-instructions to layer extras on top of the default unless you need a full replacement.
Dynamic model selection
Override the workflow's default model on a per-issue or per-comment basis by including /model provider/model-name in the trigger text:
@infer /model deepseek/deepseek-v4-flash please analyze this bug and suggest a fixThe override is parsed by the action's trigger-detection step and exported as INFER_AGENT_MODEL for that run only.
Result comment
When the run finishes, the action updates its comment with a result footer. Alongside the status, model, exit code, and job link, the footer reports:
- Duration - wall-clock time of the agent run, formatted human-readably (
0s,1m 0s,1h 1m 1s) and shown as—when unavailable. The measured window is spawn-to-exit of theinfer headlesschild, so it excludes CLI install and Node setup time. The raw millisecond value is also exposed as therun-duration-msoutput. - Tokens - prompt / completion / total token usage, plus the request count.
- Cost - per-session input / output / total cost, when the CLI reports pricing.
- Tool calls - the total number of tool calls the agent made, with the run's success rate. The rate is
succeeded / total(wheresucceeded = total - failed), so a run with failures reads its failures in proportion. Any failed calls are listed in a collapsed section just below.
The Tool calls line is only rendered when the agent made at least one tool call:
## ✅ Infer Result: Success
**Model:** `anthropic/claude-opus-4-8` · **Exit Code:** `0` · **Duration:** 1m 0s · [View Job](...)
**Tokens:** 18,432 in · 2,106 out · 20,538 total (7 requests)
**Cost:** $0.04 in · $0.02 out · $0.06 total
**Tool calls:** 12 total · 83% success rate
<details><summary>⚠️ 2 failed tool call(s)</summary>
...
</details>The total and failed tool-call counts are also exposed as the total-tool-calls-count and failed-tool-calls-count outputs for use in downstream steps.
Inputs
The table below covers the most commonly used inputs. The complete, always-current list - including newer inputs such as skills, plugins, agents, upload-artifacts, and debug - lives in action.yml in the action repository.
| Input | Required | Default | Description |
|---|---|---|---|
github-token | Yes | - | Token used for posting comments, creating branches, and opening PRs. |
github-app-slug | No | '' | Slug of the GitHub App whose bot identity authors the agent's commits (e.g. infer-bot); resolved via GET /users/{slug}[bot]. Falls back to github-actions[bot] when empty or on failure. |
model | Yes | - | Model identifier in provider/model-name form (e.g. anthropic/claude-opus-4-8). |
trigger-phrase | No | @infer | Phrase that activates the agent. Case-sensitive. |
direct-prompt | No | '' | Free-text task to run directly, bypassing issue/comment triggers. When set, the agent runs against this text under workflow_dispatch (or any event), commits to a new branch, and opens a PR; the result and PR link go to the job summary. See Direct prompt. |
version | No | pinned in action.yml | infer CLI version to install inside the runner. Defaults to the CLI release the action pins; override to install a specific version. |
apt | No | '' | Newline- or space-separated list of apt packages to install before the agent runs. Installed via DEBIAN_FRONTEND=noninteractive sudo apt-get install -y. Debian/Ubuntu runners only - on other OS families the step prints an error and fails. Packages are cached across runs (keyed on OS, architecture, and the package list); a cache failure falls back to a plain install. A plain run: step before the action is the portable alternative. |
languages | No | '' | Newline- or space-separated list of languages whose toolchains to install before the agent runs. Supported: go, rust, node (alias typescript), python. Empty = no language setup steps. See Language toolchains. |
max-turns | No | 150 | Maximum agent iterations - acts as a runaway-cost guard. |
custom-instructions | No | '' | Extra instructions appended to the default system prompt (does not replace the defaults). |
system-prompt-issue | No | '' | Overrides the action's bundled system prompt for issue-driven runs. Substitutes . See System prompt override. |
system-prompt-pr | No | '' | Overrides the action's bundled system prompt for PR-driven runs (non-fork). Substitutes , . See System prompt override. |
system-prompt-pr-fork | No | '' | Overrides the action's bundled system prompt for fork PR runs (view-only). Substitutes , , , . See System prompt override. |
system-prompt-direct | No | '' | Overrides the action's bundled system prompt for direct-prompt (workflow_dispatch) runs. No variables. See System prompt override. |
bash-allow-append | No | '' | Comma/newline-separated Go regex entries appended to the agent's read-only bash allow-list, each anchored to the whole command (e.g. npm( .*)?,go test( .*)?). Read-only git/gh come baseline-included; with git operations enabled the action additionally allow-lists the writes its PR workflow needs. |
web-fetch-domains | No | '' | Domains the WebFetch tool may use; passed as INFER_TOOLS_WEB_FETCH_ALLOWED_DOMAINS. Empty = github.com,raw.githubusercontent.com,api.github.com plus the githubusercontent.com hosts that github.com/user-attachments image links redirect to. |
vision-model | No | '' | provider/model of a vision-capable model for the CLI's image annotator. Sets INFER_VISION_ANNOTATOR_ENABLED=true + INFER_VISION_ANNOTATOR_MODEL, registering the ImageDecode tool so the agent can read screenshots/diagrams embedded in issues and PRs. The session model does NOT need vision - ImageDecode side-calls the annotator model through the gateway. Requires CLI >= v0.159.0. Empty = image understanding off. See Working with images. |
image-model | No | '' | provider/model for the ImageGeneration/ImageEdit/ImageVariation tools' one-off /v1/images/* requests. Empty = CLI default (openai/gpt-image-2, needs OPENAI_API_KEY). See Working with images. |
enable-git-operations | No | true | When false, the agent runs in comment-only mode - git/gh are not allow-listed and no PRs are created. |
enable-git-hooks | No | true | Runs git config core.hooksPath <hooks-path> on the action's workspace before the agent starts, so the agent's commits trigger the repo-defined pre-commit hook. Skipped silently when the hooks directory does not exist. See Repo git hooks. |
hooks-path | No | .githooks | Hooks directory wired via git config core.hooksPath when enable-git-hooks is true. Applies only to the action's workspace, not globally. Override for repos that keep hooks elsewhere. See Repo git hooks. |
mirror-agent-logs | No | '' | Controls whether the agent's full stdout/stderr transcript (tool inputs, tool outputs, file contents it read, web-fetch payloads, intermediate text) is mirrored to the Actions run log. Empty (the default) follows the debug input - a debug run mirrors, a normal run stays quiet. true mirrors always; false suppresses even with debug: true. Stderr (crashes, panics) and /tmp/agent-output.txt are always written. See Agent log mirroring. |
memory-repo | No | '' | Git remote URL backing the agent's persistent cross-run memory (ssh or https, e.g. git@github.com:my-org/agent-memory.git). Enables the CLI's memory git backend: pull on run start, commit + push when a fact changes. Empty = feature off. See Persistent Agent Memory. |
memory-branch | No | '' | Branch of memory-repo to sync (INFER_MEMORY_BACKEND_GIT_BRANCH). Empty = CLI default (main). |
memory-sync-on-start | No | '' | Pull memory at run start: pull or off (INFER_MEMORY_BACKEND_GIT_SYNC_ON_START). Empty = CLI default (pull). |
memory-sync-on-finish | No | '' | Push memory changes at run finish: push or off (INFER_MEMORY_BACKEND_GIT_SYNC_ON_FINISH). Empty = CLI default (push). |
memory-deploy-key | No | '' | SSH private key (e.g. a deploy key with write access) authenticating an ssh memory-repo. Secret, auto-masked. See Persistent Agent Memory. |
memory-token | No | '' | Token authenticating an https memory-repo (scoped git insteadOf rewrite). Secret, auto-masked. Empty on a same-instance https URL = falls back to github-token. |
review-inline | No | false | When true and the run is in review mode, post findings as a real GitHub pull request review with inline, line-anchored comments (including suggestion blocks the PR author can apply with one click), instead of a single conversation comment. See Inline PR review. |
reminders-config | No | '' | Verbatim reminders YAML passed to the CLI via INFER_REMINDERS_CONFIG, replacing the action's composed default. Lets a power user take full control of the CLI's native reminders (hooks, triggers, cadences). A supplied config replaces the action's default, so built-in behaviour must be re-covered if desired. Requires Infer CLI >= v0.130.0. See Native reminders and the CLI reminders docs. |
dry-run | No | false | Plan-only local-testing mode (e.g. with act): forces the bundled mock agent, simulates every GitHub mutation ([dry-run] would ...), and prints the resolved system/task/reminder prompts and tool allow-lists. Reads still run. See Local testing with act. |
mock-agent-scenario | No | happy | Which scenario the bundled mock agent runs under dry-run: happy, failures, no-todos, empty, incomplete, no-git, commit-no-push, or hang. |
record-demo | No | false | When true, a run that asks for a demo can record one: the action starts a virtual display with a terminal on tmux session demo, runs RecordStart / RecordStop without approval (60 s cap), and allow-lists tmux send-keys / tmux capture-pane and ffmpeg; the CLI's built-in demo skill does the rest and the GIF is embedded in the result comment. Recording happens only when the request says "demo" / "demonstrate". Requires CLI >= v0.215.0; the default version pin satisfies this. See Demo recordings. |
anthropic-api-key | No* | - | Required when using an Anthropic model. |
openai-api-key | No* | - | Required when using an OpenAI model. |
google-api-key | No* | - | Required when using a Google/Gemini model. |
deepseek-api-key | No* | - | Required when using a DeepSeek model. |
groq-api-key | No* | - | Required when using a Groq model. |
mistral-api-key | No* | - | Required when using a Mistral model. |
cloudflare-api-key | No* | - | Required when using a Cloudflare Workers AI model. |
cohere-api-key | No* | - | Required when using a Cohere model. |
ollama-cloud-api-key | No* | - | Required when using Ollama Cloud. |
moonshot-api-key | No* | - | Required when using a Moonshot (Kimi) model. |
minimax-api-key | No* | - | Required when using a MiniMax model. |
nvidia-api-key | No* | - | Required when using an NVIDIA model. |
zai-api-key | No* | - | Required when using a ZAI model. |
* Provide the key matching the provider of the chosen model. Multiple keys can be supplied so the same workflow handles overrides to different providers.
The action also accepts five opt-in OpenTelemetry inputs (otel-*) for exporting run telemetry to an OTLP collector. They are disabled by default and change nothing for existing workflows - see OpenTelemetry export.
Runner setup
Install system apt packages before the agent runs:
- uses: inference-gateway/infer-action@main
with:
model: anthropic/claude-opus-5
anthropic-api-key: ${{ secrets.ANTHROPIC_API_KEY }}
github-token: ${{ secrets.GITHUB_TOKEN }}
apt: |
libxml2-dev
libpq-devThe apt input is Debian/Ubuntu-only - on other OS families the step prints an error and fails. Downloaded packages are cached across runs (keyed on OS, architecture, and the package list); a cache failure falls back to a plain install. A plain run: sudo apt-get install -y ... step before the action remains the portable alternative for non-Debian runners.
Language toolchains
The languages input installs language toolchains before the agent runs, replacing the need for separate actions/setup-go / setup-node / setup-python / rust-toolchain steps. Each language resolves to a well-known action:
| Language | Action | Default version | Version-file fallback |
|---|---|---|---|
go | actions/setup-go@v7.0.0 | stable | go.mod |
rust | dtolnay/rust-toolchain@stable | stable | - |
node / typescript | actions/setup-node@v7.0.0 | lts/* | .nvmrc |
python | actions/setup-python@v7.0.0 | 3.x | .python-version |
An unsupported value fails the action with an error listing the supported values.
- uses: inference-gateway/infer-action@main
with:
github-token: ${{ secrets.GITHUB_TOKEN }}
model: anthropic/claude-opus-5
anthropic-api-key: ${{ secrets.ANTHROPIC_API_KEY }}
languages: |
go
typescriptThe version-file resolution (go.mod, .nvmrc, .python-version) works just like the upstream action: it detects the file at the repository root and reads the version from it. If the file is absent, the default version listed above is used. Per-language version pinning is not supported - use your own setup-* step if you need a specific minor or patch version.
Repo git hooks
By default (enable-git-hooks: true) the action's Configure Git step runs git config core.hooksPath <hooks-path> on its workspace before the agent starts, so the agent's commits trigger the repo-defined pre-commit hook (for example the prettier hook in inference-gateway/desktop).
The no-op guarantee: when the configured hooks directory does not exist, the step skips with a single info line - repos without hooks see no behavior change. hooks-path overrides the directory for repos that keep hooks outside .githooks; it applies only to the action's workspace, not globally.
A hook failure never strands the agent's work: the salvage step commits with --no-verify, and the agent is instructed to fall back to --no-verify rather than lose unpushed work. To skip repo hooks entirely:
- uses: inference-gateway/infer-action@main
with:
github-token: ${{ secrets.GITHUB_TOKEN }}
model: anthropic/claude-opus-4-8
hooks-path: .hooks # override for repos keeping hooks outside .githooks
# enable-git-hooks: false # skip repo hooks entirelyRequires a reasonably recent inference-gateway/infer-action release; referencing @main as every example on this page does always qualifies.
Outputs
| Output | Description |
|---|---|
result | Human-readable summary of the agent execution. |
exit-code | Exit code returned by infer - non-zero means the agent failed. |
pr-url | URL of the pull request the agent opened (empty if none). Populated for direct-prompt runs and any run that opens a PR. |
run-duration-ms | Wall-clock duration of the agent run in milliseconds (0 if unavailable). |
stopped-early | true when the agent stopped before finishing its plan (unfinished todos, uncommitted work left behind, or a job-timeout stop). |
timed-out | true when the agent was stopped because the job hit its timeout; the remaining work is recovered into a draft PR and reported as a warning, not a failure. |
failed-tool-calls-count | Number of failed tool calls detected in the agent output. |
total-tool-calls-count | Total number of tool calls the agent made during the run. |
Reference outputs in downstream steps via ${{ steps.<id>.outputs.result }}.
Persistent Agent Memory
infer-action supports the Infer CLI's persistent-memory git backend (shipped in CLI v0.127.0). When enabled, the agent's memory is pulled from a git remote at run start and committed + pushed when a fact changes, giving the agent cross-run memory in CI.
Opt-in and inert by default. When memory-repo is empty (the default), no memory environment variables are set and nothing changes for existing workflows. The feature requires Infer CLI >= v0.127.0; the action's default version pin already satisfies this.
Inputs
| Input | Env mapping | Default |
|---|---|---|
memory-repo | INFER_MEMORY_ENABLED=true, INFER_MEMORY_BACKEND_TYPE=git, INFER_MEMORY_BACKEND_GIT_REPO | '' (off) |
memory-branch | INFER_MEMORY_BACKEND_GIT_BRANCH | '' (CLI default: main) |
memory-sync-on-start | INFER_MEMORY_BACKEND_GIT_SYNC_ON_START (pull or off) | '' (CLI default: pull) |
memory-sync-on-finish | INFER_MEMORY_BACKEND_GIT_SYNC_ON_FINISH (push or off) | '' (CLI default: push) |
memory-deploy-key | SSH private key for an ssh memory-repo (secret, auto-masked) | '' |
memory-token | Token for an https memory-repo (secret, auto-masked) | '' |
The action maps these inputs to INFER_MEMORY_* environment variables via $GITHUB_ENV, writing only non-empty values so a consumer's own .infer/memory.yaml is never clobbered. The memory-sync-on-start and memory-sync-on-finish values are validated up front (pull or off only).
Auth options
The action configures authentication for the memory remote based on the URL scheme of memory-repo:
1. SSH repo + deploy key
Use an SSH remote with a dedicated deploy key. The key is written to ~/.ssh/infer-memory-deploy-key (mode 600), the host is keyscanned, and it is wired via core.sshCommand with IdentitiesOnly=yes.
- uses: inference-gateway/infer-action@main
with:
github-token: ${{ secrets.GITHUB_TOKEN }}
model: anthropic/claude-opus-4-8
anthropic-api-key: ${{ secrets.ANTHROPIC_API_KEY }}
memory-repo: git@github.com:my-org/agent-memory.git
memory-deploy-key: ${{ secrets.MEMORY_DEPLOY_KEY }}Store the deploy key as a repository secret (MEMORY_DEPLOY_KEY) and grant it write access to the memory repository.
2. HTTPS repo + token
Use an HTTPS remote with a personal access token or GitHub App installation token. The token is applied as a git insteadOf rewrite scoped to the memory repo URL (never persisted in the memory clone's .git/config).
- uses: inference-gateway/infer-action@main
with:
github-token: ${{ secrets.GITHUB_TOKEN }}
model: anthropic/claude-opus-4-8
anthropic-api-key: ${{ secrets.ANTHROPIC_API_KEY }}
memory-repo: https://github.com/my-org/agent-memory
memory-token: ${{ secrets.MEMORY_TOKEN }}Store the token as a repository secret (MEMORY_TOKEN). Both memory-deploy-key and memory-token are auto-masked in logs and redacted from the cooking comment.
3. Workflow repo branch (no extra secret)
When memory-repo is an HTTPS URL on the same GitHub instance and neither memory-token nor memory-deploy-key is set, the action falls back to github-token. This enables the lightest setup: a memory branch of the workflow repository itself, with no extra secret required. The contents: write permission (already needed for PR creation) covers the memory pushes too.
name: Infer Agent (with memory)
on:
issues:
types: [opened, edited]
issue_comment:
types: [created]
permissions:
issues: write
contents: write
pull-requests: write
jobs:
infer:
runs-on: ubuntu-24.04
steps:
- uses: actions/checkout@v7.0.1
- uses: inference-gateway/infer-action@main
with:
github-token: ${{ secrets.GITHUB_TOKEN }}
model: anthropic/claude-opus-4-8
anthropic-api-key: ${{ secrets.ANTHROPIC_API_KEY }}
memory-repo: https://github.com/${{ github.repository }}
memory-branch: agent-memoryThe action emits a ::notice:: when falling back to github-token so the behaviour is transparent.
Identity
Memory commits are attributed to the same bot identity as the agent's code commits. The action's Configure Git step exports the resolved bot identity as GIT_AUTHOR_NAME, GIT_AUTHOR_EMAIL, GIT_COMMITTER_NAME, and GIT_COMMITTER_EMAIL:
<github-app-slug>[bot]when a GitHub App token is usedgithub-actions[bot]otherwise
These environment variables are set on $GITHUB_ENV so the CLI's memory git backend picks them up for commits made outside the workspace clone.
CLI-owned defaults
The following defaults live in the Infer CLI and apply when the corresponding input is left empty:
| Setting | CLI default |
|---|---|
| Branch | main |
| Sync on start | pull |
| Sync on finish | push |
| Per-git-op timeout | 60s |
| Commit message | chore(memory): sync |
The memory-branch, memory-sync-on-start, and memory-sync-on-finish inputs override these only when set to a non-empty value.
Behaviour notes
- Best-effort. Memory sync never fails the run. If the remote is unreachable, push fails, or a rebase conflict occurs, the agent run continues and the result comment is posted as usual.
- Concurrent runs. When multiple workflow runs push to the same memory remote simultaneously, the CLI reconciles them with a push -> pull-rebase -> retry loop. Conflicting facts from the most recent push win.
- Independent of
enable-git-operations. Memory sync works regardless of whether the agent is allowed to create branches and PRs. A comment-only workflow (enable-git-operations: false) can still persist and retrieve memory. - Requires Infer CLI >= v0.127.0. The CLI version the action installs by default already satisfies this requirement.
Source references
- Example workflow:
examples/with-memory.ymlin the action repository
Direct prompt (manual runs)
Normally the action reads its task from the issue or comment that contains the trigger phrase. To run the agent against a free-text task with no issue or comment - for example from a manual workflow_dispatch form - pass the text through direct-prompt:
name: Infer (manual)
on:
workflow_dispatch:
inputs:
prompt:
description: 'Task for the agent to work on'
required: true
type: string
permissions:
contents: write
pull-requests: write
jobs:
infer:
runs-on: ubuntu-24.04
steps:
- uses: actions/checkout@v7.0.1
- uses: inference-gateway/infer-action@main
with:
github-token: ${{ secrets.GITHUB_TOKEN }}
model: anthropic/claude-opus-4-8
anthropic-api-key: ${{ secrets.ANTHROPIC_API_KEY }}
direct-prompt: ${{ inputs.prompt }}Trigger it from the Actions tab: pick the workflow, choose Run workflow, and type a task. No issue, comment, or trigger phrase is needed.
When direct-prompt is non-empty:
- The agent runs against that text instead of an issue or comment body, so no
issues/issue_commentevent is required - the action works underworkflow_dispatch(or any event). - There is no issue/PR thread to reply to, so the agent commits its work to a new branch and opens a pull request. The run's result and the PR link are written to the workflow job summary, and the PR URL is exposed as the
pr-urloutput. - All other inputs (
model,skills,max-turns,bash-allow-append, provider keys, ...) compose as usual. A/modeloverride embedded in the prompt text is honoured, just as in event-driven mode. - With
enable-git-operations: false, direct-prompt runs in advisory mode: the agent only writes its findings to the job summary, with no branch or PR.
Leave direct-prompt empty (the default) and event-driven behaviour is unchanged.
Working with images
Two independent image capabilities, both powered by the Infer CLI (requires CLI >= v0.159.0).
Reading images (screenshots in issues/PRs)
Opt in by setting vision-model to any vision-capable model your gateway serves. The session model does not need vision - the ImageDecode tool side-calls the annotator model through the gateway:
- uses: inference-gateway/infer-action@main
with:
model: deepseek/deepseek-chat # the session model does NOT need vision
vision-model: anthropic/claude-haiku-4-5-20251001
anthropic-api-key: ${{ secrets.ANTHROPIC_API_KEY }}
deepseek-api-key: ${{ secrets.DEEPSEEK_API_KEY }}This enables the CLI's ImageDecode tool and adds guidance to the system prompt: when the issue or PR text embeds an image (), the agent downloads it with WebFetch and reads it with ImageDecode before starting work. The default web-fetch-domains already include the githubusercontent.com hosts that github.com/user-attachments links redirect to.
Limitation: attachments on private repositories redirect to signed URLs that WebFetch requests unauthenticated, so private-repo screenshots may fail to download; the agent is instructed to say so and continue with the text. Public repositories work.
Generating images
The CLI's ImageGeneration/ImageEdit/ImageVariation tools are on by default and call the gateway's /v1/images/* endpoints with openai/gpt-image-2 - so image generation already works if you pass openai-api-key. Set image-model to use a different image model. Generated images land in .infer/tmp and surface through the run-artifacts upload and result-comment embedding.
Demo recordings
Set record-demo: "true" and a run can end with a short GIF of the change working. The recording procedure itself lives in the CLI's built-in demo skill (seeded by infer init like bug and tmux; requires CLI >= v0.215.0) - the action only makes recording possible; it ships no procedure of its own.
Asking for a demo
A run records only when its request says "demo" or "demonstrate" - including as the tail of a bigger ask ("implement X and demo it"). @infer /demo on a pull request demos that PR as it is, without changing code. With record-demo: "true" but no demo request, nothing is recorded and the run behaves exactly as before.
- uses: inference-gateway/infer-action@main
with:
github-token: ${{ secrets.GITHUB_TOKEN }}
model: anthropic/claude-opus-4-8
anthropic-api-key: ${{ secrets.ANTHROPIC_API_KEY }}
record-demo: 'true'What the action sets up
- A dedicated demo terminal. A virtual display (
:99, 1280x720) shows one xterm attached to the tmux sessiondemoin the repository checkout, and~/.infer/artifactsis created up front. - Recording without approval.
RecordStart/RecordStoprun without approval, capped at 60 seconds per take, and the bash allow-list gainstmux send-keys/tmux capture-paneandffmpeg. - A first-turn pointer. The
infer-action-record-demoreminder fires on the first turn and points the agent at thedemoskill: plan the demo as the last todo, rehearse unrecorded, record one take, convert it.
Where the GIF shows up
The skill converts its single good take to ~/.infer/artifacts/demo.gif (converting again replaces it), and the run-artifacts flow embeds everything in that directory in the result comment - so the GIF appears in the comment and recordings never enter the repository. Keep the upload-artifacts input enabled or there is nothing to embed.
Security
A demo run types into its recorded terminal via tmux send-keys, which means an unrestricted shell on the runner: whatever is typed there executes. Only enable record-demo on workflows whose triggers you trust. The org reusable workflow in inference-gateway/.github turns recording on only when the triggering user has write access to the repository (the "Resolve demo recording" step checks with the App token and fails closed), because the caller-side gate does not check who commented.
Local testing with act
Set dry-run: true to exercise the whole workflow in a plan-only mode - ideal for trying a workflow locally with act before it runs for real. In dry-run the action:
- Forces the bundled mock agent - no real CLI install and no provider token, so it composes with any
modelwithout spending anything. (This replaces the formeruse-mock-agentinput; usedry-run: trueinstead.) - Simulates every GitHub mutation. Instead of creating or updating a comment, the
:eyes:reaction, the "I'm cooking..." comment, comment zones, or the spinner, it logs a[dry-run] would ...line. Secret values are still redacted in the printed bodies. - Prints a DRY RUN banner with the exact system / task / reminder prompts and the resolved bash allow-list and web-fetch domains the agent would receive.
- Keeps GitHub reads real so you see the actual target issue or PR. Reads fail soft when no token is available - a public-repo read still works unauthenticated; otherwise the run warns and continues.
mock-agent-scenario selects which scripted run the mock agent performs under dry-run: happy (the default - TodoWrite passes, a read, and a commit on a fix/ branch), failures (the happy path with interspersed tool-call failures), no-todos (work without any TodoWrite calls), empty (exit immediately with no tool calls), incomplete (cut off mid-plan), no-git (edits files but never commits or pushes), commit-no-push (commits but never pushes), or hang (wedges until the job timeout).
The infer-action repo ships ready-to-run local workflows under examples/local/ that run the working-tree action (uses: ./) in dry-run, driven through act by Taskfile helpers:
task test:issue # issues event
task test:comment # issue_comment event
task test:direct # workflow_dispatch / direct-prompt mode
task test:all # all threeNo .env, token, or provider key is required - mutations are simulated and reads fail soft. Pass a token to resolve real reads:
task test:issue -- -s GITHUB_TOKEN=$(gh auth token)Recipes
PR review on push
Run the agent against every push to a PR branch and have it leave review comments without modifying the code.
name: AI PR Review
on:
pull_request:
types: [opened, synchronize, reopened]
permissions:
pull-requests: write
contents: read
jobs:
review:
runs-on: ubuntu-24.04
steps:
- uses: actions/checkout@v7.0.1
with:
fetch-depth: 0
- uses: inference-gateway/infer-action@main
with:
github-token: ${{ secrets.GITHUB_TOKEN }}
model: anthropic/claude-opus-4-8
anthropic-api-key: ${{ secrets.ANTHROPIC_API_KEY }}
trigger-phrase: '@review'
enable-git-operations: false
custom-instructions: |
- Focus on correctness, security, and performance.
- Quote specific files and line numbers.
- Do not modify code; post review feedback as a comment only.Trigger by commenting @review on the PR.
Inline PR review
By default, when the action runs in review mode, findings are posted as a single conversation comment on the issue or PR thread. Set review-inline: true to post findings as a real GitHub pull request review with inline, line-anchored comments on the specific lines of code the agent is commenting on. Each comment can include a suggestion block that the PR author can apply with one click.
name: AI PR Review (inline)
on:
pull_request:
types: [opened, synchronize, reopened]
permissions:
pull-requests: write
contents: read
jobs:
review:
runs-on: ubuntu-24.04
steps:
- uses: actions/checkout@v7.0.1
with:
fetch-depth: 0
- uses: inference-gateway/infer-action@main
with:
github-token: ${{ secrets.GITHUB_TOKEN }}
model: deepseek/deepseek-v4-flash
deepseek-api-key: ${{ secrets.DEEPSEEK_API_KEY }}
trigger-phrase: '@review'
enable-git-operations: false
review-inline: trueTrigger by commenting @review on the PR. The agent posts its findings as a formal PR review with inline comments anchored to the relevant lines and suggestion blocks the author can apply directly.
Scheduled summary / drift report
Run a daily agent that reads recent commits and posts a summary to a tracking issue.
name: Daily Drift Summary
on:
schedule:
- cron: '0 7 * * *' # 07:00 UTC daily
workflow_dispatch:
permissions:
issues: write
contents: read
jobs:
summary:
runs-on: ubuntu-24.04
steps:
- uses: actions/checkout@v7.0.1
- uses: inference-gateway/infer-action@main
with:
github-token: ${{ secrets.GITHUB_TOKEN }}
model: deepseek/deepseek-v4-flash
deepseek-api-key: ${{ secrets.DEEPSEEK_API_KEY }}
enable-git-operations: false
max-turns: 20
custom-instructions: |
- Read the last 24 hours of commits on main.
- Post a Markdown summary as a comment on issue #1 (the drift tracker).
- Group changes by area (api, sdks, docs, infra).Pair with a tracker issue containing the trigger phrase in its body so the schedule has something to fire against, or use workflow_dispatch to invoke the agent against a freshly created tracker issue.
Agent-driven release notes
On every tag push, generate release notes by running the agent over the commits since the previous tag.
name: Release Notes
on:
push:
tags:
- 'v*'
permissions:
contents: write
issues: write
jobs:
notes:
runs-on: ubuntu-24.04
steps:
- uses: actions/checkout@v7.0.1
with:
fetch-depth: 0
- uses: inference-gateway/infer-action@main
with:
github-token: ${{ secrets.GITHUB_TOKEN }}
model: anthropic/claude-opus-4-8
anthropic-api-key: ${{ secrets.ANTHROPIC_API_KEY }}
bash-allow-append: 'gh release( .*)?'
custom-instructions: |
- Diff the current tag against the previous tag.
- Categorise commits using Conventional Commits prefixes.
- Publish the result via `gh release edit <tag> --notes-file ...`.Extending the bash allow-list
The default bash allow-list is intentionally narrow (read-only commands plus read-only gh; see Default gh allowed-list). Add what your project needs.
The action exposes an input that appends to the agent's allow-list:
- uses: inference-gateway/infer-action@main
with:
github-token: ${{ secrets.GITHUB_TOKEN }}
model: anthropic/claude-opus-4-8
anthropic-api-key: ${{ secrets.ANTHROPIC_API_KEY }}
bash-allow-append: 'npm( .*)?,yarn( .*)?,pnpm( .*)?,node( .*)?,python3( .*)?,pytest( .*)?'The added entries are appended to the defaults - they do not replace them.
You can also append at the CLI layer with the INFER_TOOLS_BASH_ALLOW_APPEND environment variable (comma- or newline-separated), which merges onto the every-mode mode.all baseline. Handy for adding a couple of commands - for example letting a release agent commit and push without shipping the unrestricted .* sentinel:
- uses: inference-gateway/infer-action@main
env:
INFER_TOOLS_BASH_ALLOW_APPEND: 'git commit,git push'
with:
github-token: ${{ secrets.GITHUB_TOKEN }}
model: anthropic/claude-opus-4-8
anthropic-api-key: ${{ secrets.ANTHROPIC_API_KEY }}Controlled-autonomy CI profile
A headless infer headless is secure-by-default: in CI there is no interactive approver, so any off-list or mutating action is blocked rather than auto-run. To let an unattended agent edit files and run a curated command set without prompting, combine a block approval behaviour with a relaxed write gate and a curated allow-list - set entirely through environment variables on the step:
- uses: inference-gateway/infer-action@main
env:
INFER_TOOLS_SAFETY_APPROVAL_BEHAVIOUR: block # reject anything that would otherwise prompt
INFER_TOOLS_WRITE_REQUIRE_APPROVAL: 'false' # ...but let the agent write/edit files
INFER_TOOLS_BASH_ALLOW_APPEND: 'git add,git commit,git push'
with:
github-token: ${{ secrets.GITHUB_TOKEN }}
model: anthropic/claude-opus-4-8
anthropic-api-key: ${{ secrets.ANTHROPIC_API_KEY }}This is the recommended shape for an autonomous CI agent: explicit about exactly what may run unattended, with everything else hard-blocked rather than silently auto-approved.
Agent log mirroring
The mirror-agent-logs input controls whether the agent's full stdout/stderr transcript is mirrored to the GitHub Actions run log. By default it follows the debug input - a debug run mirrors, a normal run stays quiet - so every tool input, tool output, file content the agent reads, web-fetch payload, and intermediate text appears in the live log when mirrored. Useful for debugging and understanding what the agent did.
Why suppress the log?
GitHub Actions run logs are persisted with the workflow run, downloadable as raw logs, and visible to everyone with read access to the repository. For public repositories that means the whole world. While ::add-mask:: redacts known secret values, it cannot catch everything - a private source file, customer data in a fetched page, or a secret printed by a tool could leak into the log. mirror-agent-logs: false is a hard off-switch for the entire transcript.
What is still visible when logs are suppressed?
/tmp/agent-output.txt- the full, unredacted transcript is still written to this file on the runner. It is not uploaded or persisted beyond the job.- Cooking-comment footer - the result comment (status, model, token usage, cost, tool-call stats) is posted to the issue/PR as normal.
- Step summary - the job summary (
$GITHUB_STEP_SUMMARY) renders in full. - Minimal heartbeat - ticker updates and the final exit code still print, so the step is not completely silent.
Example
- uses: inference-gateway/infer-action@main
with:
model: anthropic/claude-opus-4-8
github-token: ${{ secrets.GITHUB_TOKEN }}
mirror-agent-logs: false # keep the full agent transcript out of the Actions run logOpenTelemetry export
infer-action can export OpenTelemetry telemetry about each agent run to any OTLP-compatible collector (an OpenTelemetry Collector, Grafana Alloy, Honeycomb, Tempo, Jaeger, and so on). The feature is opt-in and disabled by default - with otel-exporter-otlp-endpoint empty, nothing is exported and the action behaves exactly as it did before.
Export is best-effort and runs after the user-visible result comment is posted: it never blocks the result, and a slow or unreachable collector never fails the run. Resource attributes and metric / span names follow the OpenTelemetry GenAI semantic conventions, so the data lines up with other GenAI instrumentation in your backend.
Inputs
Each input maps to the standard OpenTelemetry environment variable of the same name, so the underlying exporter honours it directly.
| Input | Default | Env var | Description |
|---|---|---|---|
otel-exporter-otlp-endpoint | '' (disabled) | OTEL_EXPORTER_OTLP_ENDPOINT | OTLP HTTP endpoint, e.g. http://localhost:4318. Empty (the default) disables all export. |
otel-exporter-otlp-headers | '' | OTEL_EXPORTER_OTLP_HEADERS | Comma-separated key=value headers, e.g. Authorization=Bearer my-token. Treated as secret and auto-masked. |
otel-service-name | infer-action | OTEL_SERVICE_NAME | Value for the service.name resource attribute on exported telemetry. |
otel-resource-attributes | '' | OTEL_RESOURCE_ATTRIBUTES | Extra resource attributes in key=val,key2=val2 form, appended to the standard set. |
otel-collector | true | - | Deploy a temporary OpenTelemetry collector (Docker) for the job; it fans spans/metrics/logs back to the CLI's local receiver (so the footer's infer traces view stays complete) and, when otel-exporter-otlp-endpoint is set, forwards everything to that remote with the configured headers. Best-effort: a missing Docker or a failed start warns and the run continues. Set false to skip. |
The standard resource attributes attached to every export are service.name, service.version, gen_ai.provider.name, and CI context (cicd.pipeline.*, vcs.repository.*, github.*); otel-resource-attributes appends to that set. Export is OTLP over HTTP - point the endpoint at the collector's HTTP port (4318 on a standard collector), not the gRPC port (4317).
Example
- uses: inference-gateway/infer-action@main
with:
github-token: ${{ secrets.GITHUB_TOKEN }}
model: anthropic/claude-opus-4-8
anthropic-api-key: ${{ secrets.ANTHROPIC_API_KEY }}
otel-exporter-otlp-endpoint: https://otel-collector.example.com:4318
otel-exporter-otlp-headers: ${{ secrets.OTEL_EXPORTER_OTLP_HEADERS }} # e.g. Authorization=Bearer ...Leave otel-exporter-otlp-endpoint unset (the default) and the block above is inert - existing workflows need no changes.
Secrets and least-privilege
Three principles to follow:
Never hardcode API keys. Store them in GitHub Secrets and reference them via
${{ secrets.NAME }}. The action only reads provider keys frominputs, which are injected as environment variables for the agent step.Grant the minimum workflow permissions for the mode you want. For a PR-creating agent:
yamlpermissions: issues: write contents: write pull-requests: writeFor comment-only mode (
enable-git-operations: false):yamlpermissions: issues: write contents: readPrefer a GitHub App over the default
GITHUB_TOKENfor cross-repo or higher-trust workflows. The CLI's/install-opentaskshortcut wires the workflow up for you (see the CLI docs). Using an App token lets PRs created by the agent trigger downstream CI - PRs opened with the defaultGITHUB_TOKENdo not, by GitHub design.
Additional hardening:
- Cap
max-turnsto prevent runaway loops and bound cost. The default is150;30-50covers most issues. - Set
enable-git-operations: falsewhenever the agent only needs to read and comment. - Keep the bash allow-list narrow - only add commands you trust the agent to invoke.
- Reference
inference-gateway/infer-action@mainso every run uses the latest action; if you need reproducible runs, pin to a tagged release from the releases page.
CLI integration
The CLI ships the /install-opentask shortcut, which hands the install to the agent: it writes .github/workflows/tasks.yml from the opentask skill's canonical workflow on a dedicated branch and opens a pull request for review. Use it for first-time setup; come back to this page when you need to customise inputs, write recipes by hand, or harden secrets handling beyond the defaults.
Related
- CLI - the
inferbinary thatinfer-actioninstalls and drives. - Configuration - environment variables understood by the gateway and CLI.
infer-actionrepository - source, releases, and issue tracker.
