Queue and Retry
The queue namespace decides where review jobs wait, how many run at once,
how fast they can call each provider, and how failures are retried. The default
is an in-memory queue; for production you should switch to the durable SQLite
queue so jobs survive restarts.
For Redis, queue.redis.url_env names the environment variable containing the
connection URL. Use rediss://username:password@host:port/db for TLS and ACL
authentication; percent-encode credentials. The queue and automatic batch store
preserve the URL for the driver. For a private CA, set NODE_EXTRA_CA_CERTS to
its PEM file before starting Node; certificate verification stays enabled.
queue: kind: sqlite # memory (default) | sqlite | redis
workers: concurrency: 4 per_workspace_concurrency: 1 lock_ttl_seconds: 1800
rate_limit: per_provider_rps: gitea-internal: 5
retry: attempts: 3 backoff: kind: exponential base_ms: 2000 max_ms: 60000 jitter: trueAutomatic commit schedules
Section titled “Automatic commit schedules”Git push, P4 change-commit, and SVN post-commit events wait 300 seconds by
default. review.auto_commit controls this delay, allowed weekly windows, and
source exclusions. Set it globally, in workspaces.defaults.review, or in
workspaces.instances.<id>.review. The nearest schedule or exclude_sources
replaces the inherited value as a whole.
review: auto_commit: delay_seconds: 300 queued_timeout_hours: 72 schedule: timezone: Asia/Shanghai rules: - days: [mon, tue, wed, thu, fri] windows: - { start: "00:00", end: "13:00" } - { start: "18:00", end: "24:00" } - days: [sat, sun] windows: - { start: "00:00", end: "24:00" } exclude_sources: - id: ci-client vcs: p4 match: client: { glob: "ci-*" }Rules are combined by union. Windows include their start and exclude their
end; overnight windows belong to their starting weekday. Omitted schedules or
schedule.rules: [] allow all times. The default timezone is UTC. Running
reviews may finish after the window closes. The window also gates asynchronous
pull-request, issue, and comment processing: outside it the first attempt and
every retry wait for the next window, logged as trigger processing deferred by execution window. exclude_sources: [] clears inherited exclusions; rules
use OR, fields within a rule use AND, and each field accepts either glob or
RE2 regex.
include_branches is a receive-side branch allowlist for the same automatic
commit events: when the resolved list is non-empty, pushes to unlisted branches
are ignored at receive time without persisting a receipt. The nearest layer
wins wholesale and [] clears back to all branches. PR/MR, comment, and issue
flows are never filtered, and branchless P4/SVN hooks bypass the check.
Use exact, case-sensitive branch names such as main or release/1.x, without
the refs/heads/ prefix; glob patterns and regular expressions are not expanded.
GitLab Push Hook events use this same filter and persistent queue. Branch
creation/deletion notifications with an all-zero before/after SHA are ignored.
queued_timeout_hours bounds how long a queued automatic commit may wait
before it is terminally skipped: pending entries and queued batches
(dispatch_pending/queued/retry_wait) older than the bound are marked
queued_timeout, their in-flight run rows flip to timeout, and the Events
decision flips from queued to timeout, so a stuck queue never accumulates
silently and a long downtime never replays ancient work. Age counts from the
member’s first receipt admission, including time spent waiting for metadata
and the initial delay. A batch inherits its oldest member’s admission time.
Every automatic recovery path
checks it before acting — the startup sweep (before dispatching begins),
expired-lease reclaim, interrupted-batch recovery, legacy dead-batch requeue,
and the dispatch claim itself: an over-aged task is always terminated, never
resurrected. The dashboard’s manual Retry starts a fresh waiting period while
preserving publication checkpoints. The default is
72; 0 disables the sweep and the maximum is 8760 (365 days). The nearest
explicitly set layer wins, like the other fields above. Running batches stay
under the lease/retry lifecycle and are not timed out by the periodic sweep;
to stop one early use the IM aicr cancel command, Queue’s Cancel action,
or Runs’ Terminate action. Expired or interrupted executions are checked
before automatic recovery; a live execution with a valid lease is preserved.
Completed receipts are never relabeled as queue timeouts.
The same resolved timeout applies to waiting IM review requests, measured from
request acceptance. Expired requests close with im.queued_timeout and notify
their original conversation through the notification outbox; running requests
are preserved. Queue Requeue clears a pending batch’s retry delay, and Runs
Re-review creates a new run using the stored event and current configuration.
Stored events preserve PR/MR targets, forks and commit ranges. Historical rows
without enough event data return an explicit error instead of reviewing a
different target.
PR/MR events waiting for an execution window also expire under this timeout, including during startup recovery. Bot cancellation removes matching waiting deferrals durably before they can resume.
Pull request schedules
Section titled “Pull request schedules”PR/MR analysis can use its own weekly window with the same shape as
review.auto_commit.schedule, under review.pull_request.schedule at any of
the three layers (global, workspaces.defaults.review,
workspaces.instances.<id>.review; the nearest schedule replaces the
inherited one as a whole):
review: pull_request: schedule: timezone: Asia/Shanghai rules: - days: [mon, tue, wed, thu, fri] windows: - { start: "00:00", end: "13:00" } - { start: "18:00", end: "24:00" } - days: [sat, sun] windows: - { start: "00:00", end: "24:00" }When no layer sets review.pull_request.schedule, pull-request events fall
back to the resolved review.auto_commit.schedule; when neither is set,
every instant is allowed. Only automatic pull-request events and
comment-triggered review commands are gated — a deferred comment command
posts a reply on the PR/MR stating the scheduled start. There is no
first-receive delay or commit batching for pull requests.
review.pull_request.include_target_branches restricts PR/MR analysis to the
listed target (base) branches — for example [main] analyzes only PRs/MRs
that merge into main. It follows the same three-layer wholesale replacement
and [] clears back to all branches. Events whose target branch is unknown
(a failed PR-detail fetch for a comment command) are allowed through; push,
issue, and manual flows are never filtered.
Target branch names also use exact, case-sensitive matching. A missing ref is
allowed; a supplied empty or non-string webhook ref is rejected as invalid.
These lists apply at reception and do not re-filter previously accepted work.
workspaces: instances: atframe-utils: review: pull_request: include_target_branches: [main]Outside the window the event is held by the deferral registry instead of a
bare timer. With the observability store configured (storage.database plus
admin, see the dashboard page), deferrals persist and resume after a
restart; repeated events for the same target replace the stored envelope so
only the newest state is reviewed, and the resume instant never moves
earlier. Without the store, deferrals fall back to process memory and a
restart drops them, matching the rest of the asynchronous trigger path.
Waiting targets do not hold a running-review deduplication slot. Both timer
scheduling and the actual attempt check the window, including after a process
pause or clock change. Already running analyses can finish outside the window.
A Git push is one complete review unit, including mixed authors and merge
commits. Metadata pages and the 50-member coalescing bound cannot split it.
Across pushes, continuous due units from one raw author name + email may merge
up to 50 members; a mixed-author push or a push containing a merge/rewrite stays
isolated. Two pushes containing 40 and 20 commits produce two complete batches.
A single 550-commit push produces one batch. The atomic store bound is 4,096
members: a larger push fails with push_batch_too_large, without executing a
prefix. An internal exclusion gap fails with exclusion_scope_conflict; excluded
prefixes or tails can be omitted when the remaining range is continuous.
Receipts distinguish repository and branch. Moving already reviewed commits to
another watched branch creates new work for that branch. Queue age starts at
event admission, independent of commit dates. Duplicate deliveries preserve that
age and sealed membership. P4/SVN retain continuous same-source grouping, up to
50 members, keyed by User + Client or svn:author; hooks cover only their named
revision. Missing exclusion evidence uses bounded retries, then fails the member.
Receipts, batch membership, and execution checkpoints use the queue.kind
backend. Memory state is lost on restart. A completed checkpoint recovers local
result accounting without rerunning analysis or publication. Pending publication
retains analysis, per-channel receipts and remote operation IDs; see
remote reconciliation.
A terminal failure gets one automatic recovery; exhaustion skips the batch and
releases the stream. Admin Queue Retry uses current configuration and preserves
an existing remote journal. Old checkpoints without saved identities and memory
storage after restart cannot prevent duplicate publication.
Automatic batches, ordinary queue workers, PR/MR, issue, comment and manual
analysis share one execution budget per server process:
queue.workers.concurrency defaults to 4, and
queue.workers.per_workspace_concurrency defaults to 1. A running P4 workspace
does not prevent a different GitHub workspace from using a free slot. Claims
skip workspaces at their limit before applying the candidate bound. Scheduler
scans do not wait for analysis to finish; metadata preparation is bounded
separately per workspace. A single stream still executes its sealed batches in
order, even when the workspace limit is greater than 1.
Retries acquire a slot for each attempt; backoff and execution-window waits hold no analysis slot. Each actual start rechecks its window. Lowering concurrency does not cancel admitted work; it delays new starts until usage falls below the new limit. Raising it allows waiting work to start. Batch-store leases also enforce batch limits across consumers sharing that store; the shared budget for all entry points is process-local, not a cluster-wide quota.
Recent Runs, Events and terminal Queue history use the configurable storage retention policy.
queue.kind
Section titled “queue.kind”| Value | Description |
|---|---|
memory (default) |
In-process queue. Jobs are lost on restart. Fine for single-instance dev. |
sqlite |
Durable queue that survives restarts (single process or multiple processes sharing the same file). Recommended for production. |
redis |
Durable queue backed by Redis, for multi-instance deployments. Options below. |
rabbitmq |
Reserved — not implemented; setting it logs a warning and falls back to memory. |
queue.redis — Redis queue options
Section titled “queue.redis — Redis queue options”Redis queue connection fields are accepted as passthrough keys:
| Field | Type | Default | Description |
|---|---|---|---|
url_env |
string | – | Name of the env var holding the Redis URL. |
url |
string | – | Redis URL directly (or use host / port / password / db). |
tls |
bool | false |
Connect over TLS. |
key_prefix |
string | "aicr:" |
Key prefix for the queue. Use a unique value per environment when sharing Redis. |
queue.workers
Section titled “queue.workers”| Field | Type | Default | Description |
|---|---|---|---|
concurrency |
int > 0 | 4 |
Total simultaneous analyses across scheduler entry points in one process. |
per_workspace_concurrency |
int > 0 | 1 |
Simultaneous analyses per workspace across those entry points. |
lock_ttl_seconds |
int > 0 | 1800 |
Reserved; lock expiry is configured by the queue backend. |
queue.sqlite — durable queue options
Section titled “queue.sqlite — durable queue options”| Field | Type | Default | Description |
|---|---|---|---|
path |
string | data/queue.sqlite |
SQLite database file for the queue. |
lock_ttl_seconds |
int > 0 | 300 |
Stale-running reclaim TTL. A running job whose lock is older than this is treated as crashed and reclaimed. |
How the SQLite durable queue works
Section titled “How the SQLite durable queue works”The SQLite queue is built on better-sqlite3 and is safe for either a single process or multiple processes sharing the same file. Its key properties:
- Atomic claim via
UPDATE ... RETURNING. A worker claims the next queued job and marks itrunningin a single statement, so two workers can never grab the same job. - Stale-job reclaim after the lock TTL. A background sweep requeues any
runningjob whose lock is older thanlock_ttl_seconds, so a crashed worker’s job is eventually retried by another worker. - WAL +
busy_timeoutfor cross-process safety. The queue opens withPRAGMA journal_mode = WALandPRAGMA busy_timeout = 5000, so concurrent writers from different processes cooperate instead of erroring.
With database configuration enabled, concurrency updates apply at the next claim without cancelling running jobs. Queued jobs retain their accepted config version; the queue version ledger survives restart with a durable backend.
queue.rate_limit
Section titled “queue.rate_limit”Limits use actual LLM provider IDs, including retries and fallback requests. Publishing new rates keeps accumulated token-bucket state. Native agent CLIs are limited per launch; their internal HTTP requests remain under CLI control.
| Field | Type | Description |
|---|---|---|
per_provider_rps |
map<string, number> | Per-provider requests-per-second cap, keyed by provider id. |
queue: rate_limit: per_provider_rps: openai-prod: 5 # max 5 rps to this LLM provider idqueue.retry — use attempts + backoff
Section titled “queue.retry — use attempts + backoff”Canonical fields
The canonical retry fields are attempts and backoff. The legacy
max_attempts / backoff_seconds pair is still accepted and normalized, but
deprecated — migrate to attempts + backoff.
| Field | Type | Default | Description |
|---|---|---|---|
attempts |
int > 0 | 3 |
Total attempts including the first try. 1 = no retry. |
backoff.kind |
enum | exponential |
exponential, linear, or constant. |
backoff.base_ms |
number > 0 | 5000 |
First/backoff base delay in ms. |
backoff.max_ms |
number > 0 | 60000 |
Cap on a single backoff delay. |
backoff.jitter |
bool | true |
Add random jitter. |
Trigger-level retry exists to absorb transient IO failures (timeouts, connection
resets, DNS blips, HTTP 408/5xx). Only errors classified as transient are retried;
deterministic failures — most notably context_overflow — are never retried,
regardless of attempts. Finer-grained retries also happen one layer down: LLM
provider calls, output-channel fetches, VCS CLI network operations, GitHub App
token exchange, and the issue-triage API client each retry transient IO errors up
to 3 times with a short exponential backoff before an error can fail the whole
trigger run. At the output and triage layers only idempotent methods retry;
non-idempotent POSTs do not blindly retry. Journaled automatic-batch mutations
use the remote reconciliation protocol for retries, including Feishu UUID
deduplication. HTTP 429 is left to the LLM gateway, which honors Retry-After.
queue: retry: attempts: 3 # transient-failure retries (1 = no retry) backoff: kind: exponential base_ms: 5000 max_ms: 60000 jitter: trueLegacy fields (deprecated, normalized)
Section titled “Legacy fields (deprecated, normalized)”For backward compatibility the loader still reads these and normalizes them, but new configs should not use them:
| Legacy field | Normalized to |
|---|---|
max_attempts |
attempts (floor of the value). |
backoff_seconds |
a constant backoff with base_ms = max_ms = backoff_seconds * 1000, jitter: false. |
attempts / backoff always take precedence when both are present.
queue.dead_letter — reserved, no effect yet
Section titled “queue.dead_letter — reserved, no effect yet”The schema accepts dead_letter.enabled and dead_letter.max_age_hours, but
the runtime does not consume them today: jobs that exhaust their retries are
marked failed and recorded in the run history — there is no separate parking
area. Both fields are reserved for a future release; setting them now changes
nothing.