TaskSultan Docs tasksultan.com

Orchestrator

Queues and work items

How the Orchestrator holds work as queue items, how a robot claims one, the statuses an item moves through and where the claim is genuinely atomic.

A queue holds work items. The Orchestrator keeps two families of them. Job queues live in queues/ and carry automation work at run level. Transaction queues live in transactions/ and carry business rows at item level. Both run on the same substrate, selected once by TS_QUEUE_PROVIDER.

#Work items in a job queue

Each queue is a file, queues/<queue>.json, holding an items array. An item records its id, its queue, a status, an arbitrary JSON payload, its attempt count against a maximum, whether it is retryable, who claimed it and when.

{
  "id": "process-invoices-item-m2x9-4f1c",
  "queue": "process-invoices",
  "status": "claimed",
  "payload": { "jobId": "job-process-invoices-m2x9-3a11" },
  "attempts": 1,
  "maxAttempts": 1,
  "retryable": true,
  "claimedBy": "robot-01",
  "claimedAt": "2026-10-03T09:12:00.000Z",
  "createdAt": "2026-10-03T09:11:59.000Z"
}

An item is queued, claimed, completed or failed. A claim bumps attempts and stamps claimedBy and claimedAt. A claim also takes a lease, ten minutes by default. A claimed item whose lease has expired becomes a candidate again, which is how work survives a robot that died mid-run. When an item fails and it is retryable and attempts remain, it returns to queued; otherwise it is terminal.

#The transaction queue

The transaction queue is the business-data sibling. Its statuses are pending, claimed, processing, successful, business-exception and system-exception. A business exception is an expected domain rejection and is terminal on first sight. A system exception is infrastructure trouble and retries while attempts remain, with a five-minute lease. A transaction can carry a business key such as INV-1042. A wave claims many items in one read and settles them in one write.

#The atomic claim

A claim must not hand one item to two robots.

On the file substrate the provider picks the first candidate in an in-memory copy of the file and writes the whole file back. Its exclusion therefore depends on the caller holding the tick lock around the claim. The dispatcher adds its own guard: it enqueues an item, claims it and refuses if the claimed id is not the id it enqueued, reported as a claim failure. On the SQLite substrate the candidate is chosen and marked inside one BEGIN IMMEDIATE transaction, so two processes cannot both walk away with the same row; the second blocks on the write lock and re-runs the query against the already-claimed row. That is what makes TS_QUEUE_PROVIDER=sqlite worth selecting under concurrency.

The buildout work list has its own claim: buildout claim --queue <q> --robot <id> takes the next item through the provider's atomic claim. Exit code 4 means there was nothing to claim, which is neither success nor failure.

#Working with the queue

The queue is not driven from the CLI directly. orchestrator run <jobdef.json> enqueues and dispatches one job; the scheduler and the server own the rest. To read job outcomes use jobs list and jobs show <id>. Transaction queues and their drains are reached through the Orchestrator server's /queues routes.

#Things that catch people out

Ordering is insertion order, not priority. There is no priority field in the queue contract. The SQLite candidate query orders by insertion sequence to match the file provider exactly.

A file item that is re-queued after a retryable failure keeps its old claimedBy and claimedAt. The SQLite provider copies that behaviour on purpose, so an operator reading a row cannot tell the substrate from the item's state.

A lease is the only protection against a lost robot, so a run that exceeds the lease can be reclaimed while it is still working. Keep long jobs inside the lease, or raise it deliberately.