← LAB / Articles

Technical companion

Saved, Accepted, Safe to Use: How Work Survives a Restart

The records, checks and theory that let a task outlive the process or session that started it — worked through one small task that is interrupted, given a corrected input and made to repeat a side effect.

By Loki

This page is the deeper companion to the article Work That Outlives Its Worker and to the interactive explainer, where you can switch the protections in the worked example on and off and watch the consequences. Here the same example is taken apart precisely: the definitions, the theory each mechanism comes from and what that theory assumes, a full trace with its arithmetic, the trade-offs, the failures the example does and does not cover, and the systems that already implement these ideas.

1. The problem, stated narrowly

Long-running agent work regularly outlives the session doing it. Anthropic's November 2025 account of long-running agent harnesses describes the symptom plainly: each new session starts with no memory of the last, and an agent that ran out of context mid-feature left the next session to guess what had happened. Its remedy was a progress file, a feature list and Git commits for the next session to read.25 That is continuity of context. This page is about something narrower and older: continuity of work.

Strip the model away and the question is one that transaction processing, workflow engines and distributed systems have been answering for decades:

When the process doing a unit of work stops — or merely stops answering — and the input it was working from may change while it is gone, what must already have been written down, and checked by whom, for someone else to continue the work correctly?

“Correctly” has four separate meanings, and each fails in its own way:

  • Progress is not lost, and a half-done step is not mistaken for a done one. Otherwise work is redone at cost, or an unfinished step is treated as finished.
  • An effect on the outside world happens once. Otherwise a customer gets two emails, a card is charged twice, a feed shows two announcements.
  • Only the current owner writes. Otherwise a worker presumed dead wakes up and overwrites its successor.
  • An output is only relied on for the input it was judged against. Otherwise something accepted against an old input is used as if it were current.

None of these is new. What agent work changes is the economics and the shape of the steps: a model call is slow, costs money and does not reliably return the same text twice, so recovery should restore recorded results rather than recompute them, and a result that does get recomputed is a new result that has to be judged again.

2. Definitions

These are the terms as used on this page and in the worked example. Several of them are ordinary engineering vocabulary given one fixed meaning here; where a term is a policy of this example rather than a general law, that is said.

Work identity
A durable name for the job, independent of who is doing it: here digest/q3-report. It survives any number of attempts.
Attempt
One admitted try at the work, numbered #1, #2, … An attempt is not a session or a process; one attempt could span several of either, and a session could be replaced without the work restarting.
Owner token (epoch)
A number issued with each attempt, strictly larger than any issued before. The attempt carries it on every write; protected stores refuse writes carrying an older one.
Pinned input
The exact input version an attempt is working from, named by content hash and read from an immutable snapshot of that version.
Persisted
A step's result is durably recorded, keyed by what it was computed from. This is progress: it lets a successor restore the step instead of redoing it. It says nothing about whether the result is any good.
Finished (candidate)
The worker has assembled an output it considers complete. Nobody has judged it.
Accepted
A candidate passed an assessment and a transition recorded that fact for a named input version and assessment contract. In this example the assessment is done outside the producing worker; that is a design choice, not part of what durability means. Acceptances are appended, never edited; a later one supersedes an earlier one.
Dependable (for a given consumer)
Another task may rely on it now: it is accepted and it was accepted for the input version that consumer requires. A result can remain accepted, as history, while ceasing to be dependable for a consumer that requires the latest input.
Operation key
An identifier for an intended external effect, derived from what is being done (this announcement, for this accepted digest) rather than who is doing it. The receiver uses it to recognise a repeat.
UNKNOWN
The status of an attempt that has stopped answering. It is neither success nor failure: a silent attempt may have crashed, may be paused, or may have finished its step without recording it.

Written as a predicate, with d a digest, c a consumer and t the moment of admission:

dependable(d, c, t) ⇔ accepted(d) ∧ builtFrom(d) = required(c, t)(1)

In the worked example the newsletter's policy is required(newsletter, t) = the latest report version at t. A historical consumer could just as legitimately require an older version; the definition is relative to the consumer's declared requirement, not to “latest” in general.

3. Five records, five checkers

Durability is not one save. It is a handful of separate records, each with its own job and, crucially, its own checker — the party that must look at the record for it to protect anything.

What each record is for, who checks it, and what goes wrong without it in the worked example
RecordJobChecked byWithout it, in the example
Attempt admission: identity, pinned input, owner token — written before dispatchTurns silence into a specific question (“did attempt #1's step 3 finish?”) instead of a guessThe controller, when an attempt goes quietNothing can tell a lost attempt from a finished one
Progress journal: each finished step with the hash of what it was computed fromLets a successor restore finished steps and refuse records computed from an input that has since changedThe successor, before reusingWith version checks off, a digest quoting 4.1 million is accepted and used while the report says 4.7 million
Owner token checked by the storesMakes a worker that only seemed dead harmless when it wakesEvery protected storeWith fencing off, 3 late writes from the old worker land after takeover; with no input change, two different digests are accepted
Operation key on each external effectLets a resend after a lost acknowledgement land onceThe receiver of the effectA key per attempt, or none: the announcement is posted twice
Acceptance naming the input version and contract it was judged againstKeeps “accepted” and “safe to depend on” apart when the input movesThe dependant's admission ruleA newsletter admitted on an accepted but superseded digest quotes 4.1 million

The rest of this page takes the mechanisms one at a time, then runs them together.

4. Intent before work: what WAL does and does not give

Write-ahead logging is the oldest idea here. ARIES, the reference recovery method for write-ahead-logged databases, states the rule exactly: log records describing changes to some data must already be on stable storage before the changed data is allowed to replace its previous version on non-volatile storage.1 On restart ARIES analyses the log from the last checkpoint, then repeats history by redoing logged updates, and only then undoes the work of transactions that did not commit. The redo is conditional on a test that tells it whether an update's effect is already present on the page:

redo(r) ⇔ pageLSN(page(r)) < LSN(r)(2)

Two ideas carry over to durable work; the algorithm itself does not.

  1. Record before acting, so the record bounds what may have happened. In the example the controller writes attempt #1's admission — identity, pinned input, owner token 41 — before dispatching any work. This is WAL-inspired ordering, not ARIES's rule: ARIES orders a log write before a page write; here an attempt record is ordered before dispatch. After a failure the record does not say whether the work ran. It says what has to be reconciled.
  2. Make redo conditional on a test. ARIES's test compares sequence numbers. The example's test is “is there a journal record for this step whose input key matches mine?” (section 6). For an external effect, the test has to live where the effect lives: at the receiver (section 8).

What does not carry over: ARIES redoes deterministic page updates. A model call is not a deterministic update, so a journal for agent work stores the step's result and restores it, rather than storing an operation and executing it again.

The rollback-recovery literature adds the companion rule for effects that leave the system. Elnozahy and colleagues call it the output commit problem: the outside world cannot be relied on to roll back — a printer cannot unprint, an ATM cannot recover dispensed cash — so before output is sent, the state from which it is sent must be recoverable despite any future failure. They also note that inputs from the outside world may not be reproducible during recovery, and that a common approach is to save each input to stable storage before processing it.5 The example applies both: the report version is recorded and pinned before work, and the announcement is sent only from a durably recorded state. It additionally requires that state to be an accepted digest; that is this example's policy, not part of the output-commit rule.

5. Restore, redo and the determinism boundary

It is tempting to say that durable execution needs deterministic work and agent work is not deterministic, so durable execution does not apply. That is wrong in a precise way. Replay-based engines require determinism of the orchestration code — the part that decides what to do next — and record the results of everything else.

  • Temporal's documentation states that workflow code must be deterministic to support replay, and that non-deterministic operations such as API calls, LLM/AI invocations and database queries belong in Activities, which execute outside the replay path.9
  • Azure Durable Functions replays an orchestrator from the start against its execution history; when the history shows an activity already ran, the framework replays that activity's recorded result rather than running it again. Orchestrators must therefore be deterministic, and must take time and GUIDs from the orchestration context.1415
  • LangGraph's Functional API restores completed task results from the checkpointer instead of recomputing them, “including for long-running or non-deterministic task outputs”, and notes that different runs of a workflow can produce different results while resuming a specific thread replays the same persisted results.16
  • Restate's ctx.run wraps a non-deterministic operation and stores its result in the execution log.18

Schneider's state-machine tutorial states the property underneath: a state machine's outputs are completely determined by the sequence of requests it processes, independent of time and other activity.6 Replay works for the part that has that property, provided every non-deterministic input to it has been logged — the piecewise-deterministic assumption of the rollback-recovery survey.5

For agent work, then: a model call belongs on the recorded-result side of the boundary. Restoring it reuses exactly what was recorded. Redoing it is a new execution which may produce different output — not must — and whatever it produces has to be assessed in its own right. The worked example makes this visible by stipulation: it gives every computed step a schematic result identifier derived from the attempt that computed it, so a restored step keeps its old identifier and a redone step gets a new one. The model does not generate summary text and does not claim real reruns differ.

6. Keying progress by its inputs

A journal that records “step 2 is done” is not enough once inputs can change. The record has to say what step 2 was computed from. Content addressing is the standard way to do that. Git is, in its own documentation's words, a content-addressable filesystem whose core is a key-value store keyed by a hash of the content.22 Bazel's remote cache keys each action by a hash of its declared inputs, command line and environment, and stores outputs in a content-addressable store.23 The example applies the same idea to each step:

key(s) = H( contract ‖ declaredInputs(s) ) restore(s) ⇔ ∃ record r : r.step = s ∧ r.key = key(s) otherwise redo(s)(3)

In the example each section summary declares one input — its own section text — and the assessment contract is held constant, so the key reduces to the section's content hash. When the report changes in section 2 only, a successor restores sections 1, 3 and 4 and redoes section 2.

Reuse by key is sound under two conditions, and Bazel's own documentation lists a known issue for each:

  1. The key covers everything the output depends on. Bazel does not track tools outside the workspace, so two users with different compilers can wrongly share cache hits: the outputs differ but the action hash is the same.23 In the example, if a section summary quietly read another section, section-level reuse would be unsound.
  2. The bytes the step read are the bytes the key names. Bazel warns that when an input file is modified during a build, it might upload invalid results to the remote cache.23 Hashing a mutable file at admission and later reading it does not guarantee that. The example assumes each attempt reads an immutable snapshot of its pinned version; a real system needs that snapshot, or equivalent isolation. Rechecking a version label the worker reports about itself is not enough.

The example exercises version selection and stale reuse across two versions. It does not model a file changing during a read.

7. Leases, suspicion and fencing

A controller cannot see inside a worker. It sees heartbeats, or their absence. Gray and Cheriton defined a lease as a contract giving its holder specified rights over property for a limited time. Because the term is bounded, a grantor that loses contact with a holder only has to wait for the lease to run out, not for the holder to be found.3 Their paper is explicit that leases depend on well-behaved clocks: a server clock that runs fast can allow a write before a previous holder's lease has expired at that holder, and a client clock that runs slow can keep using a lease the server considers expired. The minimum requirement is a known bound on clock drift.

Even with good clocks, expiry is only suspicion. Kleppmann's essay on distributed locking gives the failure that matters: a lease holder pauses (a garbage-collection pause, a page fault, a stopped process), its lease expires, another client takes over, and the first resumes and writes, believing it still holds the lock. Checking the lease just before writing does not fix this, because the pause can fall between the check and the write. The fix is a fencing token, a number that increases every time the lock is granted, sent with every write; the storage remembers the highest token it has processed and rejects a write with a lower one. In his words, this requires the storage to take an active role in checking tokens.19

The example uses a stronger arrangement that is worth stating exactly, because the difference matters:

model: apply(w) ⇔ w.token = epoch (one shared epoch, compared atomically with the write) Kleppmann: apply(w) ⇔ w.token ≥ max token this store has seen(4)

In the model every protected store — the journal, the candidate store, the effect records and the acceptance transition — compares against one authoritative owner epoch, and the controller publishes the new epoch before the successor is admitted. A store that only remembers tokens it has seen on writes may still accept token 41 after 42 has been issued but before any write carrying 42 has reached it. A real system with separate stores needs the corresponding synchronisation: conditional writes against a shared epoch, or the successor's first write reaching each store before the old owner can, and monotonic epochs that survive the controller's own restarts.

Three limits follow, and they are the ones worth remembering:

  • Fencing does not prove the old worker stopped. In the example the woken attempt #1 is still alive after its write is rejected. Fencing removes its write authority; it says nothing about its health.
  • Only stores that check tokens are protected. The external feed in the example checks idempotency keys, not owner tokens.
  • A false suspicion wastes work. If the worker was merely slow, its in-flight step is lost and redone. Fencing keeps that from corrupting anything; it cannot make it free.

8. Effects: at least once, applied once

Here is the hardest window in the example. The worker records that it intends to post the announcement, sends it, the feed applies it, and the worker dies before it records the acknowledgement. Its successor finds an intent with no confirmation. That local record is compatible with two different histories: the post never arrived, or it arrived and the acknowledgement was lost. Nothing on the worker's side can tell them apart.

Saltzer, Reed and Clark made this point in 1984: reliably detecting a crash is problematic, because the problem may be a lost or delayed acknowledgement, in which case the retried request is a duplicate that only the application can discover. Duplicate suppression must be done by the application with knowledge of its own requests.4 Distributed commit has a related result — Bernstein, Hadzilacos and Goodman show that a participant that fails during its uncertainty period cannot reach a decision alone and that no atomic commitment protocol guarantees independent recovery2 — but that theorem is about voting protocols, which the example does not have. It is an analogy for why someone outside the worker must be asked, not the reason in this case. The reason here is simpler: only the receiver saw whether the effect happened.

So the practical pattern is: send at least once, and make the effect idempotent at the receiver. Temporal's documentation defines it: performing an operation multiple times has the same result as performing it once. For an operation k acting on receiver state S:

applyk( applyk(S) ) = applyk(S)(5)

A receiver achieves that for an effect that is not naturally idempotent by remembering, for each key, the request it was given and the answer it gave. In the example the receiver is the feed, and the post lives in the feed's own store. From the worker's side the feed is external and nothing the worker does can roll a post back; from the feed's side, the post and the feed's record of it can be written in one transaction:

receive(key, req):
  transaction, serialized per key:
    # nothing else on this key interleaves
    saved = remembered.get(key)
    if saved is not None:
      if saved.params != req.params:
        return reject("different request")
      return saved.result     # same answer
    outcome = apply(req)      # the post
    remembered[key] = (req.params, outcome)
  return outcome

Everything inside transaction runs as one unit that no other request for the same key can interleave with. A per-key lock, a serializable transaction, or a unique constraint on the key that aborts the losing commit before its post persists would each do. The lookup has to be inside that unit. If it sits outside, two copies of the same request arriving together can both find nothing remembered and both post, and the second answer overwrites the first. Each part of the recipe answers a requirement in the sources:

  • The receiver enforces it. Temporal states that idempotency keys are enforced by the service the activity calls, not by the activity itself, and recommends a key built from the Workflow Run ID and Activity ID because that is consistent across retry attempts but unique among executions.8
  • The lookup, the effect and the record are atomic and isolated. The AWS Builders' Library article on idempotent APIs requires “the process that combines recording the idempotent token and all mutating operations related to servicing the request” to meet the properties of an atomic, consistent, isolated and durable (ACID) operation, and its example flow begins with the service checking whether it has seen the identifier before.20 Atomicity stops the token being recorded without the resource, or the resource created without the token. Isolation is what stops two concurrent copies from both missing the lookup. AWS does not spell out that concurrent case. Stripe does: a request that conflicts with another request executing concurrently is not saved, and can be retried.21
  • The key names the operation, not the attempt, and is checked against the stored request. If each attempt generates its own key, a successor's resend looks new. AWS stores the parameters of the first request with its client request identifier and returns a validation error for a parameter mismatch, on the reasoning that the customer may have intended a different outcome. Stripe compares incoming parameters with the original request's and errors if they differ, “to prevent accidental misuse”.2021 The recipe stores req.params beside the result and rejects a mismatch.
  • A repeat gets the stored answer. AWS returns a semantically equivalent response to a request it has already seen. Stripe saves the status code and body of the first request for a key “regardless of whether it succeeds or fails” and returns them for later requests with that key, including errors.2021 In this recipe a failed post rolls the whole transaction back and nothing is remembered, so a retry tries again. Remembering the failure, as Stripe does, is the other defensible choice.
  • Keys expire. Stripe allows keys to be removed once they are at least 24 hours old, after which a reused key creates a new request.21 A retry that arrives after the receiver forgot the key is a new effect.

With the lookup, the parameter check, the post and its record serialized in one transaction, and the key kept for as long as a retry can arrive, the guarantee is one application as far as the receiver can tell, not exactly-once execution. Temporal is careful about the same distinction: with retries, an activity is observed as completed exactly once, yet may be executed multiple times and may even partially complete more than once.8 Where a receiver offers neither keys nor a way to ask whether the effect happened, the remaining tool is compensation — a saga in Garcia-Molina and Salem's sense, a sequence of transactions each with a compensating transaction to amend a partial execution. Compensation amends rather than prevents: a sent email can be followed by a correction, not unsent. Their paper also notes the case where compensation itself cannot complete and the system is stuck.7

9. Three lifecycles, not one

Mature engines name every lifecycle state rather than inferring it; Airflow's task instances, for example, move through explicit states such as queued, running, success, failed, up_for_retry and deferred, and tasks whose heartbeats time out are detected and failed or retried.24 The example keeps three small machines and keeps them separate, because they answer different questions:

ATTEMPT — is this worker allowed to act?
  admitted ─dispatch─▶ running
  running ─missed heartbeats─▶ UNKNOWN
      (never success, never failure)
  running | UNKNOWN ─write rejected─▶ fenced
      (stale token; may still be alive)
  running ─candidate refused─▶ refused
  running ─post confirmed─▶ done

ARTIFACT — what is the status of this output?
  candidate ─assessed, recorded─▶ accepted
      (for a named version and contract)
  candidate ─assessed, refused─▶ refused
  accepted ─input superseded─▶ accepted
      (history; not dependable for “latest”)

EFFECT — did the outside world change?
  intent recorded ─sent─▶ unconfirmed
  unconfirmed ─receiver answers─▶ confirmed

Keeping them apart is what lets the example say true things at the same moment that a single status field would contradict. At t13 in the main trace, attempt #2 is UNKNOWN; its digest's acceptance at t10 still stands; the announcement is unconfirmed; and the newsletter has already started on the accepted digest. Silence also gets its own state. A system that maps “process gone” to failed throws away work that may have finished; one that maps it to success invents completion.

10. Accepted versus dependable

Acceptance is a judgement recorded for an input version and a contract. Two policy choices in the example deserve to be read as policies:

  • Acceptances are append-only. The example never edits or withdraws one; a later acceptance supersedes an earlier one. This keeps history honest. It is not a law of durable systems: a flawed result can be replaced by a new assessment of the same input, recorded as a new disposition, without pretending the first judgement was never made. (Atomic commitment protocols do make individual commit decisions irrevocable,2 but that is about protocol decisions, not about whether a semantic assessment can later be revised.)
  • At most one accepted digest per report version, under one fixed contract. The example holds the assessment contract constant and never changes it.

The dependant's admission rule is where accepted and dependable separate. The example offers the newsletter three rules: take the latest finished candidate, take the latest accepted digest, or take an accepted digest for the current report (equation 1). Only the third prevents the superseded-input failure.

One case needs care. If the records carry no version at all, the third rule cannot be checked. The reader-facing model deliberately makes that rule fail open in that case — it behaves as “take the latest accepted” — to show what an unchecked freshness requirement does. A rule that fails closed would wait instead. Across the 144 input combinations where versions are off and the strict rule is selected, failing open lets the newsletter use a wrong digest in 72; failing closed prevents every one of those, but the newsletter is then never admitted in any of the 144 (it can never prove freshness), and the wrong digest is still accepted upstream in 78 of them either way. Refusing to proceed protects the consumer; it does not repair the producer.

A conservative alternative avoids mixing versions altogether: pin a run to its original input and start a new execution for a new input. AWS Step Functions' redrive continues a failed Standard execution from the unsuccessful step using the same input and the same state-machine definition version, preserving the results of successful steps, and requires a new execution for an updated definition.12 Two qualifications from the same documentation: preserved successful work has documented exceptions — after a States.DataLimitExceeded failure, a Parallel state, an Inline Map state or a Distributed Map state is rerun including its successful branches, iterations or child workflows — and redrive is only available for 14 days after the execution ends. Redrive does not itself provide the example's cross-version reuse of unchanged sections; whether a new execution reuses earlier results is an application-level choice, not a consequence of pinning.


11. The worked example

A teaching model, not a measurement. The task, report, timings and identifiers below are an original construction. Counts of model calls, ticks and posts are properties of this small model, not measurements of any system, and nothing here depicts a real run or data. The interactive explainer computes exactly the same model.

Setup

Task digest/q3-report summarises a four-section report into a digest, one model call per section, then assembles a candidate digest, has it assessed and accepted by the controller's transition, and posts one announcement to an external feed. A newsletter, ready at t12, depends on the digest and quotes its revenue figure.

The two report versions. Hashes are the first 8 hex characters of SHA-256 over the text.
Sectionv1 textv1 hashv2 hash
s1Overview. Q3 ran from July to September and the team shipped two releases.45a1f4c045a1f4c0
s2Revenue. Revenue for the quarter was 4.1 million. (v2: 4.7 million)43d57587ad64f04a
s3Costs. Operating costs held at 3.2 million.2f08201d2f08201d
s4Outlook. Q4 plans a third release and no new hires.7dac6c4a7dac6c4a
ReportSections joined by newlinesba0184f5cf47dcd5

Fixed timings, in ticks: the first owner token is 41; an attempt is presumed lost after 2 missed heartbeats; attempt #1 goes silent as it starts s3 at t3; if it was only paused it wakes at t8; the report is corrected at t4 (“during the outage”) or at t12 (“after acceptance”). The protections are on: per-step checkpoints, fencing, version pins, an operation key on the post and the strict newsletter rule. Result and digest identifiers such as 3af8da4e are schematic labels, as explained in section 5.

Main trace: interrupted, corrected, retried

Main trace. Old worker paused; report corrected during the outage; the posting worker dies after the feed applies the post.
tRecordEvent
t0attemptController admits #1, token 41, pinned to report v1 ba0184f5, recorded before dispatch.
t1journal#1 computes s1 from v1 → result 3af8da4e, recorded.
t2journal#1 computes s2 from v1 → result 8dc2b770, recorded.
t3attempt#1 starts s3 (the model call is made), then goes silent: no heartbeat, nothing recorded.
t4inputReport v2 cf47dcd5 recorded: section 2 corrected from 4.1 to 4.7 million; sections 1, 3 and 4 unchanged.
t5attempt#1 has missed 2 heartbeats: outcome UNKNOWN (s3 in flight). Crashed or paused cannot be told apart.
t5attemptController admits #2, token 42, pinned to v2 cf47dcd5.
t5journal#2 restores s1 from 3af8da4e (section hash identical in v2); must redo s2 (recorded for v1, section changed); must do s3 and s4 (never recorded).
t6journal#2 computes s2 from v2 → 2b1a6af0, recorded.
t7journal#2 computes s3 from v2 → 59977556, recorded.
t8attempt#1 wakes (it was paused, not dead) and tries to record s3. The store rejects it: token 41 < current 42. #1 stops — alive, without write authority.
t8journal#2 computes s4 from v2 → 3b74fabd, recorded.
t9artifact#2 assembles candidate digest 9723ab9c (claims v2; quotes 4.7 million): finished, not yet accepted.
t10artifactController accepts 9723ab9c for v2 under the fixed contract; its part hashes match v2.
t10effect#2 records intent to post; the feed applies post P1 with key announce:digest/q3-report:9723ab9c; #2 dies before recording it.
t12dependantNewsletter admitted on 9723ab9c under the strict rule; quotes 4.7 million.
t13attempt#2 has missed 2 heartbeats: UNKNOWN (post confirmation in flight). The t10 acceptance is unaffected. Controller admits #3, token 43, to reconcile the post.
t14effect#3 resends with the same key; the feed answers “already seen” and returns P1, posting nothing new. #3 records P1 as confirmed.

Outcome: 6 model calls against a minimum of 4 (#1 made 3, #2 made 3, #3 made none); 1 step restored; 1 stale write rejected and none applied; 1 post delivered and no duplicates; the digest first accepted for the current report at t10; the newsletter started at t12 quoting 4.7 million; no invariant broken. Note that the newsletter starts at t12 although the announcement is not confirmed until t14: dependability rests on acceptance, not on the announcement.

The cost decomposes cleanly. With a checkpoint after every step,

calls = (calls started by interrupted attempts) + (steps the successor cannot restore) main trace: 6 = 3 + 3 (#1: s1, s2, s3 in flight; #2: s2 changed, s3, s4 never recorded) end-only saves: 8 = 4 + 4 (#1's s1, s2 never reached the journal; the paused #1 also spends a call on s4)(6)

With checkpoints only at the end, attempt #1's two finished steps were never written, so #2 redoes all four; and because the woken #1's first write is now postponed until it assembles, it spends one more call before it is fenced. The digest is accepted at t11 instead of t10. Persisting more often is a cost lever in this model, never a correctness one — which matches LangGraph's description of its exit, async and sync durability modes as a performance-versus-durability trade.17

Second trace: accepted, then superseded

Same protections, but the report is corrected only at t12, after a digest was accepted for v1.

Second trace. Input change after acceptance.
tRecordEvent
t0–t3journalAs before: #1 (token 41, v1) records s1 and s2, then goes silent starting s3.
t5attempt#1 UNKNOWN; #2 (token 42) pinned to v1 restores s1 and s2 (same version) and must do s3 and s4.
t6–t7journal#2 computes s3 → 59977556 and s4 → 3b74fabd.
t8attemptThe woken #1's write is rejected (41 < 42). #2 assembles 5573113e (v1; 4.1 million).
t9artifactAccept 5573113e for v1 — correct for its input. #2 posts P1 and dies before recording it.
t12inputReport v2 recorded. The v1 acceptance stays on record; it is no longer dependable for a consumer that needs the latest report. #2 UNKNOWN; #3 (token 43) admitted to reconcile the unconfirmed post. The newsletter is ready but the strict rule holds it: no digest is accepted for v2 yet.
t13effect#3 resends with the same key; the feed returns P1; confirmed.
t14journal#4 (token 44) pinned to v2 restores s1, s3 and s4 (section hashes identical) and redoes only s2.
t15–t16artifact#4 computes s2 → ee372748; assembles 43fc86ed (v2; 4.7 million).
t17dependantAccept 43fc86ed for v2; the newsletter is admitted on it (4.7 million); #4 posts P2, a new operation with a new key, and confirms it.

Outcome: 6 model calls (#1 3, #2 2, #3 0, #4 1); two distinct announcements and no duplicates; the newsletter waits from t12 to t17; no invariant broken. Because only section 2 changed, the rebuild after the correction costs 1 model call. Rebuilding everything for v2 would have cost 4. That comparison is between two stipulated policies in this model, not a measured property of any engine. With the newsletter rule relaxed to “latest accepted” or “latest finished”, the newsletter instead starts at t12 on the v1 digest and quotes 4.1 million while the report says 4.7 million.

The failure path: version checks off

Repeat the main trace with version pins off: records keep no input hashes, so reuse is by step name alone and acceptance cannot compare versions.

Failure trace. Identical to the main trace until t5.
tRecordEvent
t5journal#2, pinned to v2, restores s1 and s2 from the v1 record 8dc2b770 — the section changed, but versions are not compared. It computes only s3 and s4.
t8artifactThe woken #1 is still fenced. #2 assembles 5573113e — the same schematic identifier as the legitimate v1 digest, because it was assembled from the same recorded results — claiming v2 while quoting 4.1 million.
t9artifactAccepted, with no version recorded: one tick earlier than the correct path, after 5 model calls instead of 6.
t12dependantThe newsletter asked for the current version, but the records carry none; the rule fails open and takes the accepted digest: it quotes 4.1 million while the report says 4.7 million.
t12–t13effect#3 reconciles the post; the feed returns P1.

Cheaper, faster and wrong. Two invariants break: a digest built from a superseded input was accepted, and the dependant consumed it. With a fail-closed rule the newsletter would wait instead; the wrong digest would still be accepted.

One change at a time

Single and paired changes from the main trace (all computed by the same model)
ChangeCallsResult
Old worker crashed instead of paused6Nothing wakes, so no write is rejected; otherwise identical — which is the point: the controller could not tell.
No input change5#2 restores s1 and s2; 5573113e (4.1 million) accepted at t9 — correct, since v1 is current.
Input change after acceptance6The second trace: newsletter waits until t17; 1-call rebuild.
No crash after the post6Two attempts; one post, confirmed directly.
Fencing off73 stale writes applied after takeover (s3, s4, the candidate). The stale v1 candidate is refused at t11 because it was built for v1 while v2 is current.
Fencing off, no input change6Two different digests accepted for v1: 5573113e at t9 and the stale worker's 6025f0d0 at t11.
Fencing off, newsletter takes latest finished7The newsletter takes the stale, never-accepted 6025f0d0 quoting 4.1 million.
Version pins off5The failure path above: 4.1 million accepted and used.
Post key per attempt, or no key6#3's resend looks new: the announcement is posted twice.
Newsletter takes latest accepted, or latest finished6No difference here: the latest candidate is already the accepted v2 digest.
…the same, with the input changed after acceptance6Newsletter admitted at t12 on the v1 digest: 4.1 million.
Checkpoint only at the end8#1 makes 4 calls and #2 makes 4; accepted at t11; no invariant broken. Cost, not correctness.

All 864 combinations

The interactive exposes eight inputs: what the silent worker really did (paused or crashed), fencing, checkpoint frequency, when the input changes (during the outage, never, after acceptance), version pins, the newsletter rule (current accepted, any accepted, latest finished), whether the posting worker dies, and the post key (per operation, per attempt, none). That is 2 × 2 × 2 × 3 × 2 × 3 × 2 × 3 = 864 combinations, and the model was run on every one.

  • With all four protections on — fencing, pins, a per-operation key, the strict newsletter rule — all 24 runs (12 event combinations × 2 checkpoint modes) break no invariant and finish: an accepted digest for the final input, its announcement confirmed. They cost 5–6 model calls (mean 5.67) with per-step checkpoints and 7–9 (mean 7.83) with end-only checkpoints.
  • Weakening exactly one protection, with per-step checkpoints, over the 12 event combinations:
Necessity within this model: one protection removed, the others kept
RemovedBrokenWhich cases
Fencing6 / 12Every case where the old worker was paused: stale writes applied; also two digests accepted for one version unless the input changed during the outage.
Version pins8 / 12Every case with an input change: superseded digest accepted and used (fail-open rule).
Operation key (none or per attempt)6 / 12Every case where the posting worker dies: duplicate announcement.
Strict newsletter rule (any accepted, or latest finished)4 / 12Every case where the input changes after acceptance: newsletter quotes the old figure.

“Necessary” here means necessary within this model and these exact switch definitions. It is exhaustive over the configured combinations, not over arbitrary interleavings, crash points, clock behaviour or storage failures. It also shows that fencing and version pins guard different things: with fencing off and pins on, stale writes pollute the record but no wrong digest is accepted; with pins off and fencing on, no stale write lands but a wrong digest is.

12. Trade-offs

ChoiceBuysCosts and risks
Persist each step, or only at the endLess redone work; a paused worker is fenced at its first writeA write per step and its latency. In the example, 6 calls versus 8.
Restore recorded results, or recomputeThe same result as before; no extra model callsStorage; a restored result keeps whatever flaws it had until it is assessed
Reuse keyed by declared inputsA precise, cheap rebuild after a partial change (1 call versus 4)Unsound if a step reads an undeclared input or a mutable one; hashing discipline everywhere
Pin a run to its input; new input means new executionSimple and conservative: never mixes versionsReuse across inputs has to be provided separately, if at all
Lease plus fencingA bounded time to replace a stuck worker; stale writes harmless wherever tokens are checkedA slow worker falsely suspected loses its work; every protected store must check; lease timing assumes bounded clock drift
At-least-once delivery plus receiver deduplicationThe effect is applied once as far as the receiver can tellNeeds a receiver that records keys atomically with effects and keeps them long enough
Assessment outside the producer, recorded by a transitionThe producer cannot accept its own work; acceptance names version and contractAnother boundary to build; assessors are fallible and can be lenient26
Strict rule for dependantsNo dependant built on a superseded inputDependants wait after a change (t12 to t17 in the example); without provenance a fail-closed rule never proceeds

13. Failure cases outside the model

The model demonstrates five failures: a duplicate effect without a stable key; a stale owner writing after it was replaced; acceptance of output built from a superseded input; a dependant using a superseded or never-accepted output; and two outputs accepted for one input. These it does not cover:

  1. A worker paused inside the effect. If an owner had already passed acceptance and was paused mid-post, its late send reaches the feed, which checks keys but not tokens. A per-operation key deduplicates it; anything else duplicates. Fencing cannot protect a resource that does not check tokens.
  2. A receiver with no idempotency. If the receiver offers no key and no way to ask whether an effect happened, at-least-once means possible duplicates. Compensation or a human check remain.
  3. Key expiry and late arrivals. A retry arriving after the receiver has pruned the key is a new request.
  4. Clocks outside their assumed bound. Lease timing is wrong; fencing prevents corruption but not wasted work.
  5. Undeclared or mutable inputs. Reuse by hash is unsound if a step read something outside its key, or bytes other than the ones hashed.
  6. A fallible assessor. An acceptance records what was judged, not that the judgement is right. In the example the assessor checks completeness and form, not every figure.
  7. Journal and effect are never one transaction. The window between recording intent and the receiver applying the effect always exists; the design has to assume it.
  8. Orchestration code that changes under running work. In replay engines, changing workflow code under running executions causes non-determinism errors9 — a code-version cousin of the changed-input problem.
  9. The controller itself restarting. The example assumes the controller survives and its records are on reliable storage. A restarted controller would need, at least, durable attempt records, the owner epoch, unconfirmed effect intents and acceptances; it would have to rebuild its suspicion timers rather than trust them, and keep the epoch monotonic across its own restart. That is a requirement, not a trace this model runs.
  10. Losing the journal. All of this assumes stable storage. Durable work is only as durable as its records.

14. Prior art and current engines

Nothing in the example is new. It is assembled from mechanisms the following systems document, and it is not offered as an alternative to them or as better or worse than any of them. The table states what each page said when it was read in October 2026; vendor documentation changes, and each statement is about that system's own semantics.

Documentation read October 2026
SystemWhat its documentation statesScope to keep
TemporalEvent history is an append-only, durably persisted log; workflow code is replayed and must be deterministic; LLM calls and other external interactions go in activities; activities retry by default; idempotency keys are enforced by the called service; a key from Run ID plus Activity ID is stable across retries.8910“Observed as completed exactly once”, yet possibly executed more than once
AWS Step FunctionsStandard workflows follow an exactly-once model in which tasks and states are never run more than once “unless you have specified Retry behavior”; asynchronous Express workflows are at-least-once and synchronous Express at-most-once. Redrive resumes a failed Standard execution from the failed step with the same input and definition. StartExecution is idempotent for a running Standard execution with the same name and input.111213Orchestration semantics; an external effect retried by a Retry clause still needs the receiver's cooperation. Redrive exceptions noted in section 10.
Azure Durable FunctionsOrchestrations are event-sourced; replay returns recorded activity results; orchestrators must be deterministic and take time and GUIDs from the context.1415Determinism constraints apply to orchestrators, not activities
LangGraphCompleted task results are restored, not recomputed, including non-deterministic ones; a task that started but did not finish may run again, so side effects should be idempotent; three durability modes; per-task pending writes survive a failed sibling.1617Different runs may produce different results; resuming a thread replays persisted ones
RestateA journal of steps and results; ctx.run stores a non-deterministic operation's result in the execution log and retries failures unless a terminal error is thrown.18The pages read do not state what happens to an external effect inside ctx.run if the process fails after the effect but before the journal entry; no claim is made either way
Apache AirflowExplicit task-instance states; tasks whose heartbeat times out are detected and failed or retried.24A scheduler's lifecycle; retained artifacts are a separate question
Bazel, GitContent-addressed storage; actions keyed by declared inputs, with documented known issues for undeclared tools and inputs modified during a build.2223Reuse only as sound as the declared inputs

Two other families are worth naming. Whole-process checkpoint and restore is simpler, but the rollback-recovery survey judges checkpoint-based recovery ill suited to applications that interact frequently with the outside world.5 And agent harnesses that hand over through progress files and Git commits, with a separate evaluator judging when work is done, address continuity of context and quality of judgement.2526 Anthropic's March 2026 follow-up reports that a separate evaluator was a strong lever but needed tuning against leniency, and that its full harness cost over twenty times as much as a solo run in that comparison. These complement the records described here; they do not replace them.

The honest summary of “why not just use an engine?” is: for many systems, do. The engines above implement most of these records. What none of them can do on your behalf is make a third party's effect idempotent, know which inputs your agent actually read, or decide what “accepted” means for your output.

15. A dated design example: Era 3

Loki's Era 3 is a proposed design for running agent work from durable factory records rather than from any one session. As of its draft architecture documents of September 2026, it is a design, not a running system: no Era 3 core, factory engine or session-capture runtime has been implemented. Its two state-walkthrough teaching applications are real, but they deliberately keep every input fixed and leave controller recovery and changed inputs out.

The draft requires most of the properties on this page and, for most of them, has not yet chosen a mechanism:

What the September 2026 drafts require, and whether a mechanism is selected
PropertyDraft requirement (paraphrased)Mechanism selected?
IdentityAn execution attempt is distinct from its session and process; work identity survives a change of sessionLinkage still to be defined
Intent before workAttempt identity committed before dispatch; admission alone does not prove executionStorage technology undecided
SilenceA process disappearing does not establish completion—
Progress versus acceptanceA committed checkpoint is work history, not acceptanceOpen
Changed inputsInvalidation after changed inputs must be designed; a stale assessment cannot accept against a changed contractNo
Ownership after restartA replaced owner's write authority and effects must be reconciled before a new writer is admitted, including if it was partitioned rather than deadNone named
EffectsPersist operation intent, collect acknowledgements, reconcile ambiguity; never claim universal exactly-oncePer operation type, undecided

The worked example on this page is a small paper rendering of the kind of proof that design asks for — interrupt a worker, replace it, continue from records without losing accepted decisions, inventing completion or repeating recorded effects — built from textbook mechanisms. It is not evidence that Era 3 does any of this. Fencing tokens, section-hash keys and per-operation keys are choices a builder might make, not choices the design has made.

16. Where the model stops

The model's assumptions, all of which the interactive shares:

  1. One model call per section summary. Result and digest identifiers are schematic: restored steps keep theirs; redone steps get new ones by stipulation. No summary text is generated or compared.
  2. A section summary depends only on its own section and the fixed contract, and each attempt reads an immutable snapshot of the version it is pinned to.
  3. The assessor checks completeness and form; version consistency is checked mechanically by comparing part hashes with the pin.
  4. One assessment contract, held constant. The newsletter requires the latest report. At most one digest is accepted per report version. These are policies of this example.
  5. Protected stores share one owner epoch and compare it atomically with each write; the new epoch is published before the successor is admitted.
  6. The feed cannot be rolled back, deduplicates only by key, keeps keys for the whole run and does not check owner tokens.
  7. A fenced worker stops when a write is rejected. It may still be alive.
  8. Discrete ticks, one step per tick, no clock skew. Missed heartbeats are suspicion, never proof.
  9. The controller never fails, and its records are on reliable storage.

Within those assumptions the model shows what each record protects and what removing it costs. It does not show that these mechanisms are sufficient for a real system, measure anything about any engine, or replace the engines that already implement them.

References

  1. C. Mohan, D. Haderle, B. Lindsay, H. Pirahesh and P. Schwarz, “ARIES: A Transaction Recovery Method Supporting Fine-Granularity Locking and Partial Rollbacks Using Write-Ahead Logging”, ACM Transactions on Database Systems 17(1), 1992. WAL rule (p. 97); conditional redo by log sequence number and restart passes (pp. 105–112).
  2. P. A. Bernstein, V. Hadzilacos and N. Goodman, “Distributed Recovery”, ch. 7 of Concurrency Control and Recovery in Database Systems, Addison-Wesley, 1987. Atomic commitment conditions, the uncertainty period and Proposition 7.2.
  3. C. G. Gray and D. R. Cheriton, “Leases: An Efficient Fault-Tolerant Mechanism for Distributed File Cache Consistency”, Proc. 12th ACM SOSP, 1989. Lease definition; §5 on clock assumptions.
  4. J. H. Saltzer, D. P. Reed and D. D. Clark, “End-to-End Arguments in System Design”, ACM Transactions on Computer Systems 2(4), 1984. Duplicate message suppression.
  5. E. N. Elnozahy, L. Alvisi, Y.-M. Wang and D. B. Johnson, “A Survey of Rollback-Recovery Protocols in Message-Passing Systems”, ACM Computing Surveys 34(3), 2002. §2.3 outside world and output commit; §3 checkpoint-based recovery.
  6. F. B. Schneider, “Implementing Fault-Tolerant Services Using the State Machine Approach: A Tutorial”, ACM Computing Surveys 22(4), 1990.
  7. H. Garcia-Molina and K. Salem, “Sagas”, Proc. ACM SIGMOD, 1987.
  8. Temporal documentation, “Activity Definition” (idempotency and retries); read October 2026.
  9. Temporal documentation, “Workflow Definition” (determinism constraints); read October 2026.
  10. Temporal documentation, “Events and Event History”; read October 2026.
  11. AWS Step Functions Developer Guide, “Choosing workflow type in Step Functions”; read October 2026.
  12. AWS Step Functions Developer Guide, “Restarting state machine executions with redrive”; read October 2026.
  13. AWS Step Functions API Reference, “StartExecution”; read October 2026.
  14. Microsoft Learn, “Durable orchestrations” (Azure Durable Functions); read October 2026.
  15. Microsoft Learn, “Orchestrator function code constraints”; read October 2026.
  16. LangChain documentation, LangGraph “Functional API”; read October 2026.
  17. LangChain documentation, LangGraph “Checkpointers” (durability modes, pending writes); read October 2026.
  18. Restate documentation, “Durable execution” and “Durable steps”; read October 2026.
  19. M. Kleppmann, “How to do distributed locking”, 8 February 2016.
  20. M. Featonby, “Making retries safe with idempotent APIs”, Amazon Builders' Library.
  21. Stripe API reference, “Idempotent requests”; read October 2026.
  22. S. Chacon and B. Straub, Pro Git, “Git Internals – Git Objects”.
  23. Bazel documentation, “Remote Caching” (including “Known issues”); read October 2026.
  24. Apache Airflow documentation, “Tasks”; read October 2026.
  25. J. Young, “Effective harnesses for long-running agents”, Anthropic Engineering, 26 November 2025.
  26. P. Rajasekaran, “Harness design for long-running application development”, Anthropic Engineering, 24 March 2026.
Authorship and contributions

Contributors: Moiré (9 October 2026: site composition under Loki’s direction).