A worker is halfway through a job when it stops. Maybe its process crashed. Maybe its machine paused it, or its session ran out of context, or the network between it and everyone else went quiet. From the outside these look the same: the heartbeats stop and nothing new is written. Somebody else now has to finish the job, and the first thing they have to decide is what has already been done.
That decision is harder than it sounds, and agent work makes it common. Anthropic's November 2025 account of harnesses for long-running agents describes the symptom plainly: each new session “begins with no memory of what came before”, and a later session would sometimes “look around, see that progress had been made, and declare the job done.” The remedy there was a progress file, a feature list and Git commits for the next session to read. By March 2026 the same company's follow-up had separated the agent doing the work from the agent judging it — calling the separation “a strong lever”, while reporting that the judge needed tuning before it stopped approving work too easily. Handing over context and judging results by a separate agent are now ordinary practice.
Watch the explanation
3:27 · narrated · optional captionsPlay when you’re ready; captions are available in the player. Open the 4K film ↗
This article is about the layer underneath: not continuity of context but continuity of work. Strip the model away and the question is one that databases, workflow engines and distributed systems have been answering for decades:
When the process doing a unit of work stops — or merely stops answering — and the input it was working from may change while it is gone, what must already have been written down, and checked by whom, for someone else to continue the work correctly?
The short answer is that durability is not one save. It is a handful of separate records, each with its own job and, crucially, its own checker: the party that has to look at the record for it to protect anything. The longer answer is easiest to see by following one small task through everything that can go wrong with it.
Saved, accepted, safe to use
Three words get used as if they meant the same thing, and most of the trouble starts there.
Persisted
A step's result is durably recorded, together with what it was computed from.
- Lets you
- restore the step instead of redoing it.
- Says nothing about
- whether the result is any good.
Accepted
A finished output passed an assessment, and that fact was recorded for a named input version and contract.
- Lets you
- say this output met this standard, for this input.
- Says nothing about
- whether that input is still the current one.
Dependable
Another task may rely on it now: it is accepted, and accepted for the input version that consumer requires.
- Lets you
- start the next piece of work on it.
- Depends on
- who is asking, and when.
A journal of finished steps is progress. A worker that says it is done has produced a candidate, which nobody has judged. An acceptance is a recorded judgement, and it is a judgement about a specific input. Whether something downstream may build on it is a third question, answered relative to what that downstream task needs. Each of these is a different record, written at a different moment, and the example below shows each one failing separately when it is missing.
One task, four steps, one input
The task has a durable name, digest/q3-report: summarise a four-section quarterly report into a short digest. That name is the work identity. It outlives any one try at the work. Each try is an attempt, numbered #1, #2 and so on, and an attempt is not the same thing as a session or a process: a session can be replaced without the work restarting, and one session can carry more than one attempt.
The work runs in a fixed order. One model call summarises each section (s1 to s4). The worker assembles the four summaries into a candidate digest. A controller, not the worker, assesses the candidate and records an acceptance. Then the worker posts one announcement to an external feed. A second task, a newsletter, becomes ready at tick 12 and quotes the digest's revenue figure.
The report exists in two versions. Section 2 of version 1 says revenue was 4.1 million; version 2 corrects it to 4.7 million. Sections 1, 3 and 4 are byte-for-byte identical in both. Each section is identified by a hash of its text:
| Section | Version 1 | Version 2 | Changed? |
|---|---|---|---|
| s1 Overview | 45a1f4c0 | 45a1f4c0 | no |
| s2 Revenue | 43d57587 · 4.1 million | ad64f04a · 4.7 million | yes |
| s3 Costs | 2f08201d | 2f08201d | no |
| s4 Outlook | 7dac6c4a | 7dac6c4a | no |
| Whole report | ba0184f5 | cf47dcd5 | yes |
Now the trouble. Attempt #1 goes silent as it starts section 3. While it is silent, the corrected report arrives. Attempt #1 was not dead; it was paused, and later it wakes up and carries on as if nothing happened. And the worker that eventually posts the announcement dies just after the feed applies it, before it can write down that it did.
Here is the whole run with every protection in place. The rest of the article takes it apart one record at a time.
- t0attemptController admits #1 with owner token 41, pinned to report v1 ba0184f5 — written down before any work is dispatched.
- t1journal#1 summarises s1 from v1 → result 3af8da4e, recorded.
- t2journal#1 summarises s2 from v1 → 8dc2b770, recorded.
- t3attempt#1 starts s3 (the model call is made), then goes silent. Nothing is recorded.
- t4inputReport v2 cf47dcd5 arrives: section 2 corrected from 4.1 to 4.7 million.
- t5attempt#1 has missed two heartbeats: unknown, not failed and not finished. Controller admits #2, token 42, pinned to v2. #2 restores s1 from the journal and must redo s2.
- t6journal#2 summarises s2 from v2 → 2b1a6af0.
- t7journal#2 summarises s3 → 59977556.
- t8attempt#1 wakes and tries to record s3. The store rejects it: token 41 is older than 42. #1 stops — alive, without write authority. Meanwhile #2 summarises s4 → 3b74fabd.
- t9artifact#2 assembles candidate digest 9723ab9c (for v2, quoting 4.7 million). Finished; not yet accepted.
- t10artifactController accepts 9723ab9c for v2. #2 records its intent to post; the feed applies post P1; #2 dies before recording that it did.
- t12dependantNewsletter admitted on 9723ab9c; it quotes 4.7 million.
- t13attempt#2 is unknown. Controller admits #3, token 43, to settle the unconfirmed post.
- t14effect#3 resends with the same operation key. The feed answers “already seen” and returns P1. Nothing new is posted; P1 is recorded as confirmed.
Write the attempt down before the work
The first record is written at t0, before anything else happens: attempt #1 exists, it is working from report version ba0184f5, and it holds owner token 41.
This is borrowed, loosely, from the oldest idea in recovery. Write-ahead logging, as formalised in ARIES, requires that the log record describing a change reach stable storage before the changed data does. The point is that after a crash, the log bounds what might have happened. The task's version of that rule is coarser — an attempt record before dispatch, rather than a log record before a page write — but the benefit is the same in kind. When #1 goes silent, the controller is not left guessing whether anyone was working on this at all. It has a specific question: attempt #1, pinned to v1, was somewhere after step 2; did step 3 finish?
The record does not answer that question. It only makes it askable. That distinction matters, because the next thing that happens is silence.
When the worker stops answering
At t3 attempt #1 starts summarising section 3, then stops sending heartbeats. By t5 it has missed two, and the controller marks it unknown.
Not failed. Not finished. Unknown. A worker that has gone quiet might have crashed before making the model call, crashed after making it, finished the step but died before recording it, or simply been paused. In this run it was paused. Switch it to crashed and the controller sees exactly the same thing up to the moment the paused worker wakes; the outcome is otherwise identical — six calls, one post, the newsletter on the corrected figure. That indistinguishability is the point. A system that maps “process gone” to failed throws away work that may have finished. One that maps it to success invents completion — which is precisely the “declare the job done” failure from the harness account above.
So silence gets its own state, and the controller does the one thing it can do safely: admit a successor, #2, with a larger owner token, 42. What #2 does first depends on the next record.
Restore what still matches, redo what doesn't
When #2 is admitted at t5, the report has already been corrected. #2 is pinned to version 2. Before doing any work it looks at the journal, which holds #1's two finished steps — and here a journal that merely says “step 1 done, step 2 done” would be useless, or worse. What makes the journal safe to reuse is that each entry records what it was computed from: the hash of the section text it summarised.
| Step | Journal entry | Needed for v2 | #2 does |
|---|---|---|---|
| s1 | 3af8da4e from section 45a1f4c0 | section 45a1f4c0 | restore |
| s2 | 8dc2b770 from section 43d57587 | section ad64f04a | redo — the section changed |
| s3 | none (in flight when #1 went silent) | section 2f08201d | do |
| s4 | none | section 7dac6c4a | do |
This is content addressing, the same idea that lets Git store objects under the hash of their contents and lets Bazel's remote cache reuse a build action whose declared inputs hash the same. A step is restored only if there is a record whose input key matches what the step needs now; otherwise it is redone. The section that changed is redone; the section that did not is reused at no cost.
Two things are worth noticing about the redo. First, it is a new result. When #2 summarises section 2 at t6 it does not reproduce #1's record 8dc2b770; it produces 2b1a6af0. In the model that difference is stipulated — every computed step gets an identifier derived from the attempt that computed it — because a model call is not a deterministic function. A rerun may produce different text, not must, and whatever it produces has to be judged on its own. That is why durable-execution engines draw the line where they do: Temporal requires the orchestration code to be deterministic and puts model calls and other external work into activities whose results are recorded, and LangGraph's Functional API restores completed task results on resume “including for long-running or non-deterministic task outputs”. Restoring reuses what was recorded. Recomputing is a fresh execution.
Second, reuse by hash is only as sound as the key. Bazel's own documentation lists a known issue for each way it breaks. If a step reads something its key does not cover — a compiler outside the workspace, in Bazel's case; in ours, a section summary that quietly looks at another section — two different situations share a key and the reuse is wrong. And if the bytes a step actually reads are not the bytes the key names — Bazel warns about files modified during a build — the key is a label, not a guarantee. The example assumes each section summary depends only on its own section, and that each attempt reads an immutable snapshot of the version it is pinned to. Those assumptions do real work, and a builder has to earn them.
The worker that wasn't dead
At t8, attempt #1 wakes up. It was paused, not crashed, and from its own point of view no time has passed. It finishes summarising section 3 and tries to write the result to the journal.
This is the failure that makes ownership after a restart hard. The controller decided #1 was gone because it stopped answering; #1 never heard that decision. Martin Kleppmann's essay on distributed locking describes exactly this shape: a lock holder pauses, its lease expires, someone else takes over, and the first holder resumes and writes, still believing it owns the lock. Checking the lease just before writing does not help, because the pause can fall between the check and the write. Leases themselves, as Gray and Cheriton defined them, bound how long a controller must wait — but they rely on clocks behaving within a known drift, and expiry is only ever suspicion.
The fix is a fencing token: a number that increases every time ownership is granted, carried on every write, and compared by the storage. As Kleppmann puts it, this “requires the storage server to take an active role in checking tokens.” In the example, every protected store — the journal, the candidate store, the effect records and the acceptance step — compares a write's token against one shared owner epoch, and the controller publishes the new epoch before admitting the successor. So when #1 arrives with token 41 after 42 has been issued, the journal refuses it, and #1 stops.
Three limits follow. Fencing does not prove the old worker stopped; in this run #1 is still alive after the rejection. It protects only the stores that check the token; the external feed in this example checks operation keys, not owner tokens. And it does not make a false suspicion free: #1's call on section 3 was wasted. Real systems with separate stores also have to arrange what the model gives itself for free — writes conditioned on a shared epoch, and an epoch that keeps increasing even if the controller itself restarts.
Turn fencing off and the cost is concrete. Attempt #1's late writes for section 3, section 4 and its own candidate digest all land after takeover: three stale writes applied, and seven model calls instead of six. In the main run its candidate is still refused at acceptance, because it was built from version 1 and version 2 is current — a different record catches it. But if the report had never changed, nothing would catch it: the model accepts #2's digest 5573113e at t9 and the stale worker's 6025f0d0 at t11, two different digests accepted for the same report, each from a worker that believed it was the owner.
Finished, then accepted — for version 2
At t9 attempt #2 assembles its candidate digest, 9723ab9c. It is finished. It is not yet accepted, and the gap between the two is the whole point of this step: a worker reporting that it is done is not the same as an output passing a check that worker did not perform.
At t10 the controller assesses the candidate and records the acceptance. The record names three things: the digest, the report version it was judged against (v2) and the assessment contract it was judged under. Before accepting, the controller checks mechanically that the candidate's part hashes match the sections of the version it is pinned to. The assessment itself checks completeness and form; it does not re-derive every figure from the source. Putting assessment outside the producing worker is a design choice rather than part of what durability means, and it is not a guarantee: the March 2026 harness account is candid that its separate evaluator found real problems and then talked itself into approving the work until it was tuned. An acceptance records what was judged. It does not certify that the judgement was right.
What the acceptance does give is a fixed point that later records can refer to. And that matters at t12, when the newsletter becomes ready. Its admission rule is strict: take a digest that was accepted for the report version that is current now. 9723ab9c qualifies, so the newsletter starts on it and quotes 4.7 million.
Notice the order. The newsletter starts at t12; the announcement is not confirmed until t14. Whether a result is safe to depend on rests on its acceptance, not on what happened to its side effects afterwards.
Once, as far as the feed can tell
The hardest moment in the run is t10. Attempt #2 writes down that it intends to post the announcement. It sends the post. The feed applies it, as P1. And #2 dies before it can record the feed's reply.
When #3 is admitted at t13 to finish the job, it finds an intent with no confirmation. That local record fits two different histories equally well: the post never arrived, or it arrived and the acknowledgement was lost. Nothing on the worker's side can tell them apart. Saltzer, Reed and Clark made the general point in 1984 in “End-to-End Arguments in System Design”: a retried request after a lost acknowledgement is a duplicate that only the application receiving it can recognise. Only the receiver saw whether the effect happened.
So the practical pattern has two halves. The sender retries until it gets an answer — at-least-once delivery. The receiver makes the effect idempotent: performing it twice has the same result as performing it once. For an effect that is not naturally idempotent, like posting an announcement, the receiver gets there by remembering keys. #3 resends with the key announce:digest/q3-report:9723ab9c; the feed has seen that key, returns P1 and posts nothing new.
Three details carry the weight, and each comes from a primary source:
- The key names the operation, not the attempt. It is built from what is being done — this announcement, for this accepted digest — not from who is doing it. Give each attempt its own key, or no key at all, and #3's resend looks new: the announcement is posted twice. Temporal's documentation recommends a key that stays the same across retry attempts for this reason, and notes that keys are enforced “by the service you are calling”, not by the caller.
- Recording the key and applying the effect are one atomic step at the receiver. The Amazon Builders' Library article on idempotent APIs requires the token and the mutations to be recorded with ACID properties; otherwise the key can exist without the effect, or the reverse.
- Receivers forget. Stripe may remove keys once they are at least 24 hours old, after which a reused key creates a new request. The example's feed keeps keys for the whole run; a real retry that arrives late enough is a new effect.
The guarantee that results is once, as far as the receiver can tell, for as long as it keeps the key — not exactly-once execution. Temporal's documentation draws the same line for its own activities, which are “observed as completed exactly once” yet “may be executed multiple times and may even partially complete more than once.” Where a receiver offers neither keys nor a way to ask whether something happened, what remains is compensation in the sense of sagas: amend the effect afterwards. A sent email can be followed by a correction; it cannot be unsent.
Try it: switch the protections off
Every protection so far has been on. The interactive below runs exactly the same model, with eight switches: what the silent worker was really doing, fencing, how often progress is saved, when the input changes, version pins, the newsletter's admission rule, whether the posting worker dies, and how the post is keyed. Its guided route walks the main run step by step; free exploration lets you break one record at a time and read the consequence beside the mechanism.
When the input changes after acceptance
Move the correction later and a different distinction comes into view. Suppose version 2 does not arrive until t12, after a digest has already been accepted.
| t | What happens |
|---|---|
| t5 | #2 (token 42) is pinned to v1, which is still current. It restores s1 and s2 and does s3 and s4. |
| t9 | Digest 5573113e (4.1 million) is accepted for v1 — correct for its input. #2 posts P1 and dies. |
| t12 | Report v2 arrives. The v1 acceptance stays on record. The newsletter becomes ready; under the strict rule it waits, because nothing has been accepted for v2. |
| t13 | #3 settles the unconfirmed P1 with the same key. |
| t14–t16 | #4 (token 44), pinned to v2, restores s1, s3 and s4 and redoes only s2; it assembles 43fc86ed (4.7 million). |
| t17 | 43fc86ed is accepted for v2; the newsletter starts on it and quotes 4.7 million; #4 posts a new announcement, P2, under a new key. |
Nothing went wrong here. The v1 digest was a correct result for the input it was judged against, and its acceptance is still true as history: the model never edits or withdraws an acceptance, only supersedes it. What changed is that it is no longer dependable for a consumer whose rule is “the latest report”. A historical consumer that asked for the Q3 report as first filed could rely on it perfectly well. Dependability is relative to what the consumer requires.
The strict rule costs time — the newsletter waits from t12 to t17. Relax it to “take the latest accepted digest” or “take the latest finished one” and the newsletter starts at t12 instead, on the v1 digest, quoting 4.1 million while the report says 4.7 million. That wait is the strict rule doing its job: acceptance and dependability come apart exactly when the input moves.
The rebuild is cheap because only one section changed: one model call, against four for rebuilding everything. That comparison is between two policies inside this model, not a property of any engine, and it only holds under the reuse assumptions above. A more conservative design avoids mixing versions altogether. AWS Step Functions' redrive, for example, restarts a failed execution with the same input and the same definition, keeping the results of steps that succeeded (with documented exceptions), and a new input means a new execution. Whether a new execution reuses earlier results is then up to the application.
Cheaper, faster and wrong
The most instructive failure is the quiet one. Go back to the main run and turn off version pins: the journal still records finished steps, but not what they were computed from, so reuse is by step name alone and the acceptance cannot compare versions.
At t5, attempt #2 is pinned to version 2. It looks in the journal, finds entries for s1 and s2, and restores both — including #1's summary of the old section 2. It computes only s3 and s4, assembles a digest at t8 that claims version 2 while quoting 4.1 million, and has it accepted at t9 with no version recorded. That is one tick earlier than the correct run, after five model calls instead of six. Then at t12 the newsletter asks for a digest accepted for the current report. The records carry no versions, so the rule cannot be checked, and in this model it is set to fail open: it takes the accepted digest. The newsletter quotes 4.1 million while the report says 4.7 million.
Everything looked better on the way to the wrong answer: fewer calls, an earlier acceptance, nothing rejected. The record that was missing is the one that says what each step was computed from, and without it the controller's acceptance and the newsletter's rule were both checking a label they could not see. A rule that failed closed would have made the newsletter wait — and, with no versions recorded anywhere, never proceed. It would have protected the consumer without repairing the producer: the wrong digest is still accepted upstream either way.
Five records, five checkers
Put the run back together and the records line up, each with the party that has to check it:
| Record | Its job | Checked by | Without it |
|---|---|---|---|
| The attempt, its pinned input and owner token, written before dispatch | Turns silence into a specific question | The controller, when a worker goes quiet | A lost attempt cannot be told from a finished one |
| A journal of finished steps, each keyed by the hash of what it was computed from | Restore what still matches; refuse what was computed from an old input | The successor, before reusing | A digest quoting 4.1 million is accepted and used |
| An owner token that the stores compare | Makes a worker that only seemed dead harmless when it wakes | Every protected store | Three late writes land; with no input change, two digests are accepted for one report |
| An operation key on each external effect | Lets a resend after a lost acknowledgement land once | The receiver | The announcement is posted twice |
| An acceptance naming the input version and contract | Keeps “accepted” and “safe to depend on” apart when the input moves | The dependant's admission rule | The newsletter quotes the superseded figure |
The model was run on every combination of its eight switches — 864 in all. With fencing, version pins, a per-operation key and the strict newsletter rule all on, all 24 runs finish with an accepted digest for the final input and a confirmed announcement, and none breaks an invariant. Remove any one of those four and some combination breaks: fencing in every case where the old worker was only paused, pins in every case where the input changes, the key in every case where the posting worker dies, the strict rule in every case where the input changes after acceptance. That is necessity within this model and these switches — not a theorem about arbitrary crash points, clocks or storage failures. It also shows that the records guard different things. With fencing off and pins on, stale writes pollute the journal but no wrong digest is accepted. With pins off and fencing on, no stale write lands but a wrong digest is.
One switch is different in kind. Save progress only at the end of the attempt instead of after every step, and #1's two finished summaries never reach the journal: #2 redoes all four, eight model calls instead of six, and the digest is accepted at t11 instead of t10. Nothing breaks. How often to persist is a cost decision, not a correctness one — the same trade LangGraph describes in its durability modes, between durability and performance overhead.
Why not just use an engine?
For many systems, do. Nothing in the example is new, and it is not offered as an alternative to the systems that already implement these records or as better than any of them. As their documentation read in October 2026 describes them: Temporal keeps an append-only event history and replays deterministic workflow code against it, recording activity results; Azure Durable Functions replays orchestrators against their history and returns recorded activity results instead of rerunning them; LangGraph restores completed task results and advises idempotent side effects for tasks that may run again; Restate stores the result of each durable step in an execution log; AWS Step Functions offers redrive and execution-level idempotency for Standard workflows. The technical companion sets out what each page states and the scope each statement needs.
What none of them can do on your behalf is make a third party's effect idempotent, know which inputs your agent actually read, or decide what “accepted” means for your output. Those three are the parts of this example that live outside any engine: the receiver that keeps keys, the honest declaration of a step's inputs, and the acceptance that names a version and a contract.
A design that asks for this
Loki's Era 3 is a proposed design for running agent work from durable factory records rather than from any one session. As of its draft architecture documents of September 2026 it is a design, not a running system: no Era 3 core, factory engine or session-capture runtime has been implemented. Its state-walkthrough teaching applications are real, but they deliberately keep every input fixed and leave controller recovery and changed inputs out. The drafts require most of the properties above — an attempt distinct from its session, intent recorded before dispatch, silence that never counts as success, progress kept separate from acceptance, reconciliation of a replaced owner's writes and effects before a new writer is admitted, and no claim of universal exactly-once effects — and for most of them have not yet chosen a mechanism. Among the evidence the design says would show it works is an experiment in which an agent is interrupted or replaced and the work continues from records without losing accepted decisions, inventing completion or repeating recorded effects. The worked example here is a small paper version of the kind of situation that experiment would face. It is not evidence that Era 3 does any of this; fencing tokens, section-hash keys and per-operation keys are choices a builder might make, not choices the design has made.
Where this stops
The example is small on purpose, and its edges are worth knowing before you borrow from it.
- The controller never fails. The example covers workers being interrupted and replaced, an ambiguous side effect and a changed input — with a controller that survives on reliable storage. A restarted controller would need, at least, durable attempt records, the owner epoch, unconfirmed effect intents and acceptances; it would have to rebuild its suspicion timers rather than trust them, and keep the epoch increasing across its own restart. That is a requirement, not something this model runs.
- The journal and the effect are never one transaction. The window between recording an intent and the receiver applying it always exists. The design has to assume it rather than wish it away.
- Fencing assumes the stores check. In the model they share one epoch. Separate real stores need conditional writes or equivalent ordering, and anything that does not check tokens is not protected by them.
- Reuse assumes honest keys. A step that reads an undeclared input, or bytes other than the ones hashed, makes restore-by-hash unsound.
- An acceptance is only as good as its assessor. Here the assessor checks completeness and form, not every figure.
- Keys expire, clocks drift and storage fails. A retry after the receiver forgot the key is a new effect; leases need bounded clock drift; and durable work is exactly as durable as its records.
- The identifiers are schematic. The model does not generate summary text. It gives redone steps new identifiers by stipulation to make the restore-versus-redo difference visible, and claims nothing about how often real reruns differ.
Within those limits the lesson is compact. A task survives the worker that started it when the records that matter were written before they were needed, and when each one has someone whose job is to check it: the controller reading an attempt record, the successor reading a journal keyed by input, the store comparing an owner token, the receiver remembering an operation key, and the dependant asking which version an acceptance was for. Saved, accepted and safe to use are three different answers. Keep them apart, and a restart becomes a question with an answer rather than a guess.
Sources
Documentation pages were read in October 2026 and may change. The technical companion gives page-level references, equations and the full traces.
- C. Mohan et al., “ARIES: A Transaction Recovery Method…”, ACM TODS 17(1), 1992.
- C. G. Gray and D. R. Cheriton, “Leases: An Efficient Fault-Tolerant Mechanism for Distributed File Cache Consistency”, SOSP 1989.
- J. H. Saltzer, D. P. Reed and D. D. Clark, “End-to-End Arguments in System Design”, ACM TOCS 2(4), 1984.
- H. Garcia-Molina and K. Salem, “Sagas”, SIGMOD 1987.
- M. Kleppmann, “How to do distributed locking”, 2016.
- Temporal documentation: Activity Definition, Workflow Definition, Events and Event History.
- AWS: Step Functions redrive; M. Featonby, “Making retries safe with idempotent APIs”, Amazon Builders' Library.
- Microsoft Learn, Durable orchestrations.
- LangGraph: Functional API, Checkpointers. Restate: Durable steps.
- Stripe, Idempotent requests; Pro Git, Git Objects; Bazel, Remote Caching.
- J. Young, “Effective harnesses for long-running agents”, Anthropic, November 2025; P. Rajasekaran, “Harness design for long-running application development”, Anthropic, March 2026.
Authorship and contributions
Contributors: Moiré (9 October 2026: site composition under Loki’s direction).