agents · Level 4

Harness state and agent memory: checkpoints, retention and deletion

Resume a small workflow, reject stale changes and make forgetting an explicit operation.

By Mickarle Wagstaff-Irons - Micky Irons

  • Level 4Frontier
  • 180 min
  • 6 chapters
  • Free PDF, no account
The Loopagents / 04

Start with the essentials

The short answer

Agent memory is stored data with a purpose and a lifecycle. In this lab, a checkpoint records progress while a separate fact record carries text and provenance. You resume a fictional workflow after closing its database connection, reject stale writes, hide expired records and delete selected rows. These local controls do not provide authentication, exactly-once external actions or erasure from backups.

What you will learn

  • Separate execution checkpoints from reusable facts.
  • Validate small versioned payloads on write and read.
  • Reject stale writes with transactional revision checks.
  • Resume a deterministic local workflow after reopening its database.
  • Apply tenant-scoped expiry, purge and revision-aware deletion.
  • Explain the boundaries of logical deletion and external-action recovery.

Who it is for

Python builders extending a small agent harness with persistent state and a clear memory lifecycle.

Before you start

  • Read Python functions, exceptions and unit tests.
  • Understand simple SELECT, INSERT and DELETE statements.
  • Recognise the plan, action and observation stages of an agent loop.

Read a sample · Chapter 01 of 06

01

Decide what memory means

Give each stored field a purpose before choosing a database.

Create a folder with store.py, demo.py and test_store.py. The complete files appear in this workbook; append the numbered sections of store.py in their printed order. Use Python 3.12 or later because connect uses the autocommit parameter introduced in 3.12. This release ran on Python 3.12.10. Run python -m unittest -v test_store and then python demo.py from that folder. The demonstration creates a temporary database containing original fictional data and removes that temporary directory on exit.

A checkpoint answers where this particular execution stopped. It might contain a validated stage number, the identifiers of completed work and enough input references to resume. A reusable fact answers something the system may consult on a later run. Combining both into one growing transcript makes expiry, inspection and correction difficult. Here they share a table but have different kinds and different payload schemas. The same key can exist independently in each kind.

The checkpoint schema is exactly schema, step and items. Schema is integer 1; step is an integer from 0 to 2; items is a list with at most eight non-empty strings. The fact schema is exactly schema, text, source and reviewed. Reviewed must be a Boolean. A fact can be false, stale or malicious even when reviewed is true. That field is an annotation, not an access grant, a truth detector or permission to execute its text.

Each text value is limited to 256 characters, and the serialised payload is limited to 2,048 UTF-8 bytes after JSON escaping. A few non-ASCII strings can reach the byte limit before their character limits. The limits are teaching choices for this small store; they are not service capacity guidance. We deliberately retain no credentials, hidden model reasoning, private transcripts or arbitrary Python objects. Provenance is a short fictional source label, not a promise that the source remains accurate.

The identity tuple is tenant, kind and key. Tenant and key use a narrow lowercase identifier format; kind is checkpoint or fact. This creates separate namespaces and makes exact operations easy to inspect. It does not authenticate the caller. A service must derive tenant from its trusted session or identity layer and enforce permission before calling these functions. Letting a browser choose any tenant string would defeat that service boundary even though the SQL includes a tenant condition.

Start store.py with the imports, Conflict exception and validation helpers below. Booleans are rejected where integers are required, despite Python treating bool as an int subclass. Extra object fields are rejected so that a later producer cannot silently add a large transcript. Validation happens on both write and read. The helpers accept already-created Python objects; they are not a general bounded HTTP body reader or a complete hostile-JSON parser.

python · 51 lines
import json
import re
import sqlite3


class Conflict(Exception):
    pass


def identity(tenant, kind, key):
    for value in (tenant, key):
        if not isinstance(value, str) or not re.fullmatch(
                r"[a-z][a-z0-9-]{0,31}", value):
            raise ValueError("invalid tenant or key")
    if kind not in ("checkpoint", "fact"):
        raise ValueError("invalid kind")


def integer(value, low, high):
    if type(value) is not int or not low <= value <= high:
        raise ValueError("integer outside bounds")


def payload(kind, value):
    if not isinstance(value, dict):
        raise ValueError("payload must be an object")
    if type(value.get("schema")) is not int or value["schema"] != 1:
        raise ValueError("unsupported schema")
    if kind == "checkpoint":
        if set(value) != {"schema", "step", "items"}:
            raise ValueError("checkpoint fields")
        integer(value["step"], 0, 2)
        items = value["items"]
        if not isinstance(items, list) or len(items) > 8:
            raise ValueError("checkpoint items")
        strings = items
    elif kind == "fact":
        if set(value) != {"schema", "text", "source", "reviewed"}:
            raise ValueError("fact fields")
        if type(value["reviewed"]) is not bool:
            raise ValueError("reviewed must be boolean")
        strings = [value["text"], value["source"]]
    else:
        raise ValueError("invalid kind")
    if any(not isinstance(s, str) or not 1 <= len(s) <= 256
           for s in strings):
        raise ValueError("text outside bounds")
    encoded = json.dumps(value, sort_keys=True, ensure_ascii=True)
    if len(encoded.encode("utf-8")) > 2048:
        raise ValueError("payload too large")
    return encoded

Try it yourself · Activity 01

15 min

Write a memory contract

Design records for the fictional Orchard workshop assistant.

  1. Separate a resumable run from a reusable preference.
  2. State the source, allowed fields and retention owner for each.
  3. Identify which layer chooses the tenant and authorises retrieval.

Could you explain why every retained field is needed to the person whose data it describes?

Worked answer

A run checkpoint can contain schema 1, step 0 and an empty items list under checkpoint/run-a. A preference can contain the fictional checklist text, source workshop-note-1 and reviewed false under fact/preference. The trusted application chooses tenant orchard. The checkpoint supports one run; the fact needs its own review and retention decision. Neither record needs a full conversation or a credential.

Read a sample · Chapter 02 of 06

02

Create a store with an explicit clock

A row can exist without being eligible for retrieval.

Append connect and get to store.py. The default database is :memory:, which disappears when the connection closes. Passing a path creates or opens a file-backed database. The demonstration uses a newly created temporary path; choose a dedicated path if experimenting with persistence. Do not point this teaching code at an existing application database. CREATE TABLE IF NOT EXISTS is initialisation, not a schema migration or compatibility check.

python · 11 lines
def connect(path=":memory:"):
    db = sqlite3.connect(path, autocommit=True, timeout=1)
    db.execute("""CREATE TABLE IF NOT EXISTS records (
        tenant TEXT, kind TEXT, key TEXT, revision INTEGER NOT NULL,
        expires INTEGER NOT NULL, body TEXT NOT NULL,
        PRIMARY KEY (tenant, kind, key))""")
    db.execute("""CREATE TABLE IF NOT EXISTS sequence (
        singleton INTEGER PRIMARY KEY CHECK(singleton=1),
        value INTEGER NOT NULL)""")
    db.execute("INSERT OR IGNORE INTO sequence VALUES (1, 0)")
    return db

python · 14 lines
def get(db, tenant, kind, key, now):
    identity(tenant, kind, key)
    integer(now, 0, 10**10)
    row = db.execute("""SELECT revision, expires, body FROM records
        WHERE tenant=? AND kind=? AND key=? AND expires>?""",
        (tenant, kind, key, now)).fetchone()
    if row is None:
        return None
    revision, expires, body = row
    if not isinstance(body, str) or len(body.encode("utf-8")) > 2048:
        raise ValueError("stored payload too large")
    decoded = json.loads(body)
    payload(kind, decoded)
    return {"revision": revision, "expires": expires, "value": decoded}

There are two tables. Records has a primary key over tenant, kind and key, plus a revision, an expiry time and JSON body. Sequence contains a single integer counter shared across records. Revision values are opaque comparison tokens, not timestamps or per-tenant activity counts. A shared counter can reveal that other writes occurred if exposed to clients. This local lab has no public API; a production interface should choose tokens and metadata disclosure deliberately.

This Python 3.12 connection uses autocommit=True. Individual SQL statements outside an explicit transaction commit through SQLite automatically. The put operation later issues SQL BEGIN IMMEDIATE, COMMIT and ROLLBACK itself. It does not call the Python connection commit method, which has different meaning under this setting. Keep transaction policy explicit when moving this code into another framework. Avoid wrapping these functions in an already-open transaction without redesigning their ownership.

Time is supplied explicitly as an integer from 0 to 10**10. The demo uses fictional ticks such as 100 and 108 so that tests are deterministic. A service must supply a trusted clock in a consistent unit. Expiry is exclusive: a record with expires 110 is visible at 109 and hidden at 110. Reads never extend the deadline. Every successful write sets a new deadline of now+ttl, where ttl is between 1 and 86,400 seconds in this exercise.

Get binds all SQL values with question-mark placeholders and filters identity and expiry in the query. If no eligible row exists, it returns None. Otherwise it checks the stored JSON size, parses it and validates the kind-specific schema again. It returns a new dictionary, so changing the returned object does not change the database until an explicit put succeeds. This is useful isolation from accidental in-memory edits, not isolation from a process with direct access to the database file.

Expiry hides a row; it does not remove it. A clock moving backwards can make a previously expired but unpurged row visible again. Trusted clock handling and a deliberate retention job matter in a real service. The caller also has to decide whether absence means create, stop or ask for new input. This lab refuses to treat an absent checkpoint as permission to restart automatically, because repeating work could be the wrong action.

Try it yourself · Activity 02

15 min

Trace a deadline

Create a fact at time 100 with ttl 10, then read it repeatedly.

  1. Predict get at times 109 and 110.
  2. Predict the stored deadline after the read at 109.
  3. Explain what remains on disk after the read at 110.

Does your current system actually run its retention job, or does it only filter old rows from search?

Worked answer

The record is visible at 109 with expires 110. Reading it does not change that value. At 110 get returns None, but the row remains in records until forget or purge removes it. A backwards clock can make an unpurged row visible again. Visibility, row retention and storage-level erasure are different questions.

Read a sample · Chapter 03 of 06

03

Reject a stale writer inside the transaction

Read a revision, propose a change and let the store decide whether it is still current.

Append put. Expected None means create only when no physical row exists. An integer expected revision means replace only that exact version. The write obtains a SQLite write transaction before reading the old revision, checks expiry and expected state, allocates a fresh revision and writes the row. SQLite allows one writer at a time. BEGIN IMMEDIATE requests that writer position early; it can fail with a busy error when another connection holds the required lock.

python · 34 lines
def put(db, tenant, kind, key, value, expected, now, ttl):
    identity(tenant, kind, key)
    integer(now, 0, 10**10)
    integer(ttl, 1, 86400)
    if expected is not None:
        integer(expected, 1, 2**63-1)
    body = payload(kind, value)
    db.execute("BEGIN IMMEDIATE")
    try:
        old = db.execute("""SELECT revision, expires FROM records
            WHERE tenant=? AND kind=? AND key=?""",
            (tenant, kind, key)).fetchone()
        if old is not None and old[1] <= now:
            raise Conflict("expired row: purge before fresh creation")
        current = None if old is None else old[0]
        if current != expected:
            raise Conflict("revision changed")
        revision = db.execute(
            "SELECT value FROM sequence WHERE singleton=1"
        ).fetchone()[0] + 1
        integer(revision, 1, 2**63-1)
        db.execute("UPDATE sequence SET value=? WHERE singleton=1",
                   (revision,))
        db.execute("""INSERT INTO records VALUES (?, ?, ?, ?, ?, ?)
            ON CONFLICT(tenant, kind, key) DO UPDATE SET
            revision=excluded.revision, expires=excluded.expires,
            body=excluded.body""",
            (tenant, kind, key, revision, now+ttl, body))
        db.execute("COMMIT")
        return revision
    except Exception:
        if db.in_transaction:
            db.execute("ROLLBACK")
        raise

Consider two workers that both read revision 7. Worker A writes its proposed value with expected 7 and receives a new revision, perhaps 8. Worker B then tries expected 7 and gets Conflict. It must not retry with expected 8 while keeping its old proposal unchanged. It should reread and recompute, merge under an explicit rule, or abandon the operation. Otherwise the revision token becomes ceremonial and a stale decision can still overwrite a better one.

The revision comparison and the row update must share the same transaction. Checking outside the transaction leaves a gap for another writer. Our global sequence advances in that transaction too. If the row write fails, the exception path rolls back the counter and the row together when a transaction remains active. The learner tests install a temporary SQL trigger that aborts an update after the counter changes, then verify both values were restored. The trigger exists only in the test database.

Why not use the old row revision plus one? Suppose revision 7 is read, deleted and later recreated as revision 7 again. A stale worker holding the old token could then match the new record. The persistent sequence prevents this delete-and-recreate reuse within this database history. Deleting a row leaves the counter intact. Rolling back a failed write may reuse its uncommitted number, which no successful caller received. Restoring an old database backup can still rewind the sequence; this lab does not solve that operational problem.

An expired row causes Conflict even when expected is None or matches the physical row. The caller must purge it and then make an explicit fresh-creation decision. This avoids silently resuming stale state. It does not revoke a worker forever: after deletion a separate create-only request can insert again. Preventing recreation after a user forget request needs a revocation or tombstone policy outside this simple store. Decide how long such a marker remains and what minimal identity it needs.

The one-second connection timeout bounds a simple lock wait, not total workflow time. The lab does not retry a busy error and does not share a connection between threads. More workers, connection pools, write contention and process crashes need separate tests. Closing and reopening a committed file proves a small persistence path; it is not a power-loss, disk-failure or distributed-consensus experiment. Never convert every exception into a retry of an external action.

Try it yourself · Activity 03

20 min

Review two competing writes

Two workers read the same live checkpoint revision.

  1. Explain why only the first replacement should succeed.
  2. Trace delete and recreate while a third worker still holds the old revision.
  3. State the recovery action after Conflict.

Which parts of your workflow can safely be recomputed, and which can cause an external effect twice?

Worked answer

A writes with the current revision and advances the persistent counter. B presents the old revision and is rejected. Deleting and recreating the record uses another fresh counter value, so the third worker cannot replace the new row with the deleted row token. After Conflict, reread and recompute or stop; copying a new revision onto a stale proposal would bypass the intended protection.

Read a sample · Chapter 04 of 06

04

Make forgetting precise

Specify exactly what was removed and what was not measured.

Append forget and purge to finish store.py. Forget removes one identity only when its revision still matches, and returns True only if a row was removed. A missing row or a revision mismatch returns False. The caller must not interpret False as proof that the requested memory is absent; a newer revision may still exist. Repeating a successful deletion returns False without changing another row. Expired rows can also be forgotten if the caller has their matching revision.

python · 7 lines
def forget(db, tenant, kind, key, expected):
    identity(tenant, kind, key)
    integer(expected, 1, 2**63-1)
    cursor = db.execute("""DELETE FROM records
        WHERE tenant=? AND kind=? AND key=? AND revision=?""",
        (tenant, kind, key, expected))
    return cursor.rowcount == 1

python · 7 lines
def purge(db, tenant, now):
    identity(tenant, "fact", "unused")
    integer(now, 0, 10**10)
    cursor = db.execute(
        "DELETE FROM records WHERE tenant=? AND expires<=?",
        (tenant, now))
    return cursor.rowcount

Purge removes all rows for one tenant whose deadlines are less than or equal to now. It includes both kinds in that tenant and returns the number removed. There is no hidden background scheduler in this example. A service needs to schedule and observe this work, including failures, retry policy and any replicas or derived stores. Purge is a deliberate maintenance operation, not something a model may invoke merely because a retrieved instruction asks it to.

Both deletion operations use parameter binding. The tenant condition is essential even when keys appear unique during development. The test gives Orchard and Harbour the same key and deadline, purges only Orchard, and verifies Harbour remains. SQL parameter binding prevents a quoted text value from being treated as a SQL fragment; it does not decide whether the caller should be allowed to delete. Application authorisation and database file permissions are separate controls.

A successful SQL DELETE establishes that the selected row is absent from this table after the statement commits. It does not establish erasure from database free pages, journals, backups, process memory, logs, exports, vector indexes or another service. SQLite documents secure_delete as an option affecting deleted table content, with limits including virtual-table shadow data. This lab does not enable it or claim media sanitisation. Whole-system deletion needs a mapped set of copies and an operational verification procedure.

Keep evidence useful without making a new copy of the material being deleted. For example, a maintenance report can record the approved operation identifier, time, affected store and count rather than echoing the fact text. Even identifiers may themselves be sensitive in another application. This lab prints only fictional progress, Boolean results and counts. No legal retention or compliance conclusion follows from these demonstration settings.

A policy may need both a maximum lifetime and a shorter idle period. Our ttl is reset by every successful write, so repeated updates can extend retention indefinitely. An absolute original-created deadline is not implemented. Likewise, source corrections do not automatically find every copied fact. A larger memory system needs source-to-derived-record links, review dates, deletion propagation and a clear answer when a worker has already loaded data into memory. Keep those requirements visible rather than hiding them behind the word memory.

Try it yourself · Activity 04

15 min

Write a deletion receipt

The demonstration forgets a checkpoint and leaves no rows in records.

  1. State the narrow observation that was verified.
  2. List three places that were not inspected for residual copies.
  3. Explain why a False result from forget is ambiguous.

Can you list all of the derived stores that would need an independent deletion action in your actual application?

Worked answer

The matching checkpoint row was removed and a count query found zero rows in records. The example did not inspect database free pages, backups or process memory, and it did not contact a vector index or another service. False means either no matching identity existed or the revision differed; reread under the current policy if you need to distinguish those cases. Do not label this result complete erasure.

Read a sample · Chapter 05 of 06

05

Resume a bounded local workflow

Prove recovery with a file reopen before connecting an external tool.

Read the memory lifecycle

A checkpoint advances through revisions 1, 2 and 3. At time 108 a fact is hidden but stored; purge removes it, then forgetting revision 3 removes the checkpoint.
A fresh fictional database follows the published three-file lab. Time values are controlled teaching ticks; the application must supply trusted identity and time. Open the full-size memory lifecycle diagram.

The checkpoint key is orchard / checkpoint / run-a. Each write uses a 60-second TTL. Closing the connection after step 1 and reopening the same file at time 102 finds step 1; the next successful advance writes step 2. This is file-reopen evidence, not a process-crash or power-loss test.

Checkpoint after each successful write
TimeStep RevisionExpires
10001160
10112161
10223162

The completed read at time 103 calls no write: revision 3 and expiry 162 stay unchanged. A stale proposal with expected revision 1 is rejected. Neither operation advances the counter. The next successful write creates orchard / fact / preference with revision 4, now=103 and ttl=5, so expires=108.

Fact visibility and physical row presence are different
EventGet result Stored row
103: createVisiblePresent
107: readVisiblePresent
108: expiryHiddenPresent
108: purgeHiddenRemoved

The read at 107 is an extra observation added for the diagram; it does not alter the demo's writes or deadlines. At 108, get returns None while the row still exists. purge(orchard, 108) removes the expired fact but keeps the checkpoint whose deadline is 162.

Recovery, rejected changes and deletion results
OperationObserved result
102: reopenStep 1 survives; advance writes step 2
103: completed readRevision 3; expiry 162 unchanged
103: stale writeExpected revision 1 rejected; current revision 3
104: other tenantHarbour read returns None; Orchard fact stays present
108: purge1 fact removed; 1 checkpoint remains
108: forget revision 3True; 0 rows remain; counter stays 4

forget with the matching checkpoint revision 3 returns True. Zero rows now remain in records, while the separate sequence counter is still 4. These revision numbers depend on this fresh database and exact write order; they are opaque comparison tokens, not timestamps.

Removing these rows does not prove erasure from free pages, journals, backups, logs, process memory or derived stores. Local checkpointing does not make an external action exactly once. Tenant filtering does not authenticate the caller. The diagram describes tested local operations, with no model call or external effect.

Save demo.py below. The workflow starts at step 0 with no items, then adds collected at step 1 and drafted at step 2. These are deterministic local markers, not a claim that a document was collected or an LLM generated a draft. Advance loads the checkpoint, changes a decoded copy and writes it back using the observed revision. A completed step-2 checkpoint returns its existing revision without refreshing its expiry.

python · 53 lines
from pathlib import Path
from tempfile import TemporaryDirectory

from store import Conflict, connect, forget, get, purge, put


def advance(db, now):
    record = get(db, "orchard", "checkpoint", "run-a", now)
    if record is None:
        raise LookupError("checkpoint absent or expired")
    value = record["value"]
    if value["step"] == 2:
        return record["revision"]
    value["step"] += 1
    value["items"].append(("collected", "drafted")[value["step"]-1])
    return put(db, "orchard", "checkpoint", "run-a", value,
               record["revision"], now, 60)


def main():
    with TemporaryDirectory() as folder:
        path = Path(folder) / "fictional.db"
        db = connect(path)
        start = {"schema": 1, "step": 0, "items": []}
        first = put(db, "orchard", "checkpoint", "run-a",
                    start, None, 100, 60)
        advance(db, 101)
        db.close()
        db = connect(path)
        print("resumed step:", get(db, "orchard", "checkpoint",
                                   "run-a", 102)["value"]["step"])
        last = advance(db, 102)
        assert advance(db, 103) == last
        try:
            put(db, "orchard", "checkpoint", "run-a",
                start, first, 103, 60)
        except Conflict:
            print("stale write: rejected")
        fact = {"schema": 1, "text": "Fictional team likes checklists.",
                "source": "workshop-note-1", "reviewed": False}
        rev = put(db, "orchard", "fact", "preference", fact, None, 103, 5)
        print("other tenant:", get(db, "harbour", "fact", "preference", 104))
        print("at expiry:", get(db, "orchard", "fact", "preference", 108))
        print("expired rows purged:", purge(db, "orchard", 108))
        print("checkpoint forgotten:", forget(
            db, "orchard", "checkpoint", "run-a", last))
        print("remaining rows:", db.execute(
            "SELECT count(*) FROM records").fetchone()[0])
        db.close()


if __name__ == "__main__":
    main()

The demo commits step 1, closes the connection and opens the same temporary file again. It reads step 1, advances to step 2 and checks that another advance leaves the revision unchanged. It then attempts a write using the original revision, creates a short-lived fact, queries another tenant, tests the exact expiry boundary and purges the expired fact. Finally it forgets the completed checkpoint and counts remaining rows. Run python demo.py for the exact output below.

text · 7 lines
resumed step: 1
stale write: rejected
other tenant: None
at expiry: None
expired rows purged: 1
checkpoint forgotten: True
remaining rows: 0

The first checkpoint is written at time 100 with deadline 160. Advancing at 101 creates deadline 161; advancing at 102 creates deadline 162. The completed read at 103 changes neither revision nor deadline. The fact written at 103 with ttl 5 expires at 108. The purge at 108 removes that fact but keeps the still-live checkpoint. Forget then removes the checkpoint. A global counter remains, even when records is empty.

Do not add an email send between the read and put and assume the revision check makes it exactly once. A crash after sending but before saving can cause a retry to send again. Two workers can also perform the same effect before one loses the checkpoint race. A real tool may accept an idempotency key, or an application may use an outbox written with its checkpoint and a separate delivery process. Those approaches still require explicit semantics and verification; none is implemented here.

Similarly, the text in a fact is untrusted data. Persisting it for a later run does not turn it into a system instruction. A model-facing retrieval layer should preserve provenance and distinguish remembered claims from governing instructions. Permission checks must run again when the fact is used. This lab never puts a fact into a model prompt or executes it as code, so its successful tests do not establish resistance to prompt injection in a future integration.

Try it yourself · Activity 05

20 min

Draw the recovery boundary

Trace the demonstration and then consider one external tool call.

  1. List the checkpoint deadlines after its three writes.
  2. Explain why repeated step 2 does not extend retention.
  3. Place a crash after an imagined external send but before put and describe the unresolved effect.

What would you need to observe to distinguish a completed tool action from a lost acknowledgement?

Worked answer

The deadlines are 160, 161 and 162. A step-2 advance returns before calling put, so its read leaves 162 unchanged. A crash after an external send can leave the checkpoint at the earlier step; restarting may repeat the send. Revision checks protect the local record from stale replacement but do not undo or deduplicate that external side effect. Add and test an explicit idempotency or delivery protocol before connecting such a tool.

Read a sample · Chapter 06 of 06

06

Test the lifecycle before adding a model

Use concrete failure cases, then state the remaining limits.

Save test_store.py and run python -m unittest -v test_store. All ten methods are printed below. They cover namespaces and tenants, exact expiry, reads that do not refresh ttl, tenant-scoped purge, competing revisions, deletion and recreation, schema and byte limits, malformed stored values, transaction rollback, two file connections and local workflow completion. Each test starts from isolated fictional data.

python · 147 lines
import json
from pathlib import Path
import sqlite3
from tempfile import TemporaryDirectory
import unittest

from demo import advance
from store import Conflict, connect, forget, get, purge, put


def fact(text="Fictional preference"):
    return {"schema": 1, "text": text, "source": "note-1", "reviewed": False}


class StoreTests(unittest.TestCase):
    def setUp(self):
        self.db = connect()
        self.addCleanup(self.db.close)

    def add(self, tenant="orchard", key="a", now=100, ttl=10):
        return put(self.db, tenant, "fact", key, fact(), None, now, ttl)

    def test_tenants_and_namespaces(self):
        a = self.add()
        b = self.add("harbour")
        self.assertIsNone(get(self.db, "third", "fact", "a", 100))
        checkpoint = {"schema": 1, "step": 0, "items": []}
        put(self.db, "orchard", "checkpoint", "a", checkpoint, None, 100, 10)
        self.assertFalse(forget(self.db, "harbour", "fact", "a", a))
        self.assertTrue(forget(self.db, "orchard", "fact", "a", a))
        record = get(self.db, "harbour", "fact", "a", 100)
        self.assertEqual(record["revision"], b)
        self.assertIsNotNone(get(self.db, "orchard", "checkpoint", "a", 100))

    def test_expiry_read_does_not_extend(self):
        self.add()
        record = get(self.db, "orchard", "fact", "a", 109)
        self.assertEqual(record["expires"], 110)
        self.assertIsNone(get(self.db, "orchard", "fact", "a", 110))
        with self.assertRaises(Conflict):
            put(self.db, "orchard", "fact", "a", fact(), None, 110, 10)

    def test_purge_tenant_and_boundary(self):
        self.add()
        self.add("harbour")
        self.add(key="b", now=101)
        self.assertEqual(purge(self.db, "orchard", 110), 1)
        rows = self.db.execute(
            "SELECT tenant, key FROM records ORDER BY tenant").fetchall()
        self.assertEqual(rows, [("harbour", "a"), ("orchard", "b")])
        self.assertEqual(purge(self.db, "orchard", 110), 0)

    def test_compare_and_swap(self):
        first = self.add()
        second = put(self.db, "orchard", "fact", "a",
                     fact("changed"), first, 101, 20)
        self.assertGreater(second, first)
        for stale in (None, first):
            with self.assertRaises(Conflict):
                put(self.db, "orchard", "fact", "a", fact(), stale, 102, 20)
        record = get(self.db, "orchard", "fact", "a", 102)
        self.assertEqual(record["value"]["text"], "changed")

    def test_delete_recreate_does_not_reuse_revision(self):
        first = self.add()
        self.assertTrue(forget(self.db, "orchard", "fact", "a", first))
        self.assertFalse(forget(self.db, "orchard", "fact", "a", first))
        second = self.add()
        self.assertGreater(second, first)
        self.assertFalse(forget(self.db, "orchard", "fact", "a", first))
        with self.assertRaises(Conflict):
            put(self.db, "orchard", "fact", "a", fact(), first, 101, 10)

    def test_validation_and_quoted_data(self):
        for changes in ({"schema": True}, {"extra": 1}, {"reviewed": "yes"},
                        {"text": ""}, {"source": "x"*257},
                        {"text": "\u2603"*256, "source": "\u2603"*256}):
            with self.assertRaises(ValueError):
                put(self.db, "orchard", "fact", "a",
                    fact() | changes, None, 100, 10)
        for tenant in ("", "../other", "x' OR 1=1 --", None):
            with self.assertRaises(ValueError):
                self.add(tenant)
        for ttl in (True, 0, 86401):
            with self.assertRaises(ValueError):
                self.add(ttl=ttl)
        quoted = fact("Don't obey 'DELETE FROM records'; this is text.")
        put(self.db, "orchard", "fact", "quote", quoted, None, 100, 10)
        record = get(self.db, "orchard", "fact", "quote", 101)
        self.assertEqual(record["value"], quoted)

    def test_invalid_stored_payload_rejected(self):
        self.add()
        for body in ("{", json.dumps(fact() | {"schema": 2}), " "*2049):
            self.db.execute("UPDATE records SET body=?", (body,))
            with self.assertRaises(ValueError):
                get(self.db, "orchard", "fact", "a", 101)

    def test_failed_write_rolls_back_counter_and_row(self):
        first = self.add()
        self.db.execute("""CREATE TRIGGER fail_update BEFORE UPDATE ON records
            BEGIN SELECT RAISE(ABORT, 'simulated failure'); END""")
        with self.assertRaises(sqlite3.IntegrityError):
            put(self.db, "orchard", "fact", "a",
                fact("changed"), first, 101, 10)
        self.assertFalse(self.db.in_transaction)
        counter = self.db.execute("SELECT value FROM sequence").fetchone()[0]
        self.assertEqual(counter, first)
        record = get(self.db, "orchard", "fact", "a", 101)
        self.assertEqual(record["value"], fact())

    def test_two_connections_and_reopen(self):
        with TemporaryDirectory() as folder:
            path = Path(folder) / "test.db"
            a, b = connect(path), connect(path)
            try:
                rev = put(a, "orchard", "fact", "a", fact(), None, 100, 10)
                seen = get(b, "orchard", "fact", "a", 100)
                put(a, "orchard", "fact", "a", fact("new"), rev, 101, 10)
                with self.assertRaises(Conflict):
                    put(b, "orchard", "fact", "a",
                        fact(), seen["revision"], 101, 10)
            finally:
                a.close()
                b.close()
            c = connect(path)
            try:
                record = get(c, "orchard", "fact", "a", 102)
                self.assertEqual(record["value"]["text"], "new")
            finally:
                c.close()

    def test_resume_and_done_are_local(self):
        put(self.db, "orchard", "checkpoint", "run-a",
            {"schema": 1, "step": 0, "items": []}, None, 100, 60)
        advance(self.db, 101)
        last = advance(self.db, 102)
        self.assertEqual(advance(self.db, 103), last)
        value = get(self.db, "orchard", "checkpoint", "run-a", 103)["value"]
        self.assertEqual(value, {"schema": 1, "step": 2,
                                 "items": ["collected", "drafted"]})
        with self.assertRaises(LookupError):
            advance(self.db, 162)


if __name__ == "__main__":
    unittest.main()

The expected runner reports Ran 10 tests followed by OK. The two-connection test deliberately reads a stale snapshot, commits a replacement through the other connection and then tries the stale write. This is a reproducible ordering test, not simultaneous-writer stress testing. The file is reopened again to inspect the committed value. The rollback test uses a simulated SQL statement failure rather than terminating a process or cutting power.

Release review also ran 600 deterministic mixed operations against a separate dictionary-based reference, comparing all six tenant/key lookups, physical rows and the sequence after each step. That produced 5,400 reference assertions across 189 writes, 217 forget attempts and 194 purges. The sequence included conflicts and expired records. This bounded cross-check is stronger than a happy-path demonstration, while still leaving schema migrations, concurrency load, hostile files and production retention operations untested.

Three separate copied implementations were deliberately broken. Removing the expiry condition exposed stale memory. Bypassing the revision comparison allowed stale writes. Removing the tenant restriction from purge deleted another tenant data. Each caused actual assertion failures, without relying on a syntax or import error. Restore the original code before continuing. Tests that only check that the functions return a value would miss these failures.

Before extending the system, write down the identity boundary, clock source, permitted payloads, correction path, absolute retention policy and every derived store. Add tests for each new promise. Authentication, file encryption, database migration, backup restoration, access logs, worker cancellation and external-action delivery are not provided by this three-file lab. Keep the small deterministic store as a reference while evaluating a framework or a larger state backend.

A useful completion report should distinguish source checks, executed local behaviour and unverified deployment behaviour. This workbook records Python 3.12.10 and the exact fictional output. It does not claim a model was trained, a human personally ran the examples, a live assistant was connected or a website was deployed. Memory becomes easier to reason about when its operations and limitations remain concrete enough to test.

Try it yourself · Activity 06

20 min

Catch and repair a tenant defect

Work in a separate copy of the three-file lab.

  1. Run the original ten tests.
  2. Replace the purge tenant predicate with a condition that accepts every tenant while keeping valid SQL.
  3. Observe the boundary test fail, restore the original and rerun the suite.

Which additional test would you need before exposing even one of these functions through a web endpoint?

Worked answer

For a local mutation, change WHERE tenant=? AND expires<=? to WHERE ? IS NOT NULL AND expires<=?. Both Orchard and Harbour expired rows can now be removed. test_purge_tenant_and_boundary must fail on the affected count or remaining rows. Restore the tenant predicate and require all ten methods to pass. This shows that the test catches this specific deletion boundary defect, not that all future authorisation paths are covered.

Keep learning

The complete workbook

Implement a bounded JSON schema and transactional compare-and-swap in SQLite. Trace a persistent revision counter across deletion and recreation, separate visibility from physical row removal, and test rollback and tenant boundaries. Three complete Python files, six worked activities and exact output support a practical review of what an agent should retain and when it should forget.

  1. 01
    Decide what memory means

    Give each stored field a purpose before choosing a database.

    Read here · 1 exercise
  2. 02
    Create a store with an explicit clock

    A row can exist without being eligible for retrieval.

    Read here · 1 exercise
  3. 03
    Reject a stale writer inside the transaction

    Read a revision, propose a change and let the store decide whether it is still current.

    Read here · 1 exercise
  4. 04
    Make forgetting precise

    Specify exactly what was removed and what was not measured.

    Read here · 1 exercise
  5. 05
    Resume a bounded local workflow

    Prove recovery with a file reopen before connecting an external tool.

    Read here · 1 exercise
  6. 06
    Test the lifecycle before adding a model

    Use concrete failure cases, then state the remaining limits.

    Read here · 1 exercise

Also inside: a 8-point checklist, a glossary of 8 terms and 10 questions and answers to test yourself. 6 hands-on exercises, each with a worked answer at the back where the workbook gives one.

No login, no card, no account. Before the download we ask you to follow Mickai (two quick links). Free to download and use for personal learning, study groups and inside your own team. Please do not resell the workbooks or republish them as your own. Link people to trust-agent.ai instead.

Test yourself

Questions and answers

Is a reviewed fact necessarily true?

No. Reviewed is a Boolean annotation, not a truth detector, instruction priority or permission grant.

Does the tenant field authenticate a caller?

No. The application must derive tenant from a trusted identity and enforce authorisation before invoking this local store.

Does reading extend expiry?

No. Reads leave expires unchanged. Each successful put sets now+ttl; completed step-2 reads do not call put.

What happens exactly at the deadline?

Get returns None when expires equals now. The physical row remains until an explicit deletion operation removes it.

What should happen after a revision conflict?

Reread and recompute under an explicit policy, merge carefully or stop. Do not attach the latest revision to the same stale proposal.

Why retain a global counter after deletion?

It avoids reusing a deleted row revision during recreation in this database history. Restoring an old backup can still rewind that history.

Does forget returning False prove absence?

No. The identity may be missing or its current revision may differ. False alone does not distinguish the two.

Does SQL DELETE prove complete erasure?

No. It removes selected rows from this table. Free pages, journals, backups, memory, logs and derived stores require separate consideration and verification.

Does checkpointing make an external send exactly once?

No. An effect can occur before its checkpoint commits, or two workers can act before one loses the revision race. External idempotency or a delivery protocol needs its own implementation.

What did the release tests establish?

Ten learner methods, 600 reference operations and three caught mutations verify bounded local state behaviour. Crash durability, real authorisation and production retention were not tested.

When you have finished

Get your certificate of completion

Type your name and download a certificate for this workbook as a PDF, ready to print or to add to LinkedIn. It is made on your own device, so your name is never sent to us. It is a self-declared certificate, not an accredited qualification.

Learn the language

Key terms

Checkpoint
A bounded record of execution progress used to decide where a specific run can resume.
Memory fact
A reusable stored claim with provenance and its own review and retention policy.
Revision
An opaque version token used to compare a proposed change with the current stored record.
Compare-and-swap
A conditional update that succeeds only when the observed version still matches.
Time to live
A duration used here to set expires to now plus ttl on each successful write.
Purge
An explicit operation that removes expired rows for a selected tenant.

6 of the workbook's 8 terms. The complete glossary is in the workbook.

Follow the evidence

Sources and checks

Facts last checked: .

Examples in this workbook were run on: Python 3.12.10 on Windows 11, standard-library sqlite3/json/unittest. File close/reopen, two-connection stale snapshots and transactional rollback executed using fictional records. No model, provider, account, package installation or network is required. (2026-09-27).

These workbooks use AI assistance. See how the workbooks are made.

  1. Python 3.12 sqlite3, transaction control and parameter bindingPython Software Foundation
  2. SQLite transaction behaviourSQLite
  3. SQLite PRAGMA secure_delete and its limitsSQLite
  4. Python 3.12 JSON encoding and decodingPython Software Foundation
  5. Python 3.12 unittestPython Software Foundation

Created by Mickarle Wagstaff-Irons - Micky Irons with the Mickai team. Published by Mickai LTD. Last updated 27 September 2026.

NextKeep going

Where to go next