retrieval · Level 4

Permission-aware retrieval and private knowledge bases

Keep tenant and document permissions in the retrieval path, then test what can cross the boundary.

By Mickarle Wagstaff-Irons - Micky Irons

  • Level 4Frontier
  • 180 min
  • 6 chapters
  • Free PDF, no account
The Indexretrieval / 04

Start with the essentials

The short answer

Permission-aware retrieval limits the documents a caller can receive before they become model context, snippets or citations. Authentication supplies identity; authorisation decides access to each resource. This course implements a small tenant-and-group policy with fictional data, tests denied paths and explores revocation, caches and audit boundaries. It is a local teaching lab, not a production authentication service.

What you will learn

  • Separate authenticated identity, document eligibility and relevance ranking.
  • Implement a tenant-and-group policy that denies missing matches.
  • Apply permission checks before ranking limits and downstream context selection.
  • Test tenant boundaries, empty permissions, malformed input and group removal.
  • Describe cache, citation and audit risks across a retrieval pipeline.
  • Prepare a synthetic integration matrix with explicit acceptance evidence.

Who it is for

Builders who understand RAG and basic Python, and need to reason about private document retrieval without connecting real private data.

Before you start

  • Understand retrieval-augmented generation and document chunks.
  • Read Python functions, sets and simple classes.
  • Run scripts and unit tests in a local folder.

Read a sample · Chapter 01 of 06

01

Put the permission boundary before the answer

A relevant document is not automatically an eligible document. Decide who can read it before asking a model to use it.

Imagine two organisations sharing a document service. Both have a group called ops and a document called launch. Alice belongs to north/ops, while Cara belongs to south/ops. Their group names match, but their organisations do not. A search for release instructions must not join those two worlds just because the text is similar. This workbook makes that difference visible with five short, fictional documents.

Authentication establishes an identity. Authorisation decides which action that identity may perform on a resource. Here the action is reading a document. The lab receives a Principal object representing a scope that a trusted server would already have resolved. It does not verify a password, session, token, issuer or group directory. Anyone running the script can construct a different Principal; the Python class is not evidence of identity.

The policy is deliberately small: a document must belong to the same tenant, and at least one of its reader groups must appear in the principal’s groups. Empty groups or an empty document reader set grant nothing. There is no anonymous-public mode, administrator override, owner exception, deny-rule precedence or group inheritance. Those would be separate policy decisions with separate tests. Write this rule in plain language before implementing it.

Draw the intended boundary as identity resolution, eligible documents, relevance scoring, bounded results, and then any answer-producing service. A hidden search result can still leak through a title, snippet, citation, cached answer or logged context. For this lab, downstream code receives only the returned Document objects. The in-memory corpus itself is trusted backend input, not a file to send to a browser or model.

Use a new local folder with Python 3.12. Create access.py, fixtures.py, demo.py and test_access.py as shown in the next chapters. No package install, account or GPU is needed. Work only with the fictional fixture; its document strings are examples chosen to expose ranking and permission mistakes. Nothing here reports Mickai infrastructure, customer data or a measured security assessment.

Try it yourself · Activity 01

15 min

Write the policy before the code

Create a small permission table for the fictional service.

  1. Give Alice north/ops and Cara south/ops.
  2. Give a north document the reader group ops.
  3. Predict each caller’s decision, then repeat with an empty reader set.

Which component in your eventual application will establish tenant and group membership?

Worked answer

Alice may read the north/ops document. Cara may not, because matching group names do not replace the tenant check. With no reader groups, neither caller matches the policy. A document owner who wants public access would need a separately defined public rule; an empty field cannot silently mean public.

Read a sample · Chapter 02 of 06

02

Represent trusted scope and document identity

Make invalid states visible and use a compound identity when document names repeat across tenants.

Save the first part below as access.py. It validates the simple identifier grammar, requires immutable group sets and bounds the amount of fixture text. Principal and Document are frozen dataclasses, which prevent ordinary field reassignment. Their type hints are not runtime validation by themselves; the post-initialisation checks enforce the conventions used by this example.

python · 45 lines
"""Local teaching policy, not an authentication service."""
from dataclasses import dataclass
import re


def identifier(value):
    if type(value) is not str or not re.fullmatch(
            r"[a-z][a-z0-9-]{0,31}", value):
        raise ValueError("Invalid identifier")


def group_set(value):
    if type(value) is not frozenset or len(value) > 100:
        raise ValueError("Expected at most 100 group IDs")
    for item in value:
        identifier(item)


@dataclass(frozen=True)
class Principal:
    tenant: str
    user: str
    groups: frozenset

    def __post_init__(self):
        identifier(self.tenant)
        identifier(self.user)
        group_set(self.groups)


@dataclass(frozen=True)
class Document:
    tenant: str
    doc_id: str
    text: str
    readers: frozenset

    def __post_init__(self):
        identifier(self.tenant)
        identifier(self.doc_id)
        group_set(self.readers)
        if type(self.text) is not str or not self.text.strip():
            raise ValueError("Expected non-empty document text")
        if len(self.text) > 10000:
            raise ValueError("Document too large for this lab")

The tenant and document ID together identify a record. north/launch and south/launch are different records. Two north/launch entries in the same corpus would be ambiguous, even if their text differs, so the retrieval function will reject that input. In a real index, the key format, chunk identifiers and source-document version must preserve the same distinction. Do not let one tenant overwrite another tenant’s chunk through a shared unqualified key.

Group names in this lab are exact, case-sensitive identifiers with a restricted ASCII grammar. A string such as ops OR true is not a group identifier here and is rejected. This is not a substitute for safe query construction in a database or search service. Use the service’s supported parameter or filter-building mechanism and keep identity attributes separate from free-text query input. The lab never assembles a database expression.

Frozen objects and frozenset values make the example easier to reason about during one call. They are not a sandbox, cryptographic proof or protection against malicious code running in the same process. The corpus and policy objects come from trusted application code in this exercise. A production adapter must validate stored permission metadata and resolve caller scope through its real authentication and policy services.

Save fixtures.py alongside access.py. The two launch records deliberately share a document ID. The unassigned record has excellent query overlap but no readers; it must remain inaccessible. The shared record grants either ops or finance within north, rather than requiring both. Every value is synthetic.

python · 20 lines
"""Every person, tenant, group and document here is fictional."""
from access import Document, Principal

DOCS = (
    Document("north", "handover", "release checklist",
             frozenset({"ops"})),
    Document("north", "launch", "release checklist deployment rollback",
             frozenset({"finance"})),
    Document("north", "shared", "release checklist deployment notes",
             frozenset({"ops", "finance"})),
    Document("south", "launch", "release checklist deployment",
             frozenset({"ops"})),
    Document("north", "unassigned", "release checklist deployment rollback",
             frozenset()),
)
ALICE = Principal("north", "alice", frozenset({"ops"}))
BOB = Principal("north", "bob", frozenset({"finance"}))
CARA = Principal("south", "cara", frozenset({"ops"}))
VISITOR = Principal("north", "visitor", frozenset())
QUERY = "release checklist deployment rollback"

Try it yourself · Activity 02

20 min

Inspect the five document identities

List each tenant/document pair and its reader groups.

  1. Explain why the two launch IDs are valid.
  2. Predict which documents Alice and Bob may read before considering query words.
  3. Explain what happens when a group set is a normal mutable set instead of frozenset.

What would a chunk need to retain when one source document is split into many searchable records?

Worked answer

The five identities are north/handover, north/launch, north/shared, south/launch and north/unassigned. Alice is eligible for north/handover and north/shared. Bob is eligible for north/launch and north/shared. The constructor rejects a mutable set under this lab’s contract. The restriction supports stable local state; it does not authenticate membership.

Read a sample · Chapter 03 of 06

03

Filter eligible records before relevance and limits

A result limit should select the best allowed matches, rather than letting denied records consume the shortlist.

Append this second part to access.py after the Document class. Keep the functions at the left margin. retrieve validates the call, checks duplicate compound identities, selects eligible records, counts distinct query-token overlaps and sorts matches. Equal scores use document ID order within the single permitted tenant. It then applies the result limit. This is a transparent keyword toy, not an embedding model or production search engine.

python · 31 lines
def allowed(doc, principal):
    return (doc.tenant == principal.tenant
            and bool(doc.readers & principal.groups))


def retrieve(documents, principal, query, *, limit=3):
    if type(principal) is not Principal:
        raise ValueError("Expected a server-resolved principal")
    if type(documents) is not tuple or len(documents) > 10000:
        raise ValueError("Expected a bounded document tuple")
    if any(type(doc) is not Document for doc in documents):
        raise ValueError("Expected validated documents")
    keys = [(doc.tenant, doc.doc_id) for doc in documents]
    if len(keys) != len(set(keys)):
        raise ValueError("Duplicate document identity")
    if type(limit) is not int or not 1 <= limit <= 50:
        raise ValueError("Invalid result limit")
    if type(query) is not str or not 1 <= len(query) <= 500:
        raise ValueError("Invalid query")
    terms = set(re.findall(r"[a-z0-9]+", query.lower()))
    if not terms:
        raise ValueError("Query needs an ASCII word or number")
    eligible = [d for d in documents if allowed(d, principal)]
    ranked = []
    for doc in eligible:
        words = set(re.findall(r"[a-z0-9]+", doc.text.lower()))
        score = len(terms & words)
        if score:
            ranked.append((score, doc.doc_id, doc))
    ranked.sort(key=lambda row: (-row[0], row[1]))
    return tuple(doc for _, _, doc in ranked[:limit])

The query release checklist deployment rollback contains four distinct terms. north/launch and north/unassigned match all four but Alice may read neither. north/shared matches three and north/handover matches two. With limit=1, Alice should receive north/shared. Ranking every record, keeping only the best one and then removing denied records could return an empty list instead. That loses an available permitted result even when the final response does not leak text.

Filtering after sending text to a reranker or answer model is a different and more serious boundary error: that service has already received the material. In this exercise all data stays in one local backend process, and scoring only inspects eligible document text. In a distributed application, identify each component allowed to process restricted content. A final output scrub cannot undo a transfer to an unauthorised downstream component.

This function returns full allowed Document objects because the next local stage might need their text. An HTTP response adapter would usually select the fields the caller needs rather than serialising every internal field, including reader groups. A citation or direct-download endpoint must perform its own access decision when called; possession of a document ID in an old result is not a lasting permission grant.

Empty eligible results and a query with no lexical match both produce an empty tuple. Malformed calls raise ValueError before a result is returned. The demo does not map exceptions to HTTP responses or implement confidential error messaging. In a service, decide what the user may learn from failures, protect detailed diagnostics and never turn a failed permission lookup into an unrestricted search.

Try it yourself · Activity 03

20 min

Follow a one-result request

Trace Alice’s query with limit=1.

  1. List eligible documents and their overlap counts.
  2. Apply the rank order and the limit.
  3. Describe the result of ranking first, limiting to one and filtering afterwards.

Which intermediate lists in your own pipeline can contain titles, snippets or full text?

Worked answer

Only north/shared and north/handover are eligible; their scores are three and two. Alice receives north/shared. A global top-one selection can choose the denied north/launch record and then discard it, leaving no result. The missing result is a sequencing defect. Sending that denied record to a separate model before filtering would also cross the intended permission boundary.

Read a sample · Chapter 04 of 06

04

Run the same query under different scopes

Hold the query constant so the effect of permissions is observable.

Permission comes first

Fictional Alice belongs to north and ops. Two of five documents pass the tenant and group checks; ranking returns north/shared then north/handover.
Follow Alice through the course lab. Labels identify every keep/deny decision; the diagram does not rely on gold alone. Open the full-size permission diagram.

The fixed query is release checklist deployment rollback. Alice has tenant north and group ops. A document is eligible only when its tenant matches and its reader groups overlap the principal’s groups. An empty reader list grants no access. The two documents named launch belong to different tenants; their full identities are north/launch and south/launch.

All five fictional records: Alice’s eligibility decisions
DocumentReaders DecisionMatching words
north/handoveropsKeep: both checks pass2
north/launchfinanceDeny: no shared groupNot scored
north/sharedfinance, opsKeep: both checks pass3
south/launchopsDeny: different tenantNot scored
north/unassignedNoneDeny: no reader grantsNot scored

Only eligible documents are scored. north/shared matches three distinct query words; north/handover matches two. Sort by that count, then by document ID for a tie, and apply limit=3. Alice receives these two results in that order. The counts are word overlap, not probabilities. Denied documents are not scored or used to fill the remaining result slot.

Same query and limit, different fictional scopes
PrincipalTenant GroupsResults in order
Alicenorthopsnorth/shared, north/handover
Bobnorthfinancenorth/launch, north/shared
Carasouthopssouth/launch
VisitornorthNoneNo results
Alice after group removalnorthNoneNo results

These identities, groups and documents are invented teaching fixtures. The illustrated scope is assumed to come from a trusted server; the local Principal class validates its shape and does not authenticate anyone. The final row creates a new principal with no groups; it does not prove distributed revocation. This is not a production security certification. A real service must not expose its denied records as this public fictional lesson does.

Save demo.py beside the other files and run python demo.py from that folder. Alice and Bob share a tenant but have different groups. Cara has Alice’s group name in another tenant. Visitor has no groups. The demo prints only returned document labels, then constructs a new scope for Alice after all of her group memberships have been removed.

python · 11 lines
from dataclasses import replace
from access import retrieve
from fixtures import DOCS, ALICE, BOB, CARA, VISITOR, QUERY

for principal in (ALICE, BOB, CARA, VISITOR):
    hits = retrieve(DOCS, principal, QUERY)
    labels = [f"{d.tenant}/{d.doc_id}" for d in hits]
    print(f"{principal.user}: {', '.join(labels) or '(empty)'}")

revoked = replace(ALICE, groups=frozenset())
print("alice after group removal:", retrieve(DOCS, revoked, QUERY))

Expected output from the delivered Python example:

text · 5 lines
alice: north/shared, north/handover
bob: north/launch, north/shared
cara: south/launch
visitor: (empty)
alice after group removal: ()

The group-removal example changes the principal supplied to the next function call. No result cache exists, so the next call returns nothing. This demonstrates fresh-scope behaviour within the lab; it does not implement directory synchronisation, token expiry, distributed invalidation or a verified policy for in-flight requests. Reusing the old ALICE object still represents the old groups. An application must define how quickly a real membership change reaches every request.

Consider an answer cache keyed only by the query. Bob could populate it with finance material and Alice could later receive that answer for the identical text. A tenant key alone would not separate these two callers. Cache design needs an equivalent effective authorisation scope, relevant policy/content versions and a revocation strategy, or an access check before reuse. Adding a user ID alone is still insufficient when that user loses a group. The safest initial teaching implementation is the one here: no answer cache.

Revocation also affects stored conversations, exported summaries, download links and previously generated citations. You cannot make a reader forget an answer already delivered. You can specify what happens on future reads, how an in-flight response observes a policy change and which retained copies remain available. Record those semantics before claiming immediate revocation. Use synthetic documents to exercise the full request chain before introducing private content.

Try it yourself · Activity 04

15 min

Design a revocation timeline

Describe Bob’s group being removed between two identical searches.

  1. Mark the identity lookup, permission snapshot, retrieval and response events.
  2. State which request may use the old snapshot.
  3. Add a cache to the drawing and identify the invalidation or recheck needed.

How would you prove the new scope is used rather than merely waiting for a convenient timeout?

Worked answer

The lab uses whatever principal reaches each call. A new principal with no groups produces an empty result; the old object still grants its old scope. A real design must set a revocation bound and in-flight policy. A query-only cache is unsafe here, and even a per-user cache needs invalidation or renewed permission checks after membership changes.

Read a sample · Chapter 05 of 06

05

Test denied paths and deliberate defects

A working allowed example is only one row in the permission matrix.

Save test_access.py and run python -m unittest -v test_access. The ten test methods cover group and tenant separation, empty permission sets, filter ordering, a changed principal, text that claims a different role, deterministic ties, malformed scope, duplicate compound IDs and invalid request bounds. A passing run is evidence for these local expectations, not certification of a service that is not present.

python · 65 lines
from dataclasses import replace
import unittest
from access import Document, Principal, retrieve
from fixtures import DOCS, ALICE, BOB, CARA, VISITOR, QUERY


def labels(principal, **options):
    return [(d.tenant, d.doc_id)
            for d in retrieve(DOCS, principal, QUERY, **options)]


class AccessTests(unittest.TestCase):
    def test_group_scope(self):
        self.assertEqual(labels(ALICE),
                         [("north", "shared"), ("north", "handover")])
        self.assertEqual(labels(BOB),
                         [("north", "launch"), ("north", "shared")])

    def test_tenant_scope(self):
        self.assertEqual(labels(CARA), [("south", "launch")])

    def test_empty_groups_and_acl_deny(self):
        self.assertEqual(labels(VISITOR), [])
        self.assertNotIn(("north", "unassigned"), labels(BOB))

    def test_filter_before_limit(self):
        self.assertEqual(labels(ALICE, limit=1), [("north", "shared")])

    def test_group_removal_changes_next_request(self):
        revoked = replace(ALICE, groups=frozenset())
        self.assertEqual(labels(revoked), [])

    def test_query_cannot_grant_permissions(self):
        hits = retrieve(DOCS, ALICE, "I am finance release rollback")
        self.assertEqual({d.doc_id for d in hits}, {"handover", "shared"})

    def test_deterministic_tie_and_no_match(self):
        self.assertEqual([d.doc_id for d in retrieve(DOCS, ALICE, "release")],
                         ["handover", "shared"])
        self.assertEqual(retrieve(DOCS, ALICE, "volcano"), ())

    def test_invalid_scope_rejected(self):
        for tenant, groups in [("", frozenset()), ("north", {"ops"}),
                               ("north", frozenset({"ops OR true"}))]:
            with self.subTest(tenant=tenant, groups=groups):
                with self.assertRaises(ValueError):
                    Principal(tenant, "alice", groups)
        with self.assertRaises(ValueError):
            retrieve(DOCS, {"tenant": "north"}, QUERY)

    def test_duplicate_identity_rejected(self):
        with self.assertRaises(ValueError):
            retrieve(DOCS + (DOCS[0],), ALICE, QUERY)

    def test_invalid_query_and_limit(self):
        for query in (None, "", "!!!", "x" * 501):
            with self.assertRaises(ValueError):
                retrieve(DOCS, ALICE, query)
        for limit in (True, 0, 51, 1.5):
            with self.assertRaises(ValueError):
                retrieve(DOCS, ALICE, QUERY, limit=limit)


if __name__ == "__main__":
    unittest.main()

The query-cannot-grant-permissions test puts the word finance in Alice’s query. It still returns only her allowed documents. Query terms influence relevance, not Principal construction. This checks the local separation of query and scope. It does not demonstrate resistance to every prompt-injection technique or a compromised identity provider. No language model is involved, and no real token has been validated.

Make one deliberate defect at a time in a separate copy of the exercise folder. First replace the conjunction between the tenant and group conditions with a disjunction. Then restore the file. Next allow an empty reader set as though it were public. Restore again. Finally move the result limit before the permission filter. Each defect should break a relevant existing test; a test that remains green has not established that boundary. Never keep the defective copy as the delivered implementation.

Extend the cases with a document whose text changes but whose denied permission remains unchanged. Alice’s returned list should stay unchanged. Also test two equally scoring allowed records in the opposite input order: the ID tie rule should preserve the output order. These properties inspect observable decisions without relying on a particular internal list implementation. Do not assert that removing a group always returns a subset of a previously limited list: a different allowed record can move into the freed result slot.

Try it yourself · Activity 05

20 min

Prove that a boundary test can fail

Run the suite against one deliberately incorrect policy in a local copy.

  1. Record the passing baseline.
  2. Change the tenant/group conjunction to a disjunction and identify a failed tenant test.
  3. Restore the source, rerun the suite and explain the observed difference.

Which denied scenario would still be missing if you tested only Alice and never Cara?

Worked answer

The disjunction allows a matching group from another tenant or an unrelated group in the same tenant. The fixture’s north and south ops scopes expose it. The original conjunction requires both conditions. A useful record contains the source change, the failing assertion and the restored passing run. Merely claiming a mutation was tried is not test evidence.

Read a sample · Chapter 06 of 06

06

Carry the boundary into a real retrieval service

The local policy is a starting point for an integration contract, with explicit gaps to close.

Before replacing the corpus tuple with a real index, write an adapter contract. State how the application resolves tenant and groups, how each document and chunk receives permission metadata, which version of that metadata the query observes and which downstream services receive text. Build a synthetic test collection whose tenant names, group names, overlapping IDs and denied high-scoring documents reproduce the local edge cases. Keep expected results independently recorded.

The Microsoft security-filter reference illustrates storing permission identifiers and applying a security filter to each query. Its filterable and retrievable field settings have different purposes; hiding an ACL field in returned documents does not itself authorise access. The provider’s supported filter semantics and vector-search execution mode need their own integration checks. This workbook uses no Azure account or API and makes no claim about verified provider latency or security.

Use an explicit release matrix for the adapter: same tenant/allowed group; same tenant/wrong group; other tenant/same group; empty ACL; changed membership; missing identity data; duplicate imported identity; denied record with the highest relevance score; and an old citation fetched after revocation. For each row, record the fixture, expected eligible IDs, actual returned IDs and the policy version. Repeat across the stages that actually read text, not only the final browser response.

An audit record can help explain a decision without copying the private material it is meant to protect. For a synthetic exercise, propose a request correlation ID, time, internal principal reference, tenant reference, policy version, operation, outcome and permitted result IDs or counts where appropriate. Treat those identifiers as potentially sensitive too. Decide who can read the log and for how long. Avoid raw document text, bearer tokens and unreviewed query strings in routine diagnostics. The lab prints synthetic result labels but does not implement a production audit service.

Record residual risks alongside the passing matrix: stale membership, partially updated chunk ACLs, mixed-version indexes, a failed directory lookup, broad backend credentials, shared caches, exported conversation history and direct file access. A permission-filtered retriever does not make an index, backup or embedding store publicly safe. Test the storage and transport boundaries separately. The outcome of this course is a reproducible local policy and an integration plan with named evidence, not a promise that private retrieval has been fully secured.

Try it yourself · Activity 06

15 min

Write the adapter acceptance record

Prepare a one-page test plan for a future private knowledge service.

  1. Choose three denied cases and one allowed case with exact expected IDs.
  2. Name every component that receives document text.
  3. Specify membership freshness, cache behaviour, citation checks and a minimal audit record.

What evidence would distinguish a successful local unit test from a verified production access boundary?

Worked answer

A concrete plan names the principal source, policy snapshot and synthetic documents; checks returned IDs before any answer call; repeats access checks for later citations/downloads; states how old cached results are invalidated or rechecked; and restricts diagnostic content. Include an explicit failure path for unavailable permission data. Do not label the integration complete until the real adapter and downstream path have been exercised.

Keep learning

The complete workbook

Implement a deny-by-default document policy and a deterministic keyword retriever in standard-library Python. Compare four fictional callers, reject malformed scope and duplicate document identities, test filtering before result limits, and plan checks for permission changes and downstream evidence.

  1. 01
    Put the permission boundary before the answer

    A relevant document is not automatically an eligible document. Decide who can read it before asking a model to use it.

    Read here · 1 exercise
  2. 02
    Represent trusted scope and document identity

    Make invalid states visible and use a compound identity when document names repeat across tenants.

    Read here · 1 exercise
  3. 03
    Filter eligible records before relevance and limits

    A result limit should select the best allowed matches, rather than letting denied records consume the shortlist.

    Read here · 1 exercise
  4. 04
    Run the same query under different scopes

    Hold the query constant so the effect of permissions is observable.

    Read here · 1 exercise
  5. 05
    Test denied paths and deliberate defects

    A working allowed example is only one row in the permission matrix.

    Read here · 1 exercise
  6. 06
    Carry the boundary into a real retrieval service

    The local policy is a starting point for an integration contract, with explicit gaps to close.

    Read here · 1 exercise

Also inside: a 8-point checklist, a glossary of 8 terms and 10 questions and answers to test yourself. 6 hands-on exercises, each with a worked answer at the back where the workbook gives one.

No login, no card, no account. Before the download we ask you to follow Mickai (two quick links). Free to download and use for personal learning, study groups and inside your own team. Please do not resell the workbooks or republish them as your own. Link people to trust-agent.ai instead.

Test yourself

Questions and answers

Does the Principal class authenticate a caller?

No. It validates the shape of a local scope object. Trusted application code must establish the real identity, tenant and memberships before constructing such a scope. Anyone running this toy script can instantiate a different principal.

What grants access in this lab?

The tenant must match and at least one reader group must intersect with the principal’s groups. Both conditions are required. Empty group sets grant nothing.

Why can two documents both be called launch?

Identity is the pair of tenant and document ID. north/launch and south/launch differ. Two entries with the same pair in one corpus are rejected.

Are frozen dataclasses a security sandbox?

No. They prevent ordinary reassignment and help maintain a stable example. They do not prove identity or defend against hostile code running in the same process.

Why filter before applying the result limit?

Denied high-scoring records should not consume a shortlist and hide available permitted matches. The lab also filters before scoring so downstream relevance code receives eligible text only.

Can a later answer scrub undo a prior text transfer?

No. If an unauthorised downstream service already received document text, removing it from the final response does not undo that transfer. Identify permission boundaries before each service receives content.

Does the demo prove immediate distributed revocation?

No. It passes a newly constructed scope to a function with no cache. Real directories, sessions, caches and in-flight requests need explicitly tested freshness and revocation rules.

Is a cache keyed only by tenant safe for Alice and Bob?

Not under this policy. They share a tenant but have different groups. Cached results must respect effective permissions and changes to policy and content, with invalidation or renewed checks.

What does the query-role test demonstrate?

Putting finance in the query does not change Alice’s supplied groups. The test checks separation of query text and local scope, not a complete prompt-injection or authentication defence.

What evidence is needed beyond the local tests?

Exercise the real identity adapter, index filters, chunk permissions, downstream text transfers, caches and later citation/download reads using an independent expected permission matrix. Record versions, failures and unresolved risks.

When you have finished

Get your certificate of completion

Type your name and download a certificate for this workbook as a PDF, ready to print or to add to LinkedIn. It is made on your own device, so your name is never sent to us. It is a self-declared certificate, not an accredited qualification.

Learn the language

Key terms

Authentication
Establishing identity through a trusted mechanism.
Authorisation
Deciding whether an identity may act on a resource.
Tenant
The organisation scope used for identity and access decisions.
ACL
Access-control list: metadata describing permitted access.
Deny by default
No matching permission grant means no access.
Security trimming
Filtering results to the caller’s eligible documents.

6 of the workbook's 8 terms. The complete glossary is in the workbook.

Follow the evidence

Sources and checks

Facts last checked: .

Examples in this workbook were run on: Python 3.12.10 on Windows 11. Four local standard-library files; no network, identity provider, private documents or model calls. Verification evidence accompanies this course. (2026-09-27).

These workbooks use AI assistance. See how the workbooks are made.

  1. Authorization Cheat SheetOWASP Cheat Sheet Series
  2. Security filter pattern for Azure AI SearchMicrosoft Learn
  3. Python 3.12 dataclassesPython Software Foundation
  4. Logging Cheat SheetOWASP Cheat Sheet Series

Created by Mickarle Wagstaff-Irons - Micky Irons with the Mickai team. Published by Mickai LTD. Last updated 27 September 2026.

NextKeep going

Where to go next