# An Index Is Not A Source: Caching Nodes Are A Separate Class With Different Rules, And It Does Not Matter Where You Start

**version** v0.33.57
**date** 9 August 2026
**from** Human (project lead)
**to** Engineering, Product, Strategy

**type** Architecture brief

*Fourth of 9 August. Resolves an objection raised in an earlier memo about pre-created mappings. Entity identifier standards are grounded and cited, and one claim in the memo is refined. Offered to be built on and challenged.*

---

## What This Is

A structural distinction that resolves an earlier objection and simplifies several things at once: **the worry that pre-created mappings and graphs are not sources of authority is correct and does not apply, because they were never meant to be sources, and the resolution is to make the distinction structural rather than implicit, so that a node is either an assertion or a pointer and a reader can always tell which; a pointer's value is not in what it contains but in making discovery of the actual source faster, which means the same rule already applied to the global entity knowledge base on 6 August generalises, namely that anchors resolve identity and never confer authority; caching nodes then acquire properties that assertion nodes do not have, since they can be wrong without being dishonest, can be regenerated by re-running the lookup that produced them, need none of the attribution and agenda apparatus specified yesterday because they assert nothing, and are safe to prune by definition, which resolves the tension left open yesterday between pruning and retaining superseded claims; the company resolution example is the right illustration and needs one refinement, because an official global identifier does exist, is free, carries ownership hierarchy and is updated daily, but covers only around two and a half million entities with barely half renewing, so the opportunity is not the absence of an authority but the gap that authority leaves, which is a sharper and more defensible position; paying for access to authoritative sources is correctly framed as part of the cost of evidence rather than an obstacle, and this is the fourth appearance of paying for trust in the corpus, now with a settlement rail that makes micropayments practical; and the most quietly important line in the memo is that it does not matter where you start, because the graph will be deep where the work is and absent elsewhere, which is not a defect to be apologised for but the property that makes the project finite.** It is the fourth document of 9 August (cross-ref: the v0.33.57 enrichment brief, the v0.33.57 attribution brief, the v0.33.54 paying-the-fact-creator brief, the v0.33.53 customised-standard brief, and the v0.33.56 402-is-the-protocol brief). New contributions: **assertion and pointer as two structurally distinct node classes, the properties that follow for caching nodes including prunability, the refinement of the company resolution claim and the resulting sharper opportunity, and starting anywhere with topical bias named as a feature.**

## The Objection, And Why It Dissolves

The memo revisits a concern raised earlier and is right to classify rather than dismiss it. The project lead: **"if or when we add to voice debriefs a set of pre-cached and pre-created mappings and graphs and relationships, that's not a good thing because they're not sources of authority."**

The objection is correct on its own terms and rests on a category error that is easy to make. A pre-created mapping is not a weak source of authority; **it is not a source at all.** It is a route.

The memo supplies the right vocabulary. The project lead: **"what are almost transient nodes on a graph, nodes that exist to resolve something, or to find something else, or find another connection, so they're nodes almost on the way to the evidence, which fundamentally you can think of as a caching node."** And the value statement. The project lead: **"their value is not in the index itself, their value is in the fact that their data points to or makes the discovery of the source of truth, or the chain of evidence, more efficient and more effective."**

That is precisely the rule applied to the global entity knowledge base on 6 August, where the conclusion was that anchors resolve identity and do not confer factual authority. **The same rule generalises**, and stating it once covers both cases.

## Two Node Classes, Structurally Distinguished

The design consequence, and it should be enforced rather than left to convention.

**A node either asserts something or points at something, and a reader must always be able to tell which.** If a reader cannot distinguish a claim from a routing hop, then every claim in the graph inherits the credibility of the weakest cache in it, and the whole trust model collapses.

| | Assertion node | Caching or pointer node |
|---|---|---|
| Claims something | Yes | No |
| Needs attribution, agenda, funding | Yes, per yesterday's brief | **No, it asserts nothing** |
| Wrong means | Somebody was mistaken or dishonest | The pointer is stale |
| Recovery | A correction, superseding the original | Re-run the lookup |
| Retained when superseded | **Yes, always** | No need |
| Safe to prune | No | **Yes, by definition** |

Two rows there do real work.

**The attribution row is a genuine simplification.** Yesterday's brief specified an apparatus of asserted-by, vouched-for-by, funded-by, published-by and cited-by, with declared interests attached. That apparatus applies to assertions. A cache mapping a company name to an identifier does not need an agenda field, because it is not taking a position.

**And the prunability row resolves a tension left open yesterday.** That brief noted that pruning and supersede-not-delete pull against each other, and proposed pruning on use rather than age while never touching a correction chain. This sharpens it considerably: **caching nodes are the prunable class**, because their contents are recoverable by re-running the lookup that produced them, while assertions and their correction chains are not.

## The Index Must Never Become The Source

The rule that protects everything above, and it is a real risk rather than a theoretical one.

A cache that people read *instead of* the source is how a pointer becomes an unaccountable authority. That is exactly the mechanism the corpus grounded on 31 July, where citation analysis found the conversion of hypothesis into fact through citation alone, and citation diversion where a work is cited for something it does not quite say.

So two interface requirements follow. **The source is always one step away and visibly so.** And **a cached value carries the timestamp of the lookup that produced it**, because a pointer without a date is indistinguishable from a claim.

The memo's own framing supports this, since the value is explicitly located in the pointing rather than in the content.

## The Company Example, Refined

The memo picks the right illustration and one detail needs correcting, because the correction makes the opportunity sharper rather than smaller. The project lead: **"how do you resolve a company name today? I don't think there is an official place where you have that, and I guess what you have is four or five places that we could use, which combined create that."**

**An official global identifier does exist.** The Legal Entity Identifier is a twenty-character code under an international standard, published by a not-for-profit foundation backed by the Group of Twenty and the Financial Stability Board, overseen by a regulatory committee. It is explicitly a public good, free of charge, with the full population publicly available for unrestricted use at all times, an open-access index updated daily, and a programmatic interface. Its first data level answers who is who, with legal name, registered address and country of formation; its second answers **who owns whom**, with direct and ultimate parent entities and ownership structure.

That last property is worth noticing. Ownership hierarchy is exactly what a graph wants and exactly what most name-resolution sources do not provide.

**And the memo's underlying point survives intact, because coverage is the problem.** Reported figures put the population at around two and a half million entities against millions of potential ones, with adoption slow outside regulatory mandates, a global renewal rate of roughly fifty-six percent, and a documented gap among smaller companies in developing economies. An entity that never registered has no identifier at all until it applies.

So the position becomes more defensible when stated precisely:

```
   HAS AN OFFICIAL IDENTIFIER        anchor to it, free, daily, with ownership
   (~2.5m entities, mostly           and DO NOT rebuild what exists
    financially regulated)
                                     note: about half of records may be
                                     unrenewed, so freshness still applies

   HAS NO OFFICIAL IDENTIFIER        THIS is the consolidation opportunity:
   (most companies in the world)      reconcile several partial registries
                                      into a resolvable pointer
```

That is a better opportunity statement than the absence of any authority, because it says exactly where the gap is and it does not invite anybody to point out that the standard exists.

The renewal figure is also a neat illustration of yesterday's freshness point: even an official, daily-updated, standards-backed index carries stale records, so a resolution result needs a date regardless of how authoritative its source.

## Paying For Trust Is A Cost Of Evidence

The memo's framing here is right and worth adopting as language. The project lead: **"some of them could be paid, there could be services that you actually have to have a subscription to access, which is okay, it's not a problem at all, because it's just part of the chain of evidence, it's part of the cost of evidence."** And the recalled position. The project lead: **"this goes back to articles we wrote a while back, where we should pay for trust."**

That framing does useful work, because it moves a subscription from an obstacle to a line item. A conclusion that rests on three paid lookups cost something to reach, and that is a fact about the conclusion rather than a complaint about the vendors.

This is now the fourth appearance of paying for trust in the corpus: fact certification and evidence-based revenue on 5 July, paying the fact creator for contextual validation on 31 July, research paid once and read many times yesterday, and now paid authoritative sources. **Four arrivals from different directions makes it a primitive.**

What is new since July is that it is buildable. The payment work of 6 August established settlement in around two hundred milliseconds at a fraction of a cent with no account required, which is exactly the shape a per-lookup charge needs and exactly what card rails cannot serve. The memo says as much. The project lead: **"even if it's a micro payment, it sort of adds up very quickly."**

One design consequence worth taking: **the chain of evidence should carry its cost.** If a node's production cost is recorded, as proposed yesterday, then a conclusion's cost is the sum along its supporting paths, which is both a pricing input and something a reader may want to see.

## It Does Not Matter Where You Start

The quietest and most important thing in the memo, and it answers the objection that kills projects of this kind before they begin. The project lead: **"people always worry about where you start. In a way, it doesn't matter where you start. What matters is that you start, because you then start navigating the graph from there."**

The failure this prevents is the belief that a complete ontology must exist before anything can be mapped, which guarantees that nothing is ever mapped. The corpus already holds the same position from the regulatory work: a customised standard **starts from nothing being relevant** and accretes provisions only as facts attach them, which inverts every tool that hands over the whole text and asks the reader to exclude.

And the memo goes further, in the part most likely to be mistaken for a weakness. The project lead: **"the graph that we're going to create is going to be very biased in terms of the content, to the topics and the areas and the stuff that we focus on, which is great, because it doesn't mean that we have to handle everything."**

**Topical bias is the feature.** A graph deep where the work is and absent everywhere else is finite, achievable, and immediately useful, whereas a graph attempting even coverage is neither. It is also the same relevance discipline the corpus applies to registers, where each role holds only what concerns it.

The consequence for planning is direct: **do not scope this by coverage.** Scope it by the questions actually being asked, and let depth follow use. The memo's closing formulation is the right one. The project lead: **"we just zoom in on that part."**

## What Publishing Our Own Indexes Buys

The memo names three reasons and they hold up.

**Trust in our own usage.** The project lead: **"at least in our usage, we have a high degree of trust of those."** That is a legitimate and modest claim: we know how they were built and when.

**A place for things that have none.** The project lead: **"they allow us to cache and index lots of other resources that might not currently have a good place."** That is the consolidation opportunity above.

**Control of publication.** The project lead: **"it allows us to be in control of what's happening, and have a place to publish our stuff a lot more effectively and a lot more visual, especially because there is still a lot of gaps in the way that a lot of this data is being published today."**

One caution against the third. Publishing resolvable indexes that others depend on makes us infrastructure, which carries a weaker version of the conflict the corpus named on 31 July when it declined to own a registry. It is weaker because **a pointer is not a claim**, so we are not adjudicating anything. But two obligations follow: an index we publish needs a freshness signal and a correction path, and it should never quietly become the thing people cite.

## What This Does Not Try To Be

- **Not a source of authority.** Caching nodes route; they do not assert.
- **Not indistinguishable from assertions.** The two classes are structurally separate and visibly so.
- **Not a replacement for existing identifiers.** Anchor to what exists; consolidate only where nothing does.
- **Not free of a freshness obligation.** Every cached value carries the date of the lookup that produced it.
- **Not scoped by coverage.** Deep where the work is, absent elsewhere, and that is the point.

## Honest Tensions

| Tension | Note |
|---------|------|
| Indexes we publish versus becoming infrastructure | A pointer is not a claim, and people will still depend on it, so freshness and correction become obligations |
| The index becoming the source | It is the documented failure mode of citation, and only interface discipline prevents it |
| Prunable caches versus rebuild cost | Recoverable does not mean free, and pruning something heavily used costs a re-lookup every time |
| Paid sources in the chain | Framing subscriptions as a cost of evidence is honest and puts a paywall between a reader and a verification |
| Topical bias as a feature | It is right and it means the graph is silent on anything outside our work, which will read as a gap to somebody |
| Anchoring to an official identifier | It is free, standard and current, and covers a minority of entities with about half of records unrenewed |

## Open Questions

| Question | Notes |
|----------|-------|
| How are the two node classes rendered differently? | A reader must never have to work out whether something is a claim or a route |
| What is the freshness policy per cache? | Different sources decay at different rates, so one interval will not serve |
| Which registries are reconciled for entities without an official identifier? | The consolidation opportunity, and it needs a named starting set |
| Does the chain of evidence display its cost? | The sum along supporting paths, and whether a reader wants to see it |
| Who pays for a paid lookup, and when? | First reader, every reader, or amortised through node access pricing |
| What is the correction path for a published index? | If we publish it and it is wrong, somebody needs a way to say so |
| What is the starting scope? | Chosen by the questions being asked rather than by any notion of coverage |

## Relationship To Previous Briefs

| Date | Document | Relationship |
|---|---|---|
| 9 Aug | `v0.33.57__arch-brief__sg-send-enrichment-and-shared-anchors-research-paid-once-wikidata-is-the-concept-layer.md` | Anchors resolve identity and not authority, which this generalises into two node classes |
| 9 Aug | `v0.33.57__arch-brief__sg-send-fact-does-not-exist-in-a-vacuum-agenda-is-context-corrections-must-propagate.md` | The attribution apparatus that caching nodes do not need, and the pruning tension this resolves |
| 31 Jul | `v0.33.54__strategy-brief__sg-send-paying-the-fact-creator-contextual-validation-not-truth-micropayments-for-correct-use.md` | Paying for trust, and the citation distortion that makes index-becomes-source a real risk |
| 28 Jul | `v0.33.53__arch-brief__sg-send-customised-standard-eu-ai-act-graph-nothing-relevant-until-facts-attach-browser-query-layer.md` | Starting from nothing relevant and accreting as facts attach, which is starting anywhere stated for a different domain |
| 6 Aug | `v0.33.56__arch-brief__sg-send-402-is-the-protocol-not-the-rail-schemes-cover-cards-and-credits-upto-solves-unknown-cost.md` | The settlement rail that makes per-lookup micropayments practical |

---

## Key Claims

| # | Claim |
|---|-------|
| 1 | The objection is right and does not apply, because a pre-created mapping is not a weak source but not a source at all |
| 2 | A node either asserts or points, and a reader must always be able to tell which |
| 3 | If claims and routes are indistinguishable, every claim inherits the credibility of the weakest cache |
| 4 | Caching nodes need no attribution or agenda apparatus, because they assert nothing |
| 5 | Caching nodes are prunable by definition, which resolves the tension left open yesterday |
| 6 | An index must never become the source, which is the documented citation failure mode |
| 7 | A cached value without the date of its lookup is indistinguishable from a claim |
| 8 | An official global company identifier does exist, is free, is updated daily and carries ownership hierarchy |
| 9 | It covers around two and a half million entities with roughly half renewing, so the gap it leaves is the opportunity |
| 10 | Paid sources are a cost of evidence rather than an obstacle, and this is the fourth appearance of paying for trust |
| 11 | It does not matter where you start, which answers the belief that a complete ontology must come first |
| 12 | Topical bias is the feature, so scope by the questions asked rather than by coverage |

---

## Sources

- The global entity identifier standard, its twenty-character structure, publication by a not-for-profit foundation backed by the Group of Twenty and the Financial Stability Board under regulatory committee oversight, its status as a public good available free of charge with the full population publicly available for unrestricted use, the open-access index updated daily, the programmatic interface, and the two data levels covering who is who and who owns whom: https://www.gleif.org/en/organizational-identity/introducing-the-legal-entity-identifier-lei/iso-17442-the-lei-code-structure and https://www.gleif.org/en/about-lei/introducing-the-legal-entity-identifier-lei and https://registry.opendata.aws/lei/
- Reported coverage and quality limitations, including a population of roughly two and a half million against millions of potential entities, slow voluntary adoption outside regulatory mandates, a global renewal rate of around fifty-six percent, and the gap among smaller companies in developing economies: https://www.digitalizetrade.org/services/global-legal-entity-identifier-foundation-gleif-global-lei-system
- Confirmation that no identifier exists for an entity that has never registered, and that records are publicly searchable and free to verify: https://rapidlei.com/knowledge-base/what-is-an-lei/
- The relationship between the identifier system and broader company registry aggregation, including quality assurance of issued data: https://knowledge.opencorporates.com/knowledge-base/lei-legal-entity-identifier/

---

This document is released under the Creative Commons Attribution 4.0 International licence (CC BY 4.0).
