Everything Is an Event, Thing, or Concept

A Minimalist Ontology for Compressed Neural Reasoning

COSAA-SST: Converged Symbolic Attention Architecture over Semantic Spacetime

Working Paper Thought Dump - Please don’t judge this is created with the help of AI putting my thoughts on paper. I have those since two years but there we are I did not come to publish them yet — February 2026

Preface: Why This Exists

There is a moment in any sufficiently long conversation about transformer architectures where someone says: “But what is the model actually doing?” Not what weights are activated. Not which attention heads fire. What is it doing — in the way that a human doing the same task would describe what they are doing.

The honest answer is: we don’t know. And we’ve built an entire civilisation-shaping technology on that ignorance.

This paper is a collection of thoughts toward an architecture where we would know. Where every computational step has a name, a type, and a reason. Where compression isn’t a post-hoc optimisation but a consequence of structural clarity. Where the model commits — irrevocably, at every layer — to saying what kind of thing it’s looking at and what kind of operation it’s performing.

The architecture is called COSAA-SST. The ontology underneath it has three types and four relations. The rest is learned.

Thought 1: Compression Is Ontologically Blind

We have a zoo of compression techniques. Pruning. Quantization. Distillation. Low-rank factorisation. Tensor decomposition. Token merging. Each removes parameters or reduces precision, and each does so without knowing what those parameters represent.

The underlying assumption: all dimensions of the representation space are equally compressible, or at minimum, compressibility can be determined by statistical properties (magnitude, variance, redundancy) alone.

This assumption is wrong. Some dimensions encode that something is an event. Others encode that something causes something else. Others encode noise. Compressing the first two degrades reasoning. Compressing the third is free. But current methods cannot distinguish between them because they have no ontological vocabulary with which to make the distinction.

The thought: What if compression knew what to preserve? What if the ontology defined the compression manifold?

Thought 2: Interpretability Is Forensic, Not Architectural

The field of mechanistic interpretability has made extraordinary progress. We now know that attention heads specialise. That induction heads exist. That superposition allows features to share dimensions. That attention patterns can implement something resembling symbolic operations.

But all of this is discovered after the fact. We train the model, then examine what it learned. We find structure, but we didn’t ask for it, and we can’t guarantee it. The interpretability is a property of the analysis, not of the architecture.

Regulatory frameworks — EU AI Act Article 13, Article 14 — don’t ask whether a model’s reasoning can be interpreted by a skilled researcher with unlimited compute. They ask whether the model’s reasoning is interpretable to the humans overseeing it. The difference is the difference between archaeology and engineering.

The thought: What if interpretability were a consequence of how the model computes, rather than a post-hoc projection of what it computed?

Thought 3: Flat Token Sequences Are the Wrong Representation

A transformer receives “The audit failure caused the remediation plan” as a flat sequence of tokens. It must discover, from context alone, that “audit failure” is an event, that “remediation plan” is a thing, that “caused” denotes a causal relationship, and that events can cause things to come into existence.

It must discover these distinctions every time, for every input, from scratch. The categories are not given. The relations are not typed. The structure is not explicit.

Human cognition doesn’t work this way. We have categories — roughly, things that happen, things that exist, and things that could be — and we have relations — roughly, causation, composition, realisation, and similarity — and we apply them automatically, instantly, pre-reflectively. We don’t rediscover that events are different from objects every time we encounter a sentence.

The thought: What if the model had categories? A very small number of them. And what if those categories did real work — constraining attention, guiding compression, structuring output?

Thought 4: These Three Problems Are One Problem

Compression is blind because it lacks ontological structure. Interpretability is forensic because computation isn’t typed. Reasoning is flat because representations lack categorical distinctions.

The common root: ontological agnosticism. The transformer refuses to commit to any representational structure beyond the uniform token embedding.

The common solution: ontological minimalism. Commit to the smallest set of distinctions that still captures the fundamental categories of reasoning. Build those distinctions into the architecture. Let everything else be learned.

Thought 5: Why Minimalism Matters

The natural instinct, when adding ontological structure to a neural architecture, is to reach for an existing heavyweight ontology. BFO (Basic Formal Ontology) has 50+ classes. Schema.org has 800+ types. Wikidata has 100,000+ relation types.

The problem is not that these ontologies are wrong. The problem is that they are too specific to be architectural. An architectural prior must be domain-general, computationally cheap, and small enough that the model can learn to use it rather than merely memorise it.

Mark Burgess’s Semantic Spacetime (SST) offers exactly this. Three types. Four relations. Everything else emerges.

Thought 6: Three Types Are Enough

TypeSymbolWhat it isWhat it does
EventETemporal, processualThings that happen
ThingTPersistent, physicalThings that exist
ConceptCAbstract, virtualThings that could be

The classification is exhaustive. Everything is an Event, a Thing, or a Concept. There is no “Unknown” category. There is no “Other.”

This is a strong commitment, and it should be. The strength of the commitment is the source of the compression. If you have 800 types, each type carries ~10 bits of information. If you have 3 types, each carries ~2 bits. The structural encoding cost drops by 5×.

But more importantly, the three-way distinction mirrors a cognitive primitive. Developmental psychology shows that infants distinguish between agents (which do things), objects (which are things), and properties (which describe things) before they acquire language. The E/T/C distinction is not an arbitrary taxonomy; it is a reflection of how minds naturally carve reality.

A question that comes up: is “democracy” an Event, a Thing, or a Concept? The answer: it’s a Concept that expresses into Events (elections, debates) and contains Concepts (representation, equality). The type system doesn’t prevent nuance; it structures it.

Thought 7: Four Relations Are Enough

RelationSymbolWhat it doesDirectionality
Leads-to→Causal/temporal sequenceDirected
Contains⊃Hierarchical groupingDirected
Expresses⇒Realisation/attributionDirected
Near~Similarity/analogySymmetric

Four relations. Each captures a fundamental mode of connection.

Leads-to is the relation of process and causation. Events cause events. Things trigger events. Events produce things. The arrow of time and the arrow of consequence.

Contains is the relation of structure and hierarchy. Things have parts. Concepts have subconcepts. Processes have subprocesses. The nesting that makes complexity manageable.

Expresses is the relation that bridges the abstract and the concrete. A contract (Thing) expresses obligations (Concepts). An election (Event) expresses democracy (Concept). A Concept realises into a Thing when an idea becomes an artefact. This is the relation that makes meaning possible.

Near is the wild card. And it is everything.

Thought 8: The Resonance-Permissive Property

Near (~) is the architectural keystone. It is symmetric and type-unrestricted: any node can be near any other node, regardless of type.

This property — resonance-permissive — is what distinguishes SST from rigid ontologies. A rigid ontology says: “Standard cannot relate to Oath. They are different types in different branches of the hierarchy.” SST says: “Of course they can relate. They’re both Concepts. They’re near each other. Now let’s explore how.”

When a compliance officer asks “How is ISO 27001 like a doctor’s oath?”, a rigid ontology has no valid path between them. In SST, both are Concepts, near is immediately valid, and the model can trace the analogy through typed operations: both express a commitment to care, both contain procedural requirements, both lead to accountability structures.

The near relation is where creativity lives. It’s the mechanism by which the model can connect a regulatory framework to a medical oath, a quantum phenomenon to a market dynamic, a childhood memory to an architectural principle. Every creative insight, at some level, is the discovery of a previously unnoticed similarity.

Formal definition: An ontology is resonance-permissive if it contains at least one relation type r such that for all type pairs (tᵢ, tⱼ) in the type system, an edge of type r between a node of type tᵢ and a node of type tⱼ is ontologically valid.

SST satisfies this through near. No connection is forbidden. But every connection must be typed. That’s not a limitation on creativity; it’s a demand for clarity.

Thought 9: The Compression Arithmetic

The numbers are striking.

  • Node type encoding: 3 types → 2 bits
  • Edge type encoding: 4 types → 2 bits
  • Total structural bits per node-edge pair: 4 bits

Compare:

  • Schema.org: 800 types → 10+ bits per node
  • Wikidata: 100,000+ relations → 17+ bits per edge
  • Open vocabulary KGs: unbounded

For a graph of 512 nodes at 10% edge density (~26,000 edges): SST structural encoding costs ~105 kilobits. A conventional KG costs ~700+ kilobits. That’s roughly 85% compression at the structural level alone, before touching the semantic embeddings.

And because the type and relation spaces are small and closed, they can be encoded as one-hot vectors and processed with simple linear operations. No large embedding tables. No hash lookups. No vocabulary management.

The ontology is the compression.

Thought 10: The Data Flow

Input Tokens
   ↓
[Embedding + Compression]     d_model=768 → d_compressed=128
   ↓
[SST Graph Construction]      Assign E/T/C types, predict →/⊃/⇒/~ edges
   ↓
[Symbolic Attention × L layers]
   ├── Tick 1: Each of H heads selects a symbol → attention flows along typed edges
   ├── Tick 2: GRU updates state → heads select new symbols based on refined context
   └── Tick K: Final refinement → symbol trace is complete
   ↓
[SST Consistency Check]       Score graph against canonical patterns
   ↓
[Dual Output]
   ├── Token logits (language modelling compatibility)
   └── Graph operation + arguments (structured reasoning)

The key architectural insight: graph structure is not a separate processing module. It is intrinsic to attention. The graph constrains where attention flows. The symbols determine how it flows. There is no GNN block followed by a Transformer block. There is one unified mechanism that is simultaneously graph-aware and attention-based.

Thought 11: Compression as Ontological Projection

The embedding stage projects from d_model = 768 to d_compressed = 128. That’s a 6× dimensionality reduction right at the input.

But this isn’t just a learned linear projection. Within the compressed space, three learned type centroids define the prototypical locations for Events, Things, and Concepts. Node type is computed as distance to these centroids. Each type has a dedicated subspace projection that extracts type-specific features.

The idea: don’t just compress. Compress into an ontologically organised space. The Event region of the embedding captures temporal dynamics. The Thing region captures persistent properties. The Concept region captures abstract relations. Compression doesn’t lose information indiscriminately — it loses information that is orthogonal to ontological structure.

This is Sebastian Jaszczur’s insight inverted: instead of learning a compressed manifold and hoping it preserves meaning, we define the meaningful manifold through the ontology and let compression operate on the residual.

Thought 12: Graph Construction as Forced Commitment

The graph constructor does three things, each involving a forced commitment:

  1. Node type classification. Every token position gets a type: E, T, or C. No hedging. Gumbel-softmax during training (soft, differentiable). Argmax during inference (hard, committed). The type clarity loss penalises entropy in the assignment — the model must be confident in its classifications.
  2. Edge existence prediction. For every pair of nodes, a scalar probability of connection. Only the top-k per node survive (k=8). The sparsity loss keeps density around 10%. The resulting graph is sparse — most nodes connect to very few others.
  3. Edge type classification. Every surviving edge gets a type: →, ⊃, ⇒, or ~. Biased by learnable type-pair priors — a 3×3×4 tensor initialised from SST’s canonical patterns. E→E prefers leads-to. T⊃T prefers contains. Everything gets a small near bias, encoding resonance-permissivity.

The output: a sparse typed adjacency tensor [batch, N, N, 4]. Four slices, one per relation type. This tensor drives everything that follows.

Thought 13: The Symbol Codebook — 48 Operations That Suffice

This is the heart.

Standard attention asks: “How much should each position attend to each other position?” It produces a continuous attention matrix via dot-product similarity.

Symbolic attention asks: “What operation should this attention head perform on this graph?” It selects a discrete symbol from a codebook, and the symbol determines the attention pattern.

The codebook has 48 entries, organised into 6 categories of 8:

Causal (0–7): CAUSE_FORWARD, CAUSE_BACKWARD, TRIGGER, PRODUCE, ENABLE, CHAIN, BLOCK, PARALLEL. These operate on → edges. They propagate, trace, chain, and interrupt causation.

Containment (8–15): DECOMPOSE, COMPOSE, ABSTRACT_UP, CONCRETE_DOWN, SCOPE_IN, SCOPE_OUT, MEMBER_OF, SUBSUMES. These operate on ⊃ edges. They take things apart, put them together, generalise, specialise.

Expression (16–23): REALIZE, MANIFEST, ATTRIBUTE, INTENTION, ROLE, QUALITY, ACTUALIZE, VIRTUALIZE. These operate on ⇒ edges. They bridge concrete and abstract — concepts becoming things, events expressing purposes.

Similarity (24–31): ANALOGIZE, CONTRAST, TRANSFER, CLUSTER, BRIDGE, DIVERGE, MATCH, ALIGN. These operate on ~ edges. They find likeness, exploit likeness, group by likeness. This is where creative reasoning lives.

Cross-Type (32–39): EVENT_TO_THING, THING_TO_EVENT, EVENT_TO_CONCEPT, CONCEPT_TO_EVENT, THING_TO_CONCEPT, CONCEPT_TO_THING, GROUND, FLOAT. These change how the model sees an entity. Reify a process. Abstract from an instance. The cognitive flexibility of re-categorisation.

Meta (40–47): ATTEND, IGNORE, REMEMBER, FORGET, COMPARE, MERGE, SPLIT, HALT. These control the reasoning process itself. Focus. Defocus. Persist. Clear. The executive functions.

Each symbol is a learnable vector. The model learns the representations of these operations through training. But the structure — which edges each symbol attends along, which type pairs each symbol expects — is given by the SST ontology. The model learns when to apply each operation. The ontology defines what each operation does.

Thought 14: How Symbol Selection Works

At each tick, each head:

  1. Encodes context. Mean-pool the current node embeddings. Project to symbol dimensionality.
  2. Scores symbols. A linear layer produces 48 logits. One per symbol.
  3. Selects one. Hard selection. Gumbel-softmax during training (differentiable). Argmax during inference (discrete).

The selected symbol then determines an attention mask:

  • The symbol’s category determines which edge types carry attention. Causal symbols → only leads-to edges. Similarity symbols → only near edges. Meta ATTEND → all edges.
  • The symbol’s type pair filters which nodes participate. CAUSE_FORWARD → only Event sources and Event targets. TRIGGER → only Thing sources and Event targets. ANALOGIZE → any types (near is unrestricted).

Attention flows only where the mask permits. The mask is not learned — it is derived from the graph structure according to the selected symbol. This is what makes the system interpretable: the symbol name tells you exactly what pattern the head is computing and why.

Thought 15: Multi-Tick Reasoning — Thinking Without Speaking

Standard chain-of-thought reasoning generates intermediate tokens. This is wasteful — the tokens cost compute, fill context windows, and are often formulaic. Multi-tick symbolic attention achieves the same purpose without generating any tokens.

Each attention block runs K=3 ticks. At each tick:

  • All heads select symbols and compute updates
  • A GRU cell integrates updates with previous state
  • Layer normalisation stabilises the result

Tick 2’s symbol selections are conditioned on Tick 1’s results. Tick 3 builds on Tick 2. The model can first identify the causal structure (Tick 1), then explore related concepts (Tick 2), then synthesise (Tick 3).

The analogy to neuroscience is apt: the thalamocortical loop doesn’t produce output after a single pass. It iterates. Percepts sharpen. Context integrates. The prefrontal cortex’s iterative refinement of working memory representations is doing something like multi-tick attention with executive control symbols.

The key difference from recurrent approaches: the number of ticks is fixed and small. There’s no vanishing gradient over long sequences. The recurrence is within each layer, not across layers. It adds depth without depth’s usual pathologies.

Thought 16: The Consistency Checker — Soft Ontological Feedback

After attention processing, the consistency checker evaluates the graph. But unlike hard ontological reasoners that reject inconsistent states, this checker produces a soft score that enters the training loss.

What it checks:

  • Canonical patterns. Is each edge’s type-pair expected? E→E via leads-to gets a high score. C→C via leads-to gets a lower (but non-zero) score. Anything via near gets a perfect score — resonance-permissive.
  • Causal cycles. Cycles in the leads-to subgraph are suspicious (but not impossible — feedback loops exist). Soft penalty.
  • Containment cycles. Cycles in the contains subgraph are more problematic (a thing can’t contain itself). Stronger soft penalty.
  • Surprises. Unusual but valid patterns are flagged — not penalised, just noted. These show up in the reasoning trace for human inspection.

The consistency score = base pattern score − 0.2 × violations − 0.05 × surprises. This enters the loss with weight λ_cons = 0.1. Light touch. Guide, don’t constrain.

Thought 17: Dual Output — Tokens for Compatibility, Graphs for Reasoning

COSAA-SST has two output heads:

  1. Token head. Standard: pool graph representation, project to vocabulary logits. This enables training on text corpora, evaluation on standard benchmarks, deployment in language tasks.
  2. Graph operation head. Novel: predict the next SST operation (one of 48 symbols) and its arguments (attention over nodes). This enables structured reasoning tasks — “what caused this?”, “what is this part of?”, “what is this similar to?” — with typed, traceable answers.

The dual output means the model can be trained on cheap, abundant text data while simultaneously developing graph reasoning capabilities that text-only models lack. The graph head doesn’t need separate supervision — the SST structure provides the inductive bias, and the consistency loss provides the gradient.

Thought 18: Glass-Box Reasoning

Every COSAA-SST prediction comes with a symbol trace. Here’s what one looks like:

INPUT: "The server failure caused data loss which triggered the incident response protocol"

GRAPH:
 [T: server] →leads_to→ [E: failure] →leads_to→ [E: data_loss]
                                                        ↓ leads_to
                                                  [E: trigger]
                                                        ↓ leads_to
                                                  [T: protocol]

TRACE:
 Layer 0, Tick 0:
   Head 0: TRIGGER        (T→E: server triggers failure)
   Head 1: CAUSE_FORWARD  (E→E: failure causes data_loss)
   Head 2: CHAIN          (E→E→E: building causal sequence)
   Head 3: ATTEND         (increasing salience on all nodes)

 Layer 0, Tick 1:
   Head 0: PRODUCE        (E→T: trigger activates protocol)
   Head 1: INTENTION      (E⇒C: inferring purpose of events)
   Head 2: ABSTRACT_UP    (→C: extracting "incident" concept)
   Head 3: REMEMBER       (persisting causal chain to next tick)

 Layer 0, Tick 2:
   Head 0: CHAIN          (completing E→E→E→T path)
   Head 1: ATTRIBUTE      (T⇒C: protocol has compliance property)
   Head 2: COMPARE        (evaluating subgraph completeness)
   Head 3: HALT           (reasoning complete signal)

CONSISTENCY: 0.97 (no violations, no surprises)

This trace is complete (every step recorded), typed (every step references SST categories), and compositional (the sequence forms a logical reasoning chain). You can read it. You can audit it. You can ask: “Why did the model think this?” and get an answer that is not a post-hoc approximation but a faithful record of the actual computation.

Thought 19: Where Creative Reasoning Happens

Watch what happens with an unexpected query:

INPUT: "How is ISO 27001 like a doctor's oath?"

Traditional ontology:
 ISO 27001 → typed as "Standard"
 Doctor's oath → typed as "Oath"
 No valid relation between Standard and Oath → query fails

SST ontology:
 ISO 27001 → typed as CONCEPT (C)
 Doctor's oath → typed as CONCEPT (C)
 Near (~) between C and C → ALWAYS valid

TRACE:
 Tick 1: ANALOGIZE   (C ~ C: these two concepts are near each other)
 Tick 2: ATTRIBUTE   (both ⇒ "commitment to protect")
 Tick 3: CONTRAST    (different domains: information vs. health)
 Tick 4: TRANSFER    (oath structure → standard structure)

The near relation opens the door. The typed operations explore what’s behind it. The model doesn’t just say “these are similar” — it traces in what way they are similar, using operations that have names and types and can be verified.

This is what “resonance-permissive” means in practice. Rigid ontologies kill analogies. SST cultivates them — but demands that they be structured.

Thought 20: Five Losses, One Objective

The training loss is a weighted sum of five terms:

LossWeightWhat it does
Language modellingλ = 1.0Learn useful representations from text
Consistencyλ = 0.1Encourage ontologically coherent graphs
Symbol diversityλ = 0.05Prevent codebook collapse
Edge sparsityλ = 0.01Keep graphs sparse (~10% density)
Type clarityλ = 0.05Encourage confident E/T/C assignments

The language modelling loss does the heavy lifting — it’s what learns representations. The other four losses shape how the model uses the SST structure. They’re guardrails, not drivers. The weights are deliberately low because the point is to guide, not to force. The ontology should emerge as the most efficient way to minimise the LM loss, not as an externally imposed constraint.

Thought 21: Temperature Annealing — From Exploration to Commitment

Both the type classifier and the symbol selector use Gumbel-softmax. The temperature τ follows a linear decay:

  • Start: τ = 1.0 (soft, exploratory — gradient flows to all types and symbols)
  • End: τ = 0.1 (hard, committed — model must choose decisively)

This mirrors how human expertise develops. A novice considers many possibilities. An expert commits quickly. The annealing schedule encodes this trajectory into the optimisation process.

Thought 22: Preventing Codebook Collapse

A known failure mode of discrete bottleneck architectures: the model finds 5–10 useful symbols and ignores the rest. The codebook collapses.

Three countermeasures:

  1. Diversity loss. Mean squared pairwise cosine similarity between symbol embeddings. High similarity → high loss. Symbols are pushed apart in embedding space.
  2. Categorical initialisation. Symbols in the same category start similar (useful structure). Symbols across categories start distant (preventing inter-category collapse). The initialisation gives the optimiser a head start.
  3. Temperature annealing. Early training uses soft selection, giving gradient signal to all symbols. Gradual hardening allows the model to specialise symbols after all have received sufficient training signal.

Thought 23: Expressiveness Is Not Sacrificed

A legitimate concern: doesn’t constraining attention to 48 discrete operations limit what the model can compute?

Three arguments that it doesn’t:

Argument from meta operators. The ATTEND symbol permits all-to-all attention with no type or edge constraints. A model that always selects ATTEND on every head at every tick reduces to standard multi-head attention. Standard attention is a special case of symbolic attention, not an excluded case.

Argument from combinatorics. With L=6 layers, K=3 ticks, H=8 heads, each step selecting from 48 symbols, the space of possible reasoning traces is 48^(6×3×8) = 48^144. That’s not a constraint on expressiveness. That’s a structured search space of incomprehensible vastness.

Argument from resonance. The near relation ensures that no node pair is ever ontologically disconnected. If the typed edges are insufficient for a particular computation, attention can always route through near edges. The constraint is not “you can’t attend there” but “you must say why you’re attending there.”

Thought 24: The Interpretability-Capability Question

Hinton’s provocation from the roundtable: “Does interpretability constrain capability?”

Possibly. But the relevant question is not whether COSAA-SST matches GPT-4 on every benchmark. The relevant question is whether the capability trade-off (if any) is worth the interpretability gain.

Consider: a model that is 5% less accurate on a language benchmark but can produce a complete, typed, auditable reasoning trace for every prediction is more useful in regulated domains (healthcare, finance, legal) than a model that is 5% more accurate but opaque. The EU AI Act doesn’t have a benchmark exception. Article 14 requires human oversight of high-risk systems regardless of their accuracy.

COSAA-SST’s bet is that the interpretability gain is not just regulatory compliance but also a capability enhancement in its own right. A model that must commit to typed operations may learn better reasoning strategies precisely because it cannot rely on opaque shortcuts. The discrete bottleneck is a regulariser that favours structured computation.

Thought 25: The Biological Analogy Is Not Accidental

The architecture mirrors biological cognition at multiple levels:

COSAA-SST componentBiological analogueShared property
Compressed embedding space with type centroidsNeocortical categorical representationsType-organised encoding
Graph-structured attention with typed edgesHippocampal relational memoryRelational encoding with typed links
Multi-tick recurrenceThalamocortical loopIterative refinement before output
Meta operators (ATTEND, IGNORE, REMEMBER, FORGET)Prefrontal executive controlResource allocation and working memory management
Dual output (tokens + graph ops)Verbal and spatial reasoning systemsMultiple output modalities

We don’t claim COSAA-SST implements these biological mechanisms. We claim the structural parallels are evidence that the architectural choices are well-motivated. Evolution has had a long time to optimise cognitive architectures. If our design converges on similar structures, that’s a signal worth attending to.

Thought 26: The Type Commitment Problem

Is “software” a Thing or a Concept? Is “inflation” an Event or a Concept? Is “the boundary between France and Germany” a Thing?

SST’s three-way classification is exhaustive by design, but borderline cases are real. The current architecture handles this through soft type assignments during training, allowing the model to hedge. At inference, the hard assignment might be wrong.

Possible mitigations: multi-type assignments (with a primary and secondary type), type uncertainty propagation (maintaining a distribution over types through computation), or domain-specific type refinement (extending E/T/C with subtypes for specific applications). We leave these to future work.

Thought 27: Does the Structure Emerge from Unsupervised Training?

COSAA-SST trains on text data with a language modelling objective. The SST structure must be discovered from this signal, augmented only by the soft consistency loss and the type clarity loss.

Will it? There’s theoretical reason to hope so — the E/T/C distinction is mirrored in linguistic categories (verbs/events, nouns/things, adjectives-and-abstractions/concepts), and the four relations have direct linguistic analogues (causal connectives, possessives, attributives, similies). The model has linguistic scaffolding to lean on.

But this is an empirical question. Phase 1 of implementation will reveal whether the inductive biases are sufficient, or whether some supervised ontological annotation is needed to bootstrap the structure.

Thought 28: The 48-Symbol Question

Why 48? Why not 24, or 96, or 256?

Honest answer: 48 is a design choice, not a derivation. The number comes from 6 categories × 8 operations per category, where both 6 and 8 are chosen to provide coverage without redundancy. But the architecture treats codebook size as a hyperparameter. The right number is whatever maximises the compression-reasoning-interpretability trade-off for a given domain and scale.

We suspect 48 is near-optimal for the general case — small enough to prevent collapse, large enough to cover the operation space, structured enough to be interpretable. But “we suspect” is not “we prove.” Empirical validation is required.

Thought 29: What Does This Architecture Do Worse?

Hinton’s other provocation: “Be explicit about what you’re giving up.”

Speed. The graph construction step adds overhead. Computing typed adjacency tensors is not free. For tasks where flat token processing suffices (simple text generation, formatting, translation), COSAA-SST is likely slower than a standard transformer of comparable size.

Simplicity. The architecture has more moving parts: embedding compression, type classification, edge prediction, edge typing, symbol selection, multi-tick recurrence, consistency checking. More parts means more things that can break, more hyperparameters to tune, more complex debugging.

Benchmark compatibility. Standard benchmarks evaluate flat text-to-text performance. COSAA-SST’s distinctive capabilities — structured reasoning, typed traces, ontological coherence — are not measured by MMLU or HellaSwag. The model may underperform on benchmarks that don’t test what it’s good at.

These are real costs. We believe they are outweighed by the benefits in domains that require interpretability, structured reasoning, and trustworthy AI. But we acknowledge that COSAA-SST is not a universal replacement for standard transformers. It is a specialised architecture for tasks where knowing why matters as much as knowing what.

Thought 32: The Principle

The principle behind COSAA-SST is simple enough to state in one sentence: A model that commits to the structure of reasoning can compress better, explain itself, and reason more coherently than a model that treats structure as optional.

Three types. Four relations. Forty-eight operations. Everything else is learned.

This is a bet on minimalism. A bet that the right architectural prior is not “give the model maximum flexibility” but “give the model the minimum structure that reasoning requires.” That ontological commitment is not a constraint on capability but a foundation for it.

Thought 33: The Name

Everything Is an Event, Thing, or Concept. We chose this title because it captures the radical commitment of the approach. Not “some things might usefully be classified as events, things, or concepts.” Not “in certain domains, a tripartite typology can be helpful.” Everything. Always. No exceptions.

That’s either a profound insight about the structure of reasoning, or a naïve oversimplification that will shatter on contact with real data.

We suspect the former. We’ll find out.

Thought 34: What If It Works?

If the architectural hypothesis proves correct — if minimal ontological commitment improves compression, interpretability, and reasoning simultaneously — the implications extend beyond transformer design.

It would suggest that the path to better AI is not through larger, more flexible models but through models that commit to the right structural priors. That the three-types-four-relations structure of SST captures something genuine about the organisation of knowledge. That interpretability by design is not merely possible but efficient — that glass-box systems can be as capable as black boxes because the glass box is the compression.

And it would offer something the field urgently needs: AI systems whose reasoning can be read, audited, questioned, and trusted. Not because we’ve learned to approximate their internals after the fact, but because they were built, from the ground up, to show their work.

The rest is implementation.

References

  • [1] Tang, Y., Wang, Y., et al. (2024). A Survey on Transformer Compression. arXiv:2402.05964.
  • [2] Jaszczur, S. & Stefaniak, M. (2025). Trainable Projections for Efficient Transformer Compression. ICLR 2026 submission.
  • [3] Zhang, X. et al. (2024). VTrans: Accelerating Transformer Compression with Variational Information Bottleneck. arXiv:2406.05276.
  • [4] Chen, J. et al. (2024). DSFormer: Dense-Sparse Weight Factorization. arXiv:2312.13211.
  • [5] Ge, T. et al. (2024). InfoTok: Adaptive Discrete Video Tokenizer. arXiv:2512.16975.
  • [6] Curtis, A. et al. (2025). TEMPEST: Transformers from Compressed Representations. arXiv:2510.23665.
  • [7] Griffiths, T.L. & Cohen, J.D. et al. (2025). Interpreting Roles of Attention Heads. arXiv:2505.13737.
  • [8] Li, Q. et al. (2025). SIGNNet: Similarity-Aware GNN with Transformer. arXiv:2504.02615.
  • [9] Zhang, L. et al. (2025). TSIformer: Multi-Scale Dilation Transformer. IEEE.
  • [10] Zhuang, B. et al. (2024). Fibottention: Diverse Attention Across Heads. arXiv:2406.19391.
  • [11] Wu, Y. et al. (2025). Graph-Aware Isomorphic Attention. arXiv:2501.02393.
  • [12] Guo, W. & Wang, X. (2025). OL-KGC: Ontology-Enhanced KG Completion. arXiv:2507.20643.
  • [13] Jain, N. & Naumann, F. (2024). ReasonKGE: Ontological Reasoning for KG Embeddings. HPI.
  • [14] Adeel, A. et al. (2025). Co4: Brain-Inspired Cognitive Computation Architecture.
  • [15] Burgess, M. (2015). Semantic Spacetime. Promise Theory and the Structure of Knowledge.
  • [16] Anthropic Research. (2025). Progress on Attention. Transformer Circuits Thread.
  • [17] Fan, J. et al. (2024). MemoryFormer. arXiv:2411.12992.
  • [18] Chen, Y. et al. (2023). Modular Transformers. arXiv:2306.02379.
  • [19] Li, Z. et al. (2023). Input Compression with Positional Consistency. arXiv:2312.12385.
  • [20] Wang, H. et al. (2025). Adaptive Semantic Retrieval Framework. Nature Scientific Reports.
  • [21] Ying, C. et al. (2024). Transformer Graph Reasoning. arXiv:2405.18512.
  • [22] Zhou, R. & Angizi, S. (2025). DeepCompress-ViT. CVPR 2025.
  • [23] DeepSeek Team. (2025). Multi-head Latent Attention. DeepSeek-R1 Technical Report.
  • [24] Bengio, Y., Hinton, G., & LeCun, Y. (2024). Deep Learning Foundations. Joint Keynote.

Originally published on LinkedIn.