Published: 2026-10-01

Table of contents
- What is PSSA?
- Why Rust, and why that is not the point
- How PSSA works: the architecture
- The benchmark numbers I ran myself
- Inference speed: 12x is not a typo
- How to clone, build, and run PSSA in under 2 minutes
- What PSSA still lacks
- FAQ
Every six months or so a project appears on Hacker News that makes you stop scrolling. PSSA hit the front page this week with a headline that sounds almost too good: a non-transformer language model, written entirely from scratch in Rust, with no PyTorch, no TensorFlow, no ML framework of any kind, that reportedly learns 0.45 nats lower cross-entropy than a matched-parameter transformer and generates text 12 times faster on the same CPU.
I cloned the repo, compiled it on an Apple M5, ran the benchmark and the evaluate command against the bundled corpus, and spent several hours reading the source. What follows is what I actually observed: where the numbers hold, where they caveat, and what a working developer should take from PSSA right now.
What is PSSA?
The name stands for Plastic State-Space Architecture. It is a small language model that combines three ideas the transformer does not use: a selective recurrent state-space layer that processes one token at a time, a bounded episodic memory bank stored in hyperbolic space, and online per-token weight updates that fold back into the base transition matrix during inference.
The Rust crate is named oxide_ai_pssa (v0.4.0 at time of writing). The binary is called oxide. The only runtime dependencies are ureq for corpus downloads, tokenizers for byte-level BPE, wgpu for WebGPU compute, and rayon for multi-threaded CPU fallback. No Python interpreter, no CUDA toolkit required to build — CUDA support is an optional feature flag.
The project is an independent research implementation by a developer under the handle Sparticle62ops, published on GitHub with an open license. The results have not been independently reproduced. Those two facts matter and I will return to them.
Why Rust, and why that is not the point
The README is explicit on this: PSSA did not use Rust for speed points. The author needed per-token weight updates, a memory bank written during the forward pass, and a scalar reference path every batched kernel could be differentiated against. Expressing that inside an autograd framework meant fighting the framework at every step. Writing the linear algebra directly made the plastic parts straightforward and the gradients checkable against a reference to around 3e-8 precision.
The architecture is the claim. The implementation language is a detail. A Python port is stated as welcome.
That framing is refreshingly honest. What it is demonstrating is whether a particular combination of selective SSM recurrence, hyperbolic memory reads, and online plastic updates can outperform a transformer at small scale — not whether Rust beats Python.
Rust’s ownership model does offer a genuine advantage here: the custom linear algebra passes Clippy checks and gradient twin-checks on every commit without a GIL or garbage collector in the path. But those are implementation-quality benefits, not architectural ones.
How PSSA works: the architecture
Each token passes through one PSSA layer with four components.
Selective recurrent SSM. This is the same family as S4 and Mamba. The recurrence is selective because the per-channel step size (delta), the input map (B), and the output map (C) are all derived from the current token, not fixed. The transition matrix A is kept negative by construction (via softplus) so the recurrence stays numerically stable. The default is 256 channels (d_m) and 16 states per channel (d_s). Each channel starts with 16 rates on log-spaced timescales from 1.5 to 200 tokens — inspired by the HiPPO initialization — so a single channel simultaneously holds recent context and longer-range structure from the start of training.
Hyperbolic episodic memory read. This is the novel part of PSSA. After the SSM output, a query is formed from both the current token and the current recurrent state, so retrieval is conditioned on where the sequence has got to, not just the token in hand. That query is projected into the Poincare ball (hyperbolic space) via a diffeomorphic map. The read is bounded at 4 nearest slots, weighted by a softmax over hyperbolic distance. Hyperbolic geometry keeps slots for general context and slots for specific episodes separable without widening the read, because distances grow toward the boundary of the ball. Cost per token is fixed at 4 lookups regardless of how much the bank holds.
Learned gate. A per-channel sigmoid gate decides how much of the memory read enters the residual stream. The read does not land unconditionally.
Plastic weight updates. This is the write path. A slot is inserted when the incoming state is novel against what the bank holds. Each slot has a refractory counter that rate-limits overwrites, protecting slots where repeated evidence has accumulated. Periodically, fast weights in the external memory are folded back into the base transition matrix by closed-form ridge regression:
A_base <- A_base + (H^T H + lambda I)^-1 H^T dH
This prevents the bank from becoming the only place long-range structure lives.
The memory module lives in src/memory.rs (291 lines). The full PSSA layer is in src/pssa.rs (1,587 lines). The linalg primitives in src/linalg.rs are 666 lines of allocation-conscious Rust with no external BLAS dependency on CPU.
| Module | Lines | Purpose |
|---|---|---|
| src/pssa.rs | 1,587 | Forward pass, plasticity, gradient checks |
| src/linalg.rs | 666 | Vectors, matrices, AdamW, RNG |
| src/memory.rs | 291 | Hyperbolic bank, read/write, refractory gate |
| src/adapter.rs | ~200 | Low-rank modular adapter projections |
| src/backend.rs | ~300 | CPU/WebGPU/CUDA dispatch |
The benchmark numbers I ran myself
I compiled PSSA on an Apple M5 MacBook. The build completed in 1 minute 25 seconds with zero errors (one style warning in numeric kernels, left intentional).
Running oxide benchmark on the built-in smoke corpus:
device cpu, 10 threads
vocabulary 8
width latent 16 / state 4 / depth 1
memory 8 slots, key width 8
schedule 120 epochs, 360 updates
epoch 1/120 loss=1.917764
epoch 4/120 loss=0.024783
epoch 7/120 loss=0.000006
epoch 10/120 loss=0.000000
wall time 15.7s
tokens 2,520
throughput 161 tokens/second
benchmark_pass ce=0.000000 completion=observes the bright moon.
The tiny benchmark corpus converges to zero loss in 7 epochs. Expected: it is a 24-token smoke test. The meaningful numbers are from the author’s WikiText-103 head-to-head, which requires a GPU session of several hours to run in full. The source is auditable and the experimental setup is documented.
From the PSSA README, head-to-head on 12.7M tokens of cleaned WikiText-103, same parameter count, same optimizer schedule, same seed:
| Metric | PSSA | Transformer |
|---|---|---|
| Training cross-entropy | 3.98 | 4.43 |
| Held-out cross-entropy (198,939 tokens) | 3.997 | 4.429 |
| Held-out perplexity | 54.4 | 83.8 |
| Next-token accuracy | 24.1% | 18.0% |
| Training throughput (CPU) | 1,716 tok/s | 415 tok/s |
The held-out gap (0.43 nats) matches the training gap. The two curves never cross across all 64 PSSA checkpoints and 43 transformer checkpoints logged during the run. PSSA is not memorizing harder — it is generalizing better, or at least that is what the numbers show at this scale.
Inference speed: 12x is not a typo
Generating 200 new tokens on the same 2-vCPU machine, same prompt, same sampler (from the PSSA README, measured results):
| PSSA | Transformer | |
|---|---|---|
| 200 tokens | 226 ms | 2,735 ms |
| Relative speed | 12x faster | baseline |
The reason is structural. A transformer re-reads its entire context window at every step, so cost per token grows with sequence length. PSSA carries a fixed-size state. Adding more tokens does not increase the per-token cost. This is the same linear-vs-quadratic argument that makes Mamba interesting, and PSSA inherits it from the SSM backbone.
On training throughput, the 4.13x advantage (1,716 vs 415 tokens per second) is less dramatic than the 12x inference gap because the transformer’s quadratic cost matters most at inference time.
How to clone, build, and run PSSA in under 2 minutes
You need Rust installed (stable toolchain, 1.75 or later). On macOS with Homebrew:
brew install rust
Then:
git clone https://github.com/Sparticle62ops/pssa
cd pssa
cargo build --release
./target/release/oxide_ai_pssa
The home screen lists commands and shows any checkpoints and corpora in the working directory. The repo ships with a bundled corpus (data/downloaded.txt, 3.2 MB) and two checkpoint files.
Run the built-in smoke test to verify your build:
./target/release/oxide_ai_pssa benchmark
You should see loss converge to 0.000000 within a few seconds and a benchmark_pass line with a sample completion.
Train on the bundled corpus for 200,000 tokens (a few minutes on CPU):
./target/release/oxide_ai_pssa train data/downloaded.txt \
-o data/my_model.pssa --max-tokens 200000 -e 1
Evaluate a checkpoint:
./target/release/oxide_ai_pssa evaluate data/heldout.txt \
--model data/my_model.pssa
The output is one JSON line: {"cross_entropy": ..., "perplexity": ..., "next_token_accuracy": ...}.
PSSA detects WebGPU acceleration automatically when an adapter is available. On my Apple M5 it picked up Metal immediately: gpu adapter: Apple M5 (backend=Metal, type=IntegratedGpu). CUDA is an optional compile feature: cargo build --release --features cuda.
The gradient check (cargo run --release --example twin_check) compares batched and scalar implementations and verifies agreement to approximately 3e-8 on every commit. That rigour is unusual in a solo research repo.
What PSSA still lacks
The author is clear about scope. Open gaps after reading the source:
Scale has not been tested. All results are at approximately 3M parameters on a small WikiText-103 slice, trained on a free hosted notebook GPU across chained 200,000-token sessions. Whether the 0.45-nat gap holds at 10x or 100x these parameters on a diverse web-text corpus is unknown.
No independent reproduction. The numbers come from one run by one developer. The experimental setup is documented and the code is public, but external reproduction has not happened yet. That is not a criticism — it is a statement of where this project sits right now.
No standard reasoning benchmarks. PSSA reports cross-entropy and next-token accuracy on WikiText-103. It does not report scores on ARC, HellaSwag, or MMLU. The step from “lower perplexity on WikiText” to “better reasoning” is not yet demonstrated.
No Python or ONNX export. The .pssa checkpoint format is not an interchange format. Loading a checkpoint requires building the Rust binary. A Python port is stated as welcome but not yet available.
These are not fatal objections. This is an early research implementation testing a set of architectural claims at small scale. On that framing, it succeeds. The claims that remain to be tested are the interesting ones.
What this means for developers
The transformer is not going away. But a project like this matters because it demonstrates three things in concrete, runnable code.
You can build a complete language model — forward pass, backward pass, custom autograd, tokenizer, checkpoint format, CLI, WebGPU backend — without a framework, in under 10,000 lines of a systems language. That is an architectural education worth having.
The post-transformer design space is genuinely open. S4 and Mamba established that SSMs can be competitive with attention. PSSA adds hyperbolic memory and online plasticity to the mix and gets results that warrant further investigation. Whether any of these ideas scale to the sizes that matter is still the open question.
The project also asks a direct question: if you have compute to grant, can this architecture beat a transformer at 10x parameters? That is a question the community can answer if someone runs the experiment.
I will be watching the repo. The gradient checks are rigorous, the code is readable, and the numbers are specific enough to verify or falsify. That puts this project in a different tier from most “LLM from scratch” projects on GitHub.
Related reading: LangGraph Tutorial: 5 Proven Steps to Fix a Fragile Agent, How to jailbreak PS5: 5 proven, painful facts about Relapse, System Design Interview: 7 Proven Fixes for a Painful Round.
Related reading: Cursor vs Claude Code: 7 Proven Tests Reveal Painful Gaps.
FAQ
What does PSSA stand for?
PSSA stands for Plastic State-Space Architecture. The plasticity refers to the online per-token weight updates that fold back into the base transition matrix during the forward pass, allowing the model to consolidate fast episodic memory into its core weights.
Is PSSA a transformer?
No. PSSA processes tokens with a selective recurrent state-space layer, not with self-attention. It does not compute pairwise token scores. Its per-token cost is constant in sequence length, unlike the quadratic cost of a standard transformer.
Is PSSA ready for production use?
Not yet. PSSA is an early research implementation tested only at very small scale on a single corpus. The checkpoint format is Rust-only, there are no standard benchmark scores, and the results have not been independently reproduced. Monitor the repo; if the scale experiments happen, the picture could change significantly.
How does PSSA compare to Mamba?
Both use selective state-space recurrences. PSSA adds two things Mamba does not have: a bounded read from an episodic memory bank in hyperbolic space, and online plastic weight updates that consolidate back into the transition matrix. The SSM backbone is described by the author as “standard selective-SSM machinery” — the novel claims are specifically the hyperbolic read and the plasticity mechanism.
GPU requirements for PSSA
PSSA runs on CPU with no GPU required, falls back gracefully when CUDA is absent, and on macOS uses WebGPU via Metal automatically. The benchmark smoke test ran in 15.7 seconds on an Apple M5 with no discrete GPU.
What corpus was used for the head-to-head comparison?
WikiText-103, cleaned via the built-in clean-wikitext command. The run used 12.7 million tokens for training and a separate 198,939-token slice for held-out evaluation. Both models used the same tokenizer, optimizer, learning-rate schedule, and random seed.
Where to find the PSSA source code
The repository is at github.com/Sparticle62ops/pssa. The full PSSA layer is in src/pssa.rs (1,587 lines). The memory bank is in src/memory.rs. The project also has a discussion thread on Hacker News.









