Noumena Research
Papers
RDEP: The NVLink Domain as a Sparse Training Computer
A domain-native MoE architecture that keeps the active dense program replica-local, owns sparse capacity once across the NVLink domain, and lets the realized route graph drive exact execution.
The MoE Atlas: A Candidate Shared Boundary for Pretraining Readiness and Geometric Forgetting
A candidate operational boundary for asking when MoE pretraining is ready for post-training—and when adaptation has forgotten what the pretrained model supplied.
Same Confidence, Opposite Truth: An Atlas of Hidden Answer Geometry
Why apparent confidence can remain nearly fixed while agreement rotates from truth into error, and how hidden answer geometry exposes the missing coordinates.
Research notes
Same Confidence, Opposite Truth
post-training barely moved the shape of agreement while moving it from gold answers into wrong ones
Let the Speedrun Search Itself
Eval-gated config-only autoresearch on the canonical super fp8 lane
Reproducing Canon, mHC, and Engram
Three architectural bets, one physics harness, and the question of which deserves a larger run
The Rack Is the Computer
RDEP makes the NVLink domain—not the rank mesh—the unit of sparse training
Do MoE Experts Need Different Learning Rates?
Why Moonlet's old 15x expert-LR rule overshoots in bf16 AdamW
One Boundary, Two Crossings
the MoE Atlas as a candidate contract for pretraining readiness and geometric forgetting
I Built Super-4096 to Fail
I gave the router 4,096 experts. One batch came back looking like fourteen.
The NVFP4 Fix Was Fighting Itself
I split the old two-knob rescue. The embedding gain helped; the logits gain gave it back.
The Experiment Had Two Model Sizes
A literal #420 transfer made the denominator part of the result
Fewer Tokens, More Time
What our first dense–MoE speedrun measured
The Loss Curve Was Fine
What a 4,096-expert router collapse taught me to measure
The Stale Batch
A one-batch resume bug and the state a trustworthy run must preserve
Why Training MoEs is So Hard
Three failure modes that make frontier MoE training qualitatively different