ferrite-lithic
Hardware you write
in Rust, as real gates.
An embedded hardware DSL, a cycle simulator and a Verilog emitter. Hardcaml's model, reimplemented natively — not a wrapper, not a binding, not a transpiler. A Rust library you call from Rust, and the only way to know it works is a corpus of twenty-one real algorithms checked three ways.
- tests
- 810
- verified designs
- 21
- crates
- 10
- Verilator
- 100%
753 unit and integration, 57 doctests
in seven tiers, each checked three ways
nine tools and one corpus
of designs cosimulated cycle by cycle
A library, not a language
Most hardware is described in a language that is not yours. This one is a Rust library: you call a builder, it hands you back a Design, and that design is a graph of nodes and edges with widths — the same thing a simulator needs and the same thing an emitter needs.
One graph, two backends, and an equivalence check between them on every single corpus design. That check is the point. A simulator and an emitter that both produce plausible answers will agree on every test anybody bothers to write — until one of them is wrong in a corner.
use ferrite_lithic::Design; let design = Design::new(); let clk = design.clock(); let rst = design.reset(); // a bit of state and one combinational path let state = design.reg(&next, &clk, &rst, &hold)?; let out = design.eq(&state, &target)?; // the two backends, from the same graph ferrite_lithic_sim::Sim::new(&design); // cycle-accurate design.emit_verilog("top"); // structural
One graph, two backends, one arbiter
The pipeline below is animated on a real clock. Data flows left to right; the equivalence check at the end compares the two right-hand branches cycle by cycle over a real stimulus, and reports the first cycle they disagree on.
Plan holds the stimulus and the column set that make the two comparable at all — without it the comparison is not between two machines but between two different experiments.Why the equivalence check fails the build when it is skipped
Verilator and Icarus are looked for rather than required, so cargo test works without them — and both print what they skipped. The CI job installs both and then fails if anything skipped, because a green run in which every equivalence test took its skip path is green and meaningless.
Read one cycle
Four things are true about a clock edge, and they are easier to see than to read. Each paragraph below moves the diagram.
- the edge
A cycle starts at a rising clock edge. Inside the cycle, nothing changes state: every register holds what it captured at the last edge, and the only thing moving is combinational logic settling.
- the budget
The clock period is the entire budget. Everything between this edge and the next has to settle inside it — which is why a design with a long combinational path is a design that will not meet timing, and why `pipeline` exists.
- before the edge
`d` has to be stable for a while before the edge, not just correct at it. That lead time is setup time, and the bracket in the hero diagram is measuring it. Get it wrong and the register captures whatever was in flight.
- after the edge
At the edge, every register simultaneously takes the value its input had. So `q` becomes the previous `d`. A register is not a wire that is slow — it is a value from one cycle ago, which is why the diagram always trails the input by one.
d is the input bus and q is the register, so q always shows what d held one cycle ago.Where the design came from
The architecture is a port of Hardcaml, Jane Street's OCaml hardware DSL, and that is worth understanding before reading any of the nine crates — several choices that look arbitrary are Hardcaml's choices that survived the translation.
Hardcaml designs are OCaml programs that build a data structure describing a circuit. Rust does not check widths at compile time here, and it cannot make combinational loops impossible; what it gives instead is memory safety without a garbage collector, handles that cannot outlive their graph, and Result on every builder operation. The honest accounting is four pages long and worth the read.
Ten crates, in dependency order
Each one is small on purpose. The number is its test count, and the tier is the phase that produced it.
- ferrite-lithic-bits96
Fixed-width bitvectors over u64 words, with runtime widths.
- ferrite-lithic-ir116
The graph: nodes, widths, and nothing else. No behaviour, no types.
- ferrite-lithic58
The front end. Design is an arena of nodes; signals are ids.
- ferrite-lithic-sim52
A cycle simulator. Sequential on the clock edge, combinational within it.
- ferrite-lithic-rtl52
The Verilog emitter: one structural always block for the whole design.
- ferrite-lithic-derive39
Port lists by derive, so a port is a struct field and not a string.
- ferrite-lithic-wave42
Cycle-indexed values asserted on directly, plus a VCD rendering.
- ferrite-lithic-cosim33
Verilator equivalence checking against the simulator, cycle by cycle.
- ferrite-lithic-tb17
Coroutine step testbenches, where every .await is one clock edge.
- ferrite-lithic-corpus305
Twenty-one verified designs. The tier that finds the tools' bugs.
Why a corpus, and what it cost
Every other crate is a tool, and a tool gets tested with inputs chosen by whoever wrote it — so its tests drift toward what the tool was built to do. A corpus entry is a real algorithm written the way hardware would want it, checked three ways: against the original crate as the golden model, through the step testbench, and against Verilator.
It paid for itself immediately. crc3 — a three-bit LFSR, one byte per cycle, the smallest design with a register and a bit stream — found four defects that 471 tests of the tools had not, including Design::sll and Design::srl being exactly swapped for the whole of A1 and A2. Every crate agreed with every other crate about it, because they were all wrong the same way.
Seven tiers, twenty-one designs
- LFSR & checksums
crc3crc32lfsrghash4 - Block ciphers
sha256chacha20aes3 - Codecs
hexbase642 - Automata
memchraho_corasickdfa3 - Data infrastructure
bitpackrleroaringhammingsortingsketch6 - Packet
packet1 - Entropy coding
huffmanfsedeflate3
The honest benchmark: the CPU wins memchr, by 4.2×
The automaton design is a real design, and it still loses. Measured per byte against the memchr crate at an assumed 1 GHz: the design retires one byte per clock edge, the crate retires 32 bytes per AVX2 instruction. That is the entire gap, and the argument for hardware is not that it wins here.
Across the corpus, two of five designs beat their reference implementation: crc3 by 2.2× and hex by 1.4×. Both win by doing less — a three-bit state and a lookup table have no generality to pay for. The ones that lose lose because the CPU has more parallelism available than a 31-node circuit uses. Why hardware loses here, and when it wins →
It is in the corpus anyway, because it is the baseline the other two automata are measured against, and because a design that claimed to beat AVX2 at single-byte search would have been dishonest. The surviving argument for the tier is the multi-pattern search, where no SIMD form exists and a 19-state automaton shares one ROM across lanes.
Four bugs, one shape
Every real defect this corpus has found looks the same: a constant that looks right, builds a graph of exactly the right width, and means the opposite of what it says. Here are two you can watch happen.
FNV-1a's offset basis
The offset basis is a published constant. It was mistyped by one hex digit — and the test that should have caught it used the same wrong constant, so it agreed with itself. Only the vectors checked against an external source disagreed.
DEFLATE peels from the bottom
RFC 1951 packs data into bytes in order of increasing bit number. A design that peels bits off the top produces a perfectly plausible bit stream that no DEFLATE decoder has ever accepted — so this is the single thing worth watching.
Nine bits in, one symbol out
A Huffman code is written most significant bit first into a stream that arrives least significant bit first. The decoder peeks nine bits, reverses them into a table index, and gets back a symbol and the number of bits that symbol cost.
The honest gap
There is no initialised-memory node in the IR. Design::mem is zero-filled in both backends, so every lookup table in the corpus is a case or a multiplexer tree. That is close to the truth for a nibble LUT and a lie for a 256-entry alphabet map.
Design::rom(address, &table, width) builds exactly what a human would write with case: one arm per entry, emitting always @* case. Watch a 256-entry table grow. A synthesis tool would infer block RAM from the same source in seconds — which is the point being recorded rather than hidden.What is not here
- No initialised memory. The highest-value addition left to the toolchain.
Design::romdocuments the cost instead of hiding it. - No dynamic Huffman tree.
Tableis built from code lengths at elaboration time. Parsing a dynamic block header — nineteen code lengths, a seven-bit code-length code, the 16/17/18 run codes — is a different module. - No LZ77 window. The DEFLATE entry is the bit layer and nothing above it: no block structure, no error detection, no back-references.
- Huffman is not one symbol per cycle. A decode spends
code_lenbits and the host supplies one bit per cycle, so the rate is one symbol per code length — seven, for DEFLATE's shortest fixed code. FSE can, because an FSE transition is allowed to cost zero bits. That difference is the whole distinction between the two decoders. - BLAKE3 and the FFT are placeholders. Files with no implementation behind them, not even declared as modules.