ferrite-lithic811 tests · 21 designs

ferrite-lithic

Hardware you write
in Rust, as real gates.

One graph, two backends, one clock. The simulator and the emitted Verilog are compared at exactly this edge, on every corpus design, cycle by cycle.

An embedded hardware DSL, a cycle simulator and a Verilog emitter. Hardcaml's model, reimplemented natively — not a wrapper, not a binding, not a transpiler. A Rust library you call from Rust, and the only way to know it works is a corpus of twenty-one real algorithms checked three ways.

tests
810

753 unit and integration, 57 doctests

verified designs
21

in seven tiers, each checked three ways

crates
10

nine tools and one corpus

Verilator
100%

of designs cosimulated cycle by cycle

A library, not a language

Most hardware is described in a language that is not yours. This one is a Rust library: you call a builder, it hands you back a Design, and that design is a graph of nodes and edges with widths — the same thing a simulator needs and the same thing an emitter needs.

One graph, two backends, and an equivalence check between them on every single corpus design. That check is the point. A simulator and an emitter that both produce plausible answers will agree on every test anybody bothers to write — until one of them is wrong in a corner.

the whole shape of a design
use ferrite_lithic::Design;

let design = Design::new();
let clk  = design.clock();
let rst  = design.reset();

// a bit of state and one combinational path
let state = design.reg(&next, &clk, &rst, &hold)?;
let out   = design.eq(&state, &target)?;

// the two backends, from the same graph
ferrite_lithic_sim::Sim::new(&design);   // cycle-accurate
design.emit_verilog("top");              // structural

One graph, two backends, one arbiter

The pipeline below is animated on a real clock. Data flows left to right; the equivalence check at the end compares the two right-hand branches cycle by cycle over a real stimulus, and reports the first cycle they disagree on.

The builder produces one graph. Both backends consume it, and Plan holds the stimulus and the column set that make the two comparable at all — without it the comparison is not between two machines but between two different experiments.

Why the equivalence check fails the build when it is skipped

Verilator and Icarus are looked for rather than required, so cargo test works without them — and both print what they skipped. The CI job installs both and then fails if anything skipped, because a green run in which every equivalence test took its skip path is green and meaningless.

Read one cycle

Four things are true about a clock edge, and they are easier to see than to read. Each paragraph below moves the diagram.

  1. the edge

    A cycle starts at a rising clock edge. Inside the cycle, nothing changes state: every register holds what it captured at the last edge, and the only thing moving is combinational logic settling.

  2. the budget

    The clock period is the entire budget. Everything between this edge and the next has to settle inside it — which is why a design with a long combinational path is a design that will not meet timing, and why `pipeline` exists.

  3. before the edge

    `d` has to be stable for a while before the edge, not just correct at it. That lead time is setup time, and the bracket in the hero diagram is measuring it. Get it wrong and the register captures whatever was in flight.

  4. after the edge

    At the edge, every register simultaneously takes the value its input had. So `q` becomes the previous `d`. A register is not a wire that is slow — it is a value from one cycle ago, which is why the diagram always trails the input by one.

Cycle 0 of 8. d is the input bus and q is the register, so q always shows what d held one cycle ago.

Where the design came from

The architecture is a port of Hardcaml, Jane Street's OCaml hardware DSL, and that is worth understanding before reading any of the nine crates — several choices that look arbitrary are Hardcaml's choices that survived the translation.

Hardcaml designs are OCaml programs that build a data structure describing a circuit. Rust does not check widths at compile time here, and it cannot make combinational loops impossible; what it gives instead is memory safety without a garbage collector, handles that cannot outlive their graph, and Result on every builder operation. The honest accounting is four pages long and worth the read.

Ten crates, in dependency order

Each one is small on purpose. The number is its test count, and the tier is the phase that produced it.

  • ferrite-lithic-bits96

    Fixed-width bitvectors over u64 words, with runtime widths.

    phase A1walkthrough
  • ferrite-lithic-ir116

    The graph: nodes, widths, and nothing else. No behaviour, no types.

    phase A1walkthrough
  • ferrite-lithic58

    The front end. Design is an arena of nodes; signals are ids.

    phase A2walkthrough
  • ferrite-lithic-sim52

    A cycle simulator. Sequential on the clock edge, combinational within it.

    phase A2walkthrough
  • ferrite-lithic-rtl52

    The Verilog emitter: one structural always block for the whole design.

    phase A2walkthrough
  • ferrite-lithic-derive39

    Port lists by derive, so a port is a struct field and not a string.

    phase A3walkthrough
  • ferrite-lithic-wave42

    Cycle-indexed values asserted on directly, plus a VCD rendering.

    phase A3walkthrough
  • ferrite-lithic-cosim33

    Verilator equivalence checking against the simulator, cycle by cycle.

    phase A3walkthrough
  • ferrite-lithic-tb17

    Coroutine step testbenches, where every .await is one clock edge.

    phase A3walkthrough
  • ferrite-lithic-corpus305

    Twenty-one verified designs. The tier that finds the tools' bugs.

Why a corpus, and what it cost

Every other crate is a tool, and a tool gets tested with inputs chosen by whoever wrote it — so its tests drift toward what the tool was built to do. A corpus entry is a real algorithm written the way hardware would want it, checked three ways: against the original crate as the golden model, through the step testbench, and against Verilator.

It paid for itself immediately. crc3 — a three-bit LFSR, one byte per cycle, the smallest design with a register and a bit stream — found four defects that 471 tests of the tools had not, including Design::sll and Design::srl being exactly swapped for the whole of A1 and A2. Every crate agreed with every other crate about it, because they were all wrong the same way.

Seven tiers, twenty-one designs

  • LFSR & checksumscrc3crc32lfsrghash4
  • Block cipherssha256chacha20aes3
  • Codecshexbase642
  • Automatamemchraho_corasickdfa3
  • Data infrastructurebitpackrleroaringhammingsortingsketch6
  • Packetpacket1
  • Entropy codinghuffmanfsedeflate3

The honest benchmark: the CPU wins memchr, by 4.2×

The automaton design is a real design, and it still loses. Measured per byte against the memchr crate at an assumed 1 GHz: the design retires one byte per clock edge, the crate retires 32 bytes per AVX2 instruction. That is the entire gap, and the argument for hardware is not that it wins here.

crc3 design1.00 cyc/byte
crc crate2.24 ns/byte
memchr design1.01 cyc/byte
memchr crate0.24 ns/byte

One design that beats its reference implementation and one that does not, in the units a ratio needs: clock edges against wall clock, read at a stated 1 GHz. The bottom row is the software being beaten — a general-purpose CRC library paying for a three-bit result.

Across the corpus, two of five designs beat their reference implementation: crc3 by 2.2× and hex by 1.4×. Both win by doing less — a three-bit state and a lookup table have no generality to pay for. The ones that lose lose because the CPU has more parallelism available than a 31-node circuit uses. Why hardware loses here, and when it wins →

It is in the corpus anyway, because it is the baseline the other two automata are measured against, and because a design that claimed to beat AVX2 at single-byte search would have been dishonest. The surviving argument for the tier is the multi-pattern search, where no SIMD form exists and a 19-state automaton shares one ROM across lanes.

Four bugs, one shape

Every real defect this corpus has found looks the same: a constant that looks right, builds a graph of exactly the right width, and means the opposite of what it says. Here are two you can watch happen.

one digit

FNV-1a's offset basis

The offset basis is a published constant. It was mistyped by one hex digit — and the test that should have caught it used the same wrong constant, so it agreed with itself. Only the vectors checked against an external source disagreed.

A constant copied into both the code and its own test is not checked twice. That is the whole lesson, and it cost one digit to learn.
bit order

DEFLATE peels from the bottom

RFC 1951 packs data into bytes in order of increasing bit number. A design that peels bits off the top produces a perfectly plausible bit stream that no DEFLATE decoder has ever accepted — so this is the single thing worth watching.

Bit 0 of each byte goes first, then bit 1, up to bit 7, then the next byte's bit 0. Watch the marker: nine cycles per byte, because this design will not load a byte over bits it has not emitted yet.
a table lookup

Nine bits in, one symbol out

A Huffman code is written most significant bit first into a stream that arrives least significant bit first. The decoder peeks nine bits, reverses them into a table index, and gets back a symbol and the number of bits that symbol cost.

The reversal is a wire permutation and costs no gates in silicon. Here it is nine one-bit selects, because the DSL has no way to say "the same bits, in the other order". Run against DEFLATE's fixed literal/length code — the RFC's own published numbers.

The honest gap

There is no initialised-memory node in the IR. Design::mem is zero-filled in both backends, so every lookup table in the corpus is a case or a multiplexer tree. That is close to the truth for a nibble LUT and a lie for a 256-entry alphabet map.

Design::rom(address, &table, width) builds exactly what a human would write with case: one arm per entry, emitting always @* case. Watch a 256-entry table grow. A synthesis tool would infer block RAM from the same source in seconds — which is the point being recorded rather than hidden.

What is not here

  • No initialised memory. The highest-value addition left to the toolchain. Design::rom documents the cost instead of hiding it.
  • No dynamic Huffman tree. Table is built from code lengths at elaboration time. Parsing a dynamic block header — nineteen code lengths, a seven-bit code-length code, the 16/17/18 run codes — is a different module.
  • No LZ77 window. The DEFLATE entry is the bit layer and nothing above it: no block structure, no error detection, no back-references.
  • Huffman is not one symbol per cycle. A decode spends code_len bits and the host supplies one bit per cycle, so the rate is one symbol per code length — seven, for DEFLATE's shortest fixed code. FSE can, because an FSE transition is allowed to cost zero bits. That difference is the whole distinction between the two decoders.
  • BLAKE3 and the FFT are placeholders. Files with no implementation behind them, not even declared as modules.