JEnt v2.2.0 LFSR Modeling

Paper: Joshua E. Hill (KeyPair Consulting) and Yvonne Cliff (Teron Labs), JEnt v2.2.0 LFSR Conditioning Analysis, Document Version v1.0, September 15, 2026.

Summary

Our paper is a mathematical and empirical treatment of the 64-bit LFSR that JEnt v2.2.0 uses as its final, non-vetted conditioning component. JEnt v2.2.0 (2019) was the first release that targeted 90B, sits behind roughly two dozen ESV certificates and many Linux kernel trees, and its design document offers only a random-map (Output_Entropy) argument plus the claim that a primitive-polynomial LFSR does not diminish entropy. We point out that claim cannot hold here: 64 × osr symbols of 64 bits each go in for every 64 bits that come out, so the overall map is compressive regardless of the polynomial.

Section 2 builds an explicit linear model. We verify the model against the actual JEnt code and describe a variety of mathy features of this conditioning.

Section 3 turns the model into a min-entropy heuristic. Under stated assumptions (each non-stuck symbol carries at least 1/osr bits; the LFSR spreads each input bit across about half the state; per-bit contributions combine via the piling-up lemma), the worst case (osr = 1) yields 42.3 bits of min entropy per 64-bit block, rising only to about 43.4 as osr grows; this is a defensible engineering estimate, not a bound.

The same section examines the stuck filter, which in JEnt v2.2.0-v3.6.3 discards raw samples whose first, second, or third timer difference is zero and feeds only survivors to the conditioner. The central argument is that min entropy is a property of the generating system, not the data, so data-dependent filtering changes the system. Measurements on twelve hardware/version combinations show the filter removes from 0.1% (Cascade Lake, v3.6.3) to 44.5% (ARM Cortex-A53, v2.2.0) of the data, driven mostly by run persistence on the lossy platforms, and show that filtering can move the 90B tool's assessment in either direction. The conclusion is that only direct assessment of the post-stuck-filtered stream is reliable in this case, and osr should be set from it.

Section 4 maps this onto the 90B non-vetted conditioning process, where h_out is the minimum of Output_Entropy, 0.999 × n_out, h' × n_out, and (per IG D.K Resolution 5) the mathematically justified level. Here, the stuck filter must be treated as the first non-vetted conditioning stage, which means that within 90B it can never be credited with improving the stream but can degrade it. Output_Entropy is shown never to be the controlling term for the LFSR.

The bulk of the section then characterizes the statistical h' assessment itself: one million pseudorandom 10^6-byte datasets run through the 90B bitstring tests give a median of 0.904 and a 95% band of [0.8758, 0.9336] (56.05-59.75 bits per block), with the compression estimator controlling 89% of the time. This assessment applies to all non-vetted conditioners whose output appears to be pseudorandom (most non-vetted conditioning functions).

Along the way we find the t-Tuple estimator produces only 56 distinct values across a million runs: for 8 × 10^6-bit strings the longest substring repeating at least 35 times is almost always 19 bits long, and the estimate then depends on a single small integer count. A Poisson model of substring occurrences predicts both the length and the resulting "atoms," ten of which account for about 8% of all final assessments. Real JEnt v2.2.0 output on four architectures reproduces the same distribution and the same top atom (0.935706).

Finally, we show that raising the symbols-per-block from 64 × osr to 126-159 × osr lifts the heuristic above the statistical result in most or all validations (mean h_out of about 57.5-57.9 bits) at a 31-45% throughput cost, with no assessment benefit beyond 159 × osr.

Impact on the ESV Program

Two findings bear directly on existing certificates. First, the conditioned-output claims. Most JEnt v2.2.0 certificates credit 56.22-59.89 bits per block, which is in the range our paper shows the h' test produces on any pseudorandom output. In other words, those claims appear to have been set by statistical testing alone, with no LFSR-specific mathematical evidence; the top of that range coincides with our observed Q[19] = 35 atom (0.935706 × 64, about 59.885). IG D.K Resolution 5 caps h_out at the mathematically justified level, so if this heuristic (or anything like it) is accepted as that evidence, the supportable figure is 42.3 bits, roughly 25-30% lower. Our paper does not claim the outputs are unsafe; a module seeding a DRBG would simply need correspondingly more blocks. But it does mean that around twenty certificates carry claims that appear to rest on statistical testing alone. The three certificates with lower claims (#E315, #E328, #E344) may already reflect some such reduction.

Second, and broader, the osr settings. Every JEnt certificate we know of (as of 20260915) covers v2.2.0-v3.6.3, and every assessment procedure we know of, including the CMUF working group's opportunistic-IID heuristic for v3.4-v3.6, set osr from the raw stream rather than the post-stuck-filter stream the conditioner actually consumes. Because the filter can remove anywhere from 0.1% to 44.5% of samples and can move the assessed per-symbol entropy either way, the osr on essentially all existing JEnt certificates, including v3.x certificates that use the vetted hash-based conditioner, was likely chosen using the wrong data.

At the process level, our paper effectively proposes a new validation workflow for these versions: treat the stuck filter as the first non-vetted stage, run a full entropy assessment and an h' bitstring assessment on the post-filter stream (Theseus provides the tooling), and use k = 1 conservatively. It also gives NIST reviewers and labs a sanity check we did not have before: h' results above 0.935706 should occur less than once in 3,500 validations, and identical h' values recurring across unrelated validations are expected rather than suspicious. The remediation path is concrete and cheap: stop discarding stuck samples (a two-line change that v3.7.0+ already made), set osr from post-filter data, and optionally raise w to support higher block claims.

Two caveats temper the impact: the heuristic rests on load-bearing assumptions and is explicitly not a bound, and our paper itself notes that enforcement of the mathematical-evidence requirement has been inconsistent, with contradictory statuses on the relevant SHALL statements, so whether CMVP requires re-assessment of affected certificates, grandfathers them, or issues new guidance is an open question.

Presentations

The findings were presented most comprehensively at the CMUF Entropy Working Group using these slides (deck version 20260901, posted 20260902).

Artifacts

The JEntLFSRModel repository contains the code used to generate the matrices in the paper, to verify the linear model against the JEnt v2.2.0 implementation, and to model every raw noise symbol in a range and produce 1 million bytes of conditioned output per symbol by repeated LFSR processing (the pseudorandomness comparison in Section 3.2). It also contains filtermass-testing.tar.xz, the modified JEnt v2.2.0 and v3.6.3 code used in the filter-mass testing of Section 3.3.2. The raw results for the h′ testing of Section 4.2.2 are in the hprime-results.tar.xz file.

Public Paper Revision History

Various versions of this paper have been publicly distributed in various venues. For reference, they are archived here:

Presentation Revision History


Mail josh@keypair.us with comments, questions, etc.

Page last updated September 20, 2026.