HARDWARE SECURITY

RTL Anti-Tampering DesignLesson 1 / 16

RTL Anti-Tampering Design Lesson 1: Why Correct RTL Can Still Be Unsafe

A correct signature result can change before it reaches the CPU. Follow the stored decision to accepted fetches, state a fault model, and separate timely blocking from later alerts.

19 min read

A chip is checking new firmware. Its signature fails, so the CPU should keep waiting. Success, failure, reset and timeout tests match the specification. Those results alone do not establish that the release path is secure.

The CPU does not recalculate the signature. It reads the stored result. If a fault changes that result from 0 to 1, does the CPU stay blocked? Follow the value to its consumer before choosing an encoding scheme.

RTL Anti-Tampering Design starts with registers, control signals and FSMs, then moves to netlist analysis and physical validation. This lesson separates two goals: computing the correct answer in normal operation and withholding unauthorized operations under faults. Each requires its own evidence.

Allow about 25–35 minutes for reading and exercises. You should know synchronous registers, basic FSMs and valid/ready interfaces. For the signature policy, see the Secure Boot lesson. The chip below is a fictional teaching design. Its RTL fragments and Python model illustrate counterexamples; they have not undergone RTL compilation, synthesis, formal proof or physical fault-injection testing.

1. The signature fails, but the CPU starts running

Imagine a library room that requires approval before it releases material. Your application fails, but the librarian still finishes the check. The record has separate checked and approved fields: checked_q and verified_q. A reader’s request, a desk ready to serve it, and permission to release correspond to valid, ready, and grant. Recording the handover corresponds to accepted_commit.

The librarian can reach the right decision while someone changes the saved approval from 0 to 1. The desk may then release material incorrectly. This story follows a saved result to its use; it does not add a second librarian or require a fresh check at every handover.

Follow one boot attempt. A checker examines the firmware image and decides whether it is allowed to execute. A register keeps that decision, and the logic that consumes the result decides whether to grant permission. That consumer might be a reset controller, a bus firewall or the CPU’s fetch interface.

Instruction fetch means the CPU reading a program instruction. We describe it through a simplified interface: the CPU requests an instruction and the interface is ready to accept the request, but execution permission is still required. The question therefore extends beyond whether the signature computation is correct: was a fetch from the unauthorized image actually accepted?

INTERACTIVE LAB · TEACHING MODEL

A failed signature—can the CPU still start?

Run the failed image without faults, then flip its stored result. Follow the checker, register and accepting edge. The full article and original diagrams remain below.

Predict first: if the stored result becomes 1, will the next fetch be accepted?

01 · Checkerauth_ok_i = 0Fail
02 · Result registerverified_q = 0checked_q = 0
03 · Release gategrant = 0checked_q && verified_q
04 · Fetch acceptanceaccepted = 0Not yet sampled at E2
E0 · reset

Reset leaves checked_q=0. The previous result is invalid; no fetch is accepted here.

Single-bit flagNot yet sampled at E2grant=0 · accepted=0
Exact four-bit comparisonNot yet sampled at E2grant=0 · accepted=0
Trusted-reference comparisonNot yet sampled at E2grant=0 · accepted=0

A failed image without faults receives no grant from any path. Try the stored-result upset next and follow how the same consumer reads Q.

Independent monitor retains the original resultref_auth_complete=0 · ref_auth_pass=0

The monitor and final gates are trusted for this model. This is a verification comparison, not implemented chip protection.

Codeword lab: how many bits separate 1001 from 0110?

Toggle a bit or choose one of sixteen words. Exact comparison decodes only 0110 as True. 1001 is legal False; the other fourteen words are neither teaching representation.

Legal False · rejectChanges from 1001:0 · Distance to 0110:4

A match establishes only that the word is 0110. The source-decision case above shows how legal True can originate from a corrupted decision.

Timing lab: was the operation accepted before the alert?

This separate illustration assumes a detector. The t0 fault has raised grant incorrectly. The first possible accepting edge is t1; reset arrives at t5. Change local-block and alert times, then inspect the first acceptance.

  1. t0FaultAcceptance history:None
  2. t1First acceptanceAcceptance history:Yes
  3. t2AlertAcceptance history:Yes
  4. t3WaitingAcceptance history:Yes
  5. t4WaitingAcceptance history:Yes
  6. t5ResetAcceptance history:Yes

One unauthorized acceptance has already occurred. A later alert, local block or reset does not erase that history.

Move the alert to t1 while leaving local blocking off: acceptance still occurs. Notification and grant are separate paths; an alert in the same cycle does not automatically change grant. These are teaching delays, not a detector, reset-controller or timing-margin simulation.

Two-state values only. One fault changes one target; a stored upset persists until reset. Original checker evidence, monitor, clock, reset and gates are trusted. The four-bit path fixes phase_ok=1 and local_fault=0. No X states, glitches, synthesis or physical injection are simulated.

Firmware checker, result register and release gate lead to an accepted CPU fetch; faults can target the entire permission path.
Figure 1. Follow permission to the point where execution is accepted. Open the diagram for its full size.

This deliberately vulnerable synchronous fragment stores the checker’s auth_ok_i result in verified_q. Everything belongs to one clock domain, and reset starts a new boot attempt.

// Deliberately vulnerable teaching fragment.
always_ff @(posedge clk_i or negedge rst_ni) begin
  if (!rst_ni) begin
    checked_q  <= 1'b0;
    verified_q <= 1'b0;
  end else if (verify_done_i && !checked_q) begin
    checked_q  <= 1'b1;
    verified_q <= auth_ok_i;
  end
end

assign exec_grant_o     = checked_q && verified_q;
assign accepted_commit = fetch_valid_i && fetch_ready_i
                       && exec_grant_o;
Two enabled flops store completion and pass results. A fault in the stored verified_q bit can open two AND stages and allow a fetch.
Figure 2. Logical view of the RTL. EN controls updates, the CLK triangle marks a rising edge, and the Rn bubble marks asynchronous active-low reset. This is not a synthesized netlist. Open to enlarge.

Read the code along the wires. The upper AND decides whether to write, and both flops share its enable. Once the check finishes, checked_q=1 stops further normal writes. The grant logic continues reading Q. A fault in the stored result can therefore change grant without another write.

A failed signature normally leaves checked_q=1 and verified_q=0, closing the gate. Flip verified_q to 1 after completion and the gate opens. If valid and ready are high, the fetch is accepted. The cryptographic core may have computed the correct answer throughout.

The two flags have different jobs. checked_q records that the check has finished; verified_q stores whether it passed. A failed image can set the first flag because the work was completed. Once a fault changes the second flag, the two ones seen by the consumer no longer faithfully represent the verifier’s decision.

Closing normal writes does not add integrity protection to stored state. Enable determines when RTL updates the register; this fault changes the retained value directly. A functional test for “no writes after verification” does not exercise that counterexample.

This is not OpenTitan RTL or a claim about a shipping chip. It illustrates how a correct result can change on its way to the consumer. A review must follow that path beyond the point where auth_ok_i is produced.

Start with one short path

Follow the same application through E0–E2. E0 clears the record; after E1 it says checked and refused. Someone changes the approval between bells, so the desk hands over material at E2. The independent review ledger still says refused, making that handover unauthorized. Bells stand for ordering only. E1 is this lesson’s local completion edge; Lesson 2 completes at edge 2 and uses a different schedule.

In the laboratory above, first inspect an unauthorized image without a fault. Then flip the stored result. Leave four-bit encoding and alert delays for later. Change one setting at a time so you can trace acceptance to that change.

TimeStored value and referenceAcceptance
E0 resetchecked=0, result=0No grant
Before E1Completion has not been capturedDo not use the post-E1 value
After E1checked=1, result=0; reference still says failureNo request accepted yet
Between E1 and E2Change the one-bit result to 1; keep checked=1The saved answer changes
E2 samplingvalid=ready=1, grant=1; reference pass=0accepted=1: unauthorized acceptance

This table fixes the order of completion, injection and request. Acceptance at E2 uses values settled before that edge. A post-E2 state change cannot rewrite an acceptance already recorded at E2. The laboratory is a two-state teaching model. E0–E2 are local labels for this short example, not Lesson 2’s edge schedule.

2. State which faults are allowed

The library experiment allows one change to the saved approval during one application. Everyone else, including the clock, follows the rules for this test. Two collection requests share that one opportunity; arriving at the desk does not renew the budget. If no collection occurs before the deadline, the record only shows no unauthorized handover in that schedule. It does not show that anyone detected or stopped the alteration. Trust defines this experiment’s scope, not permanent immunity.

When a specification says “resists fault injection,” the verification team still needs to know what to test. Which register can the attacker target? How many locations can change at once? Does a changed value last one cycle or remain stored? Those choices affect the result. Together, they form the fault model.

The first counterexample uses a narrow scope: one bit changes after the verification result has been stored.

ItemAssumption in this lesson
Protected goalAn unauthorized image receives no accepted fetch
Targetverified_q, after the checker completes
EffectOne 0→1 upset, retained until overwrite or reset
BudgetAt most one fault at one location per boot attempt
Excluded targetsChecker, monitor, clock, reset, grant gate and downstream logic are trusted in this experiment
Observation windowReset release through the first fetch, or a declared test deadline

The experiment isolates one question: can corrupting result storage release the CPU? Excluded targets are held trusted for this run, not declared immune. Adding the comparator, next-state logic or consumer requires another counterexample search under the wider model.

The observation deadline matters too. A request arriving after it may expose corruption that the test never saw. Report “no acceptance observed in this window,” rather than inferring that the fault was detected or blocked.

Voltage, clock, EM and laser perturbations are physical techniques. Bit flips, stuck-at values and altered combinational outputs are logical analysis models. Measurements and timing analysis must connect the two. Changing a Python value does not show that a laser can produce that effect on silicon.

Keep safety and availability separate as well. This lesson permits a fault to stop execution. An authorized image must still boot without faults, which needs its own checks. Tying grant permanently low satisfies the rejection property while producing an unusable product.

3. Registers and FSMs share a permission problem

The same application exposes a flow problem. The desk displays waiting, checking, and ready to release, like WAIT, CHECK, and RELEASE. A forged success condition makes it follow an allowed route to the release sign, although the application failed. Both the sign and the route look valid. A four-position approval code or duplicate records can check specified alterations; their shape or number alone cannot establish real authorization.

Widening verified_q to four bits or adding default: ERROR to an FSM can help with specific faults. Each change protects part of the path.

Design measureWhat it can addressCounterexample to investigate
Multi-bit result with an exact True comparisonRejects certain upsets that produce invalid codewordsA single upstream auth_ok flips before being encoded as legal True
Sparse states with an illegal-state error responseDetects modeled faults that produce invalid statesCorrupted cmp_pass selects the legal transition into RELEASE
Two registers with continuous comparisonDetects some mismatched storage errorsShared inputs, clocks or checker logic fail together; synthesis merges copies
An alert on detectionReports the event and starts a responseThe consumer accepts a protected operation before that response takes effect

Consider WAIT → CHECK → RELEASE. In CHECK, cmp_pass selects RELEASE or ERROR. A fault changes a failed comparison into success, and the controller follows the existing CHECK→RELEASE edge. The state is legal and the edge exists, but the authorization condition is false.

A conceptual FSM chooses ERROR or RELEASE from CHECK using cmp_pass. A corrupted condition can cause unauthorized release through a legal state and edge.
Figure 3. Conceptual FSM. The highlighted callout marks a fault in the transition condition. State encoding and the full controller transition set are unspecified.

The problem sits at the branch: which path did cmp_pass select? Checking that RELEASE belongs to the legal state set misses why the controller entered it. Trace the highlighted path, then ask which original decision should have blocked it.

Read the FSM at three levels. Check whether its current encoding belongs to a legal state, whether the conditions for this transition were satisfied, and whether the work required before arriving here actually happened. The first two checks do not remove the need for the third.

For example, completing a digest calculation gives us a digest; it does not establish that signature verification has completed and passed. Likewise, done reports completion of a task, while good represents a separate result to check. A controller that treats these meanings as interchangeable can grant permission at the wrong time.

Later lessons will calculate Hamming distance, select register encodings and examine FSM implementations. A wider Boolean representation alone does not settle these questions.

4. Watch the accepted commit, not just the alert

Material reaches the reader at t+1, the report arrives at t+2, and the room closes at t+5. Closing prevents later handovers but cannot erase the earlier one. Read accepted_commit in the handover ledger beside alert and reset. An approval light with no request is not yet a handover; the analogy must not count every light as a release of material.

Judge an authorization failure at the accepting interface. An internal flag rising does not by itself show whether the asset was delivered. We call acceptance at this simplified interface the accepted commit, giving the waveform a precise observation point.

In the simplified fetch interface, fetch_valid indicates a valid request, fetch_ready indicates that the interface is ready to accept it, and exec_grant supplies execution permission. All three must be high at the agreed acceptance edge to record fetch_valid && fetch_ready && exec_grant. Here, commit means interface acceptance; it does not describe the CPU’s complete instruction-execution sequence.

A secret-read interface might instead use the handshake that delivers sensitive data. Each design needs to identify the event where the asset or permission is actually handed over. A real system also needs to cover prefetch, DMA and other bus masters; guarding CPU reset alone may leave another path open.

A fault at t causes an accepted fetch at t plus one; alert at t plus two and reset at t plus five occur too late. Local blocking must precede acceptance.
Figure 4. A teaching timeline, not measured product latency. The operation can happen before the alert arrives.

Suppose a fault occurs at t, a fetch is accepted at t+1, the alert rises at t+2 and reset arrives at t+5. An “eventually alerts” test passes. Yet execution permission was already granted, and reset cannot withdraw that fetch.

Specify both local blocking and system response. Under the stated fault model, the local mechanism must protect the grant. The alert handler can then report, escalate or clean up. These actions have different deadlines and need separate properties.

For the failed image, the question is whether any unauthorized fetch was accepted. A later alert in the same trace cannot change a “yes” into a “no.” Record acceptance and alerts separately to establish whether the defense acted in time.

5. Express the contract in RTL and SVA

The desk now reads an entire four-position pass, accepting only 0110. It also requires a completed check, an allowed release phase, and no local error: result_code_q, checked_q, phase_ok_i, and local_fault_i. An auditor separately checks the original review ledger at every handover; that ledger represents the independent reference. A written audit rule does not build another door. SVA specifies a check rather than adding a blocking circuit.

Start with the consumer’s conditions. The teaching codewords AUTH_TRUE=0110 and AUTH_FALSE=1001 have Hamming distance 4; this is not a complete OpenTitan primitive. result_code_q becomes valid only after the current boot attempt’s check completes. phase_ok_i indicates that the consumer may grant execution in its current phase, and local_fault_i reports a local error.

localparam logic [3:0] AUTH_TRUE  = 4'b0110;
localparam logic [3:0] AUTH_FALSE = 4'b1001;

assign exec_grant_o = checked_q
                    && (result_code_q == AUTH_TRUE)
                    && phase_ok_i
                    && !local_fault_i;

assign accepted_commit = fetch_valid_i && fetch_ready_i
                       && exec_grant_o;
An exact four-bit comparison against 0110 joins checked_q, phase_ok_i and inverted local_fault_i to form grant, then combines with fetch valid and ready for acceptance.
Figure 5. Combinational logic from the second RTL fragment. The diagram starts at the stored-code input because the fragment does not supply result_code_q write logic.

The comparator asks whether the received codeword is 0110. The grant gate also checks completion, phase and local-error conditions. Neither the codeword nor the comparison proves how that value was produced. Its origin needs protection and verification of its own.

Exact decoding prevents an invalid codeword from passing merely because one bit is high. It cannot reject a legal True produced from a corrupted source decision. Storage integrity and source authorization therefore need separate analysis.

For two-state inputs, a single-bit change to FALSE cannot produce TRUE, so the exact comparison rejects it. The checker, comparator, grant gate and fetch interface still have fault surfaces. This fragment defines consumer conditions, not a finished tamper-resistant implementation.

In four-state simulation, == can produce X. Check for unknown grant and control values and fail the test on them. Do not interpret a simulator X as a specific physical fault. If the design also drives active-low reset, cpu_rst_n=1 means reset is released. A production reset controller must handle synchronization and glitches; the combinational grant above is not a complete reset controller.

Use SVA to express the requirements. The first property checks the wiring contract: when an operation is accepted, are all the conditions required by the grant logic present? It can catch omitted or incorrectly connected conditions, but the conditions themselves may still be corrupted.

The second property compares against an independent trusted observation: did this boot attempt actually complete and pass verification for the image behind the accepted operation? The third property covers a reachable authorized acceptance, so permanently refusing all work does not make a rejection property look sufficient.

// Interface contract: useful, but mostly checks the gate's wiring.
assert property (@(posedge clk_i) disable iff (!rst_ni)
  accepted_commit |->
    checked_q && (result_code_q == AUTH_TRUE)
    && phase_ok_i && !local_fault_i);

// Security goal: ref_* belongs to a trusted verification monitor.
assert property (@(posedge clk_i) disable iff (!rst_ni)
  accepted_commit |-> ref_auth_complete && ref_auth_pass);

// Avoid a vacuous result: authorized work must be reachable.
cover property (@(posedge clk_i) disable iff (!rst_ni)
  ref_auth_complete && ref_auth_pass ##[1:8] accepted_commit);
The trusted verification monitor supplies independent completion and pass evidence; a sampled assertion compares it with the faultable DUT accepted_commit output.
Figure 6. Verification setup for the security assertion. This monitor belongs to the testbench or formal environment; it is not an extra security circuit implemented in the chip.

In the one-bit counterexample, the DUT can report accepted_commit=1 while the independent reference still says the image failed. That disagreement makes the security assertion fail. Copying verified_q into the reference would erase the disagreement and hide the counterexample.

The consequent of |-> is checked at the same sampled edge. Here the interface accepts a fetch on the rising edge, with control inputs stable beforehand. If your interface commits in another cycle, change the monitor and properties accordingly.

The trusted testbench or formal monitor must establish ref_auth_complete and ref_auth_pass independently from the current image and specification, retaining them across this boot attempt. Do not rename the fault target verified_q as ref_auth_pass. The corrupted flag and its supposed evidence would both become 1, allowing the assertion to pass. The monitor is trusted within this experiment. A real independent hardware checker needs separate implementation, common-path and cost analysis.

disable iff excludes intervals with asserted reset, so these properties do not validate reset glitches or attacks on reset itself. The cover checks that a successful path is reachable; it is not a complete liveness proof.

Calculate the distance, then read the two assertions

Flipping bit 1 of refusal code 1001 gives 1011, which an exact comparison against 0110 rejects. But a forged source decision can make the printer produce a perfect 0110. The audit asking whether the desk followed its inputs may pass; the audit asking whether the application was actually approved fails. These are the wiring and reference assertions. Four differing positions do not represent four independent checks or measured resistance to forgery.

Hamming distance counts differing positions in two equal-length bit strings. Compute 1001 XOR 0110 = 1111 and count four ones: the distance is 4. Flipping stored bit 1 instead gives 1001 XOR 0010 = 1011. An exact comparison with 0110 rejects it. This example trusts the pre-encoding decision, comparator and acceptance endpoint.

SVA means SystemVerilog Assertions: properties evaluated at specified sampling events. DUT means design under test. The reference independently records the expected decision in the testbench; it is not another name for the DUT’s saved result.

Now introduce a separate pre-encoding source fault. Reset is released and all inputs settle before E2: checked=1, result_code=0110, phase_ok=1, local_fault=0 and valid=ready=1. The DUT produces grant=1, so accepted=1. The reference has complete=1 and pass=0. A property requiring the DUT grant conditions on acceptance passes; a property requiring reference authorization fails. The first checks wiring, the second checks permission. A corrupted source can still be encoded as legal 0110. This is not a one-bit storage upset crossing distance 4.

The cover delay ##[1:8] permits acceptance at one of the next one to eight sampling edges and seeks at least one matching path. It does not guarantee completion of every authorized job or impose an eight-cycle product deadline. The real harness must enforce reset exclusions, fault budgets and trusted-reference assumptions. No formal proof has run in this lesson.

6. Run a reproducible fault exercise

In the exercise, first alter the saved approval for a rejected application. Next change one position of a coded pass, then forge the decision before printing it. Start a new application for each run and retain the independent refusal decision. Compare the stored field, the value seen at the desk, and the handover. Reproducing this schedule in Python or the lab supports the specified model; it is not an actual pass-forgery experiment, RTL run, or chip test.

Download lesson01_fault_model.py and run it with Python 3:

python lesson01_fault_model.py

The model uses the standard library, two-state values and specified discrete timing. Expect six counterexample/control scenarios followed by PASS: 16 codewords + 4 single-bit faults + 6 scenarios. PASS means the model reproduced the expected outcomes, including successful bypasses of deliberately vulnerable logic. It does not mean a chip passed security validation.

ScenarioExpected observation
Authorized image, no faultBoth the vulnerable gate and the conditional reference gate allow execution
Unauthorized image, no faultBoth reject it
Stored result bit flipsVulnerable logic releases; the comparison gate with a trusted reference rejects
Upstream decision flips before encodingA legal True codeword appears; exact decoding cannot identify its false origin
Wrong transition into a legal stateState membership passes; history/authorization checking fails
Alert follows commitThe unauthorized operation was accepted; a later alert cannot repair that property

First, list the trusted boundary for each scenario without changing the program. Then check all 16 codewords: only 0110 should decode as True. Calculate the four single-bit changes to 1001. Finally, expand the model: what remains guaranteed if a fault can also target the reference checker or final grant gate? Write the new counterexample rather than changing the expected result to PASS.

With an RTL simulator, you can wrap the teaching fragment in a testbench, inject a result-register fault after failed verification, and record commit, alert and reset times. force/release, backdoor writes and an injection mux can retain faults differently. State the injector’s behavior in the model. This RTL experiment has not been run for this lesson.

7. Use each source for the question it answers

A library flowchart identifies the handover point, a policy defines who qualifies, and a particular exercise shows what happened in that exercise. Read architecture, RTL documentation, and fault analysis with the same separation. A similar door at another library does not transfer its validation to your ledger. The cited technical sources likewise do not certify this fictional gate.

SYNFI addresses faults in netlists. It uses SAT to analyze synthesized circuits, with OpenTitan case studies. Read the subcircuit boundaries, input/state setup, fault budget and definition of an effective fault before interpreting a result. Its conclusions do not certify a whole chip or every physical injection technique. SYNFI paper

SCFI addresses FSM control flow. It hardens next-state computation using control information and execution history, matching our legal-state counterexample in Section 3. This lesson does not reproduce its circuit or experimental results. SCFI paper

The Hot Chips Titan talk explains the system handoff. Slide 19 shows host firmware authentication followed by system-reset release; slides 33–34 describe Titan’s own staged verified boot. This architecture does not establish physical fault resistance for our teaching gate. Official Titan slides

For an inspectable implementation, OpenTitan rom_ctrl sends separate done and multi-bit good signals and forwards the ROM digest to the key manager. Its documentation describes good as an additional check: ROM trust also involves key derivation and lifecycle policy. Our simplified signature-before-fetch rule is not the complete boot specification for every OpenTitan lifecycle state. rom_ctrl documentation

Together, these sources help formulate review questions about result storage, consumption, execution history and verification after implementation. They address different abstraction levels.

8. Draw a permission table for your design

Build the permission table by tracing a handover back through the desk, saved record, and review source. A side entrance that releases the same material matters even if the front door stays shut; it represents another consumer, such as debug or DMA. The story helps list paths. It does not prove that every entrance has been checked or that one four-position encoding protects every field.

Start a design review with one action that grants authority: CPU fetch, debug unlock, OTP programming or secret read. Find its acceptance event, then trace backward: which logic grants permission, which result does it read, where is that result stored, and who produces it? This gives you a path to examine one stage at a time without first inventorying every flop in the chip.

Field to recordExample in this lesson
Protected action and acceptance eventCPU fetch, accepted_commit
Trusted authorization evidenceCurrent image satisfies independent signature and version policies
Producer, storage and consumerChecker → result register → fetch gate
Fault targets along the pathComparator output, register, next-state logic, decode, grant
Detection/blocking deadlineLocal block takes effect before the next possible accepting edge
Shared failure sourcesClock, reset, supply, input, decoder and synthesized shared logic
Verification evidenceFault model, counterexample traces, assertions, covers, revisions and tool limits

At each stage ask what actually completed before “success” appeared. If the only answer is another bit from the same source, include it in the next fault campaign. Check alternate paths too: can debug or DMA read secrets while the CPU remains in reset?

The table is not an instruction to apply one encoding everywhere. Result storage needs integrity analysis; an FSM needs transition-condition and history analysis; the accepting interface needs a blocking deadline. Identify how each stage can fail before selecting its countermeasure.

Completion target: explain the one-bit bypass, state a bounded fault model, and write the security property at the accepted operation. The quiz below checks those skills.

9. Engineering extension: change the assumptions

Allowing changes to the desk decision or the checked field needs a new experiment contract. The old saved-approval result cannot cover those targets. A reader might also infer private information from rejected responses, creating a separate leakage question. These correspond to added fault targets and confidentiality observations. No unauthorized handover answers the handover contract; it does not establish secrecy.

For RTL and DV engineers: extend the campaign

Establish a fault-free baseline first. For each injection, record target, effect, duration, count, initial state and observation endpoint. Count detected, blocked before commit and unauthorized commit separately. Keep timeouts and uncovered cases unresolved.

Next, move faults to checked_q, the comparator output, FSM conditions and grant logic. Then examine repeated injections and shared sources. If you add redundant flops or sparse states, inspect the synthesized netlist to confirm that the intended structure survives. Two RTL copies do not necessarily become two gate-level copies. OpenTitan’s guidelines flag these synthesis risks. Hardware design guidelines

These checks do not cover every confidentiality problem. Faulty ciphertext may expose a key. Blocked or ineffective faults can also leak information through statistical selection. The later DFA/SIFA lesson will distinguish unauthorized commits from information leakage.

This rejection property asks whether unauthorized operations were accepted. Protecting secrets needs an additional leakage goal and observation model. One PASS label cannot establish both kinds of protection.

10. The series: from registers to verification evidence

The library investigation expands along the path: the stored pass, the flow signs, the source, the handover deadline, and eventually the built implementation. This is the series’ progression from registers to FSMs and evidence. Lessons 1–10 are published and six more are planned. Imagining a complete library policy does not mean the planned lessons or chip validation already exist.

The series plans 16 lessons; Lessons 1–10 are published; six further lessons are planned. The RTL Anti-Tampering Design course page maintains the outline, lesson goals and publication status.

The course moves from registers, FSMs and permission handoffs to fault campaigns, physical-model calibration and evidence packages. Each lesson closes with a wrap-up of the threat model, causes, defenses, validation and limits.

11. References and reading assignments

This original teaching synthesis was source-checked on 2026-10-03. Claude and Grok provided independent read-only research; Codex owns the editorial decisions, source verification, implementation and website acceptance. The code, tables and diagrams are teaching models, not reproductions of paper experiments.

  1. Pascal Nasahl et al., SYNFI: Pre-Silicon Fault Analysis of an Open-Source Secure Element, TCHES 2022(4), 56–87. Author version / DOI. Assignment: identify subcircuit boundaries and effective-fault definitions; name two excluded targets.
  2. Pascal Nasahl et al., SCFI: State Machine Control-Flow Hardening Against Fault Attacks, DATE 2023; preprint 2022. Author version / DOI. Assignment: use Section 3’s FSM to distinguish state integrity from control-flow integrity.
  3. Victor Arribas, Felix Wegener, Amir Moradi and Svetla Nikova, Cryptographic Fault Diagnosis using VerFI, IEEE HOST 2020, 229–240. Author-institution abstract / DOI. The bibliography and abstract were checked; the full experiment was not reproduced. Save this gate-level diagnosis reading for Lesson 14.
  4. Scott Johnson et al., Titan: enabling a transparent silicon root of trust for Cloud, Hot Chips 30, 2018. Official slides. Assignment: compare host-reset release with Titan’s own verified boot using slides 19 and 33–34.
  5. lowRISC, OpenTitan ROM Controller: Theory of Operation. Official documentation. Assignment: trace the different consumers of done, good and the digest. Live master documentation changes; pin a revision for project validation.
  6. lowRISC, Secure Hardware Design Guidelines. Official documentation. Assignment: identify synthesis risks to redundancy and check whether your delivery process includes netlist validation.

MY ACADEMY · LESSON FILM

Lesson video

The film explains this lesson’s data path. After a section, return to the interactive exercise and change the input or fault conditions. The animation presents a teaching model; it does not replace RTL simulation.

Download MP4 · Captions VTT

Narration uses a synthetic voice. Both the interaction and animation have model boundaries; interpret results using this lesson’s sources and validation scope.

Wrap-up: take this lesson into a design review

Threat model and assumptions

Return to the opening chip. The new firmware has failed signature verification, so the CPU should keep waiting. The action we protect is instruction fetch: the CPU reading a program instruction. In this simplified model, even one accepted fetch from an unauthorized image violates the stated protection goal.

The first counterexample lets the attacker change one location: flip verified_q from 0 to 1 after verification finishes. At most one fault occurs per boot attempt, and the changed value remains until overwrite or reset. The test observes from reset release to the first fetch or a deadline declared in advance.

For this experiment, the checker, observation monitor, clock, reset, final grant gate and downstream interface are trusted. This deliberately narrow scope lets us isolate what happens when one result register is corrupted. Later scenarios that target other locations need their own assumptions; they cannot inherit this result unchanged.

Why the design fails

The verifier may compute the correct answer throughout, but the CPU's release logic reads the stored result. If that result changes during storage or transfer, the consumer may interpret failure as success. A review therefore has to follow the result to where it is used, rather than stopping at the verifier's output.

An FSM can fail in a similar way. The controller may take an existing path into a legally encoded RELEASE state even though this attempt never earned permission to take that path. A multi-bit result also requires checks on both sides of its encoding: the upstream single-bit decision, then the decoder and release gate downstream.

Defenses

First identify the point where the downstream interface actually accepts the operation. We call that event the accepted commit: at the agreed clock edge, the fetch request is valid, the interface is ready and execution permission is asserted. The defense must block an unauthorized operation before this event. An alert arriving afterward cannot change the fact that the operation was accepted.

Then examine the reasons for granting permission: did this attempt complete verification, does the result match the full codeword, is the current phase permitted, and is a local fault already present? Authorization history may need checking too. Assess how each condition obtains trusted evidence and whether the conditions share a vulnerable source. Adding more AND terms does not automatically establish a complete security guarantee.

Validation and checks to perform

Start by rerunning the six counterexample and control scenarios. Do not rely on the final PASS label alone: it means the teaching program reproduced its expected outcomes, including successful bypasses of deliberately vulnerable logic. Follow the timeline instead: when did the fault occur, when was the fetch accepted, and when did blocking and the alert take effect?

The completed demonstration is the finite Python model. If the next campaign allows faults in the final grant gate, checker or shared sources, search for counterexamples again. Do not carry over an earlier pass record as evidence for the expanded scope. Each result should state which components it trusts, which targets it excludes and when observation ends.

Limits and unverified claims

The model changes values at specified locations to explain failures along the permission path. It does not establish that a physical perturbation can produce the same effect on silicon. The lesson has also not undergone RTL simulation, synthesis, formal proof or silicon validation. Its demonstration cannot support a claim that an IP has passed anti-tampering validation.

Availability needs attention as well. Permanently disabling grant prevents unauthorized releases, but it also prevents legitimate firmware from booting. Separate checks must show that authorized firmware works without faults and that faults trigger blocking or recovery under a defined policy. This lesson uses accepted instruction fetch as its protection boundary; it makes no claim that the complete CPU or SoC is secure.

Try a changed assumption

Change just one assumption: the attacker can now target the final grant gate as well as the result register. Draw the original blocking path. If the last gate's output is corrupted, can the earlier checks still prevent that fetch? List the components that remain trusted, then write a possible bypass or a property that needs renewed validation.

This wrap-up summarizes the lesson’s teaching cases, references and experiment scope. Checks not reported as completed remain future work.

Learning guide

RTL Anti-Tampering Design

0 / 16

Open the course outline → · Progress counts published lessons only

Prerequisites

  • Synchronous registers, basic FSMs and valid/ready interfaces

What I learned

  • Trace a permission path to its accepted operation
  • State the fault model and trusted boundary
  • Distinguish state legality, history and authorization
  • Separate local blocking from delayed alerts

Key terms

Open glossary →

Further reading

Knowledge check

1. What does a fault-free functional test establish?
2. A legal FSM edge reaches RELEASE after cmp_pass flips. What failed?
3. Commit is at t+1 and alert at t+2. Was the operation prevented?
4. Where should ref_auth_pass come from?
5. An upstream false decision becomes a legal True codeword. Does exact decoding detect its origin?

Thanks for reading.

Take the concept with you, not just the terminology.

#RTL#Fault Injection#FSM#Secure Boot#Hardware Security#SYNFI#SCFI