Week 22 · Experiment: Fault injection makes hidden assumptions visible
Session 3 of 4 · Integrate the system · Phase 6
Plan about 15 minutes for explanation, 30 minutes for practical work and 10–15 minutes for documentation. A longer build may continue into the next session: stop safely, commit the current state and record the next check. Desktop simulations count as software evidence; label them clearly and record physical validation separately.
Engineering challenge
Can every module cooperate without weakening the stop and measurement contracts? This session focuses on fault injection makes hidden assumptions visible.
Before you start
The previous week’s recorded baseline and Week 21, Improve. For later sessions this week, retain the preceding session’s files and predictions.
Equipment: Desktop Python, editor, paper and ruler; for physical work, the configured 3pi+ 2040, clear floor mat and hardware checklist. Week 3 additionally uses the separate low-voltage LED circuit described in its procedure.
For any motion, verify the stop button, short time limit and clear floor area first. Keep the wheels raised for a new device program until its commands and stop behaviour are checked. A hazard or uncertain input is a reason to stop and document, not to force the trial to finish.
Theory and mathematics
Fault injection makes hidden assumptions visible
A fault-injection test supplies missing sensor data, an impossible count jump, a blocked path or a stuck encoder to see how the system responds. Start with synthetic inputs; do not damage the device or create unsafe physical faults. A timeout catches some stalled motion but does not identify its cause. Preserve the last valid estimate for diagnosis while marking it invalid for continued control. Decide which faults require manual reset. Automatic retries need a count/time budget; otherwise a recovery loop can become another unbounded motion command.
Worked example — illustrative values
Consider a sample with goal_distance = 0.02 m and bumper_pressed = True. Even if 0.02 m is inside the arrival tolerance, hazard priority selects STOP_HAZARD. With a 60 s mission budget, an elapsed time of 60 s selects STOP_TIMEOUT. A lower-priority target must not overwrite either decision.
Write the calculation in your notebook before running code. State which values you measured, which you assumed and which the program calculates. A correct numerical calculation cannot rescue an incorrect physical assumption.
Run and explain the model
The following is desktop Python, not a ready-to-run motor program. Download this week’s example, save it in your student repository and run python3 code/w22.py from the repository root. The same small model is reused across the week so you can learn it, build with it, test it and revise it.
# Desktop Python teaching example. Numerical inputs are illustrative.
def arbitrate(*, fresh, hazard, elapsed_s, arrived):
if hazard:
return "STOP_HAZARD"
if not fresh:
return "STOP_INVALID"
if elapsed_s >= 60:
return "STOP_TIMEOUT"
if arrived:
return "STOP_ARRIVED"
return "FOLLOW"
for case in [(True, True, 2, True), (False, False, 2, False),
(True, False, 60, False), (True, False, 2, True)]:
fresh, hazard, elapsed, arrived = case
print(arbitrate(fresh=fresh, hazard=hazard, elapsed_s=elapsed, arrived=arrived))
Use your own recorded data or named test fixtures instead of the illustrative inputs. Save expected and actual values side by side. If an exception appears, read its final line, identify the input or assumption that caused it and make the smallest explained correction. Do not delete validation merely to obtain output.
Understanding the model and its limits
The arbitration function gives hazard priority, followed by invalid data, timeout and arrival. Thus simultaneous arrival and contact reports STOP_HAZARD. The sample returns decision strings; the device adapter’s single motor owner converts any stop decision into motors.off and prevents stale targets from being reused. Test combinations, not only one condition at a time. A later navigation write can undo a stop unless motor ownership is explicit. Record the input sample and selected reason so replay can identify the exact decision boundary where the system diverged.
Practical instructions
- Inject stale data, impossible jumps, blocked route and timeout in replay/simulation.
- Check that no later movement command follows a latched stop without manual rearm.
- Repeat a known input trace and compare decisions exactly.
- Collect one short physical segment only after all software guards pass.
Experiment
Run at least eight named normal/fault cases and inspect the first stopping sample and all later actions.
Before testing, record your prediction, changed factor, measured response, fixed conditions and stopping rule. Save every attempted run, including failures, with a condition and source version. If hardware is unavailable, use an explicitly labelled synthetic/replay dataset and list the physical question it cannot answer. Do not invent completed trials.
Deliverable
Eight deterministic fault outcomes and no unauthorized post-stop movement.
Save notebook/w22-s3.md, the relevant code revision, raw CSV or test-case records, and one labelled diagram/plot/table. Link the files relatively from your notebook. Use the entry template and report guide.
Completion criteria
- Explain fault injection makes hidden assumptions visible in your own words using this session’s example and its units/assumptions.
- Produce the specific evidence above: Eight deterministic fault outcomes and no unauthorized post-stop movement.
- Keep predictions and raw outcomes, distinguish observations from interpretation, and explain one limitation or unresolved failure.
- Review the Git diff, commit the session’s intended files and state the next experiment or safe continuation point.
A documented failed prediction can meet the learning criteria. A missing physical trial must remain marked untested; software success alone does not validate the robot.
Reading and video
- Focused reading: Robotics Lab: device interfaces and bounded motion. Study task: Trace hazard priority through the one motor-command owner.
- Video/lecture option: MIT OpenCourseWare: Macro ME robot demonstration. Study task: Compare its sensors with your robot; do not copy hardware assumptions. Watch a relevant 5–10 minute excerpt or use the linked notes if video is inaccessible. This is supporting conceptual material; hardware in a demonstration may differ from yours.
- Practical reference: Engineering handbook and hardware setup. Manufacturer/API references and video metadata were checked on 2026-10-09; recheck the actual firmware before transferring code.
Reflection and next step
Which assumption most affected your result? Point to one observation that supports your explanation and one alternative explanation the evidence has not ruled out. Write a specific next test with a changed factor and measurable outcome, then proceed through the week’s Learn → Build → Experiment → Improve cycle.