RL Robotics LabSearch ↗
Lesson / authored

Week 14 · Experiment: False positives and false negatives

Session 3 of 4 · Detect obstacles · Phase 4

Plan about 15 minutes for explanation, 30 minutes for practical work and 10–15 minutes for documentation. A longer build may continue into the next session: stop safely, commit the current state and record the next check. Desktop simulations count as software evidence; label them clearly and record physical validation separately.

Engineering challenge

Can your detector distinguish a hazard while leaving enough time to stop? This session focuses on false positives and false negatives.

Before you start

The previous week’s recorded baseline and Week 13, Improve. For later sessions this week, retain the preceding session’s files and predictions.

Equipment: Desktop Python, editor, paper and ruler; for physical work, the configured 3pi+ 2040, clear floor mat and hardware checklist. Week 3 additionally uses the separate low-voltage LED circuit described in its procedure.

For any motion, verify the stop button, short time limit and clear floor area first. Keep the wheels raised for a new device program until its commands and stop behaviour are checked. A hazard or uncertain input is a reason to stop and document, not to force the trial to finish.

Theory and mathematics

False positives and false negatives

A true positive correctly detects a labelled hazard. A false positive calls a clear condition hazardous. A false negative misses a real hazard; a true negative correctly identifies clear. Precision is true positives divided by all positive detections; recall is true positives divided by all real hazards. These ratios can be undefined when the denominator is zero. Overall accuracy can look excellent if almost every test is easy clear floor. Include hazards and difficult near-threshold conditions, and explain which type of mistake is more costly for the mission.

Worked example — illustrative values

With fictional TP=8, FP=2, FN=2, TN=8: precision=8/10=0.8, recall=8/10=0.8, accuracy=16/20=0.8. At 0.10 m/s and a 0.20 s processing/reaction delay, the robot travels 0.020 m before additional stopping distance. These examples are not measured safety margins.

Write the calculation in your notebook before running code. State which values you measured, which you assumed and which the program calculates. A correct numerical calculation cannot rescue an incorrect physical assumption.

Run and explain the model

The following is desktop Python, not a ready-to-run motor program. Download this week’s example, save it in your student repository and run python3 code/w14.py from the repository root. The same small model is reused across the week so you can learn it, build with it, test it and revise it.

# Desktop Python teaching example. Numerical inputs are illustrative.
truth = [True, True, False, False]
readings = [810, 440, 500, 125]  # fictional
predicted = [value > 470 for value in readings]
for actual, estimate in zip(truth, predicted):
    print("TP" if actual and estimate else "FN" if actual else "FP" if estimate else "TN")

Use your own recorded data or named test fixtures instead of the illustrative inputs. Save expected and actual values side by side. If an exception appears, read its final line, identify the input or assumption that caused it and make the smallest explained correction. Do not delete validation merely to obtain output.

Understanding the model and its limits

The two Boolean lists represent independent ground truth and detector output. Their paired outcomes give one true positive, one false negative, one false positive and one true negative. In this four-case toy example, hazard recall is 1/(1+1)=0.5 and false-positive fraction among safe cases is 1/(1+1)=0.5. These are illustrative counts, not a sensor performance claim. Choose a threshold on development labels, then evaluate fresh labels. A missed hazard and an unnecessary stop have different consequences; record both rather than choosing a threshold solely by total accuracy.

Practical instructions

  1. Use ten known boundary and ten known clear positions as validation fixtures.
  2. Freeze the threshold before collecting readings and independently label each fixture.
  3. Record confusion counts and detection transitions for plain threshold and hysteresis.
  4. Measure added release delay on a replayed trace, retaining the original raw samples.

Experiment

Compare two detectors on the same 20 labelled held-out positions and report TP/FP/FN/TN with denominator-aware metrics.

Before testing, record your prediction, changed factor, measured response, fixed conditions and stopping rule. Save every attempted run, including failures, with a condition and source version. If hardware is unavailable, use an explicitly labelled synthetic/replay dataset and list the physical question it cannot answer. Do not invent completed trials.

Deliverable

20 validation labels, confusion counts and an explanation of the precision/recall tradeoff.

Save notebook/w14-s3.md, the relevant code revision, raw CSV or test-case records, and one labelled diagram/plot/table. Link the files relatively from your notebook. Use the entry template and report guide.

Completion criteria

A documented failed prediction can meet the learning criteria. A missing physical trial must remain marked untested; software success alone does not validate the robot.

Reading and video

Reflection and next step

Which assumption most affected your result? Point to one observation that supports your explanation and one alternative explanation the evidence has not ruled out. Write a specific next test with a changed factor and measurable outcome, then proceed through the week’s Learn → Build → Experiment → Improve cycle.

← Previous sessionCourse roadmapNext session →