RL Robotics LabSearch ↗
Lesson / authored

Week 23 · Improve: A change creates a new validation version

Session 4 of 4 · Test under pressure · Phase 6

Plan about 15 minutes for explanation, 30 minutes for practical work and 10–15 minutes for documentation. A longer build may continue into the next session: stop safely, commit the current state and record the next check. Desktop simulations count as software evidence; label them clearly and record physical validation separately.

Engineering challenge

How reliable is your system within its claimed conditions, including the failures? This session focuses on a change creates a new validation version.

Before you start

The previous week’s recorded baseline and Week 22, Improve. For later sessions this week, retain the preceding session’s files and predictions.

Equipment: Desktop Python, editor, paper and ruler; for physical work, the configured 3pi+ 2040, clear floor mat and hardware checklist. Week 3 additionally uses the separate low-voltage LED circuit described in its procedure.

For any motion, verify the stop button, short time limit and clear floor area first. Keep the wheels raised for a new device program until its commands and stop behaviour are checked. A hazard or uncertain input is a reason to stop and document, not to force the trial to finish.

Theory and mathematics

A change creates a new validation version

After a fix, restart validation with the new version instead of averaging old/new outcomes. Compare versions on the same diagnostic fixtures, then use fresh trials for the final claim. Report endpoint error conditional on arrival separately from overall completion fraction; a controller that stops early on difficult trials can appear very accurate if failures vanish from the error plot. Keep safety-stop rates, navigation success and measurement errors visible together. A qualified negative result is an engineering result worth presenting.

Worked example — illustrative values

In ten fictional attempts, eight arrive and two stop on hazards. Navigation success is 8/10 = 80%; safe stopping on the two hazard cases is a separate outcome. If eight arrived errors total 0.48 m, their mean is 0.06 m, conditional on arrival. It is misleading to report only that mean and hide the two failed missions.

Write the calculation in your notebook before running code. State which values you measured, which you assumed and which the program calculates. A correct numerical calculation cannot rescue an incorrect physical assumption.

Run and explain the model

The following is desktop Python, not a ready-to-run motor program. Download this week’s example, save it in your student repository and run python3 code/w23.py from the repository root. The same small model is reused across the week so you can learn it, build with it, test it and revise it.

# Desktop Python teaching example. Numerical inputs are illustrative.
from statistics import mean
trials = [
    {"arrived": True, "error_m": 0.04, "reason": "arrived"},
    {"arrived": True, "error_m": 0.08, "reason": "arrived"},
    {"arrived": False, "error_m": None, "reason": "hazard"},
]
errors = [t["error_m"] for t in trials if t["arrived"]]
print("attempts", len(trials), "arrivals", len(errors))
print("arrival_fraction", len(errors) / len(trials))
print("mean_error_given_arrival_m", mean(errors) if errors else None)

Compare the unchanged baseline and your proposed change on the same inputs. Keep the original files so another reader can reproduce the comparison. If an exception appears, read its final line, identify the input or assumption that caused it and make the smallest explained correction. Do not delete validation merely to obtain output.

Understanding the model and its limits

The arrivals list filters successful outcomes for the conditional endpoint metric, but the denominator of arrival_fraction remains every attempt. This preserves the stopped failure. Missing endpoint measurement is represented by None rather than zero. In the three-row example, arrivals are 2/3 and mean error conditional on arrival is 0.06 m. A real report also groups conditions and versions; one overall fraction can hide a systematically failing route. A causal diagnosis needs the first divergent log sample or a controlled reproduction, not just a final failure label.

Practical instructions

  1. Choose one justified repair, tag it as a new version and rerun the fault matrix.
  2. Conduct a fresh validation batch with the same acceptance rules if time permits.
  3. Compare versions without pooling incompatible results.
  4. Write the failure report with operating envelope, residual issues and next-test priorities.

Experiment

Replay the failure before/after repair and run fresh held-out trials; if unavailable, state that the repaired version is not yet validated.

Before testing, record your prediction, changed factor, measured response, fixed conditions and stopping rule. Save every attempted run, including failures, with a condition and source version. If hardware is unavailable, use an explicitly labelled synthetic/replay dataset and list the physical question it cannot answer. Do not invent completed trials.

Deliverable

A version-separated repair report and evidence-backed limitations, including unfinished physical validation if applicable.

Save notebook/w23-s4.md, the relevant code revision, raw CSV or test-case records, and one labelled diagram/plot/table. Link the files relatively from your notebook. Use the entry template and report guide.

Completion criteria

A documented failed prediction can meet the learning criteria. A missing physical trial must remain marked untested; software success alone does not validate the robot.

Reading and video

Reflection and next step

Which assumption most affected your result? Point to one observation that supports your explanation and one alternative explanation the evidence has not ruled out. Write a specific next test with a changed factor and measurable outcome, then proceed through the week’s Learn → Build → Experiment → Improve cycle.

← Previous sessionCourse roadmapNext session →