Week 12 · Learn: Training and validation are different jobs
Session 1 of 4 · Tune and validate · Phase 3
Plan about 15 minutes for explanation, 30 minutes for practical work and 10–15 minutes for documentation. A longer build may continue into the next session: stop safely, commit the current state and record the next check. Desktop simulations count as software evidence; label them clearly and record physical validation separately.
Engineering challenge
Does the chosen controller work on a case it was not tuned to pass? This session focuses on training and validation are different jobs.
Before you start
The previous week’s recorded baseline and Week 11, Improve. For later sessions this week, retain the preceding session’s files and predictions.
Equipment: Desktop Python, editor, paper and ruler; for physical work, the configured 3pi+ 2040, clear floor mat and hardware checklist. Week 3 additionally uses the separate low-voltage LED circuit described in its procedure.
For any motion, verify the stop button, short time limit and clear floor area first. Keep the wheels raised for a new device program until its commands and stop behaviour are checked. A hazard or uncertain input is a reason to stop and document, not to force the trial to finish.
Theory and mathematics
Training and validation are different jobs
Tuning uses data to choose gains. Validation tests the frozen choice on conditions not used for tuning. If you retune during validation, that run becomes more training data. Select an attainable reference and a modest disturbance as training cases, then reserve another reference or surface for validation. A motor plant with lag, caps and noise is a model with assumptions, not a digital twin. Use it to investigate controller reasoning; hardware trials still need independent measurements and conservative limits.
Worked example — illustrative values
For fictional post-transient errors 0.01, −0.02, 0.00 and 0.01 m/s, mean absolute error is 0.01 m/s. If peak speed is 0.12 m/s for a 0.10 target, overshoot is 100×(0.12−0.10)/0.10 = 20%. Define the window and nonzero target before using either metric.
Write the calculation in your notebook before running code. State which values you measured, which you assumed and which the program calculates. A correct numerical calculation cannot rescue an incorrect physical assumption.
Run and explain the model
The following is desktop Python, not a ready-to-run motor program. Download this week’s example, save it in your student repository and run python3 code/w12.py from the repository root. The same small model is reused across the week so you can learn it, build with it, test it and revise it.
# Desktop Python teaching example. Numerical inputs are illustrative.
errors = [0.01, -0.02, 0.0, 0.01] # fictional
mae = sum(abs(error) for error in errors) / len(errors)
reference, peak = 0.10, 0.12
overshoot_percent = max(0, peak - reference) / reference * 100
print("mae_m_s", mae, "overshoot_percent", overshoot_percent)
Predict the example’s output by hand. Mark the inputs, units and assumptions; explain where this model could fail. If an exception appears, read its final line, identify the input or assumption that caused it and make the smallest explained correction. Do not delete validation merely to obtain output.
Understanding the model and its limits
The metric code illustrates two different questions: mean absolute tracking error and peak overshoot relative to a reference. It does not run a controller; use the traces saved in Weeks 9–11. Overshoot is zero if the peak does not exceed the target. A zero reference needs a different metric because the percentage formula would divide by zero. Declare a settling band and duration before examining your final trace. Count stopped or invalid runs separately rather than assigning them an attractive small error. The chosen controller must pass the held-out route with its settings unchanged.
Practical instructions
- Write three controller requirements including accuracy and stop guardrails.
- Separate two training cases from two validation cases in a test table.
- Define transient exclusion time, error metric and overshoot formula.
- Predict the effect of noise and an unreachable target before running.
Experiment
Calculate metrics for a trace that meets accuracy but exceeds command jitter; decide pass/fail under the full set of requirements.
Before testing, record your prediction, changed factor, measured response, fixed conditions and stopping rule. Save every attempted run, including failures, with a condition and source version. If hardware is unavailable, use an explicitly labelled synthetic/replay dataset and list the physical question it cannot answer. Do not invent completed trials.
Deliverable
Three predeclared measurable requirements, training/validation split and unambiguous metric definitions.
Save notebook/w12-s1.md, the relevant code revision, raw CSV or test-case records, and one labelled diagram/plot/table. Link the files relatively from your notebook. Use the entry template and report guide.
Completion criteria
- Explain training and validation are different jobs in your own words using this session’s example and its units/assumptions.
- Produce the specific evidence above: Three predeclared measurable requirements, training/validation split and unambiguous metric definitions.
- Keep predictions and raw outcomes, distinguish observations from interpretation, and explain one limitation or unresolved failure.
- Review the Git diff, commit the session’s intended files and state the next experiment or safe continuation point.
A documented failed prediction can meet the learning criteria. A missing physical trial must remain marked untested; software success alone does not validate the robot.
Reading and video
- Focused reading: Robotics Lab: scientific testing guide. Study task: Write a tuning/validation split and predeclare a settling criterion.
- Video/lecture option: MathWorks: What is PID control?. Study task: Draw the closed-loop signal path and label the measured quantity. Watch a relevant 5–10 minute excerpt or use the linked notes if video is inaccessible. This is supporting conceptual material; hardware in a demonstration may differ from yours.
- Practical reference: Engineering handbook and hardware setup. Manufacturer/API references and video metadata were checked on 2026-10-09; recheck the actual firmware before transferring code.
Reflection and next step
Which assumption most affected your result? Point to one observation that supports your explanation and one alternative explanation the evidence has not ruled out. Write a specific next test with a changed factor and measurable outcome, then proceed through the week’s Learn → Build → Experiment → Improve cycle.