RL Robotics LabSearch ↗
Lesson / authored

Week 21 · Experiment: A test plan distinguishes development from evidence

Session 3 of 4 · Define the mission · Phase 6

Plan about 15 minutes for explanation, 30 minutes for practical work and 10–15 minutes for documentation. A longer build may continue into the next session: stop safely, commit the current state and record the next check. Desktop simulations count as software evidence; label them clearly and record physical validation separately.

Engineering challenge

What exactly will your robot prove, and how will someone else check it? This session focuses on a test plan distinguishes development from evidence.

Before you start

The previous week’s recorded baseline and Week 20, Improve. For later sessions this week, retain the preceding session’s files and predictions.

Equipment: Desktop Python, editor, paper and ruler; for physical work, the configured 3pi+ 2040, clear floor mat and hardware checklist. Week 3 additionally uses the separate low-voltage LED circuit described in its procedure.

For any motion, verify the stop button, short time limit and clear floor area first. Keep the wheels raised for a new device program until its commands and stop behaviour are checked. A hazard or uncertain input is a reason to stop and document, not to force the trial to finish.

Theory and mathematics

A test plan distinguishes development from evidence

Development trials guide changes; validation trials estimate behaviour with the design frozen. Create a requirements-to-tests table before either. Each test needs a starting setup, input condition, expected result and measurement method. Include stopping and no-route cases alongside successful navigation. Decide how to record an interrupted trial before running it. If software changes during validation, start a new batch. This preserves the meaning of the result and prevents quietly mixing different systems.

Worked example — illustrative values

An illustrative mission requires endpoint error ≤0.10 m, elapsed time ≤60 s and zero boundary breaches on a 1.2 m mapped route. A run ending 0.07 m away in 42 s passes those two numerical checks, but fails overall if it crosses a boundary. A safe timeout stop is a successful guard test and a failed navigation run.

Write the calculation in your notebook before running code. State which values you measured, which you assumed and which the program calculates. A correct numerical calculation cannot rescue an incorrect physical assumption.

Run and explain the model

The following is desktop Python, not a ready-to-run motor program. Download this week’s example, save it in your student repository and run python3 code/w21.py from the repository root. The same small model is reused across the week so you can learn it, build with it, test it and revise it.

# Desktop Python teaching example. Numerical inputs are illustrative.
def evaluate(error_m, elapsed_s, boundary_breaches, arrived):
    navigation_pass = arrived and error_m <= 0.10 and elapsed_s <= 60 and boundary_breaches == 0
    return {"navigation_pass": navigation_pass,
            "error_m": error_m, "elapsed_s": elapsed_s,
            "boundary_breaches": boundary_breaches}

print(evaluate(0.07, 42, 0, True))
print(evaluate(0.07, 42, 1, True))
print(evaluate(0.07, 60, 0, False))

Use your own recorded data or named test fixtures instead of the illustrative inputs. Save expected and actual values side by side. If an exception appears, read its final line, identify the input or assumption that caused it and make the smallest explained correction. Do not delete validation merely to obtain output.

Understanding the model and its limits

The acceptance function combines arrived state, independent endpoint error, elapsed time and boundary count with logical AND. Every required condition must hold. A timeout can be a correct guard response while failing navigation because arrived is false. Separate guard-test outcomes from the mission success table. Before coding this rule, classify several fictional rows manually and resolve ambiguities such as missing measurements. A missing endpoint must not silently become zero error. The final criteria should reflect the surveyed course and your demonstrated capabilities, rather than a target chosen after seeing the results.

Practical instructions

  1. Pilot the measurement procedure without changing final success rules.
  2. Create a requirements-to-tests table with at least six normal/fault conditions.
  3. Have another person measure the same stopped pose and compare ruler readings.
  4. Set the final batch size and distinguish development data from held-out evaluation.

Experiment

Compare two observers on three stopped positions and note disagreement in the measurement procedure.

Before testing, record your prediction, changed factor, measured response, fixed conditions and stopping rule. Save every attempted run, including failures, with a condition and source version. If hardware is unavailable, use an explicitly labelled synthetic/replay dataset and list the physical question it cannot answer. Do not invent completed trials.

Deliverable

A six-condition test plan and a reproducible independent measurement method.

Save notebook/w21-s3.md, the relevant code revision, raw CSV or test-case records, and one labelled diagram/plot/table. Link the files relatively from your notebook. Use the entry template and report guide.

Completion criteria

A documented failed prediction can meet the learning criteria. A missing physical trial must remain marked untested; software success alone does not validate the robot.

Reading and video

Reflection and next step

Which assumption most affected your result? Point to one observation that supports your explanation and one alternative explanation the evidence has not ruled out. Write a specific next test with a changed factor and measurable outcome, then proceed through the week’s Learn → Build → Experiment → Improve cycle.

← Previous sessionCourse roadmapNext session →