Conversation
Data resources are keyed by full topic id (/sensors/scan, not scan) and the reading is nested under .data, so sections 5-7 printed null for every field. Section 8 read .value from the configurations list, which never carries a value; read each parameter's own detail endpoint instead. The fault collection has no entity_id/code fields, so the snapshot/bulk-data walkthrough always fell back to "skip" - resolve the owning App from the fault's reporting_sources instead. Also drop the Components/Apps columns that read fields the API never returns (area, namespace) for ones that do (description, component id), and fix the diagnostic_bridge id typo in the README (live id is hyphenated). Covered by a new smoke_test.sh section that runs check-demo.sh live against a fault and asserts no null fields.
check-entities.sh and check-faults.sh read fields the SOVD entity and fault responses never carry (area, category, is_located_on, hosted_by, code, reporter_id, message, timestamp), so every labeled field printed null. Read the real fields instead (fault_code, severity_label, reporting_sources, x-medkit.component_id, ...), matching the shape moveit_pick_place/check-faults.sh already uses. setup-triggers.sh and watch-triggers.sh both watched apps/diagnostic-bridge, which reports nothing for this demo - faults arrive from apps/anomaly-detector. Fixed both to the live entity id. The nav-failure inject script never brings a fault to CONFIRMED state under the gateway's own OnChange notification path (verified live: inject-localization-failure fires reliably, inject-nav-failure never does across repeated attempts), so the hinted inject script and README example now point at inject-localization-failure.sh. Covered by a new smoke_test_turtlebot3.sh section that runs both scripts against a live fault and asserts no null fields, and a new smoke_test_navigation.sh section that drives setup-triggers.sh / watch-triggers.sh / inject-localization-failure.sh and asserts the SSE stream delivers an event.
run-demo.sh's text called the simulated robot a TurtleBot3. The image builds the Robotnik RB-Theron description (Dockerfile.gateway), not TurtleBot3.
move-arm.sh always printed a success line and exited 0, even when the controller aborted the goal (the pick-and-place loop competes for the same action). It now parses the action's own final status, exits non-zero and prints a failure line when the goal did not succeed, and ./move-arm.sh demo runs every step regardless of earlier failures while still exiting non-zero overall. The local-vs-container branch also checked only that `ros2 node list` succeeded, which is true even on an empty, disconnected graph; it now confirms the target action is actually listed. The container exec no longer requests a TTY, so the script also works from a pipe or CI. check-entities.sh displayed several fields (component area, app category, app location, function category/host, fault code/reporter) that the API never populates for this demo, always printing null. Those columns are dropped or replaced with the real equivalents (app-to-component links via x-medkit, real fault field names) so the explorer only prints real data. README.md documents the new goal-preemption behavior and corrects the manipulation-monitor entity id in the triggers section (the demo scripts already use the hyphenated id; the docs still had the old underscored one).
The container scripts changed parameters with the ros2 CLI, whose first invocation in a container races an unstarted ros2 daemon and fails with "Node not found". Switch every inject and restore-normal script to the gateway's configuration API (PUT .../configurations/<param>), which has no daemon to warm up, matching how sensor_diagnostics already does this. path-planner blocks its single-threaded executor for the full injected planning delay on every cycle, so its parameter service can stay busy long after the delay was set. restore-normal on the planning ECU retries its writes there until they land. The smoke test verifies the injection took effect through the resulting fault instead of a live parameter read-back, since a read racing that busy window does not reliably recover within any bounded wait, while a later write does. Move the container_scripts COPY in the Dockerfile below the colcon build step: the scripts are not a build input, and putting them last means a script-only change no longer invalidates the compiled packages. Add an end-to-end section to the multi-ECU smoke test that runs every inject and restore-normal script on a fresh stack, verifies the parameters they touch actually change and reset, checks that all injected faults are gone after restore, and drives one real write failure to prove it is reported by name.
… direct reads The check-demo.sh test only looked for null fields, so it passed when sections 5-8 printed nothing, dropped parameters or printed fixed values. It now compares sections 5-7 with direct reads of the same topics, within 8 sigma of the live noise configuration, and requires a second run a second later to print a new sample. Section 8 must list exactly the parameters the configurations endpoint lists, each with the value and ROS type its detail endpoint returns. check-demo.sh printed the constant type "parameter" for every LiDAR parameter. It now prints the ROS type from the detail endpoint. The test strips ANSI colours by piping the script output into sed, which needs no shellcheck disable.
…ion check A LiDAR read carries 360 ranges and 360 intensities, which buried the printed values in the failure line.
The script output is piped straight into sed instead of being captured first and fed back through a here-string, which shellcheck flags as SC2001.
The trigger test ran inject-localization-failure.sh directly and passed on any event, so a wrong inject hint in setup-triggers.sh or an event for another fault went unnoticed. It now runs the command setup-triggers.sh prints and requires an event carrying LOCALIZATION_UNCERTAINTY. The inject leaves AMCL with a uniform particle cloud, and the cleanup only deleted fault records, so a second run on the same stack failed its localization checks. The test now sets the AMCL pose from the spawn point plus odometry, runs restore-normal.sh, forces AMCL updates for two detector intervals without LOCALIZATION_UNCERTAINTY coming back, and drives both goals again, which returns the robot to the spawn point. The README notes that restore-normal.sh does not re-localize AMCL.
…t severity The localization check read LOCALIZATION_UNCERTAINTY as CONFIRMED at ERROR severity from the fault list. A fault keeps the highest severity it ever had, also when a FAILED event reactivates it after a clear, so once the injected localization failure raised it at ERROR, every later WARN crossing during a healthy drive read as ERROR and failed the check. The check now compares AMCL's published position spread with the detector's configured ERROR threshold, which is what the detector compares when it reports at ERROR severity.
The EXIT trap ran resume_pick_place_loop before print_summary, and print_summary takes the script's exit status from $?. The resume call always succeeds, so a run that ended early with a non-zero status and no recorded failure (the gateway never became healthy, for example) was reported as "All smoke tests passed!" and exited 0. The trap now saves $? on entry and restores it right before print_summary. errexit is turned off inside the trap, because under set -e restoring a non-zero status would end the trap before the summary runs.
…lues
The move-arm.sh checks decided the real outcome of a goal from the
script's own exit code, so a script that swapped success and failure
still passed. They now read the status the action client printed itself
("Goal finished with status: ...") and require: SUCCEEDED gives exit 0
and exactly the success line; any other status gives a non-zero exit and
exactly one failure line that names that status.
For ./move-arm.sh demo, each of pick, place and ready (home) must have
exactly one result line, in its own step, matching that step's real
status, and the command must exit non-zero when a step really failed.
The goals are now preempted on purpose instead of by racing two
processes: a probe inside the container waits until the arm controller
and MoveGroup are idle, then answers the next goal the controller starts
with a competing goal. Pausing pick_place_loop now also waits for the arm
to go idle, since a MoveGroup goal sent before the pause still runs.
The check-entities.sh check passed on output that held only the headings.
It now compares what the script prints with direct API reads: every
known component with its name and description, every app and its
component, and each active fault's code, severity, status and sources.
The check-entities.sh and check-faults.sh section injected inject-nav-failure.sh as soon as the gateway answered. On a fresh stack Nav2 is not active yet, bt_navigator rejects the goal and NAVIGATION_GOAL_ABORTED never appears, so the section failed on every fresh start, which is how CI runs it. It now waits up to 180 s for bt-navigator and planner-server to report the active lifecycle state. The file header no longer claims the test injects no fault.
…mulated pose inject-localization-failure re-initialises AMCL to a uniform particle cloud, and restore-normal only cancelled goals, reset velocity limits and cleared faults. AMCL stayed on the scattered cloud, so the fault came back and navigation ran on a wrong pose estimate. restore-normal now reads the robot's pose from the running Gazebo simulation and sets it through AMCL's set_initial_pose operation before it clears the faults. The map frame of the demo is the Gazebo world frame, so no spawn pose is assumed. The script fails when it cannot re-localize. The inject script's comment said the goal is typically rejected; Nav2 accepts it and drives from the scattered estimate. The README describes the new restore step.
…uild The container scripts are not build inputs, but they were copied before colcon build, so any script change rebuilt the workspace layer.
…CL against the simulation The navigation test set the AMCL pose itself before running restore-normal.sh, so it proved the test could recover the demo, not that the demo recovers itself. It now relies on restore-normal.sh alone and compares AMCL's estimate with the robot's pose in Gazebo, within 0.2 m and 0.2 rad, since a cloud that converged on the wrong place is no longer reported as uncertain.
…s it The arm controller can fail to deliver the goal response to a freshly started `ros2 action send_goal` whose endpoints are not yet discovered. It logs "Failed to send goal response ... client will not receive response" and drops the goal without running it, and the CLI then waits for that response forever. move-arm.sh hung there with no result. Each CLI run is now limited to 30 s by `timeout` (inside the container on the docker branch, since killing `docker exec` leaves the CLI running). When a run got no goal response at all, the goal never ran, so the script sends it again, up to three times. An accepted or rejected goal is never sent twice; if its result does not arrive in time the script reports it as failed. PYTHONUNBUFFERED keeps the CLI's lines when the timeout ends it. The smoke test simulates the lost response with a ros2 on PATH that drops the first goal and hands the next to the real CLI in the container, and requires exactly one resend and the real SUCCEEDED result. The unreachable-ros2 check now reads its output from a file instead of piping a large string into `grep -q` under pipefail.
…delay path-planner waited out planning_delay_ms inside its timer callback on a single-threaded executor, so every parameter request queued behind the delayed cycle. The gateway timed those requests out and then reported the node unavailable for its negative-cache window, which made restore-normal after inject-planning-delay take more than a minute. Run the planning timer in its own callback group on a two-thread executor, so the parameter services answer while a cycle waits. The delay still holds back each path and still raises the SLOW_PLANNING diagnostic. A new planning_delay_ms also ends the wait of the cycle in progress, so a restore takes effect at once instead of after one more delayed cycle whose stale path keeps behavior-planner reporting faults after they were cleared.
All six container scripts write parameters through one put_config helper. A refused write is named on stderr together with its HTTP status. The Scripts API reports stderr as the error message of a failed execution and drops stdout, so the failure lines the scripts printed to stdout never reached the caller, who only saw "Script exited with code 1". path-planner now answers parameter requests during an injected delay, so the planning restore-normal no longer retries its writes for minutes. README: document how the scripts change parameters and report a failed write, and drop the troubleshooting note about a slow planning restore.
…essage Read the parameter writes of each container script from its put_config calls and check all of them through the gateway with plain GETs: the written value after each inject, the launch value (ECU params file, else the node declaration) after each restore-normal. Static checks fail when a script writes a parameter in a form the test does not read, when an inject does not change its parameter, when restore-normal does not write it back, and when a container script is not executed by the test. For the planning delay, check the PATH_PLANNER fault, that paths come one delay apart while it is injected and at the planning rate after restore, and that restore-normal completes within 15 s. Run the host wrappers: inject-cascade-failure.sh, then restore-normal.sh must exit 0 within 30 s and leave every parameter at its launch value and no fault. Replace the write-failure check, which ran its own curl snippet in the container, with a committed script run through the Scripts API: another client locks gripper-controller's configurations, actuation restore-normal must fail and its error message must name exactly the two refused writes. The lock is released and the ECU restored afterwards.
One read of the Gazebo pose topic returned the model twice. restore-normal would then post two JSON documents to set_initial_pose and fail, and the navigation test could not parse the simulated pose. Both now use the first match only.
…, time the bound Before every restore-normal whose result is checked, move each parameter it writes away from its launch value through the gateway (a bool flips, a double moves by 0.001, an integer by 1; injected values stay), and check that it is away. The launch-value checks after the restore then fail for any write the script skips, not only for the injected parameters. The write-failure test runs one script. Check that every container script carries the same put_config and the same ERRORS exit check as that script, initialises ERRORS once and makes all its writes before the check. The write-failure test also checks that the refused writes left their parameters unchanged and that the other writes landed, and the test ends by checking that no fault is left. Time the planning restore-normal from before the POST until its completed status is read, and require that time to be within 15 s. The wait runs to four times the bound so a slow restore reports its time. The host restore-normal.sh bound is measured in milliseconds the same way.
…y goal
The demo check counted a step with no final action status as a failed
step and accepted any "Failed: <label> (" line for it, and it discarded
the preempt probe's exit code. A move-arm.sh that printed three failure
lines without sending a goal passed. Every checked goal must now reach a
final status the action client printed itself. The probe must exit 0
after it saw the goal it preempted end with a non-SUCCEEDED status, and
that status must be the one move-arm.sh's own goal reports (the ready
goal, and the demo's pick step).
The fault setup waited for the arm controller's action server and for
each goal response and result with no limit, inside `|| true`, so a
healthy gateway with a missing controller hung the test forever and left
the probe running in the container. The fault setup is now a mode of the
same probe, and every wait in it is bounded: 30 s for the server, 60 s
for an idle arm, 10 s for a goal response or result, and a 240 s
`timeout` around the whole process in the container. A probe that runs
out records a FAIL, and the pick-and-place loop is resumed after it.
move-arm.sh runs are limited to 400 s, check-entities.sh to 120 s, and
the real CLI behind the fake ros2 to 60 s. The fault setup also waits for
a fault that manipulation_monitor reports, not any fault.
…emo.sh /health answers about a second before the gateway links the sensor nodes. check-demo.sh started in that window printed null for every LiDAR, IMU and GPS field and still reported the demonstration complete. It now waits up to 30 s for the sensor data and the LiDAR configurations it prints, and stops with a message and exit code 1 when they do not appear.
…e ranges at default noise A first check-demo.sh run starts as soon as /health answers, polled every 0.2 s, and must print no null and either real values or the stop message with a non-zero exit. The section 5-7 comparisons ran while the injected LiDAR noise was 0.5 m, where 8 sigma covers the whole 0.12..3.5 m range, so invented ranges passed. They now run before the fault injection, at the configured noise, and the LiDAR check fails when 8 sigma is not below a tenth of the range span. The run with the fault active keeps the section 8 and fault detail checks.
…t with direct reads The test only required no null and the injected fault code, so a script printing wrong components or a made-up severity passed. It now compares the printed component ids, each app's component and, for every active fault, the code, severity label, status, sources and occurrence count with direct API reads. The fault list is read before and after each script run and the printed faults must equal one of the two reads.
…live AMCL updates The inject now starts from a parked pose. The test asserts in Gazebo that the parked pose is farther from spawn than the agreement tolerance, so a restore-normal.sh that assumed the spawn pose fails the agreement check. The forced AMCL updates discarded every response. Each update must now return 200 and be followed by an AMCL pose with a newer stamp, or the check that LOCALIZATION_UNCERTAINTY stays absent fails.
The header pointed at a comment in smoke_test_turtlebot3.sh that no longer exists, and that test now injects a navigation failure itself. The reason that holds is that the thresholds count FAILED and PASSED events, and only a direct report sends an exact number of them.
turtlebot3 check-faults.sh and sensor_diagnostics check-demo.sh read /faults without checking the response. While the fault manager is not available the gateway answers 503 with an error body, and both scripts then reported no active faults and exited 0. They now name the failed read with its HTTP status and exit 1.
The smoke tests pause the demo's fault manager with SIGSTOP, so GET /faults answers 503 as it does before the fault manager is up, and run check-faults.sh (turtlebot3) or check-demo.sh (sensor_diagnostics). The script must exit non-zero with the failed-read message and must not claim there are no faults. The fault manager is then resumed and /faults must answer again within 30 s. The EXIT trap also resumes it and keeps the script's exit status.
| printf '%s\n' "${output}" | ||
| # An accepted or rejected goal has its answer. Only a goal that got | ||
| # no response at all never ran, so only that one is sent again. | ||
| if grep -qE '^(Goal accepted with ID|Goal was rejected)' <<< "${output}"; then |
There was a problem hiding this comment.
Everything without a "Goal accepted/rejected" line is treated as a lost goal, including docker: command not found, a wrong CONTAINER and The passed action type is invalid; those fail instantly, yet the loop prints "No goal response within 30 s, sending the goal again" twice and ends with status: UNKNOWN. Resend only when the output shows the CLI got as far as Sending goal:, and report any other output as the failure it is.
| # demo looks identical to being inside the container. Checking that the | ||
| # action itself is listed avoids that false positive. | ||
| can_reach_action_locally() { | ||
| command -v ros2 &> /dev/null && ros2 action list 2> /dev/null | grep -qFx "${ACTION}" |
There was a problem hiding this comment.
Inside the container this hits the cold-start effect from #77 item 2: with no ros2 daemon up, the first ros2 action list answers from a direct node after a short discovery spin and can be empty, so the script falls through to docker exec, which does not exist in the container, while README line 76 still promises it works from inside. Retry the listing once, treat a missing docker as local, or drop the in-container claim.
| } | ||
|
|
||
| if [ "$EARLY_RC" -ne 0 ]; then | ||
| if grep -q "Sensor data not available" <<< "$EARLY_PLAIN"; then |
There was a problem hiding this comment.
This branch also passes when check-demo.sh gives up after its 30 s wait, which is exactly the first-run failure #77 reports, so a gateway that links in 31 s keeps the check green. wait_for_runtime_linking right below proves linking completes, so require EARLY_RC -eq 0 here, or record the linking time and fail when it exceeds DATA_WAIT_SEC.
| if [ "$waited" -ge "$DATA_WAIT_SEC" ]; then | ||
| echo_error "Sensor data not available at ${GATEWAY_URL} after ${DATA_WAIT_SEC}s." | ||
| echo " Check that the sensor nodes are running, then retry." | ||
| exit 1 |
There was a problem hiding this comment.
After inject-failure.sh the IMU stops publishing on purpose, so this gate times out and the script exits before the fault sections, which is the one moment a user runs it to see the fault. Report the missing sensor and continue to the fault sections instead of exiting.
| FIRST_FAULT=$(echo "$FAULTS_JSON" | jq -r '.items[0].fault_code') | ||
| REPORTING_SOURCE=$(echo "$FAULTS_JSON" | jq -r '.items[0].reporting_sources[0] // empty') | ||
| FIRST_ENTITY=$(curl -s "${API_BASE}/apps" | jq -r --arg node "$REPORTING_SOURCE" \ | ||
| '.items[] | select(.["x-medkit"].ros2.node == $node) | .id' | head -n 1) |
There was a problem hiding this comment.
The anomaly detector reports with source_id = "/processing/anomaly_detector/" + source (anomaly_detector_node.cpp:213), e.g. /processing/anomaly_detector/imu_sim, while the app's node path is /processing/anomaly_detector, so this exact match fails for every directly reported IMU/GPS fault and sections 10-12 are skipped when one is first in the list. Match on a node-path prefix with a / boundary ($node == $src or ($src | startswith($node + "/"))).
| angle_max: .data.angle_max, | ||
| range_min: .data.range_min, | ||
| range_max: .data.range_max, | ||
| sample_ranges: .data.ranges[:5] |
There was a problem hiding this comment.
The || echo fallback right below is dead: jq exits 0 on a 404 body, so before the scan is linked this section prints "angle_min": null and the "Gazebo may still be starting" hint never appears. Use jq -e '.data | select(. != null) | {...}' or check %{http_code} so the fallback actually runs.
|
|
||
| echo_step "6. Faults" | ||
| curl -s "${API_BASE}/faults" | jq '.items[] | {code: .code, severity: .severity, reporter: .reporter_id}' | ||
| curl -s "${API_BASE}/faults" | jq '.items[] | {code: .fault_code, severity: .severity_label, status: .status, sources: .reporting_sources}' |
There was a problem hiding this comment.
Same dead fallback in the joint-states block above: a 404 for data/joint_states prints "joint_names": null and the hint never runs, because jq exits 0 on the error body. Use jq -e '.data | select(. != null) | {...}'.
| ros2 param set /actuation/joint_driver failure_probability 0.0 || ERRORS=$((ERRORS + 1)) | ||
| put_config joint-driver inject_overheat false | ||
| put_config joint-driver drift_rate 0.0 | ||
| put_config joint-driver failure_probability 0.0 |
There was a problem hiding this comment.
The writes now fail the script, but the fault clear below still swallows failure (-f ... || true on both DELETEs), so with the fault manager unavailable (the 503 state this PR's own tests drive) the execution completes with "status": "restored" while the ECU's faults stay latched; the turtlebot3 restore in this PR reports that case. Capture %{http_code} of the final DELETE and fail on non-2xx; same in perception-ecu and planning-ecu restore-normal.
|
|
||
| # Terminal 2: Inject a fault - the trigger fires in Terminal 3! | ||
| ./inject-nav-failure.sh | ||
| ./inject-localization-failure.sh |
There was a problem hiding this comment.
"any new or updated faults reported by the anomaly detector" overstates it: faults reported under a source sub-path (NAVIGATION_GOAL_ABORTED / NAVIGATION_GOAL_CANCELED via /goal_status) produce no trigger event, so ./inject-nav-failure.sh stays silent on the watcher. Say so here and name inject-localization-failure.sh as the inject that fires.
restore-normal cleared the ECU's faults with `curl -sf ... || true` twice
and then always printed "restored". While the ECU's fault manager does not
answer, the gateway rejects DELETE /faults with 503, so the execution
completed as restored while the faults stayed.
The second clear now decides the result: a status other than 2xx exits 1
and names the clear on stderr ("FAIL: clear faults (HTTP 503)"), which the
Scripts API returns as the execution error. The first clear stays best
effort. Both clears get a 30 s curl timeout.
The smoke test stops each ECU's fault manager with SIGSTOP and checks,
through the Scripts API and with a direct run of the script, that
restore-normal fails and names the clear. It covers the fault manager
stopped for both clears and stopped only between the first and the final
clear, and that restore-normal completes again once it resumes.
…l ros2 without docker move-arm.sh sent a goal again on any output without an accepted or rejected line, so a failure that never sent the goal (no such container, no action server, an invalid action type) ran three times and printed "No goal response within 30 s" even when the command failed at once. Now only output with "Sending goal:" and no response is sent again, and the resend line names the 30 s wait only when the timeout ended the run. A goal that was never sent is reported once with the command's output and "Failed: <pose> (goal not sent: ...)", and the script exits non-zero. Without a docker CLI the script now uses the local ros2 directly, and it no longer runs `docker ps` before it knows it needs it. Inside the container a cold `ros2 action list` missed the arm action about half the time, the script then fell through to `docker exec`, and every run printed "docker: command not found". Where docker is available, the local-reach check lists the actions a second time before it falls back to `docker exec`. The smoke test runs the not-sent cases for real: a missing container from the host, and in the container no docker and no ros2, a ROS domain without the action server, and a shadowed control_msgs. Each must be reported once, non-zero, with no resend line. A copy of move-arm.sh run in the container with the ros2 daemon stopped before every run must get the goal accepted in all eight runs, with no "command not found" line. Fake ros2 CLIs cover the listing that misses once and a CLI that drops the goal at once after "Sending goal:".
check-entities.sh piped the joint_states read straight into jq. jq exits 0 on an error body, on empty input and on a reply without data, so the "not available" fallback never ran and the script printed "joint_names": null. With joint_state_broadcaster unloaded the gateway answers 200 with an empty data object, which hit exactly that. The script now prints the values only when the reply has joint names, and the hint otherwise. The smoke test unloads joint_state_broadcaster, waits until the gateway serves no joint names, and requires check-entities.sh to finish with the hint in section 5 and no null anywhere; it then loads the controller again and waits for the data to return. With data, section 5 must show the joint names the API returns. The exit trap reloads the controller if a run stops halfway.
… and sub-path fault sources The run started right after /health must exit 0 and print values in sections 5-8; a stop with "Sensor data not available" no longer passes. A new section restarts the demo, fails the IMU before anything reads it and requires check-demo.sh to name the IMU as having no data, print no null, still print LiDAR and GPS values and run sections 9-12. It times the readiness wait at the default and with DATA_WAIT_SEC 0 and 2 against the limit plus one request. Another section reports faults through the fault manager's report_fault service: one from /processing/anomaly_detector/imu_sim, which sections 10-12 must show on apps/anomaly-detector without null, and one from /processing/anomaly_detector_extra, which must not map to any App.
…sub-path fault sources check-demo.sh waited for a message from every sensor and gave up with exit 1, so after inject-failure.sh the fault sections never ran. The wait counted passes, and a read of a sensor without a message blocks for the gateway's sample timeout, so a 30 s wait took 66 s. The wait now uses wall-clock time: up to DATA_WAIT_SEC (default 30, settable) for the gateway to link each sensor node and up to 5 s for its first message, each request bounded to 3 s. It names the sensors it waits for and the ones left without data, whose sections print a message and no null fields. Sections 10-12 now pick the App whose node is the fault's reporting source or a path segment above it, so anomaly detector faults reported as /processing/anomaly_detector/<sensor> get the detail and bulk-data sections. Section 12 says so when no rosbag is listed, and section 8 when no configuration is.
…trigger across a navigation failure With the Gazebo server stopped, /scan keeps its publisher but gets no message. Before anything reads /scan, the test stops the simulator, confirms the scan read has no data, and requires check-entities.sh to print the LiDAR hint and no null. The later run with data must print the scan's angle and range limits as a direct read returns them. The run also creates the trigger with setup-triggers.sh and watches it with watch-triggers.sh across inject-nav-failure. A fault reported under the anomaly detector's own node path must reach the watcher, and no NAVIGATION_GOAL_* event may, as the README states.
…tatus The watcher that stops a fault manager between the two clears exited 1 when restore-normal ended without a pause it saw, and under errexit the bare `wait` on it aborted the suite before the setup failure was recorded, skipping the remaining ECUs and sections. The miss is now recorded as that setup failure and the run continues; the fault manager is resumed either way. The watcher prints "armed" before it polls and the script starts only after that, and it exits as soon as the script ends without a pause. The failure message check accepted any three digits, so a message naming a status the clear never got passed. The test now measures what a fault clear gets from the ECU gateway in the same state, requires it to be non-2xx, and requires the error and the stderr of a direct run to name exactly that status.
When only the second clear fails, the first one may already have removed the ECU's faults, so "the faults stay" was not always true. The README now says that the second clear decides the result, that the error names the status it got, and that the ECU may still hold faults.
jq exits 0 on an error body and on empty input, so the fallback after `|| echo` never ran and a scan read without data printed "angle_min": null. check-entities.sh now prints the scan values only when the read carries data, and the hint otherwise.
The README said the trigger fires on any new or updated fault. It fires for the localization fault inject-localization-failure.sh causes, and not for the navigation goal faults inject-nav-failure.sh causes, which only show in the fault list.
The in-container move-arm.sh runs stopped the ros2 daemon and ignored the result, so a run with the daemon still up counted as a cold run. A warm daemon lets an implementation with a single `ros2 action list` pass. Each run now records `ros2 daemon status` in a file in the container; move-arm.sh starts only if it reports the daemon not running, and a run without that proof is recorded as a failure. run_move_arm_inside runs move-arm.sh only when its setup succeeds. The check-entities.sh check with data compared only the joint names. It now also requires positions and velocities as arrays of numbers, one per joint name the API returns. The values move with the arm, so they are checked by type and count only.
…adiness wait is needed on every run The first check-demo.sh run after /health races the linking, and on a fast start a copy without any wait finds all data at its first read and passes. A new section restarts the demo, stops lidar_sim with SIGSTOP before anything reads its topic and resumes it 2 s into the run. The run must exit 0, print values in sections 5-8 and say it waits for lidar-sim. The EXIT trap resumes the node as well. The DATA_WAIT_SEC sweep adds 08, a whole number of seconds with a leading zero, which must wait at least 3 s and at most 8 s plus one request.
The digit-only check accepts 08, and bash then reads it as an invalid octal number: the deadline stayed empty and check-demo.sh skipped the readiness wait. The value is now converted in base 10 after the check.
…n inject and require a successful scan read The watcher was guarded only by a 2 s sleep before inject-nav-failure, and its liveness was proven after the navigation fault. It is now proven before the inject as well: a fault reported from the anomaly detector's own node path is repeated every 2 s until its event arrives. Only a watcher live before and after the inject counts for "no NAVIGATION_GOAL_* event"; otherwise that check fails as not checked. The no-scan setup discarded the HTTP status, and an error body has no .data either, so a failed read counted as the no-data state. The setup now requires a successful read with an empty data object.
… no-data setup The README says an accepted goal is never sent again and a missing result ends as "Failed: <pose> (status: UNKNOWN)", but no test reached that path: every fake lost the goal before "Goal accepted with ID", and every real accepted goal reached a final status. The fake ros2 has a mode that prints "Goal accepted with ID" and then hangs; move-arm.sh must send the goal once, print no resend line, and exit non-zero with status UNKNOWN. The joint_states no-data setup accepted any read without joint names, including a 404, a 500 or a connection failure, so the hint check could run against an error response. It now passes only when the read answers HTTP 200 with an empty data object, and it reports the status and body otherwise.
…w and resume one while the wait goes on The link wait was never engaged: natural linking takes about 2 s and the held LiDAR stays linked, so a script that waits only 5 s passed. A new section restarts the demo, holds the linked IMU, takes lidar_sim off the ROS graph and checks the gateway has unlinked it. The IMU resumes 7 s and lidar_sim returns with the launch's parameters 9 s into the run. The run must wait for the lidar-sim link, exit 0 and print values in sections 5-8, and must not report the IMU without a message once it sent data while the wait went on. The EXIT trap resumes and restores both nodes. The failed-IMU run also checks that the time printed for the IMU matches the measured wait. The SIGSTOP helpers now take the node name.
…message window A sensor was dropped once its 5 s window ended, even while the wait went on for another sensor. The summary then said it had no message for the whole wait and that its data section showed no values, and the section, which reads it again, printed values when it had started in the meantime. Such a sensor stops holding the wait and is still read on each pass, so a late first message is noticed. The summary gives the time each sensor was read for and says the data sections read every sensor again.
…ts use Under smoke_lib.sh's `set -euo pipefail`, a reader that exits before the end of its input (awk `exit`, `head`) sends SIGPIPE to the writer, the pipeline returns 141, and an assignment from it ends the suite. The sensor smoke test died this way in readiness_wait_output, an awk that stops at section 1 behind the full run output. readiness_wait_output, the reported-IMU-seconds read and the TurtleBot3 trigger id read now use one awk on a here-string, so no writer sits behind the early exit. The early readers left are in fail and echo message arguments, whose status is not used.
Description
The helper scripts a user runs by hand printed
null, failed on the first run after start, or watched the wrong entity. CI did not see this, because the smoke tests only exercised the gateway API. This PR fixes the scripts and adds checks to the existing smoke tests so each problem fails CI.sensor_diagnostics
check-demo.shreads the data resources the sensors really publish (/sensors/scan,/sensors/imu,/sensors/fix) and the message under.data. Section 8 prints each parameter with its value and type from the configuration detail./faultsread exits non-zero and does not say "no faults".multi_ecu_aggregation
PUT /apps/<app>/configurations/<param>). The firstros2 param setin a fresh container failed withNode not found.error.message.path_plannerruns its planning timer in its own callback group on a two-thread executor. Its parameter service now answers during an injected delay, so the planningrestore-normaltakes about 3 s (75 s before).container_scripts/after the workspace build, so a script change does not rebuild the workspace.moveit_pick_place
move-arm.shuses the local ROS 2 CLI only when the arm action is visible locally. Before,ros2 node listreturning 0 on an empty graph sent every host with ROS 2 sourced down the local path.docker execno longer needs a TTY.check-entities.shandcheck-faults.shshow real values, andcheck-faults.shreports a failed fault read.turtlebot3_integration
check-entities.shandcheck-faults.shshow real values and report a failed fault read. The README API examples print real values.setup-triggers.shandwatch-triggers.shwatchanomaly-detector, which reports the navigation faults. The hint is nowinject-localization-failure.sh. The gateway does not send trigger events for faults reported from a sub-path of a node (/bridge/anomaly_detector/goal_status), soinject-nav-failure.shproduces no event. That is a gateway issue and is not changed here.restore-normalputs AMCL back on the robot's pose read from Gazebo, so localization works again afterinject-localization-failure. It fails when a write is refused.container_scripts/after the workspace build.ota_nav2_sensor_fix
run-demo.shnames the robot the image builds (RB-Theron).Tests
The new checks compare script output with direct API reads, run the multi-ECU scripts as the first ones on a fresh stack, and drive real failures (a configuration lock that answers 409, a stopped fault manager that answers 503). Every new check was first run against the old scripts or a broken copy of the new ones, and failed. The EXIT traps that resume paused processes keep the real exit status, so a gateway that never starts still fails the job.
Related Issue
closes #77
Checklist