← All work

MOOS-IvP CI/CD

2026

MOOS-IvP is an open-source framework for robot autonomy. It had many ways to build software, but not one dependable pipeline for checking whether a change still behaved correctly in a running mission. I built a continuous integration and delivery (CI/CD) pipeline to automate those checks and catch regressions, bugs introduced by new changes, before they reached someone’s boat or field test.

Five sets of simulated mission tests in the MOOS-IvP testing matrix
A condensed view of the testing matrix. Each entry summarizes a set of checks performed during simulated missions.

From a build check to a behavior check

I built two complementary testing layers. Fast C++ unit tests check individual pieces of code, such as geometry calculations and configuration parsing. Mission harnesses, reusable setups for running and checking a mission, then launch the actual robot software in controlled simulations and check what happens over time.

The mission layer uses reusable starting missions with small changes for each case. A case can move a simulated vehicle, wait for messages or flags, and produce a pass or fail result with logs that make the outcome inspectable. That lets me reproduce failures that would be invisible to a unit test or a build matrix alone.

Each harness run gets its own working directory, and each case gets a copy of the mission and its own ports. That keeps one case’s generated files and processes from becoming another case’s hidden inputs. I also scoped cleanup to the processes started by that run, with failed cleanup reported and the run preserved for diagnosis.

A failure that looked like a pass

Developing new test cases has helped us find and fix many bugs already present in MOOS-IvP. One example was a behavior, a module that guides a vehicle’s decisions, intended to pause its decision-making loop for two seconds. A mission test measured closer to four: the completion function could trigger the delay twice. A single-trigger guard addressed the cause in the proposed upstream fix.

The test measures the actual pause rather than trusting the behavior’s status message, and checks it at several simulation speeds. Other cases exposed a coordinate-conversion object that failed when copied and a documented viewer setting that never reached the graphical interface at startup. Each case gives us a repeatable way to check the correction and catch the bug if it returns.

The two-second stall regression
Requested stall2 seconds
Gap between autonomy updatesAbout 4 seconds
Accepted test band1.6–2.4 seconds
CauseThe function handling completion could trigger the delay twice.
CorrectionA single-trigger guard and steadily advancing timer, checked at multiple simulation speeds.

Source revisions and reproduction steps. This is a specific regression example, not a live status display.

A workflow for other contributors

The pipeline’s ongoing role is to catch regressions before proposed changes are accepted. I developed the current testing system independently and discussed it with core MOOS-IvP developers as it grew. Other contributors have asked me to run the suite against their proposed changes, which has helped me refine both the tests and the fixes under review. The aim is for a developer to run a focused case locally, see the same result in the automated pipeline, and understand what failed without recreating an entire field setup.

The pipeline also checks native macOS and containerized Linux builds for compatibility across supported environments. Those build checks run alongside the unit and mission tests.

← All work