Why I made the workflows
I started this project because I expect coding agents to become a routine part of robotics development, and I wanted MOOS-IvP to be ready for that shift. They can already inspect a repository, change its code, and run tests, but using that freedom well depends on understanding the system they’re working in. In MOOS-IvP, a convincing C++ implementation can still miss the mission setup, vehicle messaging, or checks needed to show that the robot behaves as intended.
My aim was to make the community’s engineering knowledge reusable across projects and agents, rather than relying on someone to explain the same conventions each time. I put that knowledge into workflows for finding the right documentation, using established patterns, and checking the result. As coding tools improve, I want that guidance to help new users get started and give experienced developers more time to work on the autonomy itself.
The ten connected workflows cover apps, behaviors that guide vehicle decisions, missions, evaluation, test harnesses, maps, installation, documentation, and log analysis. They’re meant to lead an agent through the whole job. For example, it can build a normal mission first, add checks that judge the outcome, run a harness that varies the mission across test cases, and inspect the logs when the result doesn’t match its expectation.
The handoffs are a central part of the design. A mission workflow can ask for documentation before choosing a parameter, or hand a runtime problem to log analysis rather than guessing from configuration alone. Evaluation adds grading to an ordinary working mission; a harness then varies that mission across cases. Keeping those responsibilities separate gives the agent a clearer idea of what to do next and what evidence the previous step actually supplied.
The packages include reference material, reusable templates, and scripts as well as instructions. That lets a workflow provide an established launcher or a repeatable check instead of asking the model to reconstruct everything from prose. I maintain agent-neutral sources and generate self-contained packages for Codex and Claude Code, so improvements can carry across tools without maintaining two different versions of the engineering guidance.
Testing the idea
I ran a controlled comparison of agent attempts with and without the plugin across 11 MOOS-IvP tasks, four models, and repeated runs. Each result was checked for both completion and engineering practice. Those are different measures: an app can satisfy the immediate request while still missing the help text, launch structure, or validation a MOOS developer would expect.
Across that fixed benchmark, completion and engineering conformance improved overall with Skills available. One multi-case harness task completed less often in the guided condition, and some agents had access to the right workflow without applying it well. Those failures show where the guidance or its use needs further work.
The comparison contains 440 attempts: five runs per task, model, and condition, each in a fresh workspace. Both conditions had the same toolchain, public internet access, and three-hour limit. The prompts asked for the result without revealing the grading criteria. An attempt counted as complete only when every required part worked and could be checked from the delivered evidence.
For example, one task asks an agent to build an app that follows another moving vehicle, called a contact. Completion requires it to track that vehicle, update its destination point, handle missing or invalid reports, and show the operator what it’s doing. Conformance separately checks the engineering practices a MOOS-IvP developer would expect, including help text, launch structure, and validation.
| Measure | Without Skills | With Skills |
|---|---|---|
| Complete attempts | 59.5% | 75.5% |
| Mean conformance | 56.2% | 85.7% |
220 attempts per condition. Completion means satisfying the full request; conformance measures adherence to the tested engineering practices. These aggregate gains don’t hold for every task or model. Methods, per-model results, and limitations
Where it fits
I maintain the portable skill sources and package them for coding agents. The project article walks through how the workflows fit together, while the benchmark article has the full methods, results, and limitations. I see the plugin as a way to make specialized knowledge easier to use and test, not as a substitute for an engineer who knows the robot and checks what the agent actually delivered.
