Process Orchestration (s10)

Register external processes with world.recipe(...) and Elodin starts them in order, gates the sim on their readiness, and tears them down with it.

A software-in-the-loop (SITL) or hardware-in-the-loop (HITL) simulation is rarely a single process. Flight software, SITL firmware, controllers, link bridges, and video receivers all run next to the physics. Each one needs to start in the right order and find its ports, and each one must stop when the simulation stops.

s10 is Elodin's process orchestrator. You describe each external program as a recipe. When you launch the simulation, s10:

  1. starts every recipe, building Cargo recipes first,
  2. waits for each recipe's readiness probe before starting anything that depends on it, including the simulation,
  3. prefixes every line of each process's output with the recipe name,
  4. stops everything when the simulation exits or a recipe fails.
flowchart LR
  Launch["elodin editor / run / monte-carlo"] --> Plan["python main.py plan"]
  Plan --> Group["s10 group: recipes + sim"]
  Group --> Ready["Recipes start, readiness probes pass"]
  Ready --> Sim["Simulation starts"]
  Sim --> Teardown["Simulation exits: SIGTERM, 2 s, SIGKILL"]

When to use a recipe

You want to…Use
Exchange data with external code in lockstep every tickAn in-process bridge in pre_step / post_step with StepContext
Have Elodin launch and clean up an external binary next to the simAn s10 recipe (this page)
Connect a separately managed program or a hardware rig over the networkAn Impeller2 / gRPC client talking to Elodin DB

These approaches combine. A typical SITL setup launches the flight software as a recipe and talks to it through a socket opened in post_step.

Quick start

import elodin as el
from pathlib import Path

world = el.World()
# ... entities, systems, schematic ...

world.recipe(
    el.s10.PyRecipe.process(
        name="controller",
        cmd=str(Path(__file__).parent / "controller.sh"),
        ready=el.s10.Ready.tcp("127.0.0.1:9000"),
    )
)

world.run(system, simulation_rate=1000.0)
elodin editor main.py   # interactive
elodin run main.py      # headless, exits with the sim's status

The controller starts first. The simulation starts once something accepts connections on 127.0.0.1:9000. Controller output appears in the terminal prefixed with controller, and the controller is stopped when the simulation exits.

Recipe types

el.s10.PyRecipe.process(...) runs any command. el.s10.PyRecipe.cargo(...) builds a Rust crate with cargo build, then runs the resulting binary. Both accept the same options:

ParameterTypeDefaultMeaning
namestrrequiredLog prefix, and the key other recipes use in depends_on
cmd (process)strrequiredExecutable path or command name on PATH
path (cargo)strrequiredCrate directory or Cargo.toml path
package, bin (cargo)strNoneSelect a package or binary in a workspace
argslist[str][]Command-line arguments
cwdstrinheritedWorking directory
envdict[str, str]{}Extra environment variables, layered on the inherited environment
restart_policyel.s10.RestartPolicyNeverNever or Instant (respawn on every exit)
depends_onlist[str][]Recipe names that must be ready before this one starts
readyel.s10.ReadyNoneReadiness probe (see below)
ready_timeoutstr"30s"Probe deadline: "500ms", "10s", "2m", "1h"
silenceboolFalseDiscard the process's stdout and stderr

Recipes are registered with world.recipe(recipe), keyed by name. See the Python API for signatures.

Readiness and ordering

Readiness probes

A recipe without a probe counts as ready as soon as it is spawned. A probe makes s10 wait for a real signal. Probes are checked every 50 ms until they pass or ready_timeout expires. A timeout is an error and stops the whole run.

ProbePasses when
el.s10.Ready.tcp("127.0.0.1:9000")A TCP connection to the address succeeds. Use an IP literal, not a hostname.
el.s10.Ready.unix("/tmp/app.sock")A Unix socket connection succeeds
el.s10.Ready.file("/tmp/app.ready")The path exists
el.s10.Ready.delay(100)The given number of milliseconds have passed since spawn (not bounded by ready_timeout)
el.s10.Ready.log("listening")A line containing the pattern appears in the recipe's log. Only works in elodin monte-carlo campaigns, where s10 writes per-recipe log files.

For cargo recipes, the build finishes before the readiness clock starts, so a slow compile doesn't use up delay or ready_timeout.

The simulation waits for every probed recipe

When Elodin plans the run, it makes the simulation depend on every recipe that has a readiness probe. You don't list the simulation anywhere. Add probes to the services you want to wait for.

Several services in parallel

world.recipe(el.s10.PyRecipe.process(
    name="telemetry", cmd="./run-telemetry.sh",
    ready=el.s10.Ready.tcp("127.0.0.1:5000"),
))
world.recipe(el.s10.PyRecipe.process(
    name="video", cmd="./run-video.sh",
    ready=el.s10.Ready.unix("/tmp/video.sock"),
))

Both start at once, and the simulation starts after both report ready.

Services that must start one at a time

world.recipe(el.s10.PyRecipe.process(
    name="router", cmd="./run-router.sh",
    ready=el.s10.Ready.unix("/tmp/router.sock"),
))
world.recipe(el.s10.PyRecipe.process(
    name="flight-software", cmd="./run-flight-software.sh",
    depends_on=["router"],
    ready=el.s10.Ready.tcp("127.0.0.1:9005"),
))

flight-software starts only after router is ready, and the simulation starts after both. Give the earlier service a probe: without one, depends_on only waits for it to be spawned. A cargo recipe with depends_on also waits to build until its dependency is ready.

If a dependency exits or times out before it is ready, its dependents fail immediately with dependency "router" exited before "flight-software" became ready, and the run stops. s10 does not detect dependency cycles, so a cycle waits forever.

Run a setup script to completion first

Readiness means "up", not "finished". To hold the simulation until a script completes, have the script create a marker file as its last step, and probe for that file:

import tempfile

setup_done = Path(tempfile.mkdtemp(prefix="sim-setup-")) / "done"

world.recipe(el.s10.PyRecipe.process(
    name="setup",
    cmd="bash",
    args=["-c", f'"$0" && touch "{setup_done}" && sleep 1', str(Path(__file__).parent / "setup.sh")],
    ready=el.s10.Ready.file(str(setup_done)),
    ready_timeout="5m",
))
  • && means a failing setup.sh never creates the marker, so the simulation never starts and the run fails.
  • The sleep 1 gives the 50 ms probe time to see the marker. A process that exits before its probe passes counts as having died before becoming ready.
  • A fresh temporary directory prevents a marker left over from an earlier run from passing the probe immediately.

How you launch changes the behavior

LaunchWhat happens
elodin editor main.pyPlans the run, starts recipes and the sim, and watches files. Saving a file in the sim's directory restarts the sim, and saving a file in a cargo recipe's crate rebuilds and restarts that recipe. Process recipes are not restarted.
elodin run main.pySame plan, run once. Exits with the simulation's status.
elodin monte-carlo run main.py …Plans once, then patches the plan for each run with per-worker ports and environment. See SITL patterns.
python main.py runNo plan. The Python process is the simulation, and recipes run on a background thread. depends_on ordering between recipes still applies, but the simulation does not wait for readiness probes.

Prefer the elodin CLI for anything that depends on startup order.

Lifecycle and shutdown

EventEffect
The simulation exits, successfully or notEvery recipe is stopped
A recipe fails: spawn error, build error, readiness timeout, or a dependency that died before becoming readyEvery recipe and the simulation are stopped (under python main.py run, the simulation keeps running)
A recipe exits on its own after becoming readyNothing else stops (with restart_policy=Instant, it respawns)
The editor closes, you press Ctrl-C, or a campaign run times outEverything is stopped

"Stopped" means SIGTERM, up to 2 seconds to exit, then SIGKILL. On Linux, the editor and campaign runner also place the whole stack in a cgroup and clean it up at the end. That catches daemonized grandchildren that would otherwise keep ports bound.

ctx.stop_recipes() lets a step callback stop recipes before the simulation finishes. It only has an effect with python main.py run. Under the elodin CLI the simulation runs with --no-s10, so the call does nothing, and the simulation exiting already stops every recipe.

SITL patterns

Build from source in the editor, prebuilt in campaigns

Rebuilding the flight software from source is convenient while editing. It is wasteful in a campaign, where every worker would run cargo build. While the campaign runner plans, it sets ELODIN_MONTE_CARLO_PLANNING=1, so a simulation can switch to a binary built once by the campaign's [[build]] step:

controller_dir = Path(__file__).parent / "controller"
if os.environ.get("ELODIN_MONTE_CARLO_PLANNING") == "1":
    controller = el.s10.PyRecipe.process(
        name="controller",
        cmd=str(controller_dir / "target" / "release" / "controller"),
        ready=el.s10.Ready.delay(100),
    )
else:
    controller = el.s10.PyRecipe.cargo(
        name="controller",
        path=str(controller_dir),
        ready=el.s10.Ready.delay(100),
    )
world.recipe(controller)

falcon9, apollo-lander, and rc-jet all use this pattern.

Per-worker ports

Parallel campaign runs can't share fixed ports. Declare named ports in campaign.toml under [resources.ports]. The runner gives each worker its own slot and exports it as ELODIN_MC_PORT_<NAME>. There are two ways to use it:

From Python, el.monte_carlo.port("state", 9013) returns the worker's port in a campaign and the default otherwise. Pass the value to the recipe through env:

state_port = el.monte_carlo.port("state", 9013)
controller = el.s10.PyRecipe.cargo(
    name="controller",
    path=str(controller_dir),
    env={"ELODIN_MC_PORT_STATE": str(state_port)},
)

At planning time this bakes in the default port. For each run, the campaign runner overwrites ELODIN_MC_PORT_* in every recipe's environment with that worker's values, so the process always sees its own port.

Inside recipe arguments, s10 expands ${NAME} and ${NAME:-default} placeholders when it spawns the process:

el.s10.PyRecipe.process(
    name="controller",
    cmd="./controller",
    args=["--port", "${ELODIN_MC_PORT_CONTROLLER:-31337}"],
    ready=el.s10.Ready.unix("${ELODIN_MONTE_CARLO_RUN_DIR:-/tmp}/controller.sock"),
)

Placeholders are expanded in args, cwd, and probe addresses and paths. They are not expanded in cmd or in env values. $$ produces a literal $. For a tool that reads its port from an environment variable, export it through a small shell wrapper:

el.s10.PyRecipe.process(
    name="controller",
    cmd="sh",
    args=[
        "-c",
        'export CONTROLLER_PORT="${ELODIN_MC_PORT_CONTROLLER:-31337}"; exec "$@"',
        "controller-wrapper",  # becomes $0
        "./controller", "--config", "sitl.toml",
    ],
)

s10 fills in the ${...} value. "$@" has no braces, so s10 passes it through for the shell to expand.

In campaigns, every recipe also gets fail_on_error turned on, so a non-zero exit fails the run. Each recipe's output goes to runs/<run_id>/logs/<recipe>.log.

HITL patterns

In a HITL rig, the flight software runs on the vehicle hardware, not on the host. s10 manages the host-side companions: link bridges (serial, radio, or UDP), log and video receivers, and operator tools. The same readiness rules let you hold the simulation until the link to the hardware is actually up:

world.recipe(el.s10.PyRecipe.process(
    name="radio-bridge",
    cmd="./radio-bridge",
    args=["--device", "/dev/ttyACM0", "--listen", "127.0.0.1:9100"],
    ready=el.s10.Ready.tcp("127.0.0.1:9100"),
    ready_timeout="60s",
))
  • Gate on the link, not on spawn. Probe the bridge's socket so the simulation never ticks against a link that isn't up yet. If the bridge itself waits for a device node to appear, Ready.file("/dev/ttyACM0") works too.
  • Allow for slow hardware. Radios, USB enumeration, and boot sequences can take far longer than the 30 s default ready_timeout.
  • Choose a restart policy deliberately. Under the default Never, a bridge that crashes mid-run stays down while the simulation keeps running. Instant respawns it immediately, every time, including in a tight loop if the device is missing.
  • Use elodin run for repeatable rig runs. It exits with the simulation's status and stops every companion process.

For the hardware side, see crazyflie-edu, which runs the same controls in SITL or HITL, and the Aleph flight computer docs.

Gotchas

Outside campaigns, a recipe that crashes does not stop the simulation. The simulation keeps ticking against a dead process. Detect the lost connection in your bridge if you need to fail fast interactively.

  • From Python, restart_policy defaults to Never and processes are never restarted on file changes.
  • Ready.log(...) only works inside elodin monte-carlo campaigns.
  • Ready.tcp needs an IP literal such as 127.0.0.1:9000. localhost:9000 is rejected.
  • ${...} placeholders are not expanded in cmd or env values.
  • Readiness probes only hold the simulation when it is launched through the elodin CLI.
  • ctx.stop_recipes() does nothing under the elodin CLI.
  • cargo recipes run target/debug/<package name>. Passing bin= changes what is built, not what is run, so keep the binary name equal to the package name.
  • sim and render-server are reserved recipe names. Elodin adds the render server automatically when the world has sensor cameras.
  • In the editor, saving main.py restarts the simulation but reuses the original plan. Restart the editor after adding or removing world.recipe(...) calls.
  • s10 orchestration is not available on Windows.

Debugging

  • Read the plan. python main.py plan /tmp/plan writes the exact recipe group the CLI will run to /tmp/plan/s10.toml. Check that each recipe appears, and that the sim entry's depends-on lists the recipes you expect it to wait for.
  • Follow the prefixes. Each recipe's output is prefixed with its name (stdout in blue, stderr in red), and the simulation's own output with sim. A line like controller killed with code 1 shows how a recipe exited.
  • Campaign logs. In elodin monte-carlo runs, each recipe writes to runs/<run_id>/logs/<recipe>.log, and results.csv records a failure_reason naming the recipe that failed.
  • Pick the Python interpreter. Set ELODIN_PYTHON to choose the interpreter used to plan and run the simulation. Otherwise Elodin uses the active virtualenv, ./.venv, uv run, or python3, in that order.

Examples

ExampleShows
betaflight-sitlBetaflight SITL firmware as a plain process recipe, plus an optional controller
falcon9Rust flight software: cargo in the editor, prebuilt in campaigns, named ports
apollo-landerThe same split in a full Monte Carlo campaign (tutorial)
rc-jetController with an override to a prebuilt binary and a long ready_timeout
monte-carloMinimal Python controller recipe in a campaign
video-streamSeveral shell-script recipes, including one that exits immediately when unused
logstreamA C++ client built and run by a script recipe