YAML Spec API¶
The YAML spec API lets you declare a ROSE workflow as a data file instead of Python code. The spec is schema-validated when loaded — missing slots, unknown keys, and type errors are caught before any infrastructure starts. Task functions are plain Python functions with no decorators required. The same spec file works with any ROSE learner type: Active Learning, Reinforcement Learning, or UQ.
Note
The YAML spec currently supports only the four built-in learner.type values listed below:
SequentialActiveLearnerParallelActiveLearnerSequentialReinforcementLearnerParallelReinforcementLearnerSequentialUQLearnerParallelUQLearner
Custom Learner subclasses are not yet expressible in YAML — LearnerBuilder raises ValueError for any other learner.type value. Use the Python API (decorator-based, e.g. SequentialActiveLearner(asyncflow)) if you need a custom learner implementation.
Loading a Spec¶
For the common case — build a learner and run it — LearnerBuilder loads and validates the YAML itself; you never need to touch a config object:
from rose.spec.builder import LearnerBuilder
builder = LearnerBuilder("workflow.yaml", asyncflow) # validates schema on load
cfg = builder.config # typed WorkflowConfig, if you need it
Reach for load_spec instead when you need WorkflowSpec-level features: spec variants via workflow_with(), import validation, or the .workflow coroutine used by ROSE as a Service:
from rose.spec import load_spec
spec = load_spec("workflow.yaml") # validates schema on load
cfg = spec.config # typed WorkflowConfig object
Both raise ValueError with a precise message on any schema violation. Neither imports task modules by default — see Import Validation to enable that check.
Resource Specification via task_description¶
Every task — simulation, training, active_learn, environment, update, … — can carry an optional task_description. In the Python API this is a parameter on the decorated function itself, with a default value:
@acl.simulation_task(as_executable=False)
async def simulate(*args, task_description={"process_templates": [(4, {})]}, **kwargs):
...
The accepted keys depend entirely on which execution backend your asyncflow session is using — see asyncflow's Execution Backends guide for the full per-backend reference.
The YAML spec exposes the same dict as a task_description: key on the corresponding task slot — see the side-by-side example in Task Types below.
Sequential Active Learner¶
Task functions referenced by function: tasks:simulate are ordinary Python callables in a tasks.py module — no ROSE decorators required:
# tasks.py
def simulate(*args, **kwargs):
... # return simulation result
def train(sim_result, **kwargs):
... # return trained model
def active_learn(sim_result, model, **kwargs):
... # return updated dataset
def check_mse(*args, **kwargs):
... # return float metric value
Parallel Learner — Shared Tasks¶
When all parallel learners run the same task implementations, declare the task slots at the top level exactly as in the sequential case.
Each learner receives a unique learner_id kwarg (0, 1, …) in its task calls. Task functions can use it to write to isolated files or namespaces.
Parallel Learner — Distinct Tasks (learners: block)¶
When each parallel learner needs its own task implementations — different models, different executables, different command-line flags — use the learners: block. Each entry gets a label that is injected as learner_label into every task kwarg.
The builder creates a single dispatch closure per slot that routes to the correct command or function based on learner_id. The learner_id kwarg is consumed by the routing layer and never reaches the user function.
Note
All learners in a learners: block must use the same task type (python or shell) for each slot. Mixed types within a slot are rejected at load time.
Task Types¶
task_description example¶
Note
For parallel learners using the learners: block, all entries must use the same task_description for each slot — the backend registers it once at task-registration time.
Stop Criterion¶
The evaluator function follows the same *args, **kwargs convention as task functions. ROSE stops the loop when metric_value <operator> threshold is satisfied, or when max_iter is reached — whichever comes first.
Parameters¶
The parameters: block defines key-value pairs that are injected into every task's **kwargs at every iteration:
def simulate(*args, **kwargs):
dataset = kwargs["dataset"] # "my_dataset"
batch_size = kwargs["batch_size"] # 32
growing_pool = kwargs["growing_pool"] # False
iteration = kwargs["iteration"] # 0, 1, 2, ...
...
Reserved keys¶
The following keys are injected automatically by the builder and must not appear in parameters::
| Key | Available in | Description |
|---|---|---|
iteration | all tasks | Current loop counter (0-based) |
pythonpath | all tasks | Contents of remote.pythonpath as a list |
learner_id | parallel tasks | Integer index of the learner (0, 1, …) |
learner_label | parallel tasks (when label set) | Human-readable learner name from learners[].label |
Remote Config¶
remote:
pythonpath:
- /path/to/my/task/modules
- /path/to/shared/utilities
backends: [dragon] # optional — default shown; use [concurrent] for CPU-only runs
remote.pythonpath¶
remote.pythonpath entries are added to sys.path on the remote worker before any task module is imported. They are also injected into every task's kwargs["pythonpath"] (as a list), so task functions can construct file paths without duplicating the value in parameters::
def train(sim_result, **kwargs):
pythonpath = kwargs.get("pythonpath", [])
base = pythonpath[0] if pythonpath else ""
script = Path(base) / "scripts" / "train_model.py"
...
remote.pythonpath is the single edit point for the remote worker path — changing it updates both the import path and the kwarg.
remote.backends¶
remote.backends selects which Rhapsody backends are instantiated on the remote orbit endpoint session. The value is a list of backend name strings:
| Value | Backend class | When to use |
|---|---|---|
dask | DaskExecutionBackend | Dask-based execution |
concurrent | ConcurrentExecutionBackend | CPU-only runs, no Dragon needed |
dragon | DragonExecutionBackend | For highly compute intensive surrogates |
Example — switch to concurrent for a CPU-only endpoint:
backends is consumed only at engine-creation time and is not injected into task kwargs.
remote.target — bootstrapping the endpoint for rose run --remote¶
remote.target tells rose run <yaml> --remote how to launch the remote orbit endpoint itself, instead of assuming one is already running. Omit it to keep today's manual-bootstrap behavior (start the endpoint yourself, use --local/the orbit backend directly). See the Command-Line Interface page for the --local/--remote flags themselves, and rose setup for an interactive wizard that walks you through getting this working.
remote:
backends: [dragon]
target:
kind: sfapi # iri | sfapi | psij
endpoint: nersc # iri/sfapi only: nersc | olcf
resource_id: perlmutter
account: amsc007
queue_name: debug
walltime_min: 30
n_nodes: 1
constraint: cpu
home_dir: /global/u2/m/merzky
tunnel: none # none | forward | reverse
| Key | Type | Required | Description |
|---|---|---|---|
remote.target.kind | iri | sfapi | psij | yes | Bootstrap mechanism. Use sfapi for NERSC (IRI's /compute/* routes are broken server-side there); iri for OLCF; psij to submit via an already-connected login-node endpoint instead of IRI/SFAPI. |
remote.target.endpoint | string | iri/sfapi only | nersc or olcf. |
remote.target.resource_id | string | iri/sfapi only | Target compute resource, e.g. perlmutter, odo. |
remote.target.home_dir | string | iri/sfapi only | User $HOME on the target — resolves the endpoint wrapper script path. |
remote.target.login_host | string | when tunnel: forward | Login host to tunnel through. |
remote.target.edge_name | string | psij only, unless remote.embedded: true | Name of the already-connected login-node endpoint to submit through. Not needed when remote.embedded: true — PsiJ runs on the embedded broker itself. |
remote.target.executor | string | no | PsiJ executor name (default: local). |
remote.target.account | string | yes | Allocation/project account. |
remote.target.queue_name | string | no | Queue/partition. |
remote.target.walltime_min | int | no | Default 30. |
remote.target.n_nodes | int | no | Default 1. |
remote.target.constraint | string | no | Scheduler constraint (e.g. cpu). |
remote.target.reservation | string | no | Reservation name. |
remote.target.workdir | string | no | Job working directory. |
remote.target.environment | dict | no | Extra environment variables for the bootstrap job. |
remote.target.setup | list[str] | no | Shell setup lines run before the endpoint starts. |
remote.target.tunnel | none | forward | reverse | no | Default none. SSH tunnel mode for the endpoint's broker connection. |
Credentials are read from disk/env, never stored in the spec: iri reads a bearer token from ~/.amsc/token_<endpoint>; sfapi reads a client ID from $SFAPI_CLIENT_ID and a private key from ~/.amsc/sfapi_key_<endpoint>.pem.
remote.embedded — run without a standalone broker¶
remote.embedded: true hosts the ORBIT broker inside the rose run --remote process itself (EndpointRuntime/EmbeddedBroker) instead of connecting to one already running elsewhere. No separate radical-orbit-broker.py deployment is needed. Mutually exclusive with remote.broker_url.
remote:
embedded: true
target:
kind: psij
account: amsc007
queue_name: debug
# edge_name omitted — PsiJ runs on the embedded broker itself
The embedded broker is still a real broker: it needs the same operator-placed cert/key/token under ~/.radical/orbit (or their env-var redirects) as a standalone one — see radical.orbit's CLAUDE.md TOKEN_RECIPE if those aren't set up yet.
Combined with target.kind: psij, target.edge_name is not required — PsiJ loads directly on the embedded broker (this only works when the process running rose run --remote itself has batch-scheduler access, e.g. you're on a login node). target.kind: iri/sfapi work the same as without embedded — only where the broker lives changes.
Tracking¶
tracking:
backend: mlflow # mlflow | clearml | none (default)
experiment: my-exp # experiment/project name
run_name: run-01 # optional run label
See MLflow integration and ClearML integration for configuration details.
Spec Variants with workflow_with()¶
workflow_with() returns a new WorkflowSpec with selective overrides applied. The original spec is never mutated:
base_spec = load_spec("workflow.yaml")
test_spec = base_spec.workflow_with(max_iter=2, parameters={"dataset": "test_ds"})
large_spec = base_spec.workflow_with(max_iter=50, parameters={"batch_size": 128})
Accepted override keys:
parameters— merged (not replaced) into the existingparameters:block- Any
learnerfield:max_iter,parallel_learners - Any other top-level spec field
Unknown keys raise ValueError immediately.
Import Validation¶
By default, load_spec does not import task modules — this is intentional when remote.pythonpath points to paths that only exist on the remote worker. To catch function: typos locally during development:
With validate_imports=True, every module:callable string is resolved in the current environment. All failures are collected and reported in a single ValueError before any infrastructure starts. Use this during development when task files are locally accessible.
Spec Reference¶
Top-level keys¶
| Key | Type | Required | Default | Description |
|---|---|---|---|---|
learner | object | yes | — | Learner type and loop settings |
learner.type | string | yes | — | sequential_active_learner, parallel_active_learner, sequential_reinforcement_learner, uq_active_learner |
learner.max_iter | int | no | 0 | Maximum iterations |
learner.parallel_learners | int | no | 2 | Parallel learner count when learners: is absent |
simulation / training / active_learn | TaskDef | AL yes | — | Task slots for Active Learning |
environment / update | TaskDef | RL yes | — | Task slots for Reinforcement Learning |
prediction / active_learn / uncertainty | TaskDef | UQ yes | — | Task slots for UQ-based AL |
learners | list | no | — | Per-learner task definitions for heterogeneous parallel |
stop_criterion | object | yes | — | Stopping condition |
parameters | dict | no | {} | User-defined kwargs injected into all tasks |
remote.pythonpath | list[str] | no | [] | Paths added to sys.path on worker; injected as pythonpath kwarg |
remote.backends | list[str] | no | ["dragon"] | Rhapsody backends requested on the remote orbit session |
remote.broker_url | string | no | $RADICAL_ORBIT_BROKER_URL | Broker URL override for --remote (no on-disk fallback — orbit resolves URL from CLI/API/env only) |
remote.target | object | no | null | Bootstrap config for --remote — see remote.target |
remote.embedded | bool | no | false | Host the broker in-process instead of connecting to one — see remote.embedded |
tracking.backend | string | no | none | mlflow / clearml / none |
tracking.experiment | string | no | ROSE-Spec | Experiment name |
tracking.run_name | string | no | null | Run label |
TaskDef fields¶
| Key | Type | Required | Description |
|---|---|---|---|
type | python | shell | yes | Task execution mode |
function | string | if type: python | module:callable — dotted module path and function name |
command | string | if type: shell | Shell command; supports {param} placeholders filled from parameters: |
task_description | dict | no | Resource hints forwarded to the execution backend |