OpsQuest sandbox and safety¶
OpsQuest has two execution models and one optional presentation boundary:
- Linux missions use a teaching shell implemented entirely in Go over in-memory state.
- Optional Docker missions parse a deliberately small command language and translate typed actions into fixed Docker CLI arguments for attempt-owned resources on the active Docker-compatible engine.
play --webserves sanitized mission snapshots from an ephemeral IPv4 loopback port; the browser cannot submit commands or approve completion.
Neither model passes a raw player line to a host shell.
Trust boundaries¶
Editable source: trust-boundaries.excalidraw
The diagram focuses on command and storage authority. The separate loopback presentation boundary is described in Web companion boundary.
The player controls command text and navigation choices. Trusted OpsQuest code decides whether that text is a session control, a simulated-shell command, or one of the supported Docker teaching forms. Only three components normally interact with durable or external host facilities:
profile.Storereads and atomically replaces one configured profile file. Player commands cannot address this path.dockerlab.execRunnerinvokes the discovered Docker executable with arguments constructed insideinternal/dockerlab. It never invokessh -cand never forwards the original line.webapp.Serverlistens on127.0.0.1for the lifetime of oneplay --webcommand. It has no environment, runner, profile store, or command-input dependency.
Web companion boundary¶
The companion uses an ephemeral port and a cryptographically random one-time pairing token. A successful exchange creates an HTTP-only, same-site cookie and redirects away from the token-bearing URL. Exact Host and Origin checks reduce loopback confusion and DNS-rebinding exposure; restrictive CSP, frame, referrer, content-type, permissions, and cross-origin headers constrain the embedded page.
After pairing, the page may read the latest sanitized snapshot and subscribe to Server-Sent Events. It cannot submit terminal lines, address virtual files, choose Docker aliases or IDs, alter profile data, reveal a hint without the CLI control, or mark an objective complete. Static assets use no external scripts, fonts, trackers, or network service.
Snapshots are complete current projections with monotonically increasing event IDs. A slow subscriber skips obsolete queued snapshots and converges on the latest state. They exclude raw mission setup and condition objects, command text, terminal output, and internal resource identifiers. The page constructs all mission text with DOM text nodes rather than HTML interpolation.
Linux command execution¶
Editable source: command-execution-pipeline.mmd
Sandbox.Execute processes a line in explicit stages:
- Bound input: reject a command line over 64 KiB before adding it to the attempt's 100-entry history.
- Lex: recognize words, quotes, escapes, comments, variables, pipes, and
<,>, or>>. Expansion reads only the sandbox environment. - Parse: build pipeline stages and attach at most one input and output redirection to each stage.
- Expand: resolve eligible globs against the virtual filesystem and enforce expanded token and argument budgets.
- Preflight compositions: reject unsupported interactive-editor or script placement before an earlier pipeline stage can mutate state.
- Dispatch: call a Go method from the supported-command switch. Nested
find -execand scripts share dispatch budgets. - Move virtual data: pipeline output becomes the next stage's input; redirection reads or writes only virtual files.
- Return learning metadata: output, successful command names, maximum pipeline width, or a virtual editor request flow back to
game.Session. - Observe outcomes: output validators compare the returned text; state validators query the active environment.
Unsupported commands fail at dispatch. The shell implements a teaching subset, not process lookup, so a name absent from the dispatcher can never fall through to the host.
A full pipeline is not one transaction: an ordinary earlier stage may mutate virtual state before a later stage fails. Safety-sensitive operations preflight their own affected state, and compositions known to be unsupported (vi in a pipeline, or a script receiving pipeline/file input) are rejected before any stage runs.
Virtual state model¶
One sandbox.Sandbox owns:
| State | Representation | Persistence |
|---|---|---|
| Files and directories | Normalized absolute paths to typed entries with content, mode, and owner | Attempt only |
| Working directory | Virtual absolute path | Attempt only; child scripts restore caller scope |
| Environment | String map initialized with virtual HOME and USER |
Attempt only; exported child-script values restore on return |
| Processes | Mission-provided PID map and running flags | Attempt only; no host PID is visible or signalable |
| Archives | Logical archive metadata and payload entries | Attempt only; kept synchronized with virtual path mutations |
| History | Last 100 accepted command lines | Attempt only |
Paths are cleaned relative to the virtual current directory. The filesystem has its own /; it is not mounted or mapped to the host filesystem. Recursive writes, copies, archive operations, environment updates, and other amplifying operations calculate their budget before publishing the change. Virtual root/current-directory removal, file/directory type corruption, and archive extraction outside the chosen virtual destination are rejected.
The modal vi implementation is also virtual. sandbox returns an editor request, game handles terminal keys, and saving calls back into Sandbox.SaveEditorFile; no host editor or file API receives the virtual path.
Resource ceilings¶
The limits are part of the isolation model, not tuning suggestions.
| Resource | Limit |
|---|---|
| Command line | 64 KiB |
| Expanded token text | 2 MiB |
| Expanded arguments | 4,096 |
| Pipeline stages | 64 |
| Dispatches per execution, including nested work | 512 |
| One virtual file | 2 MiB |
| One command's output | 2 MiB |
| Total virtual file content | 8 MiB |
| Virtual filesystem entries | 4,096 |
| Virtual path | 4,096 bytes |
| Owner value | 256 bytes |
| Logical archive payload | 8 MiB |
| Logical archive entries | 4,096 |
| Environment entries | 256 |
| Environment data | 256 KiB |
| Script file / line | 64 KiB / 8 KiB |
| Script nesting / dispatched commands | 8 / 256 |
| Script output | 1 MiB |
vi file |
256 KiB |
Scripts are interpreted line by line through the same lexer, parser, expander, and dispatcher. They can use supported commands, variables, pipelines, and virtual redirection, but not loops, conditionals, functions, substitutions, background jobs, external programs, sh -c, stdin-fed source, or interactive editor calls.
Docker teaching boundary¶
Docker missions are opt-in and follow a narrower path:
- The mission catalog accepts only pinned image references and bounded logical fixtures.
- Availability checks the Docker executable, active Docker-compatible engine, and exact local image without creating resources or pulling images. An active
orbstackcontext is identified for provider-specific guidance only. - The factory first sweeps orphaned fixtures (see below), then generates a random session ID, creates the mission's declared networks, and creates containers. Networks and containers get generated names, five ownership labels, and two advisory owner labels (process ID and host name).
- Player text is parsed into
list,start,restart,stop,rm,inspect,logs, orhelp; flags and aliases must match the small grammar.psfilters are limited to fixedstatusandhealthvalues and are evaluated in Go over sanitized inspection data rather than passed to Docker.logs --tailaccepts only a whole number up to 10,000 orall.rmhas no--forceform and refuses a running container.docker network ls,inspect,create,rm,connect, anddisconnectaccept no options, validated logical names only, and never the built-inbridge,host,none, ordefaultnetworks. - The adapter resolves a validated logical alias to a tracked exact container ID.
exec.CommandContextreceives fixed arguments constructed by the adapter.- Observations inspect exact IDs or enumerate by the session label, then verify the complete ownership-label set.
Closeseals the attempt immediately, re-inspects ownership, removes only matching exact IDs (containers first, then networks), and retains unresolved resources for a retry. Resources the player already removed count as resolved.- Engine errors that name a tracked container are rewritten to its logical identity, and a missing container is reported as removed, so real IDs and generated names never reach the player.
Created containers use a read-only root filesystem, a bounded temporary filesystem, an unprivileged numeric user, no Linux capabilities, no-new-privileges, and explicit PID, memory, CPU, file-descriptor, restart, and stop-timeout limits. Diagnostic and logging fixtures use fixed shell programs; bounded log text, exit status, and lifetime are passed as data arguments rather than interpolated into those programs. Long-lived fixture processes exit after 24 hours. Health probes are the fixed BusyBox applets true or false, selected by a validated health enum. The restart policy is no except for a declared crash-loop fixture, which uses on-failure:50 and therefore stops relaunching on its own. Containers do not receive host bind mounts, devices, privileged mode, host networking, published ports, or a Docker socket.
Lab networks¶
A fixture without declared networks runs with --network none, exactly as before networks existed. A fixture with declared networks starts on the first one and joins the rest before it starts, with its logical alias as its network alias. Every lab network is created by OpsQuest with fixed options: the bridge driver and --internal, so it has no route to external networks. Players cannot pass network options, can create at most 4 networks per attempt, and can only attach this attempt's containers to this attempt's networks. docker network rm is refused while any container is attached, including stopped ones, because Docker would otherwise leave those containers unable to start. inspect shows logical network names and members, never engine IDs.
The internal flag is not the only control. Players cannot run commands inside containers: there is no exec, and fixtures run fixed programs. So no player-controlled traffic starts inside a lab network, whatever the engine allows between an internal network and its host gateway.
Orphaned fixtures¶
A process that is killed before Close leaves its fixtures behind. Before each Docker attempt, and on demand through opsquest doctor --cleanup, the factory lists containers and networks labeled com.opsquest.managed=true and removes an exact ID only when all of these hold:
- the managed, schema, session, mission, and alias labels are present and well formed;
- the name is the generated
opsquest-<session>-cNN(container) oropsquest-<session>-nNN(network) form for that same session; - either the owner-host label matches this machine and the recorded owner process is no longer running, or the resource is older than the 24-hour fixture lifetime.
Orphaned containers are removed before orphaned networks.
Owners on other hosts, and fixtures created before owner labels existed, can therefore only age out. A plain opsquest doctor counts orphans without removing anything. A failed automatic sweep never blocks mission setup.
Docker-specific ceilings include:
| Resource | Limit |
|---|---|
| Player Docker line | 64 KiB |
| Images per mission | 16 |
| Containers per mission | 32 |
| Declared networks per mission | 8 |
| Declared networks per container | 4 |
| Player-created networks per attempt | 4 |
| Declarative diagnostic log per container | 8 KiB |
| Captured Docker stdout and stderr | 2 MiB each |
| Normal Docker operation | 10 seconds |
| Cleanup attempt | 10 seconds |
The Docker-compatible engine remains a powerful external dependency; OpsQuest reduces exposure by constraining syntax, setup, runtime options, identity, and cleanup scope. It does not claim to turn an untrusted engine into a security boundary.
Threat-to-control map¶
| Threat | Primary controls | Evidence location |
|---|---|---|
| Player launches a host command | Closed Go dispatcher; no fallback process lookup; raw line never reaches a shell | internal/sandbox/shell.go, dispatch.go, hardening tests |
| Player reads or writes a host path | Independent in-memory root and virtual path resolver | internal/sandbox/filesystem*.go, regression tests |
| Expansion exhausts memory | Line, token, argument, output, entry, and aggregate budgets | quota tests in internal/sandbox |
| Archive escapes extraction target | Strict metadata validation and destination containment | archive and hardening tests |
| Script bypasses shell restrictions | Same parser/dispatcher, bounded nesting and steps, unsupported syntax rejection | commands_script.go and tests |
| Docker input becomes arbitrary CLI flags | Typed parser accepts only exact actions and validated logical aliases | internal/dockerlab/parser.go and tests |
| Cleanup removes another container | Exact ID plus managed/schema/session/mission/alias label verification | internal/dockerlab/cleanup.go, ownership.go and tests |
| Partial Docker setup leaks resources silently | Factory may return a partial environment; managed cleanup and retryable Close |
internal/dockerlab/factory.go, fixtures.go, environment contract tests |
| A killed process leaks running fixtures | 24-hour fixture lifetime; owner-process and age-based sweep with complete label and generated-name proof | internal/dockerlab/janitor.go and tests |
| Orphan sweep removes a live or foreign container | Live same-host owners are never swept; unknown owners only after the fixture lifetime; exact IDs only | internal/dockerlab/janitor_test.go |
| Removed containers leak engine identity | Missing-container errors mapped to the alias; tracked IDs and names redacted from engine errors | internal/dockerlab/transport.go, container_commands.go and tests |
| Lab network reaches the host network or internet | Fixed --internal bridge networks; no network options; no in-container command execution |
internal/dockerlab/network.go, network_test.go |
| Player attaches a lab container to a host or foreign network | Built-in names rejected by the parser; only tracked attempt networks resolve; exact IDs only | internal/dockerlab/parser.go, network_commands.go and tests |
| Network cleanup strands or removes the wrong network | Containers removed first; label and generated-name verification; retryable unresolved set; janitor sweeps networks after containers | internal/dockerlab/cleanup.go, janitor.go and tests |
| Persisted display text injects terminal controls | Profile names reject non-printable characters and normalize legacy values | internal/profile/profile.go and tests |
| Another site reaches the loopback companion | One-time capability pairing, exact Host/Origin checks, same-site HTTP-only cookie, no permissive CORS | internal/webapp/server.go and tests |
| Browser input reaches a mission environment | Companion exposes only read-only state and SSE routes; game.Session has no companion command callback |
internal/webapp, internal/game/companion.go |
| Browser observes private attempt internals | Dedicated sanitized snapshot type; no setup, raw condition, transcript, or resource-ID fields | internal/game/companion.go and tests |
Safety review checklist¶
Changes that affect parsing, paths, recursion, redirection, scripts, archives, quotas, Docker execution, or environment cleanup should verify all of the following:
- Player-controlled text still cannot become host shell input or an unrestricted process argument.
- Paths resolve only in the active virtual filesystem or to an exact adapter-owned resource.
- Failure paths preserve existing state where an operation promises preflight or atomic publication.
- Success, boundary, and rejection behavior have focused tests.
- Mission validators continue to test observable results rather than a command transcript.
make check-allpasses; Docker adapter changes also runmake docker-integration, ormake orbstack-integrationfor that provider, when prerequisites are available.
The authoritative implementation is the code and tests. This guide explains their intended security properties; if a diagram and a tested invariant disagree, treat the mismatch as a documentation defect.