An app is a piece of software that ORC8R installs and runs on your nodes. Instead of logging into each machine and setting things up by hand, you attach an app to a pool and ORC8R installs it on every node for you. This page explains how to use existing apps and, for the more technical reader, the full artifact.yaml reference for authoring your own — including apps that serve traffic.

How apps get onto nodes

You choose apps when you request nodes. In the Choose apps step of the request wizard (see Pools):

  1. Click Add Application.
  2. Pick an app from the catalog.
  3. Choose a version, and fill in any settings the app asks for.

Every node created for that pool then installs the apps you selected. If you add none, nodes start with just the agent. Some apps run a task once and finish (for example, installing a package); others run continuously as a service.

Configuring apps for a project or organization

You can set defaults for an app so that everyone requesting nodes gets sensible values without having to fill them in each time. Open a project and go to Project Apps (or an organization and go to Organization Apps). For each app you can:

  • Pin a version so new nodes always use a known-good release.
  • Set default values for the app's settings.
  • Lock settings so users cannot change them at request time.
  • Disable the app so it does not appear in the catalog for that project or organization.

Settings flow from broad to narrow: an organization's defaults apply to all its projects, and a project can add its own on top. If your organization disables an app, a project inside it cannot turn it back on.

Authoring your own app

If the app you need does not exist, you can build one. An app is packaged as a standard OCI artifact (the same kind of package used for container images), so it can be stored in an ordinary registry. You describe the app in a small file called artifact.yaml, then build it with the orc build command, which validates the file and tells you about mistakes while packaging rather than at deploy time.

The sections below are the complete reference for that file.

The shape of artifact.yaml

The file has five top-level keys:

KeyWhat it holds
artifactTypeAlways application/vnd.orc8r.app.v1 — this is what marks the package as an ORC8R app.
annotationsThe app's name and description, as org.opencontainers.image.title and org.opencontainers.image.description.
configEverything about how the app behaves: its settings, its endpoints, and its lifecycle. The rest of this page is config.
filesThe scripts and binaries the package ships — a list of paths. They land in the app's working directory.
platformsThe operating systems and architectures the app supports, each a {os, arch} pair — for example Linux on amd64 and arm64.

Only artifactType is required; everything else has a sensible empty default. Set org.opencontainers.image.title too, because it is what names your app in the catalog.

Shipping files

files lists what the package carries. An entry is either a plain path, relative to the recipe's own directory:

files:
  - install-apt.sh
  - bin/server

or an object, which lets you pin the bytes and fetch them from elsewhere:

FieldMeaning
pathRequired. Where the file lands in the app's working directory, and where it is read from locally.
urlOptional. Fetched when the local file is missing or does not match sha256.
sha256Optional but strongly recommended with url. The digest the bytes must have.
files:
  - path: bin/widget
    url: https://example.com/widget/2.4.0/widget-linux-amd64
    sha256: "9f2c…"

Fetching happens once, at build time, and the bytes are embedded into the package like any other file — so nodes never reach out to that URL, and a vendor that deletes the asset cannot break your deploys. A digest mismatch fails the build. The fetched bytes are also written to path in your working directory, so a second build is offline.

Paths must be relative and normalized: .. and absolute paths are rejected, and two entries may not resolve to the same path. On Unix the execute bit is captured and restored on the node, so a script committed executable arrives executable.

One package, many platforms

platforms lists the {os, arch} pairs the app supports, and vars under each one carries per-platform values that get spliced into the recipe:

files:
  - path: "bin/widget-{os}-{arch}"
config:
  start:
    command: "{startCommand}"
platforms:
  - os: linux
    arch: amd64
    vars: &posix
      startCommand: ./bin/widget-linux-amd64 --serve
  - os: linux
    arch: arm64
    vars: *posix
  - os: windows
    arch: amd64
    vars:
      startCommand: .\bin\widget-windows-amd64.exe --serve

A {name} placeholder is replaced with vars.name for the platform being built. {os} and {arch} are supplied for you; every other name must exist in that platform's vars, or the build fails with unresolved interpolation. YAML anchors (&posix / *posix) are the tidy way to share one vars map across several platforms — the built-in shell app uses exactly that to give POSIX and Windows their own start command.

Substitution reaches every string value under config, plus files[].path and files[].sha256. It does not reach files[].url or annotations, so keep those literal.

There is no escape for a literal brace. Every { starts a placeholder, which is why shell snippets in a recipe must be written $VAR rather than ${VAR} — the braces would be read as a template key and fail the build. Move anything brace-heavy into a script file, where it is never interpolated.

Leave platforms out entirely and you get a single platformless package that runs anywhere — and no placeholders, since there is no platform to resolve them against.

Settings: the params block

config.params describes the settings an operator fills in. It has two parts: required, the list of setting names that must be provided, and properties, a map from each setting name to how it behaves.

Fields appear in the request form alphabetically unless you say otherwise. To control the order, add an x-ui-order list alongside properties naming the settings in the order you want them shown — any setting you leave out follows the ordered ones. For example, x-ui-order: [repo_url, labels, runner_group] puts the repository URL first.

Property fieldMeaning
typeThe value's type, usually string.
titleThe label shown for the field in the request form.
descriptionHelp text shown under the field.
placeholderExample text shown in the empty field.
enumA fixed list of choices, offered as a dropdown instead of a free-text box.
contentMediaTypeThe value's media type, such as application/x-pem-file. It also makes the value arrive as a file rather than an environment variable (see Values the platform fills in).
sensitivetrue stores the value securely and never echoes it back in plaintext.
lifetimestartup (withdrawn once the app has started) or runtime (available for the whole run). When unset, operator secrets default to startup and everything else to runtime.
x-sourceMarks a value the platform fills in on the node rather than a person typing it — see Values the platform fills in.

Each setting reaches your commands as an environment variable with an upper-cased name — a setting called packages arrives as PACKAGES.

This is the built-in apt app, which installs Debian or Ubuntu packages. It has one required setting and runs a script during the install phase:

artifactType: application/vnd.orc8r.app.v1
annotations:
  org.opencontainers.image.title: apt
  org.opencontainers.image.description: Install Debian/Ubuntu packages via apt-get
config:
  params:
    type: object
    required: [packages]
    properties:
      packages:
        type: string
        title: Packages
        description: Space-separated package names (e.g. git curl htop)
        placeholder: "curl git vim"
  install:
    command: sh install-apt.sh
files:
  - install-apt.sh
platforms:
  - os: linux
    arch: amd64
  - os: linux
    arch: arm64

When someone adds this app to a pool and types curl git vim into the Packages field, every node runs install-apt.sh with PACKAGES set to curl git vim, installing those packages.

Lifecycle phases

The rest of config is the app's lifecycle, given as named phases. You include only the ones you need. A phase's command is either a string (sh start.sh) or an argv list (["/usr/sbin/server", "--flag"]).

PhaseRunsFields
installonce, when the app is first set up on a nodecommand, timeout
startto launch the appcommand, service, restart, gui
stopwhile the app is still running, to let it wind down safelycommand, signal, timeout, grace
stoppedafter the app has exited, however it exitedcommand, timeout
uninstallwhen the app is removed from the nodecommand, timeout

The start phase is what makes the difference between the two kinds of app. A start command that keeps running is the app's long-lived service; an app with only an install command that finishes is a one-time task. Its other fields cover apps run as a managed system service or as a desktop program: service and restart for the former, gui for the latter.

stop and stopped are the two halves of shutting an app down, and they run at different moments: stop runs while the app is still alive and is your chance to let it finish safely; stopped runs once it has exited. See Stopping an app cleanly.

Timeouts take the grammar <integer>(s|m|h) — "600s", "5m". The defaults are 600s for install, 30s for stop with a 10s grace after the signal, 60s for stopped, and 300s for uninstall.

Exit 0 means completed, not failed

For a command-mode app the runtime watches the process: failing to start, or exiting non-zero, is an app failure, and exiting 0 is a legal terminal outcome — the app is reported completed. That is the whole mechanism behind run-to-completion apps; there is no separate one-shot mode to declare. An apt install that finishes is completed, and the node moves on.

For a service-mode app the runtime watches the service manager's state rather than raw process exits. The app fails only when the service reaches a terminal failed or stopped state outside orchestration — exits the manager absorbs under its own restart policy are not failures, and they run neither stop nor stopped.

Running as a system service

Add service to start and your app becomes a managed system service instead of a child process. The name is a plain identifier — jenkins-agent, never jenkins-agent.service — and the node turns it into whatever its platform calls a service (a systemd unit on Linux, a service entry on Windows). Every phase gets the resulting name in APP_SERVICE. If you put ${APP_VERSION} in the identifier, each installed version gets its own service.

With service and a command, the node writes the service definition for you, from your config: the restart policy, the stop signal, and the grace before the kill all come straight from the phases above. You ship no unit files and no registration scripts.

With service alone, you keep ownership: your install phase registers the service under exactly that derived name and your uninstall phase removes it, and the node only starts, stops and watches it.

restart is never (the default), on-failure, or always, and it is the service manager that honors it — the runtime itself never restarts an app. Do not enable boot auto-start (systemctl enable, a Windows Automatic start type): after a reboot the node alone restores run state, so an app an operator stopped stays stopped.

Prefer a service for anything long-lived. A service keeps running when the ORC8R agent restarts or is upgraded — the node re-adopts it by name afterwards — while a plain command app is a child of the agent and is stopped and started again with it.

There is no field for running as another user: an app runs as the user the agent runs as. If yours must run as a different account, drop privileges inside the start command with setpriv --reuid=svc --regid=svc --init-groups (Linux) or sudo -u svc (macOS). Do not use su or runuser — they put your app in a new session where the stop signal can no longer reach it, and su kills its own child a couple of seconds after passing SIGTERM on, cutting your shutdown short.

Scripts, and how a phase finds one

Lifecycle logic usually ships as a script rather than an inline command, discovered by naming convention: install-<app>, start-<app>, stop-<app>, stopped-<app>, uninstall-<app>, with the extension deciding the interpreter.

ExtensionPlatformRun as
.shUnixsh <script>
.ps1Windowspowershell -File <script>
.cmdWindows fallbackcmd /C <script>

A command you declare wins; a script found by name is the fallback. That is what lets one declaration cover every platform: the built-in os-update app ships both install-os-update.sh and install-os-update.ps1 and declares an install phase carrying only a timeout — no command, so each node picks up its own script at install time. Phases with neither a command nor a script are skipped.

Commands: string or argv

A command is either a string, run through the platform shell (sh -c on Unix, powershell -Command on Windows), or an argv array run directly with no shell at all:

start:
  command: ["/opt/widget/bin/widget", "--master", "${WIDGET_URL}"]

The runtime substitutes ${VAR} tokens from the app's environment inside each element, and a reference to an unset variable is an error rather than an empty string. Prefer the argv form for new apps: it is portable across platforms and immune to quoting, which is where the string form goes wrong.

What a phase can read from its environment

Every phase runs in the app's working directory — where the files you shipped landed — with a normal login environment rather than the bare service-manager one, so profile-derived PATH and ~-relative tooling behave as they would for a human at a shell.

VariableAvailable inValue
each param, upper-casedevery phasethe operator's value — packages arrives as PACKAGES
<PARAM>_FILEevery phasethe path to a 0600 file, for file-backed params
APP_VERSIONevery phasethe resolved upstream version; unset only when the version is the implicit default
APP_INSTANCEevery phasewhich copy of the app this is — the app name, or name:version when several versions run on the node
APP_SERVICEevery phase, service appsthe platform service name derived from your service identifier
APP_PIDstop, stoppedthe app's process id
APP_STOP_REASONstop, stoppedwhy it is stopping — see below
APP_STOP_STATUSstoppedhow your stop command went: ok, failed, timeout or skipped
APP_EXIT_CODE, APP_EXIT_SIGNAL, APP_RUN_DURATIONstoppedhow the run ended — one of the first two is set, plus the seconds it ran

The APP_* names are reserved, and a runtime value beats a colliding param. APP_VERSION is the one every install script cares about — see App versions.

Stopping an app cleanly

Shutting an app down happens in four steps, in this order:

  1. Your stop command runs while the app is still alive. Use it to make the app safe to signal — take it out of rotation, unregister it from its controller, wait for in-flight work to finish. Do not kill the app here; that is the next step's job. Exiting 0 means "go ahead".
  2. The app is signalled — SIGTERM unless you named another signal — or, for a service app, stopped through its service manager. Windows has no signal to send, so on Windows the graceful exit has to come from the stop command.
  3. After grace, whatever is still alive is killed. That grace is always granted.
  4. Your stopped command runs, once the app has exited — after a clean shutdown, a kill, or a crash alike.

A slow or failing stop command never blocks the shutdown: exceeding its timeout or exiting non-zero is recorded and the flow carries on, and the stopped phase reads what happened in APP_STOP_STATUS. An operator can force a stop at any moment, which jumps straight to the kill.

stopped is where you release anything that lives outside the node — a runner registration with a CI controller, a lease, a DNS entry. Which of those you should release depends on APP_STOP_REASON:

APP_STOP_REASONWhat it meansUsually
restartthe same install is about to start again (an edited parameter, a renewed certificate)leave external registrations in place
stopthe install stays on the node and may start again laterleave external registrations in place
terminatethe node is going away for goodrelease everything external here
shutdownthe agent itself is exiting, for an upgrade or a rebootleave everything in place; the app comes back
exitthe app ended on its ownusually nothing to do

Write both commands so they can be cut short and run again: a timeout can interrupt them, and they may run after a crash in which an earlier attempt got nowhere. Neither stop nor stopped may reboot the machine — install, start and uninstall may.

Declaring endpoints

An app that serves traffic says so by declaring endpoints. An endpoint is one thing your app listens for connections on: where it listens, what it speaks there, how to tell when it is ready, and whether more than one node may answer for it. That is the whole contract a package carries. Nothing about who may reach it or which network it sits on belongs in a package — those are decisions of the deployment, made in the request composer (see Project and organization exposure). Put the other way round: your package says what you serve, and the deployment says who may reach it.

the app package (artifact.yaml)endpointpgport5432protocoltcpprobetcp connectserveprimarysays WHAT it serves+the request (one deployment)exposeorgservice namedbport5432says WHO may reach itone endpoint, two decisions, made in two places

endpoints is a map under config, keyed by endpoint name:

config:
  endpoints:
    pg:
      port: 5432
    metrics:
      port: 9187
      protocol: http
      probe:
        http: /metrics
FieldMeaning
portRequired. The listen port, 1–65535.
protocoltcp (the default), udp, http, or https. http is a plaintext HTTP/1.1+ listener; https is the same but the app terminates TLS on the port itself (it serves HTTPS, consuming tls.cert/tls.key). Both are eligible for an HTTP probe. Declare https when your app serves TLS — the platform then knows the scheme rather than guessing it, and a plaintext http endpoint is refused public exposure (Public exposure).
probeOptional readiness check — exactly one of tcp, http, or command, plus optional interval and timeout.
serveprimary (the default) or spread — how many nodes of a scaled pool answer for this endpoint. See below.

Endpoint names are DNS labels — one dot-separated piece of a name, the db in db.shop.internal: lowercase letters, digits and dashes, 1–63 characters, no leading or trailing dash. Most paths select an endpoint by port and never put its name in DNS, but SRV records — DNS entries that hand back a port alongside a name — and front doors that pick a service from the hostname a client asked for do, so the constraint is enforced up front rather than caught by a rename later.

orc build validates all of this:

endpoint "Metrics" is not a DNS label (lowercase alphanumeric and '-', 1-63 characters, no leading or trailing '-')
endpoint "pg" port must be an integer in 1..=65535
endpoint "web" probe.http requires protocol "http", not "tcp"
endpoint "api" probe must declare exactly one of tcp, http, command
endpoint "pg" serve must be one of primary, spread
endpoint "metrics" port 5432 duplicates endpoint "pg" on transport "tcp"

The last one is the rule that surprises people: collisions compare transports, and http and https both count as tcp, because they are TCP listeners. So two endpoints may not share a port number unless they sit on different transports. tcp and udp on the same number coexist legally — QUIC beside TCP on 443, DNS on 53.

Serve mode: primary or spread

A pool can run one node or twenty. serve is where your package says how many of them may answer a request of a given endpoint — the one thing about spreading traffic that the deployment cannot decide for you, because only the app knows whether a second node holds the same answer.

config:
  endpoints:
    pg:
      port: 5432          # serve: primary — the default
    web:
      port: 8080
      protocol: http
      serve: spread
ModeWho answers
primary (default)The slot-1 node and nobody else. While that node is not ready the endpoint's port refuses connections and its names answer nothing — never a stand-in.
spreadEvery node that passes this endpoint's probe.

primary is the default on purpose. A package that says nothing has not said that a replica may serve a write, so a pool scaled from one node to three keeps pointing at the one node it always pointed at until you say otherwise. Declare spread when any node can serve any request of the endpoint — a stateless HTTP tier, a read-only mirror, a metrics port.

The two modes live side by side on one app. A database can serve pg as primary and metrics as spread; each port is decided on its own terms.

One rule catches people out. A pool has one set of names but many ports, and a name is a single answer for all of them. So if any endpoint of the version serves primary, the pool's plain names answer the slot-1 node alone — db.shop.internal, its service-name aliases, and its public name, everywhere. A client that looked up the bare name may dial any port the pool declares, and only slot 1 is correct for all of them. Nothing else changes: the slot names (db-1, db-2, …) still name their own nodes, SRV records still list every ready node, and delivery on each individual port is still exactly what that port's mode says. If you want the names to spread, every endpoint of the app has to spread.

While the slot-1 node is unready, a name under this rule answers empty rather than falling through to a sibling — the same fence a slot name carries. See Networking for what each name means.

Probes

A probe is a small check ORC8R repeats to decide whether an endpoint is ready to serve traffic.

ProbePasses when
tcp: truea connection to the port is accepted — the default for tcp, http, and https endpoints
http: /patha GET on that path answers 2xx or 3xx (redirects are not followed). Valid on http and https endpoints; on https the probe dials over TLS. It does not validate the certificate — it dials the node's own addresses, where the certificate's name cannot match, and readiness is a liveness check, not the trust boundary (that is the consumer, through ca.bundle).
command: …the command exits 0

interval and timeout take the grammar <integer>(s|m|h) — "10s", "1m", "1h". Nothing finer than a second exists. The defaults are a 10-second interval and a 5-second timeout; an endpoint goes unready after 3 consecutive failures and recovers after 1 success. A udp endpoint has no default probe: without a command probe its readiness simply follows the app's.

Bind [::], not loopback — and not one address family

Make your app listen on [::], the dual-stack wildcard. An app that listens only on 127.0.0.1 never passes its probe, and an app that listens only on 0.0.0.0 fails it the moment the endpoint is exposed publicly. This is the one that costs people an afternoon.

A probe dials every address the endpoint is served on: the node's own address — the one internal consumers resolve — and, for a publicly exposed endpoint, the pool address the public traffic arrives on. It does not dial 127.0.0.1. That is deliberate: a passing probe has to mean what the platform takes it to mean — that this node is one of the ones traffic gets sent to — and a socket that is not bound where the traffic lands serves nobody.

There are two ways to get that wrong:

  • Loopback only. An app bound to 127.0.0.1 reports unready forever, its name answers nothing, and the pool looks healthy while serving no traffic.
  • One address family only. Public delivery is direct: the claiming host forwards the client's packets to your node unrewritten, so the address your app must have a socket bound to is the pool address itself, not the node's. Pool addresses are usually IPv6 — a routed v6 block is near-free, a v4 one is rented — and 0.0.0.0 is an IPv4 wildcard whatever the word "wildcard" suggests. An app bound there has no socket for what arrives, and the node's kernel answers every public connection with a reset.

So listen on [::], which accepts both families on one socket, and let the platform's default-deny policy be what keeps the port private — privacy is the deployment's job, not your listen address's. Where a runtime cannot do both families on one socket (Python's http.server is the classic — HTTPServer is AF_INET and nothing else), open one listener per family. See Project and organization exposure.

Probe failures are recorded verbatim and name the address that failed, so the status tells you which problem you have:

connect to 100.64.0.3 port 5432 failed: Connection refused
connect to 2001:db8::5 port 5432 failed: Connection refused
connect to 100.64.0.3 port 5432 timed out
http probe /healthz on 100.64.0.3 returned 503
probe command exited with exit status: 1

A refusal naming the pool address while the node's own address answers is the second failure above, and it has no other symptom: the app is running, internal consumers are served, and the internet gets a reset.

A probe command runs in the app's working directory with the app's start-time environment. One trap: parameters delivered as files with a startup lifetime are deleted once the app has started, so a probe command must not depend on a *_FILE variable. Read what you need at start, or probe over the port.

Values the platform fills in

Some parameters are not typed in by a person filling a form — a peer's name, a trust bundle, a certificate. Declare them in config.params with an x-source, and the platform fills them in on the node at deploy:

config:
  params:
    type: object
    properties:
      slot:
        type: string
        title: Runtime slot
        x-source:
          kind: pool.slot
      writer:
        type: string
        title: Writer address
        x-source:
          kind: peers.first
      tls_ca:
        type: string
        title: TLS CA bundle
        contentMediaType: application/x-pem-file
        x-source:
          kind: ca.bundle
      tls_cert:
        type: string
        title: TLS certificate
        contentMediaType: application/x-pem-file
        x-source:
          kind: tls.cert
          params:
            purpose: server_tls
      tls_key:
        type: string
        title: TLS private key
        contentMediaType: application/x-pem-file
        x-source:
          kind: tls.key
          params:
            pair: tls_cert

A slot is a pool node's stable identity, and slot 1 is the writer by convention; Networking overview has the rest. A leaf is the certificate issued to your app itself, as opposed to the chain that signed it.

kindparamsWhat arrives
pool.slot—the node's own runtime slot as a decimal string; 1 is the writer by convention
pool.name—the name of the pool the node belongs to
peers.firstpool, projectthe slot-1 node's DNS name, {pool}-1.{project}.{suffix}
peers.gossippool, projectthe pool's plain name — one name that means the pool as a whole, answering per its serve modes (above)
peers.anypool, projectthe plain name too; where it spreads, the resolver rotates its answers
peers.allpool, projectone slot name per allocated slot, newline-separated, delivered as a file
endpoint.nameendpoint: which of your endpoints — omit it when your app declares exactly onethe endpoint's canonical name: its public name where it is exposed publicly, its internal one otherwise
endpoint.portendpointthe port the endpoint actually listens on, as a decimal string — your declared port with the deployment's override applied
endpoint.urlendpoint{scheme}://{name}, plus :{port} unless the port is the scheme's default (443 for https, 80 for http). The scheme is the protocol you declared: https, http, tcp, or udp
ca.bundlescope: project (default) or orgthe trust chain, PEM
tls.certpurpose: client_mtls (default) or server_tls; optional sansa signed leaf plus its issuing chain, PEM
tls.keypair: the sibling tls.cert property's namethe matching private key, PEM — always file-backed and implicitly sensitive

Four things about these are worth internalising.

peers.* hand you names, not addresses. That is what makes them stay correct: nodes are replaced, slots are reused, and the next DNS answer absorbs the change without re-resolving the parameter or restarting your app. Resolve them at connect time; do not parse them, cache them, or split them on a delimiter.

A name that answers nothing is the fence, not an error. peers.first resolves to the slot-1 name whether or not a healthy slot-1 node exists — while none does, the name answers empty and your connection attempt fails. The plain name of a primary pool is fenced the same way. That is the writer fence working as designed. Retry; do not fall back to another name.

endpoint.* is your app's own address, not a neighbour's. The three endpoint kinds resolve within the assignment declaring them, from the endpoint surface the deployment settled — so they answer "what am I reachable at", which is what an absolute link, an OAuth callback or a root_url needs. Declaring one on an app with no endpoints, omitting endpoint on an app with several, or naming an endpoint the app does not declare are all authoring errors: the deploy fails with a message naming the fix, rather than the platform guessing a port for you — the same message wherever your app runs. Exposure is a deployment decision, so changing it reconfigures the pool and the values roll with the nodes.

File-backed delivery. A parameter arrives as an upper-cased environment variable, unless it is sensitive, has a contentMediaType, or is a tls.key/peers.all — those arrive as a 0600 file whose path is in <NAME>_FILE. Putting contentMediaType: application/x-pem-file on your certificate and bundle parameters is the idiom worth copying: PEM in an environment variable is unpleasant for everyone.

Certificates renew by restarting

There is no reload channel. Each converge pass — each run of the agent bringing the node in line with what it has been assigned — arms a single timer from the shortest-lived leaf it delivered: two thirds of the certificate's validity window, with jitter so pool nodes do not all restart at once. When it fires, the pass runs again, mints fresh material, and restarts the app through the ordinary pipeline.

Design for that: start fast, hold no unflushed state, and read your certificate at startup rather than watching the file. If a leaf cannot be parsed the agent warns tls leaf renewal NOT armed; this certificate will expire unattended and leaves the app running, which is the one case you have to notice yourself.

For purpose: server_tls, the names you are authorized to assert are your pool's plain name, your own slot name, and both forms of the endpoint's service name if one is chosen. Extra names are requested with sans; ask for anything outside that list and you get no certificate at all.

A publicly exposed endpoint is served instead by a certificate a public authority issued for that one public name (Public exposure), delivered through the same binding. Where a pool exposes several public services, say in sans which one this binding is for — the service name, or its full public name — since one binding carries one certificate. Public and internal names never share a certificate: no public authority will certify an .internal name, so a request mixing the two gets nothing.

Tracking upstream versions

One package serves every release of the software it installs. config.versions declares where those releases are published — a project's GitHub releases, a JSON version index, a fixed list — so the catalog offers them without you republishing, and config.default_version names the one selected by default. Your install script is handed the chosen version in APP_VERSION and fetches exactly that upstream artifact.

Both halves of that — declaring the versions, and installing a pinned one correctly — are on their own page: App versions.

Building and publishing

orc build reads artifact.yaml from the directory you point it at, defaulting to the current one:

$ orc build -t widget:default --push
Resolved reference: v0.orc8r.com/widget:default
sha256:2f9c…
FlagEffect
-t, --tagthe reference to publish under; repeat it, or comma-separate, for several tags
--pushpublish after building, instead of only caching locally
-o, --output <DIR>write a portable OCI layout directory instead of caching (not with --push)
--format plainwhole-blob layers, required for registries outside ORC8R

Omit -t and the reference comes from your annotations — org.opencontainers.image.title, plus org.opencontainers.image.version if you set one, else the tag default. A reference with no registry in it goes to the built-in registry; write the registry into the reference (ghcr.io/acme/widget:1.0) to publish elsewhere. External registries need --format plain, and orc build tells you so rather than failing halfway through the upload.

Publishing under the tag default is the usual choice: it is the tag an app reference with no version resolves to, and the one version discovery reads its declaration from.

Every build also attaches your artifact.yaml to the package verbatim, so orc clone recovers the exact recipe — comments, anchors and all — rather than a reconstruction. That, not a copy on your laptop, is the durable source of an app.

Build-time validation is not the whole story. orc build checks the recipe: the artifact type, path safety, digests, interpolation, and that config parses. Deeper checks on the parameter schema and endpoints run when the package is ingested. Push to a scratch tag early rather than discovering a schema complaint on the day you ship.

Packaging gotchas

files is a top-level key, a sibling of config, not something inside it. Its entries land in the app's working directory, which is also the working directory of every lifecycle command — so command: sh install-app.sh and command: ./bin/server both just work. Nested paths are fine; .. and absolute paths are rejected.

Values go under params. On the deployment side, an app assignment's values live under params. vars is a recipe's build-time template map and nothing else — writing it where params belongs would otherwise start your app with no configuration at all, so the request is rejected outright, with a message that says exactly this.

Stop the distribution from starting your service for you. On Debian and Ubuntu, installing a package runs its postinst, which starts the service immediately, on the distribution's default configuration, on the distribution's port, before any of your configuration exists. In a pool that is a port conflict at best, and a service that answers the probe while running nothing you configured at worst. Neutralise it inside your install phase:

# Refuse service starts for the duration of the install.
printf '#!/bin/sh\nexit 101\n' > /usr/sbin/policy-rc.d
chmod +x /usr/sbin/policy-rc.d

DEBIAN_FRONTEND=noninteractive apt-get install -y postgresql

rm -f /usr/sbin/policy-rc.d
# Write your configuration, then start it yourself in the start phase.
systemctl mask --now postgresql || true

The same reflex applies to the distribution's own automation: masking apt-daily.timer, apt-daily-upgrade.timer and unattended-upgrades.service is what the agent itself does before baking an image, and for the same reason — nothing should be reconfiguring a node behind the platform's back.

Wait for your dependencies in your own start script. A pool comes up in parallel and nothing sequences apps across pools for you. The pattern is a bounded retry loop around the name, not a sleep:

# Wait for the writer to exist. Until slot 1 is healthy, WRITER answers nothing.
i=0
while ! pg_isready -h "$WRITER" -q; do
  i=$((i + 1))
  [ "$i" -lt 120 ] || { echo "writer $WRITER never became ready" >&2; exit 1; }
  sleep 2
done
exec /usr/sbin/my-service --upstream "$WRITER"

Two details make this work: the loop re-resolves the name every iteration, so a replacement writer is picked up without any parameter changing; and the failure is bounded and loud, so a genuinely broken deployment fails a deploy instead of hanging forever in starting.

A complete example

A single-writer database that serves postgres and a metrics endpoint, wires itself up from platform-filled values, and branches on its slot:

artifactType: application/vnd.orc8r.app.v1
annotations:
  org.opencontainers.image.title: examplepg
  org.opencontainers.image.description: Single-writer database, slot 1 is the writer
config:
  params:
    type: object
    properties:
      slot:
        type: string
        title: Runtime slot
        x-source:
          kind: pool.slot
      writer:
        type: string
        title: Writer name
        x-source:
          kind: peers.first
  endpoints:
    pg:
      port: 5432
      probe:
        tcp: true
        interval: 5s
    metrics:
      port: 9187
      protocol: http
      probe:
        http: /metrics
  install:
    command: sh install-examplepg.sh
  start:
    command: sh start-examplepg.sh
files:
  - install-examplepg.sh
  - start-examplepg.sh
platforms:
  - os: linux
    arch: amd64

start-examplepg.sh branches on $SLOT: slot 1 initialises and serves, anything else seeds from $WRITER and follows it. The platform is told nothing about roles, and does not need to be.