When something on the network does not work, ORC8R tells you what went wrong in words. A failure shows up in one of three places: when you submit a request, on the pool, or on the endpoint. Each prints its own reason. This page is organised by what you saw, and it quotes the messages exactly, so you can search it for the text on your screen.
The public edge — DNS records, addresses, and certificates — is managed for you on ORC8R Cloud, so the failures below are the ones that are yours to fix: how your app binds, how it answers its probe, and how names resolve between your nodes. If a public name will not resolve or a certificate never appears for a correctly exposed https service, that is a platform matter — check the request was accepted, then contact support.
Where failures show up
When you submit a request. Anything ORC8R can decide from the request itself is refused the moment you submit it: an endpoint the app does not declare, a setting an endpoint does not have, two apps asking for the same port, a service name someone else has already claimed. The message names the exact field that is wrong, as a dotted path:
apps.postgres.endpoints.pg.name: service name "maindb" is already claimed by pool "db" in project "shop"
apps.postgres.endpoints.pg.weight: unknown endpoint decoration field; accepted: expose, name, port
apps.web.endpoints.api: app "web" does not declare endpoint "api"; declared: http, metrics
tcp port 8080 is claimed by both "web" endpoint "http" and "proxy" endpoint "front"; override one with an endpoint port
None of this reaches a node, and nothing is dropped quietly. An exposure is either accepted or refused; it is never accepted and then forgotten. Fix the field the message names and submit again.
On the pool. A pool that is blocked or rejected carries a state_change_reason — one line saying why. A configuration failure that would fail the same way every time blocks the pool immediately, instead of retrying into a wall:
configuration failure: <what the node reported>
3 consecutive node failures: <reason>
On the endpoint. Once your app is deployed, each combination of node, app and endpoint carries a status: ready, unready, or no report at all, which reads as unknown. A node counts as endpoint-healthy only when the node is online and the endpoint reports ready. An unready status carries the probe's own words.
An endpoint never turns ready
The status carries the probe's own words. These are the ones you will see, and what each one means:
connect to 100.64.0.3 port 5432 failed: Connection refused the app is not listening there
connect to 100.64.0.3 port 5432 timed out a firewall on the node, or the app is wedged
http probe /healthz on 100.64.0.3 returned 503 the app is up and says it is not ready
http probe /healthz on 100.64.0.3 failed: <transport error> nothing is speaking HTTP on that port
probe command exited with exit status: 1 your probe command's own verdict
probe command timed out raise `timeout`, or make the check cheaper
Every address-dialing reason names the address that answered that way, and that address is the diagnosis. A probe dials every address the endpoint is served on — the node's own, plus the pool address a public exposure delivers it on — and passes only when all of them answer.
A Connection refused that never clears is almost always a listen address:
- On the node's own address: the app bound loopback. An app on
127.0.0.1reportsunreadyforever while looking perfectly healthy over SSH. - On the pool address, while the node's own address answers: the app did not bind that family. Public delivery hands your node the client's packets unrewritten, so the app needs a socket on the pool address — usually IPv6 — and a
0.0.0.0bind is IPv4 only. Internal consumers are served throughout, which is why nothing else looks wrong.
Bind [::]. See Authoring apps.
The status does not flip on the first blip, so give it time before you read it: three consecutive failures to go unready, one success to come back, on a 10-second interval by default.
A name does not answer, or your app cannot find another service
Internal names are answered by the resolver inside the agent on the node itself. When a name gives you nothing, work down this list:
- Is the node online and endpoint-healthy? Address answers follow node liveness; SRV answers follow endpoint health. The plain name —
db.shop.internal, the one that means the pool — answers empty when no node it may offer is online. - Is it a slot name?
db-1.shop.internalanswers nothing while slot 1 is unhealthy, by design — a slot name is fenced to its own node. It never substitutes another node. If your client is happy with any node, ask for the plain name instead — but note that the plain name is fenced to slot 1 too whenever any endpoint of the app servesprimary, which is the default (Authoring apps). - Are you calling from another project? Then write the name out in full. A node adds its own project's suffix to short names, but that shortcut is not extended across projects, and the endpoint must also be exposed at
orgscope. From outside its own project, an unexposed sibling's name answersNOT_FOUND. Expose the endpoint atorgscope, and use the full name. See Project and organization exposure. - Is the resolver installed on this node? When it is, the agent logs
internal DNS resolver listeningandinternal DNS zone routed to the agent resolver, naming the mechanism it used. If instead you seeresolv.conf cannot address a resolver off port 53; internal names will not resolve, the node's operating system cannot route the internal suffix to the resolver's address, and no internal name will ever resolve on that node.
A name inside the internal suffix that names nothing answers NXDOMAIN, so a resolver walking its search list moves on rather than stopping there.
Negative answers, slot answers and SRV answers carry a 1-second TTL; spread answers 5 seconds. If a stale answer persists much longer than that, the cache holding it is not the platform's: application-level resolvers that cache forever are the usual culprit — the JVM's default is the classic — and networkaddress.cache.ttl is the usual fix.
Common traps
An identical reconfiguration is a no-op. There is no "touch" that re-runs anything. Re-submitting a byte-identical request changes nothing, and an address set that is already in place costs the agent no commands. The composer even leaves controls at their defaults out of the document it submits, precisely so that forking a pool and changing nothing stays a genuine no-op. If you want something re-done, change it, or restart the component whose state you suspect.
An endpoint goes unready with Connection refused on the pool address. The app bound one address family and the pool address is the other one — a 0.0.0.0 listener under a v6 pool address is the case that happens. Nothing else shows it: the app runs, its internal name answers, and only the public traffic — which arrives on the pool address unrewritten — is refused. Bind [::], or open a listener per family. See Authoring apps.
An endpoint setting the platform does not have is rejected, not ignored. An endpoint takes expose, name and port and nothing else; anything else is refused at submission, naming the dotted path of the key. A silently dropped setting is a misconfiguration you would find out about much later.
An unresolvable app cannot be exposed publicly. An app that does not resolve in the project contributes no endpoints to the pool surface, so its public exposure would vanish along with it. Submission refuses instead. Publish the app, or pin a version that exists.
Service names are quarantined after release. Another project cannot take a name that was recently released. The project that released it can re-claim it immediately — that asymmetry is what lets a service migrate between pools without a gap.
Related pages
- Networking overview — the model these failures are defending.
- Authoring apps — probes, and binding
[::]. - Project and organization exposure — the two internal scopes and their names.
- Public exposure — flipping an endpoint to the internet.
- Support — when a public name or certificate needs the platform's attention.