Skip to content
Runtime supervisor

Runtime supervisor

Every Anetos application runs its long-lived work (HTTP servers, queue workers, pub/sub listeners, the scheduler, background tasks) as goroutines inside one process, managed by a supervisor. It decides what runs, what happens when something fails, and how everything stops.

Components

Anything long-running is a component:

// illustrative (from supervisor/component.go)
type Component interface {
	Name() string
	Run(ctx context.Context) error
}

Run blocks until the work is done or ctx is canceled. On cancellation, the component stops taking new work, finishes or hands back what’s in flight, and returns. app.Go(name, fn) wraps a function as a component; app.Component(c) adds your own type.

Roles: one binary, many shapes

Components declare roles. A process can run all of them or only some:

./blog run                            # everything: dev and small deployments
./blog run --only=http                # web machines
./blog run --only=workers            # background machines
./blog run --only=scheduler          # one machine, or several with OnOneServer

The built-in components with roles are the web server (http), the queue’s workers (workers, from q.Work; see Queues) pub/sub listeners (listeners, from pubsub.Listen; see Pub/sub listeners) and the scheduler (scheduler, from schedule.ForApp once it has tasks; see Scheduling). Your own components choose theirs with anetos.Roles("workers").

run is the binary’s default command (app.Execute; see Commands); in code, pass roles to app.Run(ctx, "http").

  • Components without roles run in every process.
  • Asking for a role no component declares is an error, so typos in --only fail loudly.

Split processes share work through shared stores: the database or Redis for the queue, the cache (database or Redis) for the scheduler’s OnOneServer locks, and a broker (Redis, Google Pub/Sub) for pub/sub; the memory drivers stay inside one process. examples/saas runs one binary as four processes, http, workers, listeners and scheduler, and its roles_test.go follows a sign-up from one to the other.

Failure policies

PolicyOn error or panicUse for
RestartNever (default)Logged; component stays stopped; app keeps runningOne-off background tasks
RestartOnFailureRestarted after exponential backoff with jitterWorkers, listeners, pollers
StopOnFailureWhole app shuts down; Run returns the errorComponents the app can’t live without, e.g. the HTTP server

Panics are recovered and treated as failures, with the stack trace logged.

Backoff starts at Initial (default 1s), doubles each consecutive failure up to Max (default 30s), and varies by ±20% so many instances don’t restart in lockstep. A component that ran longer than Max before failing counts as healthy again, and its delay resets. If MaxRestarts (default unlimited) is exceeded, the failure escalates like StopOnFailure, so a crash-looping process exits and your orchestrator (systemd, Kubernetes…) can take over.

Staged shutdown

When the run context is canceled (usually by SIGINT or SIGTERM) or a component escalates a failure, the supervisor stops components in stages. Each stage is canceled only after the previous one has fully drained:

OrderStageWhy this order
1StageIngress (HTTP)Stop accepting requests; finish in-flight ones
2StageSchedulerStop starting scheduled runs, which produce jobs
3StageListenersStop pulling external messages, which also produce jobs
4StageWorkersFinish in-flight jobs, including ones just produced above
5StageBackgroundAd-hoc tasks (app.Go)

The rule is producers stop before consumers, so work already accepted gets done. All stages share one deadline (in an app, APP_SHUTDOWN_TIMEOUT minus the part reserved for shutdown hooks). If it passes, the remaining stages are canceled together with a short grace period, and Run returns an error naming only the components that are still running, so the process can exit anyway.

If every component finishes on its own, Run returns too. From that moment new components are refused with ErrStopping rather than being started and immediately canceled.

Warning: A component that ignores ctx can’t be stopped gracefully. It will be reported as stuck when the deadline passes.

Health

  • app.Supervisor().Ready() is true while running and not shutting down, and every component that implements Ready() bool is ready. Such a component that is waiting to restart or has failed counts as not ready. It will back the planned readiness endpoint.
  • app.Supervisor().ShutdownDeadline() is when the components’ share of APP_SHUTDOWN_TIMEOUT runs out, once shutdown has begun: queue workers use it to stop their jobs in time.
  • app.Supervisor().Status() lists each component’s state (pending, starting, running, backoff, done, failed, stopped), restart count and last error.

Coming from Laravel? This replaces running php artisan queue:work, schedule:run and Supervisor (the process manager) as separate programs. Everything runs in your app’s own binary, and --only splits it apart when you need to scale.

Related