Logs, metrics and alerts
← Docs · Run in production

Logs, metrics and alerts

Live logs, log history and retention, saved searches, log drains, access logs, OpenTelemetry export, metric alerts and the background job queue view.

Everything here is on the app page, on the Logs and Metrics tabs, and in the API. Nothing needs installing in your app, except the OpenTelemetry SDK if you want traces.

Live logs

The Logs tab streams what your containers print, as they print it. You can filter by process. railyard logs my-app --follow --process web does the same from a terminal. Download saves the recent output.

Log history and retention (beta)

Railyard also stores every line, so the Log history card can search back in time.

  • How long logs are kept: Team → Logs → Keep app logs for (days). The default is 7 days.
  • Time range: last hour, 24 hours, 7 days, 30 days, or between two dates and times.
  • Search: plain text, or /regex/ (/regex/i ignores case). You can also filter by minimum level, process, and for access logs by status and path.
  • Paging: 200 lines at a time, newest first. Older results carries on from where you stopped.
  • Download these lines saves the range as a text file.

Saved searches and log alerts

Save a search, and optionally alert on it: N matching lines within M minutes. Alerts go to notification rules with log alerts ticked, at most once per window.

Log drains

Logs tab → Log drains sends every log line to your own sink, either an HTTPS endpoint that takes JSON or a syslog server (syslog://logs.example.com:514). Access log lines arrive with source: "router".

Access logs

Every request that reaches your app through Railyard’s proxy becomes one router line:

at=info method=GET path="/cart?id=7" host=shop.example.com status=200 duration=12.3ms bytes=532 fwd="203.0.113.9" user_agent="curl/8.4.0"

5xx responses are logged at error level, so error filters and log alerts see them. Access logs show up in live logs, log search, downloads and drains. Railyard’s own uptime probes are left out. To turn them off, untick Settings → Networking → Access logs → Collect access logs. That takes effect at once, with no deploy needed.

Metrics

The Metrics tab shows requests per minute, p50/p95/p99 latency, error rate, bandwidth and the slowest endpoints.

Metric alerts

On the Metrics tab, the Alerts card lets deployers and above add rules such as p95 response time above 800 ms for 5 minutes. You can alert on:

  • p95 response time, in ms;
  • 5xx error rate, in %;
  • memory, as a % of the container’s limit;
  • CPU, in % (100 means one core).

Rules are checked every minute. A rule fires once the value has stayed above the threshold for the whole duration, and one dip below resets the clock. When it fires, a message goes to notification rules with metric alerts ticked, and an entry opens in the card’s History with the start time, how long it lasted and the peak. When the value drops back, a back to normal message goes out. You can pause and resume rules.

Background job queues (beta)

For apps running Sidekiq, Solid Queue, GoodJob, Resque or Delayed Job, the Metrics tab shows a Background jobs card. Railyard reads it from your app’s own Redis or Postgres every minute. It shows:

  • the queue depth and the oldest waiting job;
  • how many jobs are busy, processed and failed;
  • the 50 most recent failures, each with Retry and Discard;
  • a 24-hour chart.

When the queue stays backed up, rules with job backlog ticked get a message, and another when it has caught up.

It can’t read a Redis outside Railyard, Solid Queue on MySQL, or Sidekiq namespaces.

OpenTelemetry

Settings → General → OpenTelemetry. Enter an OTLP endpoint, a protocol (http/protobuf, http/json or grpc) and optional headers such as x-honeycomb-team=YOUR_KEY. Headers are stored encrypted. On the next deploy, Railyard sets the variables every OpenTelemetry SDK reads:

  • OTEL_EXPORTER_OTLP_ENDPOINT, OTEL_EXPORTER_OTLP_PROTOCOL and OTEL_EXPORTER_OTLP_HEADERS;
  • OTEL_SERVICE_NAME, which is the app’s name;
  • OTEL_RESOURCE_ATTRIBUTES, which carries the commit SHA, the environment, the app and the team.

A variable you set yourself always wins.

Tick also send Railyard’s CPU, memory and request metrics to post them to <endpoint>/v1/metrics every minute. That export always uses OTLP over HTTP (port 4318 on most collectors).

To instrument the app, add the vendor-neutral SDK and let it read those variables. For Rails that’s opentelemetry-sdk, opentelemetry-exporter-otlp and opentelemetry-instrumentation-all with c.use_all, on the http/protobuf protocol. For Node it’s @opentelemetry/auto-instrumentations-node/register, and for Python it’s opentelemetry-instrument.

API

MethodPathPurpose
GET/api/applications/:id/logsRecent container logs
GET/api/applications/:id/log_searchStored lines: q, level, process, status, path, range or from/to
GET/api/applications/:id/log_downloadThe same filters, as text
GET / PATCH / DELETE/api/applications/:id/otel_exportOpenTelemetry settings
GET / POST/api/applications/:id/metric_alertsAlert rules and their history
GET/api/applications/:id/job_queueQueue stats and failed jobs, with retry and discard