Logs, metrics and alerts
Live logs, log history and retention, saved searches, log drains, access logs, OpenTelemetry export, metric alerts and the background job queue view.
Everything here is on the app page, on the Logs and Metrics tabs, and in the API. Nothing needs installing in your app, except the OpenTelemetry SDK if you want traces.
Live logs
The Logs tab streams what your containers print, as they print it. You can filter by process. railyard logs my-app --follow --process web does the same from a terminal. Download saves the recent output.
Log history and retention (beta)
Railyard also stores every line, so the Log history card can search back in time.
- How long logs are kept: Team → Logs → Keep app logs for (days). The default is 7 days.
- Time range: last hour, 24 hours, 7 days, 30 days, or between two dates and times.
- Search: plain text, or
/regex/(/regex/iignores case). You can also filter by minimum level, process, and for access logs by status and path. - Paging: 200 lines at a time, newest first. Older results carries on from where you stopped.
- Download these lines saves the range as a text file.
Saved searches and log alerts
Save a search, and optionally alert on it: N matching lines within M minutes. Alerts go to notification rules with log alerts ticked, at most once per window.
Log drains
Logs tab → Log drains sends every log line to your own sink, either an HTTPS endpoint that takes JSON or a syslog server (syslog://logs.example.com:514). Access log lines arrive with source: "router".
Access logs
Every request that reaches your app through Railyard’s proxy becomes one router line:
at=info method=GET path="/cart?id=7" host=shop.example.com status=200 duration=12.3ms bytes=532 fwd="203.0.113.9" user_agent="curl/8.4.0"
5xx responses are logged at error level, so error filters and log alerts see them. Access logs show up in live logs, log search, downloads and drains. Railyard’s own uptime probes are left out. To turn them off, untick Settings → Networking → Access logs → Collect access logs. That takes effect at once, with no deploy needed.
Metrics
The Metrics tab shows requests per minute, p50/p95/p99 latency, error rate, bandwidth and the slowest endpoints.
Metric alerts
On the Metrics tab, the Alerts card lets deployers and above add rules such as p95 response time above 800 ms for 5 minutes. You can alert on:
- p95 response time, in ms;
- 5xx error rate, in %;
- memory, as a % of the container’s limit;
- CPU, in % (100 means one core).
Rules are checked every minute. A rule fires once the value has stayed above the threshold for the whole duration, and one dip below resets the clock. When it fires, a message goes to notification rules with metric alerts ticked, and an entry opens in the card’s History with the start time, how long it lasted and the peak. When the value drops back, a back to normal message goes out. You can pause and resume rules.
Background job queues (beta)
For apps running Sidekiq, Solid Queue, GoodJob, Resque or Delayed Job, the Metrics tab shows a Background jobs card. Railyard reads it from your app’s own Redis or Postgres every minute. It shows:
- the queue depth and the oldest waiting job;
- how many jobs are busy, processed and failed;
- the 50 most recent failures, each with Retry and Discard;
- a 24-hour chart.
When the queue stays backed up, rules with job backlog ticked get a message, and another when it has caught up.
It can’t read a Redis outside Railyard, Solid Queue on MySQL, or Sidekiq namespaces.
OpenTelemetry
Settings → General → OpenTelemetry. Enter an OTLP endpoint, a protocol (http/protobuf, http/json or grpc) and optional headers such as x-honeycomb-team=YOUR_KEY. Headers are stored encrypted. On the next deploy, Railyard sets the variables every OpenTelemetry SDK reads:
OTEL_EXPORTER_OTLP_ENDPOINT,OTEL_EXPORTER_OTLP_PROTOCOLandOTEL_EXPORTER_OTLP_HEADERS;OTEL_SERVICE_NAME, which is the app’s name;OTEL_RESOURCE_ATTRIBUTES, which carries the commit SHA, the environment, the app and the team.
A variable you set yourself always wins.
Tick also send Railyard’s CPU, memory and request metrics to post them to <endpoint>/v1/metrics every minute. That export always uses OTLP over HTTP (port 4318 on most collectors).
To instrument the app, add the vendor-neutral SDK and let it read those variables. For Rails that’s opentelemetry-sdk, opentelemetry-exporter-otlp and opentelemetry-instrumentation-all with c.use_all, on the http/protobuf protocol. For Node it’s @opentelemetry/auto-instrumentations-node/register, and for Python it’s opentelemetry-instrument.
API
| Method | Path | Purpose |
|---|---|---|
| GET | /api/applications/:id/logs | Recent container logs |
| GET | /api/applications/:id/log_search | Stored lines: q, level, process, status, path, range or from/to |
| GET | /api/applications/:id/log_download | The same filters, as text |
| GET / PATCH / DELETE | /api/applications/:id/otel_export | OpenTelemetry settings |
| GET / POST | /api/applications/:id/metric_alerts | Alert rules and their history |
| GET | /api/applications/:id/job_queue | Queue stats and failed jobs, with retry and discard |