← Back to dashboard
railwayhostingfreshReader view (for NotebookLM)

Railway Integration Guide

What is Railway?

The real model

The home for processes that outlive a request.

Railway's value is the shape it hosts, not the language it runs. Bind `process.env.PORT` on `0.0.0.0` or your health check never returns 200 and the deploy silently stalls for 300s — but bind IPv6 `::` for anything the private network must reach, which is the exact inverse rule and the classic Railway footgun. Config lives in `railway.json`, migrations belong in `deploy.preDeployCommand` (never `startCommand`), and reference variables like `${{Postgres.DATABASE_URL}}` wire services together without copy-pasting a secret. For codeAmani, this is where boda-dispatch's `webhook_jobs` worker belongs: a continuous `FOR UPDATE SKIP LOCKED` poller is a shape Vercel structurally cannot hold, while the Next.js front end stays on Vercel.

Six Railway primitives

One project, one private network — a container per job, with the database beside it.

Text
██████╗  █████╗ ██╗██╗     ██╗    ██╗ █████╗ ██╗   ██╗
██╔══██╗██╔══██╗██║██║     ██║    ██║██╔══██╗╚██╗ ██╔╝
██████╔╝███████║██║██║     ██║ █╗ ██║███████║ ╚████╔╝
██╔══██╗██╔══██║██║██║     ██║███╗██║██╔══██║  ╚██╔╝
██║  ██║██║  ██║██║███████╗╚███╔███╔╝██║  ██║   ██║
╚═╝  ╚═╝╚═╝  ╚═╝╚═╝╚══════╝ ╚══╝╚══╝ ╚═╝  ╚═╝   ╚═╝

Railway Integration Guide

Focus — Deploying long-running services, workers, cron jobs and managed databases on Railway from Claude Code: CLI, config-as-code, the public GraphQL API, and the private-network topology that keeps egress costs at zero.

Overview

Railway is a container hosting platform. You point it at a repo, it builds an OCI image (via Railpack, the successor to Nixpacks) and runs it as a long-lived process with a public HTTPS domain, TLS, and health-checked zero-downtime deploys.

The distinction that matters when choosing it:

ShapePlatformWhy
Next.js app, request/responseVercelServerless is the right fit; keep the default
Static site + edge functionsNetlify / CloudflareEdge-first
Process that must outlive a requestRailwayQueue consumers, pollers, WebSocket servers, schedulers
Managed Postgres + Redis on one private networkRailwayOne project, one private network, no egress fees between services

A Railway project contains services (each a deployed container) across environments (production, staging, PR environments). Services in the same project and environment reach each other over a private IPv6 network — traffic there is free and never leaves Railway.

Official Documentation

CLI setup

Bash
# Install (v5.49.2 at time of review)
npm install -g @railway/cli

railway login            # opens a browser; use `railway login --browserless` over SSH
railway init             # create a new project from the current directory
railway link             # or: attach this directory to an existing project

railway up               # build + deploy, streaming logs
railway up -d            # detached — don't stream
railway up --ci          # CI mode: no interactive prompts
railway up --service my-api --environment staging

Useful day-to-day commands:

Bash
railway add                        # add a service or a database (Postgres, Redis, MySQL, Mongo)
railway variables                  # list variables for the linked service
railway variables set KEY=value    # set one
railway run -- npm run dev         # run locally WITH the remote environment's variables injected
railway logs                       # tail deploy/runtime logs
railway open                       # open the project dashboard

railway run is the one to remember: it injects the live environment's variables into a local process, so local dev hits the same database and secrets as the deployed service without ever copying them into a .env file.

Config as code

Commit a railway.json (or railway.toml) next to your service. It overrides dashboard settings, so infrastructure changes ship in the same PR as the code.

JSON
{
  "$schema": "https://railway.com/railway.schema.json",
  "build": {
    "builder": "RAILPACK"
  },
  "deploy": {
    "startCommand": "node dist/worker.js",
    "preDeployCommand": "npm run db:migrate",
    "healthcheckPath": "/health",
    "healthcheckTimeout": 300,
    "restartPolicyType": "ON_FAILURE",
    "restartPolicyMaxRetries": 10
  }
}

The TOML form is equivalent:

TOML
[build]
builder = "railpack"
buildCommand = "npm run build"

[deploy]
preDeployCommand = ["npm run db:migrate"]
startCommand = "node dist/worker.js"
healthcheckPath = "/health"
healthcheckTimeout = 300
restartPolicyType = "on_failure"

Key fields:

  • build.builder — RAILPACK (default), DOCKERFILE, or NIXPACKS. Use DOCKERFILE with build.dockerfilePath when you need exact control.
  • deploy.preDeployCommand — runs to completion before the new version takes traffic. The correct home for migrations; a failure aborts the deploy.
  • deploy.healthcheckPath — Railway polls it until it returns HTTP 200, and only then swaps traffic to the new deployment. Anything else (including 503) stalls activation until healthcheckTimeout (default 300s), after which the deploy fails. Railway does not keep polling after go-live.
  • deploy.restartPolicyType — ON_FAILURE | ALWAYS | NEVER.
  • deploy.multiRegionConfig — replica counts per region, e.g. us-east4-eqdc4a, europe-west4-drams3a, asia-southeast1-eqsg3a.

railway.toml does not support volume-mount configuration — use railway.json or the dashboard for volumes.

The PORT contract

Railway injects a PORT environment variable and expects your server to bind it on 0.0.0.0. Hardcoding a port is the single most common cause of a service that builds fine and then fails its health check.

TypeScript
const port = Number(process.env.PORT) || 3000;
app.listen(port, "0.0.0.0", () => console.log(`listening on ${port}`));

If your app cannot listen on PORT (for example when using target ports), set a PORT variable explicitly so Railway probes the right one.

Variables, references and private networking

Railway variables are per-service, per-environment. Reference variables interpolate one service's value into another's, so a connection string is never copy-pasted:

Bash
# In the app service, referencing the Postgres service in the same project:
DATABASE_URL=${{Postgres.DATABASE_URL}}
REDIS_URL=${{Redis.REDIS_URL}}

# Reference another service's private address:
API_URL=http://${{api.RAILWAY_PRIVATE_DOMAIN}}:3000

Prefer the private form. Every service gets a RAILWAY_PRIVATE_DOMAIN resolvable only inside the project's IPv6 network:

  • Traffic over the private network is free — public egress is billed at $0.05/GB, and a chatty worker-to-database link over the public domain is a silent, recurring cost.
  • The database never needs a public endpoint at all.

Bind private listeners to IPv6 (::) — a server listening only on 0.0.0.0 is unreachable over the private network.

Railway also injects RAILWAY_ENVIRONMENT, RAILWAY_SERVICE_NAME, RAILWAY_PUBLIC_DOMAIN and RAILWAY_GIT_COMMIT_SHA — useful for tagging Sentry releases and structured logs.

Public GraphQL API

One endpoint: POST https://backboard.railway.com/graphql/v2.

Token typeHeaderScope
Account / Workspace / OAuthAuthorization: Bearer <TOKEN>Account or workspace-wide
Project tokenProject-Access-Token: <TOKEN>A single environment in one project
Bash
curl --request POST \
  --url https://backboard.railway.com/graphql/v2 \
  --header "Project-Access-Token: $RAILWAY_PROJECT_TOKEN" \
  --header 'Content-Type: application/json' \
  --data '{"query":"query { projectToken { projectId environmentId } }"}'

From TypeScript — note fetch, never a shell-based HTTP call:

TypeScript
const RAILWAY_API = "https://backboard.railway.com/graphql/v2";

export async function railwayQuery<T>(
  query: string,
  variables: Record<string, unknown> = {},
): Promise<T> {
  const token = process.env.RAILWAY_API_TOKEN;
  if (!token) throw new Error("RAILWAY_API_TOKEN is not set");

  const res = await fetch(RAILWAY_API, {
    method: "POST",
    headers: {
      "Content-Type": "application/json",
      Authorization: `Bearer ${token}`,
    },
    body: JSON.stringify({ query, variables }),
  });

  if (res.status === 429) {
    const retryAfter = res.headers.get("Retry-After") ?? "60";
    throw new Error(`Railway rate limit hit; retry after ${retryAfter}s`);
  }
  if (!res.ok) throw new Error(`Railway API ${res.status}: ${await res.text()}`);

  const body = (await res.json()) as { data?: T; errors?: { message: string }[] };
  if (body.errors?.length) throw new Error(body.errors.map((e) => e.message).join("; "));
  if (!body.data) throw new Error("Railway API returned no data");
  return body.data;
}

Rate limits (responses carry X-RateLimit-Limit, -Remaining, -Reset):

PlanPer hourPer second
Free100—
Hobby1,00010
Pro10,00050

Cron jobs

Set a Cron Schedule on a service and Railway runs its start command on that schedule. The rules are strict and worth internalising:

  • Standard 5-field crontab, UTC always — there is no timezone setting.
  • Minimum interval is 5 minutes.
  • The service must exit when the task finishes. A process that lingers (an open DB pool, a listening server) blocks the next run.
  • If the previous run is still going when the next fires, Railway skips the new one rather than running them concurrently.
Text
*/15 * * * *   # every 15 minutes
0 3 * * *      # 03:00 UTC daily  (= 06:00 EAT; Kenya is UTC+3 year-round)

Because the schedule is UTC and East Africa Time has no DST, an EAT-local job is a fixed −3h offset — 0 3 * * * is reliably 6am in Nairobi.

CI/CD with GitHub Actions

Deploy on green tests using a project token stored as a repository secret:

YAML
# .github/workflows/railway-deploy.yml
name: Deploy to Railway
on:
  push:
    branches: [master]

jobs:
  deploy:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4
      - uses: actions/setup-node@v4
        with:
          node-version: 24
          cache: npm
      - run: npm ci
      - run: npm test
      - name: Deploy
        run: |
          npm install -g @railway/cli
          railway up --ci --service "$RAILWAY_SERVICE"
        env:
          RAILWAY_TOKEN: ${{ secrets.RAILWAY_TOKEN }}
          RAILWAY_SERVICE: api

RAILWAY_TOKEN in the environment authenticates the CLI non-interactively — no railway login step. Scope it to a project token so a leaked CI secret can touch exactly one environment.

Pricing model

Usage-based, billed per minute of actual consumption:

ResourceRate
Memory$10 / GB / month ($0.000231 / GB / min)
vCPU$20 / vCPU / month ($0.000463 / vCPU / min)
Network egress$0.05 / GB
PlanSubscriptionIncluded usagePer-service ceiling
Hobby$5/mo$548 GB RAM, 48 vCPU, 6 replicas
Pro$20/mo$201 TB RAM, 1,000 vCPU, 42 replicas
EnterpriseCustomCustom2.4 TB RAM, 2,400 vCPU, 50 replicas

The subscription includes an equal amount of usage, so a small always-on worker on Hobby is often fully covered. Because billing tracks provisioned resources over time, an idle service still costs memory — size containers deliberately rather than leaving defaults.

codeAmani notes

Security

  • Railway tokens (RAILWAY_TOKEN, RAILWAY_API_TOKEN) are server-side only. Never expose one to the browser and never prefix it NEXT_PUBLIC_.
  • Prefer project tokens over account tokens for CI — blast radius is one environment instead of the whole workspace.
  • Store secrets in Railway's variable store (or bind them through Hazina); never commit them. railway run is the sanctioned way to use production variables locally without writing them to disk.
  • Keep databases on the private network only. A managed Postgres with a public proxy endpoint is an internet-reachable database — remove the public endpoint unless something outside Railway genuinely needs it.
  • Webhook receivers deployed here still owe signature verification: the Stripe signing secret, Svix for Clerk, and M-Pesa callback validation for Kenya-targeted projects. Railway terminating TLS proves nothing about the sender.

Where Railway fits our stack

Vercel remains the default for Next.js front ends. Railway earns its place for the part Vercel structurally cannot host — a process that outlives a request:

  • The webhook_jobs durable queue in boda-dispatch is claimed with FOR UPDATE SKIP LOCKED by a worker that polls continuously. That is a background-worker shape, not a serverless one: a Vercel function would have to be re-invoked by cron and would contend for locks on every tick.
  • M-Pesa reconciliation pollers, WhatsApp session keep-alives, and Daraja token refreshers (OAuth2 tokens expire hourly) all want one long-lived process holding state, not N cold starts.
  • Daraja requires HTTPS callbacks. A Railway service gets a public HTTPS domain immediately, which beats ngrok for a shared staging environment.

Kenya-targeted projects

  • The nearest region to East Africa is europe-west4-drams3a (Netherlands); asia-southeast1-eqsg3a is the alternative. Neither is close, so keep chatty round-trips off the user path — do M-Pesa STK Push server-side and let the callback carry the result rather than long-polling from a 2G handset.
  • Co-locate the worker and its Postgres in the same project so the hot path runs over the private network; only the user-facing hop crosses the public internet.
  • Cron runs in UTC, which maps to EAT at a fixed −3h (no DST).

Provenance

A Railway deploy is a deploy, not a downloadable artifact — per our SLSA policy that means no provenance target. Pin the GitHub Actions used in the deploy workflow to commit SHAs and document the build; there is nothing for slsa-verifier to verify. If a project additionally ships a container image or release tarball, that artifact takes Build L3 on its own track — see supply-chain/CLAUDE_CODE_INTEGRATION.md.

Troubleshooting

SymptomCauseFix
Build succeeds, deploy never activatesHealth check never returns 200Bind process.env.PORT on 0.0.0.0; confirm healthcheckPath exists and is unauthenticated
service unavailable on health checkHardcoded port, or target ports in useBind PORT, or set a PORT variable telling Railway which port to probe
Service unreachable over private networkListening on 0.0.0.0 onlyBind IPv6 (::) for private-network traffic
Cron job never fires againPrevious run never exitedClose DB pools and exit; overlapping runs are skipped, not queued
Cron fires less often than expectedInterval below the floorMinimum is 5 minutes
Unexpected egress chargesServices talking over public domainsSwitch to RAILWAY_PRIVATE_DOMAIN / reference variables
429 from the GraphQL APIPlan rate limitHonour Retry-After; batch queries
Migrations race the new deployMigration in startCommandMove it to deploy.preDeployCommand