Diff v1 → v2

v1: bot backfill · 2026-06-23 07:11:37
v2: bot legacy · 2026-07-23T08:15:50Z
  # Rails Production Playbook
  
  A short, opinionated reference for building serious Rails apps. The **principles**
  are language-agnostic — anyone can lift them — but the **gems and snippets** are
  Rails-first. Default posture: *boring, explicit, deny-by-default, observable.*
  
  > Rule of thumb: reach for a pattern when the pain it removes is real, not
  > hypothetical. Every item below has a "use when" so you don't over-engineer.
  
  ---
  
  ## Architecture & code design
  
  ### Sandi Metz discipline — lean controllers, service objects, enums
  Keep classes small and single-purpose. Controllers only orchestrate: authenticate,
  authorize, call one object, render. Push logic into POROs ("service objects").
  - **Guidelines:** ≤100 lines/class, ≤5 lines/method, ≤4 params, one job per object.
  - **Gems:** plain Ruby first; `interactor` or `dry-monads` if you want a Result type.
  - **Use when:** always for controllers; extract a service the moment an action does
    more than one thing or touches more than one model.
  - Use Rails `enum` for finite states that are *attributes*; use a **state machine**
    (below) when transitions have rules.
  
  ### Service objects + Result types (dry-monads / dry-contracts)
  Make success/failure explicit instead of raising for control flow.
  - **Gems:** `dry-monads` (`Success`/`Failure`, `Do` notation), `dry-validation`
    / `dry-contracts` (typed, composable input validation at the boundary).
  - **Use when:** multi-step operations that can fail in known ways (payments,
    imports, signups). Validate untrusted input with a contract *before* the service.
  ```ruby
  class CreateInvoice
    include Dry::Monads[:result, :do]
    def call(params)
      attrs   = yield InvoiceContract.new.(params).to_monad   # Failure stops the chain
      invoice = yield persist(attrs)
      Success(invoice)
    end
  end
  ```
  
  ### Slice / vertical-feature development
  Organize by **feature**, not by layer. A change should touch one slice
  front-to-back rather than smearing across `app/*`. Enforce boundaries in a modular
  monolith with **Packwerk** (`packs/`), split to services only when a slice needs
  independent scaling or deploys.
  - **Use when:** the app outgrows "one big `app/`" and teams step on each other.
  
  ### EventBus / pub-sub
  Decouple side effects (emails, audit, search reindex) from the core write.
  - **In-process, cheap:** `ActiveSupport::Notifications` or `wisper`.
  - **Durable / event-sourced:** `rails_event_store`.
  - **Use when:** one action fans out to ≥2 unrelated reactions. Keep the *write*
    transactional; emit events `after_commit` so listeners never see rolled-back data.
  
  ### Enums — named values, never magic strings
  Raw strings scattered through code (`"pending"`, `"active"`, `where(status: "shipped")`)
  are typo-prone and have no single source of truth. Centralize the allowed values.
  - **Rails `enum`:** maps a column to named values; gives you scopes
    (`Order.shipped`), predicates (`order.shipped?`), and setters (`order.shipped!`).
    Define the mapping **explicitly** so values are stable — never rely on positional
    integers, and prefer a string-backed column for readability in the DB.
  - **Richer enums (behavior/metadata/i18n per value):** `enumerize`, or a small PORO
    registry / value object when each value carries logic. Translate display labels
    via i18n — don't hardcode human strings either.
  - **Use when:** any attribute drawn from a fixed, known set (status, role, kind,
    priority). One definition, referenced everywhere.
  ```ruby
  class Order < ApplicationRecord
    enum :status, { draft: "draft", paid: "paid", shipped: "shipped" }, default: :draft
    # → Order.paid, order.paid?, order.shipped!, Order.statuses
  end
  ```
  - **Relationship to state machines:** an `enum` lists the *values*; a **state
    machine** governs the *transitions* between them. If moving between values has
    rules, guards, or callbacks (`draft → paid → shipped`, never `draft → shipped`),
    let the state machine own the column — it becomes the single source of truth for
    both the allowed states *and* the legal moves. Plain `enum` is the right tool only
    when any value can follow any other.
  
  ### State machines (AASM)
  Model lifecycles with explicit, guarded transitions and callbacks.
  - **Gems:** `aasm` (attribute-based, ergonomic — status lives on the model);
    `statesman` (Gusto — stores each transition as a row in its own table, giving a
    built-in audit trail and history, at the cost of more setup);
    `state_machines-activerecord` (another attribute-based option).
  - **Pick:** `aasm` for simple lifecycles on the record; `statesman` when you need the
    full transition history / audit of *how* it got to a state.
  - **Use when:** an object moves through statuses with rules (`draft → published →
    archived`), not just a flat enum.
  
  ---
  
  ## Data & persistence
  
  ### IDs: bigint internal, UUID external, friendly_id for URLs
  - **`bigint`** primary keys for internal joins (fast, compact).
  - **`uuid`** (a separate `public_id`) for anything exposed in APIs/URLs — never leak
    sequential counts or enable ID enumeration.
  - **`friendly_id`** for human-readable slugs (`/users/jane-doe`, `/articles/the-title`).
  - **Use when:** any app with a public surface. Add the UUID column from day one.
  
  ### Multi-tenancy: RLS / ActsAsTenant as first class
  Make tenant isolation a property of the data layer, not something each query
  remembers to add.
  - **`acts_as_tenant`:** sets a current tenant and auto-scopes every query.
  - **Postgres RLS (row-level security):** the strongest option — the database itself
    refuses cross-tenant rows even if app code forgets. Belt *and* suspenders.
  - **Use when:** B2B / SaaS with shared tables. Decide *before* writing schema;
    retrofitting tenancy is painful. If you don't need tenancy, don't add it.
  ```sql
  ALTER TABLE invoices ENABLE ROW LEVEL SECURITY;
  CREATE POLICY tenant_isolation ON invoices
    USING (account_id = current_setting('app.current_account_id')::bigint);
  ```
  
  ### JSONB for metadata
  Use a `jsonb` column for sparse, schemaless, or fast-evolving attributes —
  *not* as an excuse to avoid real columns for things you query or index heavily.
  - **Rails:** `store_accessor :metadata, :referrer, :utm_source`; GIN-index when you
    query inside it.
  - **Use when:** flexible per-record extras (settings, integration payloads,
    feature config). Promote a key to a real column once you filter/sort on it.
  
  ### Transactions — when & how
  Wrap writes that must succeed or fail **together**.
  - **Do:** group related `INSERT`/`UPDATE`s; use DB constraints as the real guarantee;
    put side effects in `after_commit`, never mid-transaction.
  - **Don't:** make HTTP/network calls inside a transaction (holds locks, can't roll
    back the remote side); don't span a transaction across user think-time.
  - Use advisory locks / `with_lock` for "check-then-act"; pick an isolation level
    deliberately when you have read-modify-write races.
  
  ### Auditing (paper_trail / audited / logidze)
  Track who changed what, when — for compliance, debugging, and undo.
  - **`paper_trail`:** version history with whodunnit, restore previous versions.
  - **`audited`:** lighter change log.
  - **`logidze`:** Postgres-trigger based (no app-layer overhead).
  - **Use when:** regulated data, financial records, or anything users dispute.
  
  ---
  
  ## Security
  
  ### Policies, zero trust, deny-by-default
  Authorization is a first-class layer, not scattered `if current_user.admin?`.
  - **Gems:** `pundit` or `action_policy` (policy objects per resource);
    `rack-attack` (rate limit / block); `brakeman` + `bundler-audit` in CI.
  - **Posture:** **deny by default** — no policy means no access. Authorize every
    action; never trust client-supplied IDs, scopes, or roles. Strong params always.
  - **Use when:** always. Add a default `deny` and opt routes *in*.
  
  ### Secrets vs config — credentials vs ENV vs YAML
  Three buckets, don't mix them:
  - **`.yml`-backed config objects** for non-secret, environment-shaped values:
    URLs, database names, endpoints, timeouts, feature defaults.
    Gems: `config` (`Settings.x.y`) or `anyway_config`.
  - **Rails encrypted credentials** for *secrets*: DB passwords, integration API
    keys, signing keys (`rails credentials:edit`, committed encrypted).
  - **ENV** only for *runtime-start* values that differ per host/boot and aren't
    secret by nature (e.g. which credentials key to load, `RAILS_MAX_THREADS`).
  - **Rule:** if leaking it is a breach → credentials. If it just shapes the
    environment → yml. If the platform injects it at boot → ENV.
  
  ---
  
  ## Reliability & resilience
  
  ### HTTP errors the RFC way (RFC 9457 `problem+json`)
  Return machine-readable errors with a stable shape, not ad-hoc JSON.
  - **Shape:** `type`, `title`, `status`, `detail`, `instance` (+ `errors`, `trace_id`).
  - **Rails:** `rescue_from` → render `application/problem+json`. No dedicated gem
    needed; a thin renderer is enough.
  - **Use when:** any API with external consumers (incl. your own SPA / agents).
  ```ruby
  rescue_from Pundit::NotAuthorizedError do |e|
    render json: { type: "/errors/forbidden", title: "Forbidden",
                   status: 403, detail: e.message, trace_id: request.request_id },
           status: :forbidden, content_type: "application/problem+json"
  end
  ```
  
  ### Idempotency
  Make retried writes safe — same request, same result, no duplicates.
  - **How:** accept an `Idempotency-Key` header; store key + response; replay on
    repeat; reject body mismatch (`422`). Back it with a unique index.
  - **Gems:** roll your own table, or `sidekiq-unique-jobs` for job-level dedup.
  - **Use when:** payments, order creation, anything a client/network may retry.
  
  ### Circuit breakers
  Stop hammering a failing dependency; fail fast and recover automatically.
  - **Gems:** `stoplight` (general), `semian` (Shopify — for net/http, MySQL, Redis),
    `faraday-retry` for retries with backoff.
  - **Use when:** you call flaky third parties or internal services over the network.
  
  ### Feature flags
  Decouple deploy from release; kill-switch risky paths; gradual rollout.
  - **Gem:** `flipper` (+ `flipper-active_record`, `flipper-ui`).
  - **Use when:** trunk-based dev, canary releases, per-account betas, ops kill-switch.
  
  ### `/health` — real readiness, not just "the web server is up"
  Verify the whole boot: DB reachable, migrations current, Redis/queue up,
  required integrations reachable. Separate **liveness** (am I running?) from
  **readiness** (can I serve traffic?).
  - **Rails 7.1+:** `/up` exists but only proves the process booted — extend it.
  - **Gems:** `okcomputer` or `health_check` for multi-dependency checks.
  - **Use when:** any real deployment — load balancers and orchestrators gate on it.
  
+ ### Versioning & releases
+ Every deploy should carry a real, traceable version — not just a git SHA in your head.
+ If you deploy with **Kamal** (Basecamp's own, very Rails-idiomatic tool), auto-bump a
+ semantic version from your Conventional Commits and tag the release **only on a healthy
+ deploy**, via `.kamal/hooks/` pre/post-deploy hooks. Full recipe:
+ [[kamal-auto-version|Kamal Auto-Versioning]].
+ - **Tooling:** `svu` (Conventional Commits → semver: `feat:`→minor, `fix:`→patch,
+   `!`/`BREAKING CHANGE:`→major), a `pre-deploy` hook to preview/guard, a `post-deploy`
+   hook to `git tag` after success.
+ - **Surface it:** expose the version + git SHA on your deep `/health` (above) so a running
+   instance self-reports exactly what shipped — invaluable during an incident. (In Rails,
+   read it from an ENV/initializer rather than the Go `ldflags` trick in the reference.)
+ - **Use when:** any Kamal-deployed app — so every release is tagged in git and traceable.
+ 
  ---
  
  ## Scalability
  
  - **Vertical:** bigger box first — simplest win until a single node is maxed.
  - **Horizontal:** more app servers behind a balancer; tune **Puma** workers/threads
    and the DB **connection pool** together; move work to background jobs
    (`Sidekiq` / `GoodJob` / `Solid Queue`).
  - **Database:** read replicas + Rails multi-DB (`connects_to`), then sharding only
    when a single primary is the bottleneck. Cache aggressively (`Solid Cache`,
    Russian-doll fragment caching).
  - **Use when:** scale to the *next* bottleneck you can measure — not preemptively.
  
  ---
  
  ## Observability
  
  You can't fix what you can't see. Aim for the three pillars + errors.
  - **Logs:** structured/JSON (`lograge` + custom fields, request id, tenant id).
  - **Metrics:** `yabeda` → Prometheus → Grafana; instrument with
    `ActiveSupport::Notifications`.
  - **Traces:** OpenTelemetry (`opentelemetry-ruby`) across web → job → DB → HTTP.
  - **Errors:** Sentry / Honeybadger with release + user/tenant context.
  - **Use when:** before you need it. Wire request-id + tenant-id through every log
    line so an incident is greppable end-to-end. (Plug your own stack in here.)
  
  ### Correlation IDs — one thread through everything
  A single id that follows **one logical operation** end-to-end: inbound request →
  background jobs → outbound HTTP → emitted events → every log line. It turns "what
  actually happened to this request?" from archaeology into a single grep.
  - **Vs. request id:** a request/trace id is per-hop; a **correlation id** spans the
    whole business operation across processes and async boundaries. Often you carry
    both.
  - **How:** accept it at the edge (`X-Correlation-ID` header) or mint one; stash it in
    `ActiveSupport::CurrentAttributes` so any code can read it without threading it
    through every method; then propagate it deliberately —
    - into **every structured log line** (and your metrics/trace context),
    - echoed back in the response and in `problem+json` (`trace_id`),
    - copied into **background job** arguments/metadata (e.g. a Sidekiq client
      middleware), so async work keeps the same id,
    - forwarded as a header on **outbound** service calls,
    - attached to **emitted events** (see EventBus) so reactions stay traceable.
  - **Gems/Rails:** `ActiveSupport::CurrentAttributes`, `ActionDispatch::RequestId`
    (`X-Request-Id`); a small Sidekiq middleware pair to set/read it on jobs.
  - **Use when:** anything that crosses more than one process, job, or service — i.e.
    every non-trivial system. Pairs naturally with **idempotency** keys and the
    **EventBus**.
  
  ---
  
  ## Testing
  
  ### The suite (rspec, factories, fakes, end-to-end)
  - **`rspec-rails`** as the framework; table-driven where it fits.
  - **`factory_bot`** + **`faker`** for test data; prefer factories over fixtures.
  - **Fakes/stubs for I/O:** `webmock` / `vcr` for HTTP, in-memory fakes for adapters
    — never hit real third parties in tests.
  - **End-to-end user-story tests:** drive real flows (sign up → do the thing → see
    result). `capybara` + `cuprite`, or **Playwright via MCP** for browser-level
    journeys an agent can run.
  - **Use when:** unit-test logic-heavy POROs; integration-test the slice; reserve a
    handful of slow E2E tests for critical revenue paths.
  
  ### Performance & N+1 testing
  Catch slow queries and fan-out in CI, not in production.
  - **Gems:** `prosopite` (N+1 detection — strict, low false positives; pair with
    `pg_query`), `bullet` (alternative), `test-prof` (`let_it_be`, factory profiling),
    `rspec-benchmark` / `derailed_benchmarks` for memory & boot.
  - **Use when:** any list/index endpoint, any view that loops over associations.
    Fail the build on a new N+1.
  
  ---
  
  ## User feedback
  
  Errors and guidance are product, not afterthoughts.
  - **Errors:** human-readable HTML pages for browsers, `problem+json` for APIs
    (above). Tell the user *what to do next*, surface a trace id for support.
  - **Flash / notifications:** `noticed` for multi-channel (in-app, email) notices.
  - **Onboarding / tutorials:** `shepherd.js` / `intro.js` product tours.
  - **Use when:** every user-facing failure and every first-run experience.
  
  ---
  
  ## Async & agentic engineering
  
  - **Background work:** `Sidekiq` / `GoodJob` / `Solid Queue` for anything slow,
    retryable, or fan-out. Jobs should be **idempotent** (see above) and small.
  - **Agentic subagents:** decompose AI work into bounded, single-purpose subagents
    with pre-fetched context; orchestrate sequentially or in parallel waves;
    verify outputs with an independent pass (trust-but-verify). Treat an LLM call
    like any flaky network dependency — wrap it (timeouts, retries, circuit breaker)
    and make the surrounding job idempotent.
  - **Use when:** long-running, parallelizable, or AI-driven work that shouldn't
    block the request cycle.
  
  ---
  
  ## Quick reference
  
  | Concern | Reach for | Use when |
  |---|---|---|
  | Multi-tenancy | `acts_as_tenant` + Postgres RLS | shared-table SaaS |
  | Authorization | `pundit` / `action_policy`, deny-by-default | always |
  | Secrets | Rails credentials | leaking = breach |
  | Config | `config` / `anyway_config` (yml) | non-secret env shape |
  | Runtime boot | ENV | per-host/per-boot |
  | API errors | RFC 9457 `problem+json` | any external API |
  | Idempotency | `Idempotency-Key` + unique index | retryable writes |
  | Circuit breaker | `stoplight` / `semian` | flaky dependencies |
  | Auditing | `paper_trail` / `audited` / `logidze` | regulated/disputed data |
  | Flexible attrs | `jsonb` + `store_accessor` | sparse/evolving metadata |
  | Feature flags | `flipper` | decouple deploy/release |
  | Enums | Rails `enum` / `enumerize` | fixed value set, no transition rules |
  | State machine | `aasm` / `statesman` | guarded lifecycles (owns the enum column) |
  | Result types | `dry-monads` + `dry-contracts` | fail-able multi-step ops |
  | Events | `ActiveSupport::Notifications` / `rails_event_store` | fan-out side effects |
  | Health | `okcomputer` (deep `/health`) | every deployment |
+ | Versioning | `svu` + Kamal pre/post-deploy hooks → git tag | traceable releases (see [[kamal-auto-version\|Kamal Auto-Versioning]]) |
  | N+1 / perf | `prosopite`, `test-prof` | list/index endpoints |
  | Testing | `rspec`, `factory_bot`, `vcr`, Playwright(MCP) | everywhere |
  | Observability | OTel, `yabeda`, `lograge`, Sentry | before you need it |
  | Correlation ID | `X-Correlation-ID` + `CurrentAttributes` | trace one op across processes |
  | Background/agents | `Sidekiq`/`Solid Queue`, bounded subagents | slow/parallel/AI work |
  | IDs | bigint internal · uuid external · `friendly_id` | public surfaces |
  | Scale | vertical first, then horizontal + replicas | at a measured bottleneck |