Diff v4 → v5

v4: bot legacy · 2026-08-19T06:46:24Z
v5: bot legacy · 2026-08-19T06:53:18Z
  # Rails Production Playbook
  
  A short, opinionated reference for building serious Rails apps. The **principles**
  are language-agnostic — anyone can lift them — but the **gems and snippets** are
  Rails-first. Default posture: *boring, explicit, deny-by-default, observable.*
  
  > Rule of thumb: reach for a pattern when the pain it removes is real, not
  > hypothetical. Every item below has a "use when" so you don't over-engineer.
  
  ---
  
  ## Architecture & code design
  
  ### Sandi Metz discipline — lean controllers, service objects, enums
  Keep classes small and single-purpose. Controllers only orchestrate: authenticate,
  authorize, call one object, render. Push logic into POROs ("service objects").
  - **Guidelines:** ≤100 lines/class, ≤5 lines/method, ≤4 params, one job per object.
  - **Gems:** plain Ruby first; `interactor` or `dry-monads` if you want a Result type.
  - **Use when:** always for controllers; extract a service the moment an action does
    more than one thing or touches more than one model.
  - Use Rails `enum` for finite states that are *attributes*; use a **state machine**
    (below) when transitions have rules.
  
  ### Service objects + Result types (dry-monads / dry-contracts)
  Make success/failure explicit instead of raising for control flow.
  - **Gems:** `dry-monads` (`Success`/`Failure`, `Do` notation), `dry-validation`
    / `dry-contracts` (typed, composable input validation at the boundary).
  - **Use when:** multi-step operations that can fail in known ways (payments,
    imports, signups). Validate untrusted input with a contract *before* the service.
  ```ruby
  class CreateInvoice
    include Dry::Monads[:result, :do]
    def call(params)
      attrs   = yield InvoiceContract.new.(params).to_monad   # Failure stops the chain
      invoice = yield persist(attrs)
      Success(invoice)
    end
  end
  ```
  
  ### Slice / vertical-feature development
- Organize by **feature**, not by layer. A change should touch one slice
- front-to-back rather than smearing across `app/*`.
- - **Use when:** always as a habit. The *tooling* to enforce it is a separate,
-   much later decision — see the ladder below.
+ Organize by **feature within the layer**. A change should be implementable by
+ touching one domain namespace across `app/models`, `app/services`, `app/policies`
+ — not by smearing across unrelated domains, and not by inventing a folder per
+ feature. The slice is the **namespace**, not the directory.
+ - **Use when:** always. The layout that makes it work is the next section; the
+   tooling to *enforce* it is a much later decision.
  
- ### Modular boundaries — namespaces first, Packwerk maybe never
+ ### Modular boundaries — namespace inside the layer, and stop there
  
  Boundary enforcement is the most over-adopted idea in Rails architecture. The
- honest position: **almost every app that reaches for Packwerk should have stopped
- one rung lower.** Climb only when the rung you're on has demonstrably failed.
+ position here: **layer-first directories with a domain namespace inside each one
+ is the answer, not a waypoint.** Everything below it is a cost you should need
+ evidence to pay.
  
- **Rung 1 — namespaced modules in the directories you already have.** Free, zero
- tooling, zero config, works today:
+ #### The layout: layer first, domain second
  
  ```
- app/models/billing/invoice.rb        → Billing::Invoice
- app/services/billing/charge.rb       → Billing::Charge
- app/policies/billing/invoice_policy.rb
+ app/models/billing/invoice.rb           → Billing::Invoice
+ app/services/billing/charge.rb          → Billing::Charge
+ app/policies/billing/invoice_policy.rb  → Billing::InvoicePolicy
+ app/queries/billing/unpaid.rb           → Billing::Unpaid
+ app/contracts/billing/create_contract.rb
  ```
  
- A module is a boundary. `Billing::` reaching into `Catalog::` internals is visible
- in a diff and greppable in review. Most teams never need more than this, and the
- ones that do usually needed *ownership* rather than *enforcement*.
+ Two questions, two axes, both readable straight off the path: **the directory says
+ what kind of object this is; the namespace says whose domain it belongs to.** That
+ is not a compromise between layer-first and feature-first — it is both, and it
+ costs nothing.
  
- - **Use when:** always. This is the default and the destination for most apps.
+ - **Zeitwerk namespaces it for free.** `app/models` is the autoload root, so
+   `billing/invoice.rb` resolves to `Billing::Invoice` with no configuration, no
+   initializer, no `push_dir`.
+ - **One organizing scheme for the whole app.** Every file's location is derivable
+   from two facts you already know. There is never a placement argument.
+ - **It greps both ways.** `app/services/billing/` is everything billing *does*;
+   `app/services/` is every service in the app. A domain folder gives you the first
+   and destroys the second.
+ - **A module is a boundary.** `Billing::Charge` reaching into `Catalog::Product`
+   internals is visible in a diff and greppable in review.
+ - **Use when:** always. This is the default and the destination.
  
- **Rung 2 — a domain folder under `app/`.** When a slice has enough of its own
- models, services and policies that scattering them across `app/models/billing`,
- `app/services/billing`, `app/policies/billing` costs more than it buys, invert it:
+ Vertical-slice thinking still applies — a feature should be implementable by
+ touching one namespace across the layers, not by smearing across unrelated
+ domains. That is a property of the *namespace*, and it does not require the files
+ to sit in one folder.
  
+ #### Rejected: the all-in-one domain folder
+ 
+ The inversion — `app/billing/` holding its own models, services and policies —
+ looks tidier in a directory listing and is worse in every other respect:
+ 
  ```
- app/billing/
+ app/billing/            # don't
  ├─ invoice.rb
  ├─ charge.rb
- ├─ invoice_policy.rb
- └─ contracts/create_contract.rb
+ └─ invoice_policy.rb
  ```
  
- **The gotcha that bites everyone here.** Rails autoloads *every* subdirectory of
- `app` that exists at boot (except `assets`, `javascript`, `views`), and — per the
- Rails guides — those directories "represent the root namespace: `Object`". So
- `app/billing/invoice.rb` defines `Invoice`, **not** `Billing::Invoice`. The folder
- name does nothing. Fix it explicitly:
+ - **It fights the autoloader.** Rails autoloads *every* subdirectory of `app` that
+   exists at boot (except `assets`, `javascript`, `views`), and per the Rails guides
+   those directories "represent the root namespace: `Object`". So
+   `app/billing/invoice.rb` defines `Invoice`, **not** `Billing::Invoice` — the
+   folder name does nothing. Getting the namespace you thought you already had costs
+   a per-domain initializer:
  
+   ```ruby
+   # config/initializers/autoloading.rb — the tax the layer-first layout doesn't pay
+   module Billing; end
+   Rails.autoloaders.main.push_dir("#{Rails.root}/app/billing", namespace: Billing)
+   ```
+ 
+ - **It splits `app/` into two schemes.** The moment one domain inverts, `app/` has
+   both layer-first and feature-first directories, and every new file is a
+   placement decision. Half-migrated is the normal end state and the worst one.
+ - **It hides the layer census.** "Show me every policy" and "show me every
+   service" stop being one path each — and those are the reviews that catch a
+   missing gate.
+ - **It buys nothing the namespace didn't already give you.** The boundary was
+   never the folder.
+ 
+ #### Enforcing it without Packwerk
+ 
+ Layer-first and Packwerk are **architecturally mismatched**, and this is the part
+ usually left unsaid: a Packwerk package is a *single directory tree* rooted at its
+ `package.yml`. A layer-first domain spans four or five trees, so "billing" cannot
+ be one package — it would be one package per layer, all mutually dependent, which
+ expresses nothing. If you keep the layout above, Packwerk is mostly unavailable to
+ you. That is fine; the two cheap mechanisms that *do* fit are enough:
+ 
+ **Ownership, by glob.** CODEOWNERS handles the layer-first layout natively:
+ 
+ ```
+ app/*/billing/     @acme/billing
+ app/*/catalog/     @acme/catalog
+ ```
+ 
+ **A boundary spec.** Fifteen lines, no dependency, and it fails the build — the
+ same "a rule with a test behind it, or it isn't a rule" posture as the rest of this
+ playbook:
+ 
  ```ruby
- # config/initializers/autoloading.rb
- module Billing; end
- Rails.autoloaders.main.push_dir("#{Rails.root}/app/billing", namespace: Billing)
+ # spec/architecture/boundaries_spec.rb
+ BOUNDARIES = { "Billing" => %w[Catalog Shipping], "Catalog" => %w[Billing] }.freeze
+ 
+ RSpec.describe "domain boundaries" do
+   BOUNDARIES.each do |domain, forbidden|
+     it "#{domain} does not reach into #{forbidden.join(', ')}" do
+       files = Dir["app/*/#{domain.underscore}/**/*.rb"]
+       offenders = files.flat_map { |f|
+         File.readlines(f).each_with_index.filter_map { |line, i|
+           "#{f}:#{i + 1}" if line.match?(/\b(#{forbidden.join('|')})::/)
+         }
+       }
+       expect(offenders).to be_empty, "cross-domain reference:\n#{offenders.join("\n")}"
+     end
+   end
+ end
  ```
  
- - **Use when:** a slice is big enough that the layer-first split hurts. Costs one
-   initializer line per domain and buys a real front-to-back folder.
+ Crude — it reads text, not constants, so it cannot see a dynamic `const_get` — and
+ that is the trade: it costs nothing, is obvious to everyone, and catches the
+ violation that actually happens (someone types `Catalog::Product` in a billing
+ service). Add a domain to the hash when a boundary starts mattering; delete a line
+ when it stops.
  
- **Rung 3 — Packwerk.** Static analysis of constant references that fails the build
- when a package reaches somewhere it declared it wouldn't. Actively maintained
- (v3.3.0, May 2026) and genuinely good at the one thing it does: **holding a line
- you have already drawn.** It will not draw the line for you.
+ - **Use when:** you have a boundary worth naming. Which is the same bar Packwerk
+   should be held to, at about 1% of the cost.
  
- - **Use when:** *all* of these are true — multiple teams, a repeatedly violated
-   boundary you can name, evidence of breakage from that coupling, and a person
-   who owns the migration. Miss the last one and you get a `todo` file nobody
-   burns down.
- - **Don't use when:** the motivation is "we might scale", tidiness, or a rewrite
-   that's really about disliking the current structure. Packwerk charges rent
-   immediately and pays out only after the boundary work is finished.
+ #### Packwerk — when, and why probably not
  
+ Static analysis of constant references that fails the build when a package reaches
+ somewhere it declared it wouldn't. Actively maintained (v3.3.0, May 2026) and
+ genuinely good at the one thing it does: **holding a line you have already drawn.**
+ It will not draw the line for you.
+ 
+ - **Use when** *all* of these hold: multiple teams, a repeatedly violated boundary
+   you can name, evidence of breakage from that coupling, an owner for the
+   migration — **and** you are willing to give up the layer-first layout, because a
+   package has to be one directory tree.
+ - **Don't use when** the motivation is "we might scale", tidiness, or a
+   restructure that is really about disliking the current shape. Packwerk charges
+   rent immediately and pays out only after the boundary work finishes.
+ 
  **Read Shopify's own retrospective before adopting it.** They wrote it, and their
  assessment is unusually candid:
  
  - **Privacy checks were removed in v3.0.** They "introduced several problems",
    broke Rails conventions, and turned Packwerk "into something it was never
    intended to be: an API design tool."
  - **They discussed removing Packwerk from their monolith**, "given the costs it
    incurs and the weaknesses and blind spots."
  - **It is blind to anything outside Zeitwerk.** Code loaded via `require`,
    `autoload` or `ActiveSupport::Autoload` is invisible, so "a package that is
    well-defined according to Packwerk may actually crash with name errors when its
    code is executed."
  - **It tells you a violation exists, never how to fix it.**
  - **Domain-based grouping failed; functional grouping worked.** Developers "group
    code into packages based strongly on semantic clues that in many cases have
    little relation to how the code actually runs."
  - They believe they were "likely the first Packwerk user to completely work
    through an entire package todo file, years after its initial release" — which
    tells you how often adoption actually finishes.
  
  The line worth keeping: *"It is much harder to bend this behavior to fit your
  mental models than it is to bend your mental models to fit what a codebase
  actually does."*
  
- **Rung 4 — engines, then separate services.** Engines drag routes, fixtures and
- initializers along and Packwerk "doesn't help with sorting through" any of them.
- Extract a service only when a slice needs independent scaling or deploys — not for
- modularity, which rungs 1–3 already bought.
- 
- #### `packs/` is a convention, not a requirement — don't use it
- 
- Packwerk's `package_paths` default is `**/` — meaning any directory containing a
- `package.yml` is a package. The `packs/` top-level folder comes from the
- surrounding gem ecosystem, not from Packwerk, and it is the wrong choice.
- 
- **Put packages under `app/`:**
+ **If you adopt it anyway, do not use `packs/`.** Packwerk's `package_paths`
+ default is `**/` — any directory containing a `package.yml` is a package. The
+ `packs/` top-level folder comes from the surrounding gem ecosystem, not from
+ Packwerk. A `packs/` tree sits outside `app`, so it needs autoload wiring added,
+ which is most of what the `packs-rails`-shaped gems exist to do; and it puts a
+ second Rails app in your repo, so every new file becomes a placement argument.
+ Point it at directories that already exist instead:
  
  ```yaml
  # packwerk.yml
  package_paths:
-   - .           # the root package
-   - app/*       # one package per domain folder
- ```
- 
- ```
- app/
- ├─ billing/
- │  ├─ package.yml          # dependencies: ["app/platform"]
- │  ├─ invoice.rb
- │  └─ charge.rb
- ├─ catalog/
- │  ├─ package.yml
- │  └─ product.rb
- └─ platform/
-    └─ package.yml
+   - .
+   - app/*/*        # e.g. app/models/billing, app/services/billing
  ```
  
- Three reasons `app/` wins:
- 
- 1. **Zeitwerk already autoloads it.** Every subdirectory of `app` is an autoload
-    path at boot. A `packs/` tree is outside `app`, so it needs autoload wiring
-    added — which is most of what the `packs-rails`-shaped gems exist to do. Choose
-    `app/` and that entire dependency disappears.
- 2. **It keeps one answer to "where does code live".** `packs/billing/app/models/`
-    *and* `app/models/` is two Rails apps in one repo, and every new file becomes a
-    placement argument. Rung 2's layout is the same shape with none of that.
- 3. **It survives abandoning Packwerk.** Delete the `package.yml` files and the
-    gem, and `app/billing/` is still a perfectly ordinary namespaced Rails
-    directory. Rip Packwerk out of a `packs/` layout and you own a migration.
- 
- If you'd rather not invert the layer split at all, the smallest possible version
- is a `package.yml` beside a namespace inside the standard directories —
- `app/models/billing/package.yml`. Same enforcement, zero structural change, and
- the diff to adopt or revert is a handful of files.
- 
  - **Rule:** whatever you pick, one layout for the whole app. A half-migrated
    `packs/` tree beside a populated `app/` is worse than either alone.
  
+ #### Engines, then separate services
+ 
+ Engines drag routes, fixtures and initializers along, and Packwerk "doesn't help
+ with sorting through" any of them. Extract a service only when a slice needs
+ independent scaling or deploys — never for modularity, which the namespace already
+ bought.
+ 
  ### EventBus / pub-sub
  Decouple side effects (emails, audit, search reindex) from the core write.
  - **In-process, cheap:** `ActiveSupport::Notifications` or `wisper`.
  - **Durable / event-sourced:** `rails_event_store`.
  - **Use when:** one action fans out to ≥2 unrelated reactions. Keep the *write*
    transactional; emit events `after_commit` so listeners never see rolled-back data.
  
  ### Enums — named values, never magic strings
  Raw strings scattered through code (`"pending"`, `"active"`, `where(status: "shipped")`)
  are typo-prone and have no single source of truth. Centralize the allowed values.
  - **Rails `enum`:** maps a column to named values; gives you scopes
    (`Order.shipped`), predicates (`order.shipped?`), and setters (`order.shipped!`).
    Define the mapping **explicitly** so values are stable — never rely on positional
    integers, and prefer a string-backed column for readability in the DB.
  - **Richer enums (behavior/metadata/i18n per value):** `enumerize`, or a small PORO
    registry / value object when each value carries logic. Translate display labels
    via i18n — don't hardcode human strings either.
  - **Use when:** any attribute drawn from a fixed, known set (status, role, kind,
    priority). One definition, referenced everywhere.
  ```ruby
  class Order < ApplicationRecord
    enum :status, { draft: "draft", paid: "paid", shipped: "shipped" }, default: :draft
    # → Order.paid, order.paid?, order.shipped!, Order.statuses
  end
  ```
  - **Relationship to state machines:** an `enum` lists the *values*; a **state
    machine** governs the *transitions* between them. If moving between values has
    rules, guards, or callbacks (`draft → paid → shipped`, never `draft → shipped`),
    let the state machine own the column — it becomes the single source of truth for
    both the allowed states *and* the legal moves. Plain `enum` is the right tool only
    when any value can follow any other.
  
  ### State machines (AASM)
  Model lifecycles with explicit, guarded transitions and callbacks.
  - **Gems:** `aasm` (attribute-based, ergonomic — status lives on the model);
    `statesman` (Gusto — stores each transition as a row in its own table, giving a
    built-in audit trail and history, at the cost of more setup);
    `state_machines-activerecord` (another attribute-based option).
  - **Pick:** `aasm` for simple lifecycles on the record; `statesman` when you need the
    full transition history / audit of *how* it got to a state.
  - **Use when:** an object moves through statuses with rules (`draft → published →
    archived`), not just a flat enum.
  
  ---
  
  ## Data & persistence
  
  ### IDs: bigint internal, UUID external, friendly_id for URLs
  - **`bigint`** primary keys for internal joins (fast, compact).
  - **`uuid`** (a separate `public_id`) for anything exposed in APIs/URLs — never leak
    sequential counts or enable ID enumeration.
  - **`friendly_id`** for human-readable slugs (`/users/jane-doe`, `/articles/the-title`).
  - **Use when:** any app with a public surface. Add the UUID column from day one.
  
  ### Multi-tenancy: RLS / ActsAsTenant as first class
  Make tenant isolation a property of the data layer, not something each query
  remembers to add.
  - **`acts_as_tenant`:** sets a current tenant and auto-scopes every query.
  - **Postgres RLS (row-level security):** the strongest option — the database itself
    refuses cross-tenant rows even if app code forgets. Belt *and* suspenders.
  - **Use when:** B2B / SaaS with shared tables. Decide *before* writing schema;
    retrofitting tenancy is painful. If you don't need tenancy, don't add it.
  ```sql
  ALTER TABLE invoices ENABLE ROW LEVEL SECURITY;
  CREATE POLICY tenant_isolation ON invoices
    USING (account_id = current_setting('app.current_account_id')::bigint);
  ```
  
  #### Two database roles: `app` (restricted) and `migrator` (owner)
  RLS is only as strong as the role that connects. **A table's owner, any
  `SUPERUSER`, and any role with `BYPASSRLS` silently ignore every policy** — so if
  your app connects as the same role that ran the migrations, RLS is decoration.
  Split the roles:
  
  - **`myapp_migrator`** — owns the schema, runs migrations, creates policies.
    Used by `rails db:migrate` / deploy tasks only. Never by the running app.
  - **`myapp_app`** — owns nothing, `NOSUPERUSER NOBYPASSRLS`, holds only
    `SELECT/INSERT/UPDATE/DELETE` on tables plus `USAGE` on sequences. This is the
    role Puma and Sidekiq connect as. Policies actually apply to it.
  
  ```sql
  CREATE ROLE myapp_migrator LOGIN PASSWORD '...' NOSUPERUSER NOBYPASSRLS;
  CREATE ROLE myapp_app      LOGIN PASSWORD '...' NOSUPERUSER NOBYPASSRLS;
  
  -- migrator owns the schema; app only uses it
  GRANT USAGE ON SCHEMA public TO myapp_app;
  ALTER DEFAULT PRIVILEGES FOR ROLE myapp_migrator IN SCHEMA public
    GRANT SELECT, INSERT, UPDATE, DELETE ON TABLES TO myapp_app;
  ALTER DEFAULT PRIVILEGES FOR ROLE myapp_migrator IN SCHEMA public
    GRANT USAGE, SELECT ON SEQUENCES TO myapp_app;
  
  -- belt: FORCE makes the policy apply even to the table owner
  ALTER TABLE invoices ENABLE ROW LEVEL SECURITY;
  ALTER TABLE invoices FORCE  ROW LEVEL SECURITY;
  ```
  Add `FORCE ROW LEVEL SECURITY` anyway — it closes the hole if someone ever points
  the app at the owning role by mistake.
  
  **Wire it in Rails.** Two URLs, one per role; migrations use the privileged one:
  ```yaml
  # config/database.yml
  default: &default
    adapter: postgresql
    url: <%= ENV["DATABASE_URL"] %>            # myapp_app — runtime
    migrations_paths: db/migrate
  
  production:
    primary:
      <<: *default
    primary_migrator:                           # rails db:migrate uses this
      <<: *default
      url: <%= ENV["MIGRATOR_DATABASE_URL"] %>  # myapp_migrator
      migrations_paths: db/migrate
      database_tasks: false
  ```
  Then set the tenant per request, **transaction-scoped** (`set_config(..., true)`)
  so a pooled connection can never leak one tenant's id into the next checkout:
  ```ruby
  class ApplicationRecord < ActiveRecord::Base
    def self.with_tenant(account_id)
      transaction do
        connection.exec_query(
          "SELECT set_config('app.current_account_id', $1, true)",
          "set_tenant", [account_id.to_s]
        )
        yield
      end
    end
  end
  ```
  
  #### Use the restricted role in development too
  The most common way RLS fails in production is that nobody ever exercised it.
  Locally everyone connects as their own superuser (`postgres`, or the macOS
  username Homebrew created), every policy is bypassed, and a query missing its
  tenant scope looks perfectly fine — until prod, where it returns zero rows or,
  worse, the policy was never right and prod is the first place it's tested.
  
  - **Dev and test connect as `myapp_app`, same as production.** Create the role in
    `bin/setup` / a seed task so a fresh clone gets it automatically.
  - **Only `db:migrate`, `db:prepare`, and `db:schema:load` use `myapp_migrator`** —
    in every environment, so dev and prod fail the same way.
  - **Payoff:** a forgotten `with_tenant` blows up on your laptop, in a failing
    spec, instead of in an incident. Write a spec that connects without a tenant set
    and asserts the query returns **no rows** — that test is the real proof RLS is on.
  - **Gotcha:** if dev used to run as owner/superuser, expect a wave of permission
    errors the first time you switch (missing `GRANT`s on new tables, sequences,
    extensions). That's the point — those same gaps were latent in prod.
  
  ### JSONB for metadata
  Use a `jsonb` column for sparse, schemaless, or fast-evolving attributes —
  *not* as an excuse to avoid real columns for things you query or index heavily.
  - **Rails:** `store_accessor :metadata, :referrer, :utm_source`; GIN-index when you
    query inside it.
  - **Use when:** flexible per-record extras (settings, integration payloads,
    feature config). Promote a key to a real column once you filter/sort on it.
  
  ### Transactions — when & how
  Wrap writes that must succeed or fail **together**.
  - **Do:** group related `INSERT`/`UPDATE`s; use DB constraints as the real guarantee;
    put side effects in `after_commit`, never mid-transaction.
  - **Don't:** make HTTP/network calls inside a transaction (holds locks, can't roll
    back the remote side); don't span a transaction across user think-time.
  - Use advisory locks / `with_lock` for "check-then-act"; pick an isolation level
    deliberately when you have read-modify-write races.
  
  ### Auditing (paper_trail / audited / logidze)
  Track who changed what, when — for compliance, debugging, and undo.
  - **`paper_trail`:** version history with whodunnit, restore previous versions.
  - **`audited`:** lighter change log.
  - **`logidze`:** Postgres-trigger based (no app-layer overhead).
  - **Use when:** regulated data, financial records, or anything users dispute.
  
  ---
  
  ## Security
  
  ### Policies, zero trust, deny-by-default
  Authorization is a first-class layer, not scattered `if current_user.admin?`.
  - **Gems:** `pundit` or `action_policy` (policy objects per resource);
    `rack-attack` (rate limit / block); `brakeman` + `bundler-audit` in CI.
  - **Posture:** **deny by default** — no policy means no access. Authorize every
    action; never trust client-supplied IDs, scopes, or roles. Strong params always.
  - **Use when:** always. Add a default `deny` and opt routes *in*.
  
  ### Secrets vs config — credentials vs ENV vs YAML
  Three buckets, don't mix them:
  - **`.yml`-backed config objects** for non-secret, environment-shaped values:
    URLs, database names, endpoints, timeouts, feature defaults.
    Gems: `config` (`Settings.x.y`) or `anyway_config`.
  - **Rails encrypted credentials** for *secrets*: DB passwords, integration API
    keys, signing keys (`rails credentials:edit`, committed encrypted).
  - **ENV** only for *runtime-start* values that differ per host/boot and aren't
    secret by nature (e.g. which credentials key to load, `RAILS_MAX_THREADS`).
  - **Rule:** if leaking it is a breach → credentials. If it just shapes the
    environment → yml. If the platform injects it at boot → ENV.
  
  ### Static security analysis — one command, `bin/ci`
  Security scanning that only runs on someone's laptop when they remember doesn't
  run. Put every static check behind **one entrypoint** developers and CI both call,
  so "did it pass?" has a single answer.
  
  Rails 8.1 ships this natively: `bin/ci` driven by a `config/ci.rb` DSL. On older
  Rails, write `bin/ci` as a shell script that exits non-zero on the first failure —
  the value is the single entrypoint, not the DSL.
  ```ruby
  # config/ci.rb
  CI.run do
    step "Setup",     "bin/setup --skip-server"
    step "Style: Ruby",   "bin/rubocop"
    step "Security: Gem audit",      "bin/bundler-audit"
    step "Security: Importmap audit", "bin/importmap audit"
    step "Security: Brakeman",
         "bin/brakeman --quiet --no-pager --exit-on-warn --exit-on-error"
    step "Tests: Rails",  "bin/rails test"
    step "Tests: System", "bin/rails test:system"
  end
  ```
  The GitHub Actions workflow then shrinks to one line — `run: bin/ci` — which means
  CI can't drift from what you run locally.
  
  **What to put in it:**
  - **`brakeman`** — Rails-aware SAST: SQL injection, mass assignment, unsafe
    `render`/`redirect_to`, XSS, command injection, weak crypto. Run with
    `--exit-on-warn`; triage real findings and check in `config/brakeman.ignore`
    (with a reason per entry) rather than lowering confidence thresholds.
  - **`bundler-audit`** — gem dependencies against the Ruby Advisory DB.
    `bundle audit check --update`.
  - **`ruby_audit`** — the same for the Ruby and RubyGems versions themselves,
    which `bundler-audit` doesn't cover.
  - **`bin/importmap audit`** — JS dependency CVEs when you use importmaps; use
    `yarn npm audit` / `pnpm audit` if you bundle instead.
  - **`rubocop`** + **`rubocop-rails`/`-rspec`/`-performance`** — style is not
    security, but consistent code makes the security diffs readable. Add
    **`erb_lint`** for template-level escaping issues.
  - **Secret scanning:** `gitleaks` (or `trufflehog`) over the diff *and* the
    history — a rotated key still in git history is still leaked. Pair with a
    pre-commit hook so it never lands in the first place.
  - **Workflow hardening:** `actionlint` + `zizmor` if you're on GitHub Actions —
    CI itself is an attack surface (unpinned actions, `pull_request_target`,
    injectable `${{ }}` expressions).
  - **Dependabot / Renovate** for the *upgrade* side; the audit steps above are the
    *gate*. You want both — one opens PRs, the other fails the build.
  - **`semgrep`** when you need custom org rules ("never call this helper",
    "no raw SQL outside `app/queries`") that Brakeman doesn't model.
  
  **Rules that keep it honest:**
  - **Fail the build, don't warn.** A scanner whose output is advisory is ignored
    within two sprints.
  - **Every suppression carries a reason and an owner.** Ignore files are fine;
    anonymous ignore files are technical debt with a security label.
  - **Run it on the PR, and on a schedule.** New CVEs land against unchanged code —
    a nightly `bin/ci` run catches what the PR gate structurally can't.
  - **Keep it fast.** If `bin/ci` takes 20 minutes nobody runs it locally, and the
    local run is the whole point.
  - Static analysis finds *classes* of bugs, not your bugs. It does not replace
    authorization policies, RLS, or a review of the actual threat model.
  
  ---
  
  ## Reliability & resilience
  
  ### HTTP errors the RFC way (RFC 9457 `problem+json`)
  Return machine-readable errors with a stable shape, not ad-hoc JSON.
  - **Shape:** `type`, `title`, `status`, `detail`, `instance` (+ `errors`, `trace_id`).
  - **Rails:** `rescue_from` → render `application/problem+json`. No dedicated gem
    needed; a thin renderer is enough.
  - **Use when:** any API with external consumers (incl. your own SPA / agents).
  ```ruby
  rescue_from Pundit::NotAuthorizedError do |e|
    render json: { type: "/errors/forbidden", title: "Forbidden",
                   status: 403, detail: e.message, trace_id: request.request_id },
           status: :forbidden, content_type: "application/problem+json"
  end
  ```
  
  ### Idempotency
  Make retried writes safe — same request, same result, no duplicates.
  - **How:** accept an `Idempotency-Key` header; store key + response; replay on
    repeat; reject body mismatch (`422`). Back it with a unique index.
  - **Gems:** roll your own table, or `sidekiq-unique-jobs` for job-level dedup.
  - **Use when:** payments, order creation, anything a client/network may retry.
  
  ### Circuit breakers
  Stop hammering a failing dependency; fail fast and recover automatically.
  - **Gems:** `stoplight` (general), `semian` (Shopify — for net/http, MySQL, Redis),
    `faraday-retry` for retries with backoff.
  - **Use when:** you call flaky third parties or internal services over the network.
  
  ### Feature flags
  Decouple deploy from release; kill-switch risky paths; gradual rollout.
  - **Gem:** `flipper` (+ `flipper-active_record`, `flipper-ui`).
  - **Use when:** trunk-based dev, canary releases, per-account betas, ops kill-switch.
  
  ### `/health` — real readiness, not just "the web server is up"
  Verify the whole boot: DB reachable, migrations current, Redis/queue up,
  required integrations reachable. Separate **liveness** (am I running?) from
  **readiness** (can I serve traffic?).
  - **Rails 7.1+:** `/up` exists but only proves the process booted — extend it.
  - **Gems:** `okcomputer` or `health_check` for multi-dependency checks.
  - **Use when:** any real deployment — load balancers and orchestrators gate on it.
  
  ### Versioning & releases
  Every deploy should carry a real, traceable version — not just a git SHA in your head.
  If you deploy with **Kamal** (Basecamp's own, very Rails-idiomatic tool), auto-bump a
  semantic version from your Conventional Commits and tag the release **only on a healthy
  deploy**, via `.kamal/hooks/` pre/post-deploy hooks. Full recipe:
  [[kamal-auto-version|Kamal Auto-Versioning]].
  - **Tooling:** `svu` (Conventional Commits → semver: `feat:`→minor, `fix:`→patch,
    `!`/`BREAKING CHANGE:`→major), a `pre-deploy` hook to preview/guard, a `post-deploy`
    hook to `git tag` after success.
  - **Surface it:** expose the version + git SHA on your deep `/health` (above) so a running
    instance self-reports exactly what shipped — invaluable during an incident. (In Rails,
    read it from an ENV/initializer rather than the Go `ldflags` trick in the reference.)
  - **Use when:** any Kamal-deployed app — so every release is tagged in git and traceable.
  
  ---
  
  ## Scalability
  
  - **Vertical:** bigger box first — simplest win until a single node is maxed.
  - **Horizontal:** more app servers behind a balancer; tune **Puma** workers/threads
    and the DB **connection pool** together; move work to background jobs
    (`Sidekiq` / `GoodJob` / `Solid Queue`).
  - **Database:** read replicas + Rails multi-DB (`connects_to`), then sharding only
    when a single primary is the bottleneck. Cache aggressively (`Solid Cache`,
    Russian-doll fragment caching).
  - **Use when:** scale to the *next* bottleneck you can measure — not preemptively.
  
  ---
  
  ## Observability
  
  You can't fix what you can't see. Aim for the three pillars + errors.
  - **Logs:** structured/JSON (`lograge` + custom fields, request id, tenant id).
  - **Metrics:** `yabeda` → Prometheus → Grafana; instrument with
    `ActiveSupport::Notifications`.
  - **Traces:** OpenTelemetry (`opentelemetry-ruby`) across web → job → DB → HTTP.
  - **Errors:** Sentry / Honeybadger with release + user/tenant context.
  - **Use when:** before you need it. Wire request-id + tenant-id through every log
    line so an incident is greppable end-to-end. (Plug your own stack in here.)
  
  ### Correlation IDs — one thread through everything
  A single id that follows **one logical operation** end-to-end: inbound request →
  background jobs → outbound HTTP → emitted events → every log line. It turns "what
  actually happened to this request?" from archaeology into a single grep.
  - **Vs. request id:** a request/trace id is per-hop; a **correlation id** spans the
    whole business operation across processes and async boundaries. Often you carry
    both.
  - **How:** accept it at the edge (`X-Correlation-ID` header) or mint one; stash it in
    `ActiveSupport::CurrentAttributes` so any code can read it without threading it
    through every method; then propagate it deliberately —
    - into **every structured log line** (and your metrics/trace context),
    - echoed back in the response and in `problem+json` (`trace_id`),
    - copied into **background job** arguments/metadata (e.g. a Sidekiq client
      middleware), so async work keeps the same id,
    - forwarded as a header on **outbound** service calls,
    - attached to **emitted events** (see EventBus) so reactions stay traceable.
  - **Gems/Rails:** `ActiveSupport::CurrentAttributes`, `ActionDispatch::RequestId`
    (`X-Request-Id`); a small Sidekiq middleware pair to set/read it on jobs.
  - **Use when:** anything that crosses more than one process, job, or service — i.e.
    every non-trivial system. Pairs naturally with **idempotency** keys and the
    **EventBus**.
  
  ---
  
  ## Testing
  
  ### The suite (rspec, factories, fakes, end-to-end)
  - **`rspec-rails`** as the framework; table-driven where it fits.
  - **`factory_bot`** + **`faker`** for test data; prefer factories over fixtures.
  - **Fakes/stubs for I/O:** `webmock` / `vcr` for HTTP, in-memory fakes for adapters
    — never hit real third parties in tests.
  - **End-to-end user-story tests:** drive real flows (sign up → do the thing → see
    result). `capybara` + `cuprite`, or **Playwright via MCP** for browser-level
    journeys an agent can run — same pass also yields the QoE numbers (Web Vitals,
    console errors, page weight), which no request spec will ever give you. Recipe:
    [QA and QoE Testing with Playwright MCP](https://wiki.joexbayer.dev/wiki/qa-and-qoe-testing-with-playwright-mcp);
    the case for measuring it: [QA and QoE](https://wiki.joexbayer.dev/wiki/qa-and-qoe).
  - **Use when:** unit-test logic-heavy POROs; integration-test the slice; reserve a
    handful of slow E2E tests for critical revenue paths.
  
  ### Performance & N+1 testing
  Catch slow queries and fan-out in CI, not in production.
  - **Gems:** `prosopite` (N+1 detection — strict, low false positives; pair with
    `pg_query`), `bullet` (alternative), `test-prof` (`let_it_be`, factory profiling),
    `rspec-benchmark` / `derailed_benchmarks` for memory & boot.
  - **Use when:** any list/index endpoint, any view that loops over associations.
    Fail the build on a new N+1.
  
  ---
  
  ## User feedback
  
  Errors and guidance are product, not afterthoughts.
  - **Errors:** human-readable HTML pages for browsers, `problem+json` for APIs
    (above). Tell the user *what to do next*, surface a trace id for support.
  - **Flash / notifications:** `noticed` for multi-channel (in-app, email) notices.
  - **Onboarding / tutorials:** `shepherd.js` / `intro.js` product tours.
  - **Use when:** every user-facing failure and every first-run experience.
  
  ---
  
  ## Async & agentic engineering
  
  - **Background work:** `Sidekiq` / `GoodJob` / `Solid Queue` for anything slow,
    retryable, or fan-out. Jobs should be **idempotent** (see above) and small.
  - **Agentic subagents:** decompose AI work into bounded, single-purpose subagents
    with pre-fetched context; orchestrate sequentially or in parallel waves;
    verify outputs with an independent pass (trust-but-verify). Treat an LLM call
    like any flaky network dependency — wrap it (timeouts, retries, circuit breaker)
    and make the surrounding job idempotent.
  - **Use when:** long-running, parallelizable, or AI-driven work that shouldn't
    block the request cycle.
  
  ---
  
  ## Quick reference
  
  | Concern | Reach for | Use when |
  |---|---|---|
- | Modular boundaries | namespaced modules in `app/` | always — start and usually stop here |
- | Boundary enforcement | Packwerk, packages under `app/*` (never `packs/`) | multiple teams + a named, repeatedly violated boundary + an owner |
+ | Modular boundaries | `app/<layer>/<domain>/` — layer first, domain namespace second | always; this is the destination, not a waypoint |
+ | Boundary enforcement | CODEOWNERS globs + a boundary spec | any boundary worth naming |
+ | Packwerk | only if you will abandon layer-first for it | multiple teams + named violated boundary + an owner; read the retrospective first |
  | Multi-tenancy | `acts_as_tenant` + Postgres RLS | shared-table SaaS |
  | DB roles | `app` (NOBYPASSRLS) vs `migrator` (owner) | any app using RLS — in dev too |
  | Authorization | `pundit` / `action_policy`, deny-by-default | always |
  | Static analysis | `bin/ci` → `brakeman`, `bundler-audit`, `gitleaks` | always; fail the build |
  | Secrets | Rails credentials | leaking = breach |
  | Config | `config` / `anyway_config` (yml) | non-secret env shape |
  | Runtime boot | ENV | per-host/per-boot |
  | API errors | RFC 9457 `problem+json` | any external API |
  | Idempotency | `Idempotency-Key` + unique index | retryable writes |
  | Circuit breaker | `stoplight` / `semian` | flaky dependencies |
  | Auditing | `paper_trail` / `audited` / `logidze` | regulated/disputed data |
  | Flexible attrs | `jsonb` + `store_accessor` | sparse/evolving metadata |
  | Feature flags | `flipper` | decouple deploy/release |
  | Enums | Rails `enum` / `enumerize` | fixed value set, no transition rules |
  | State machine | `aasm` / `statesman` | guarded lifecycles (owns the enum column) |
  | Result types | `dry-monads` + `dry-contracts` | fail-able multi-step ops |
  | Events | `ActiveSupport::Notifications` / `rails_event_store` | fan-out side effects |
  | Health | `okcomputer` (deep `/health`) | every deployment |
  | Versioning | `svu` + Kamal pre/post-deploy hooks → git tag | traceable releases (see [[kamal-auto-version\|Kamal Auto-Versioning]]) |
  | N+1 / perf | `prosopite`, `test-prof` | list/index endpoints |
  | Testing | `rspec`, `factory_bot`, `vcr`, Playwright(MCP) | everywhere |
  | Observability | OTel, `yabeda`, `lograge`, Sentry | before you need it |
  | Correlation ID | `X-Correlation-ID` + `CurrentAttributes` | trace one op across processes |
  | Background/agents | `Sidekiq`/`Solid Queue`, bounded subagents | slow/parallel/AI work |
  | IDs | bigint internal · uuid external · `friendly_id` | public surfaces |
  | Scale | vertical first, then horizontal + replicas | at a measured bottleneck |
  
  ---
  
  ## Sources
  
  - Shopify Engineering, [A Packwerk Retrospective](https://shopify.engineering/a-packwerk-retrospective) — the authors on what worked, what did not, and why they considered removing it.
  - [Packwerk USAGE.md](https://github.com/Shopify/packwerk/blob/main/USAGE.md) — `package_paths`, `include`/`exclude` defaults.
  - Rails Guides, [Autoloading and Reloading Constants](https://guides.rubyonrails.org/autoloading_and_reloading_constants.html) — why an `app/` subdirectory is the root namespace, and `push_dir(namespace:)`.