Part 4 — Running It in Production

P Authentication, Accounts and Access

This chapter takes a person standing in front of a login box and follows them all the way through to a database row they are allowed to read. It covers how logins actually work, which login method to pick for a wholesale apparel ERP and why, how to run multi-factor authentication and sessions for real employees on real schedules, how to invite and remove staff, and how a role matrix layers on top of the row-level security from chapter 5. It ends with a complete, runnable implementation built from the standard library — the sessions you own. Where the text below cites a managed vendor's behavior (Supabase's, mostly) it is for contrast: a worked example of the decisions you are now making yourself. The figures that will date carry that check date next to them.

In this chapter14 sections · about 93 min
  1. What you need to know first
  2. Authentication versus authorization, concretely
  3. Sessions, cookies and tokens from zero
  4. Login methods compared honestly
  5. Why the rep at a trade show needs a different login than the controller at a desk
  6. Multi-factor authentication
  7. Session management for business users
  8. Onboarding a customer
  9. Roles and permissions in practice
  10. Customer-facing access: the B2B portal
  11. Service accounts and API keys
  12. Account security operations
  13. Auditing authentication events
  14. A complete worked implementation: sessions you own

What you need to know first

Before any of the rest makes sense, a handful of ideas need to be solid. None of them are hard. They are just usually explained badly.

Statelessness, cookies and tokens

A request is stateless. When your browser asks your server for a page, that request arrives with no memory of anything that happened before. The server has no idea who you are. Every single request has to carry proof of identity with it, or the server treats you as a stranger. Everything in this chapter exists to solve that one problem: how does request number 4,000 prove it comes from the same person who typed a password an hour ago.

A cookie is a small piece of text the server asks the browser to store and send back. The server responds with a header that says Set-Cookie: session=abc123. The browser writes that down and attaches Cookie: session=abc123 to every future request to that site, automatically, until the cookie expires or is deleted. You do not write code to send it. The browser does it.

That automatic-attachment behavior is why cookies are convenient, and it is also why cross-site request forgery exists. Cross-site request forgery (CSRF) is an attack where a page on some other website makes your browser fire a request at your ERP, and your browser helpfully attaches your cookie to it. We deal with that below.

A token is a string that proves something. That is the entire definition. A token can be a random 32-byte value that means nothing on its own and only has meaning because your database has a row matching it. Or a token can be self-describing: it carries the facts inside itself and a cryptographic signature proving those facts were not edited. The first kind is called an opaque token or a reference token. The second kind is what a JWT is. Both are "tokens". They behave completely differently, and confusing them causes most auth bugs.

Hashing, DNS records and TLS

Hashing is one-way. Encryption is two-way. If you encrypt something you can decrypt it later with the key. If you hash something you can never get it back; you can only hash a new input and see whether the two hashes match.

Passwords must be hashed, never encrypted, using a slow algorithm designed for the job — bcrypt, scrypt or Argon2, so an attacker who steals your database cannot test billions of guesses per second. This book hashes with scrypt from node:crypto (the worked implementation below); managed vendors do the same with bcrypt or Argon2 behind an auth.users table you never see. Invitation tokens and API keys should also be stored hashed, for the same reason, which we build below.

DNS is a lookup table for names, and you will be editing it. Your domain registrar or DNS host (Cloudflare, Route 53, Namecheap) holds a set of records:

  • An A record maps a name to an IPv4 address.
  • A CNAME (canonical name) record maps a name to another name, so lookups follow a pointer.
  • A TXT record holds arbitrary text, and email authentication uses TXT records to publish policy.
  • An MX (mail exchanger) record says where mail for the domain goes.

You will add TXT records for SPF and DMARC, plus TXT or CNAME records for DKIM, so that your password-reset emails actually arrive. DNS changes are not instant: each record has a TTL (time to live, in seconds) and resolvers cache the old answer until it expires. Set TTL to 300 while you are experimenting.

TLS is the padlock. Turn it on everywhere. TLS (Transport Layer Security, the successor to SSL) encrypts the connection between browser and server so nobody on the coffee-shop wifi can read the cookie going past. Any cookie that grants access must be marked Secure, which tells the browser to refuse to send it over plain HTTP.

Certificates are issued automatically these days through ACME (Automatic Certificate Management Environment) — a protocol where your host proves to a certificate authority such as Let's Encrypt that it controls the domain, gets a certificate back, and renews it on a timer without anyone watching. Caddy runs ACME for you automatically (chapter O), and the managed platforms do the same for domains you attach, so you rarely touch a certificate by hand. Never disable TLS "just for testing" on a domain that also serves real sessions.

JWKS, PKCE and who vouches for whom

Two more acronyms you will meet in this chapter. JWKS (JSON Web Key Set) is a small public file your auth server publishes at a fixed URL, listing the public keys anyone can use to check a JWT's signature. PKCE (Proof Key for Code Exchange, pronounced "pixy") is a handshake that stops someone who intercepts a login link or redirect code from redeeming it, because only the browser that started the flow holds the matching secret. You implement PKCE at your callback route for OAuth (it is a hash and a compare); managed auth vendors do it for you. Either way the principle is the same: the code is worthless without the secret only the starting browser holds.

Two pieces of identity vocabulary. An identity provider (IdP) is the system that holds the truth about who someone is and vouches for them — Google Workspace, Microsoft Entra ID (formerly Azure AD), or Okta. A relying party is your app, which trusts that vouching. Enterprise single sign-on (SSO) is the arrangement where the customer's IdP does the authenticating and your app accepts the answer.

A word about scale. A wholesale apparel brand in the $10M–$80M revenue range usually has somewhere between a dozen and a couple of hundred internal users. The count varies widely with how much of the warehouse and sales force is in-house rather than outsourced. Add a few hundred more if you open a B2B portal for the retailers who buy from you. Every auth vendor's free or entry tier covers that comfortably, so choose on operational fit and support burden, not price per user.

Core principle

Authentication answers "who is this?" once, at the door. Authorization answers "may they do this?" on every single request afterwards. Getting the door right is a weekend of work. Getting the second question right, consistently, in every query, is the actual job, and chapter 5's row-level security is how you make the database enforce it instead of trusting your application code to remember.

Authentication versus authorization, concretely

Authentication is proving identity. Authorization is granting permission. People conflate them because the login screen is where both usually get set up at the same moment. Separating them in your head has a practical payoff: it tells you where each bug lives.

In your ERP the split runs like this. Maria types her email and password. Your login handler checks the scrypt hash, confirms it matches, and issues her a session. That is authentication, and it happened exactly once. Maria now clicks into an order for account "Nordstrom", changes the ship date, and saves. Before that write lands, something must decide whether Maria — a sales rep, not a sales manager — may change ship dates on an account that belongs to a different rep. That is authorization, and it happens on that request and every other one.

The failure modes differ too. An authentication failure is loud: the user cannot get in, they call you, you fix it. An authorization failure is silent, and nobody notices for six months until a rep screenshots another rep's margin sheet into a group chat.

So authorization must be enforced where it cannot be bypassed. Check permissions only in your React components and anyone who opens the network tab and replays a request walks straight past you. Check only in your API route handlers and you are one forgotten if away from a leak. Chapter 5 puts the check in Postgres row-level security, where every query — from your app, from a script, from a mistake — goes through the same gate. This chapter's job is to hand that gate an accurate, unforgeable statement of who the requester is and what role they hold.

Sessions, cookies and tokens from zero

The oldest and still one of the safest patterns works like this. You log in. The server generates a long random string — OWASP's session management guidance asks for at least 64 bits of entropy, produced by a cryptographically secure random generator of at least 128 bits, and stores a row keyed by that string containing your user id, creation time, and last-seen time. The server sends the random string back as a cookie. On each request the server looks the string up, finds the row, and knows who you are.

The random string means nothing on its own. It is a coat check ticket, and its power comes entirely from the row in your database, which means you control it completely. Delete the row and the session is dead on the very next request. Change someone's role and the next lookup sees the new role. Revocation is instant and free. The cost is a database read per request, which is irrelevant at 40 concurrent users and is why consumer-scale sites moved to JWTs.

What a JWT actually is

A JSON Web Token is three chunks of Base64URL text joined by dots: header.payload.signature. The header says which algorithm signed it. The payload is JSON containing claims — facts about the subject. The signature is a cryptographic proof that whoever holds the signing key produced exactly this header and payload.

eyJhbGciOiJFUzI1NiIsImtpZCI6IjJhNDMi...   <- header
.eyJzdWIiOiIzZmE4NS0uLi4iLCJyb2xlIjoi...  <- payload
.MEUCIQDx8kR2vT1p9nQ0hZfL4mYc2wJ...       <- signature

decoded payload:
{
  "iss": "https://erp.brandco.com/auth",
  "sub": "3fa85f64-5717-4562-b3fc-2c963f66afa6",
  "aud": "authenticated",
  "role": "authenticated",
  "email": "maria@brandco.com",
  "aal": "aal2",
  "amr": [{"method": "password", "timestamp": 1753372800},
          {"method": "totp",     "timestamp": 1753372812}],
  "session_id": "9c1b0e7a-4d2f-4a11-b8c5-7e1f0d3a2b44",
  "app_metadata": {"tenant_id": "b21c...", "role": "sales_rep"},
  "iat": 1753372812,
  "exp": 1753376412
}

Read that payload left to right:

  • iss is who issued it.
  • sub is the subject — the user's UUID, the thing you join on. A UUID is a long random identifier, unique enough that two systems will never generate the same one.
  • aud is the intended audience; you set it (say erp-sync) when you mint a sync token, and PowerSync is configured to accept that exact value.

PowerSync validates that claim: omit it and you get PSYNC_S2105 ("JWT payload is missing a required claim 'aud'"), so set it when your server mints the token.

The rest of the payload:

  • role is the Postgres role the query will run as.
  • aal is the Authenticator Assurance Level — aal1 means one factor was used, aal2 means a second factor was verified too.
  • amr (authentication methods references) lists the actual methods and when each happened.
  • session_id ties the token back to a server-side session row.
  • app_metadata is where you put claims the user must not be able to edit.
  • iat and exp are issued-at and expiry, both Unix seconds; keep the gap short — this book mints fifteen minutes for sync tokens, an hour is the common managed default.

Three things people get wrong about JWTs:

  1. Encoding is not encryption. Anyone can paste the payload into a Base64 decoder and read every claim, so never put anything secret in it.
  2. The signature only proves the token was not tampered with. It says nothing about whether the token is still appropriate for the person holding it.
  3. Most important of all: a JWT cannot be revoked. Once issued, any server holding the public key will accept it until exp passes. If you fire someone at 10:02 and their token expires at 10:59, they have 57 minutes of valid access unless you build something extra.

For a business application the revocation problem dominates. Offboarding must be immediate. Role changes must take effect now, not "within the hour". A controller who discovers a rep is quoting unauthorized discounts wants that rep locked out while they are still on the phone.

Storage is the other reason. A JWT held in browser localStorage is readable by any JavaScript on the page, so one compromised npm dependency or one cross-site scripting hole hands over a working credential. Cross-site scripting (XSS) is a bug where attacker-supplied text gets executed as code inside your page. A cookie marked HttpOnly is invisible to JavaScript: the browser still sends it, but no script can read it, so an XSS bug becomes "attacker can act while the page is open" rather than "attacker exfiltrates a token and uses it from elsewhere for the next hour".

Set-Cookie: __Host-erp-session=eyJhbGciOi...;
            Path=/;
            Secure;
            HttpOnly;
            SameSite=Lax;
            Max-Age=3600

Every attribute here earns its place:

  • __Host- is a browser-enforced prefix: the browser rejects the cookie unless it also has Secure, has Path=/, and has no Domain attribute, which stops a compromised subdomain from planting a cookie on your main site.
  • Secure means HTTPS only.
  • HttpOnly means JavaScript cannot read it.
  • SameSite=Lax means the browser sends it on normal top-level navigations to your site but not on cross-site form posts or background fetches, which blocks most CSRF.
  • Max-Age=3600 is one hour in seconds.

Use SameSite=Strict if you have no inbound links from email that need to land already-logged-in; use Lax if you do, which for an ERP with emailed order links you probably do.

One practical note on cookie naming: the __Host- prefix (which the browser only honors on a Secure, Path=/, host-only cookie) is worth adopting for your session cookie. Managed clients often skip it — Supabase writes sb-<project-ref>-auth-token, split across numbered chunks when the token is long. You control the rest of the attributes through the cookieOptions you pass when you create the server client.

The hybrid the managed vendors use, and why this book does not need it

Managed auth (Supabase's is a fair example) issues two things: a short-lived JWT access token (one hour by default) and a long-lived refresh token. The refresh token is opaque, single-use, and backed by a server-side session row. When the access token expires, the client trades the refresh token for a fresh pair.

That gives you a bounded version of the revocation problem. Kill the session server-side and the refresh fails, so the blast radius is at most one access-token lifetime. Refresh tokens carry a small reuse interval — ten seconds is customary, so a legitimate double-refresh from two server-rendered requests does not blow up the session. A genuinely stolen refresh token used outside that window trips reuse detection, and the session and its refresh tokens are revoked.

The important operational decision is where the tokens live, and this book's answer sidesteps the whole apparatus: the browser holds one opaque cookie pointing at a server-side session row, and there is no access/refresh token pair on the client to refresh, rotate or steal. A managed SDK instead parks the token pair in cookies and runs a per-request hook to refresh them — more moving parts to buy statelessness you did not ask for. Either way you get cookie security with JWT performance. Do not use the browser-only client with localStorage for an ERP.

Never trust getSession() on the server

The rule generalizes past any one vendor, and the managed docs say it bluntly: a session loaded from client storage "is loaded directly from local storage and isn't re-validated against the Auth server", so it must not be trusted in server code. In this book's design there is nothing to re-validate on the client at all — the cookie is opaque and getSession (Step 6) hits the database, which is the source of truth. If you ever do trust a JWT server-side, verify its signature against your JWKS on every call, or getUser() when you need a fresh round-trip to the auth server. Reserve getSession() for the case where you need the raw token string to forward to another service such as PowerSync.

Login methods compared honestly

Five realistic options, compared in the table below and then taken one at a time, with a recommendation at the end.

MethodUser experienceMain failure modeSupport burdenUse in a wholesale ERP
Email + passwordFamiliar; fast on repeat; needs no second channelForgotten passwords; reuse of a breached passwordMedium — resets are a steady source of tickets, but self-serveDefault for all internal staff
Magic linkNo password to remember; one clickEmail delay or spam folder blocks login entirely; corporate scanners pre-click and burn the link; link opens in the wrong browserHigh during outages, near zero otherwiseGood for the retailer B2B portal; poor at a trade show
One-time code (OTP)Type 6 digits; survives cross-browser and cross-deviceSame email dependency; code entry on a phone keyboardMediumBetter than magic link on iPads and shared devices
Social login (Google/Microsoft)One tap if already signed in to work accountPersonal vs work account confusion; account lost when employee's Google account is deletedLow — provider handles resetsStrong if the customer is a Google Workspace or Microsoft 365 shop, which many apparel brands are
Enterprise SSO (SAML / OIDC)Invisible: the IdP already knows themMetadata expiry, clock skew, certificate rollover; misconfigured attribute mappingLow daily, spiky at setup and cert renewalOnly when the customer's IT department demands it, which is usually your larger accounts

Email and password

Everyone understands it. It needs one round trip and no second channel, so it works on bad wifi, on a borrowed laptop, and on the day your transactional email provider has an incident.

The modern rules for passwords are shorter than you expect. NIST SP 800-63B revision 4 (published 26 August 2025) tells verifiers they "SHALL NOT impose other composition rules (e.g., requiring mixtures of different character types) for passwords" and "SHALL NOT require subscribers to change passwords periodically". On length it says verifiers "SHALL require passwords that are used as a single-factor authentication mechanism to be a minimum of 15 characters in length", and MAY allow a shorter password when it is only part of a multi-factor flow, but never below eight characters.

Forced changes happen only on evidence of compromise. So: a 12-character minimum with mandatory second factor for anyone who matters, no character-class rules, and leaked-password checking against the Have I Been Pwned Pwned Passwords corpus — a single range-query API call from your signup handler, free.

A magic link mails a URL that logs you in when clicked. A well-built magic link expires after about an hour, is one-time use, and is rate-limited to one request per address every 60 seconds — the shape your auth_token table (Step 4) already enforces.

The failure modes are not obvious, so learn them before you commit to this. Corporate email security appliances follow every link in every message to check for malware, and because the link is single use, the scanner consumes it, so the user clicks and gets "link expired". Corporate mail scanners cause this: "certain email providers may have spam detection or other security features that prefetch URL links from incoming emails". Its suggested fixes are to send a six-digit code instead of a link, or to build your own link that lands on a page with a button the user must press.

Some mail clients open links in an in-app browser with no cookie jar shared with Safari, so the session lands somewhere the user cannot get back to. And if your email is delayed by four minutes, your login is delayed by four minutes, with the user staring at a blank screen.

One-time code

Same email round trip, but the user types six digits instead of clicking. This solves the scanner problem (scanners cannot type), solves the wrong-browser problem, and works when the user is reading email on their phone but logging in on an iPad. Link or code is the same server flow — mint a single-use token, mail it — differing only in whether the email template embeds it in a URL or prints it as digits; the code variant sidesteps the scanner problem because there is nothing for a scanner to click. {{ .ConfirmationURL }} or {{ .Token }}.

Social login

"Sign in with Google" or "Sign in with Microsoft" delegates authentication to a provider the customer already pays for. Password resets, MFA and offboarding all become the customer's IT problem, which is where they belong. The risk is account confusion: a rep signs up with their personal Gmail on day one, and eighteen months of order history is attached to an identity you cannot control. Restrict the allowed email domains at signup and refuse anything else.

Enterprise SSO via SAML or OIDC

SAML 2.0 (Security Assertion Markup Language) and OIDC (OpenID Connect) are protocols for "let the customer's identity provider vouch for this person". OIDC is the modern one. It is built on OAuth 2.0, the standard for granting one site limited access to an account held somewhere else, and it passes identity around as JSON and JWTs. SAML is the older XML-based one, and it is still what enterprise IT departments ask for by name.

SAML is where the honest answer is: this is the one auth job worth an outside dependency. A conformant SAML 2.0 service provider is a large, security-critical spec you should not hand-roll — reach for a library (or a purpose-built SSO gateway like WorkOS) for this one integration, deliberately and on the record. Whichever you pick, the moving parts are the same. The docs state you need CLI version v1.46.4 or higher (checked July 2026).

# Register the customer's identity provider with your SSO layer,
# whichever you chose. Every one takes the same three inputs:
#   1. the IdP's metadata (a URL or an XML file they send you)
#   2. the email domain(s) that should route to it
#   3. a mapping from their attribute names onto your claims
#
# e.g. with the WorkOS CLI:
workos sso create-connection \
  --type saml \
  --metadata-url 'https://retailco.com/idp/saml/metadata' \
  --domain retailco.com

That registration tells your SSO layer: "any user whose email ends in retailco.com should be bounced to this identity provider". The metadata URL is a document the customer's IdP publishes describing its signing certificate and endpoints; using the URL rather than a file means certificate rotations are picked up automatically, which removes a common cause of SSO outages.

On the client you then call signInWithSSO({ domain: 'retailco.com' }) and the browser redirects out and back. The attribute mapping file translates the IdP's field names (which will be something like http://schemas.xmlsoap.org/ws/2005/05/identity/claims/emailaddress) into your own claim names. Note that once SSO owns a domain, the IdP owns that account's second factor too — your MFA policy defers to theirs for those users.

Cost check, every figure read off the vendor's pricing page in July 2026 and all subject to change:

  • A managed database that bundles SSO (Supabase Pro) runs about $25/month with the first ~50 SSO users included — a data point for what the convenience costs.
  • WorkOS charges $125 per SSO connection per month for the first 15 connections, dropping to $100 for connections 16–30 and lower again above that; its AuthKit user-management product is free up to 1 million monthly active users, then $2,500 per additional million.
  • Clerk's Pro plan is $25/month, includes one enterprise connection per app, and charges $75/month for each additional connection in the 2–15 range.
  • Auth0's Essentials tier starts at $150/month for 500 MAU and includes 3 enterprise connections, with extras at $100/month up to 30 total.

Per-connection pricing is the industry norm, so SSO cost scales with your number of customers, not your number of users.

Recommended default

Email and password, with mandatory TOTP multi-factor for any role that can move money or change prices, plus optional "Sign in with Microsoft / Google" for customers already on a managed work account. Add one-time email codes as the recovery path. Add SAML SSO only when a specific customer's IT department asks and will pay for it. This combination fails gracefully — if your email provider is down, staff can still log in, and it is the cheapest to support.

Why the rep at a trade show needs a different login than the controller at a desk

Consider two real days. In February your sales rep is at Coterie in New York, standing with an iPad in a booth alongside hundreds of other exhibitors on the same overloaded venue wifi. A buyer from a 12-door boutique chain is in front of her and has eleven minutes. If the app asks her to log in again, she loses the buyer.

Meanwhile your controller is at a desk with a wired connection, a password manager and a laptop that never leaves the building, releasing a payment run against factory invoices. If the app asks him to re-authenticate first, he shrugs and does it. Those two people need opposite defaults.

DimensionSales rep on iPad at a showController at a desk
NetworkCongested, intermittent, sometimes noneReliable wired or office wifi
Login friction toleranceNear zero during show hoursHigh; re-auth on sensitive actions is fine
Second factorDevice-bound (Face ID / device passcode); an SMS code is unusable with no signalTOTP app or passkey, no problem
Session lengthLong: must survive a 10-hour show day and ideally the whole showShort: 8-hour absolute, 30-minute idle
Device trustCompany-owned, MDM-enrolled, supervised iPadCompany laptop, disk encrypted
Blast radius if stolenOwn accounts only, no pricing edits, no bankingEverything financial, so lock it down hard

MDM, blast radius and two policies

MDM means mobile device management: software (Jamf, Kandji, Mosyle, Microsoft Intune) that the company installs on a device it owns, so it can push settings, force a passcode, lock the device into one app, and wipe it remotely when it goes missing. An MDM-enrolled iPad is a device you can still control after it walks out of the booth, which is why it earns a longer session than a random browser. One more phrase from that table: "blast radius" means how much damage one stolen credential can do before anyone notices.

The design that follows: log the rep in before the show, on the office wifi, on a company iPad enrolled in MDM, with a long session and a device-bound second factor. Give the rep an app that works fully offline (chapter 6), so a dropped connection never turns into a login prompt. Give the controller a shorter session, mandatory TOTP, and step-up re-authentication on payment release — step-up meaning you ask for the password or a fresh code again at the moment of the dangerous action, even though the user is already signed in. Same system, two policies, keyed off role and device.

Concretely, add a "trusted device" concept: a company iPad that has completed enrollment gets a device record, and sessions on that device get the long timebox. A browser on an unknown machine gets the short one. We store that as a device row and a claim, shown later.

Multi-factor authentication

A factor is a category of evidence: something you know (password), something you have (phone, security key), something you are (fingerprint, face). Multi-factor means evidence from two different categories. Two passwords still count as one factor, because both are things you know. A password plus a code from an app on your phone counts as two.

FactorHow it worksPhishing-resistant?Works offline?Verdict for an ERP
TOTP authenticator appShared secret plus current time generates a 6-digit code every 30sNo — a fake login page can relay the code in real timeYes, fullyBaseline. Require for privileged roles.
Passkey / WebAuthnPublic-key pair bound to your domain; unlocked by biometric or device PINYes — the browser refuses to sign for the wrong originYes (local unlock)The right direction; ecosystem support is still maturing
Push approvalProvider app shows "Approve?"; user tapsNo, and open to approval fatigue — the attacker fires prompts all night until the user taps to make them stopNo — needs dataOnly with number matching: the login page shows a number the user must find on the prompt
SMS codeCode texted to a phone numberNoNo — needs cell signalWeakest. Last resort only.
Recovery codesTen single-use strings printed at enrollmentn/aYesMandatory backstop for every enrolled user

Why SMS is the weakest option

Three independent problems:

  • SIM swap: an attacker social-engineers the carrier into moving the victim's number to their own SIM, and every code follows.
  • Number porting and carrier account takeover do the same thing.
  • And the SS7 signaling network that routes text messages between carriers has known interception weaknesses.

NIST SP 800-63B rev 4 states that "use of the PSTN for out-of-band verification is restricted", the public switched telephone network being ordinary phone and SMS delivery, and tells verifiers to weigh risk signals such as a recent device swap, SIM change or number port before relying on it.

A mundane problem too: a rep in a concrete trade-show basement has no cell signal, so the text never arrives. Offer SMS as an alternate, never as the only factor. Cost matters as well — a managed phone-MFA add-on runs on the order of $75/month plus your SMS provider's per-message fee.

TOTP as the baseline

TOTP (Time-based One-Time Password, RFC 6238) is what Google Authenticator, 1Password, Authy and Microsoft Authenticator generate. At enrollment the server generates a random secret and shows it as a QR code; thereafter both sides compute the same 6-digit code from that secret plus the current 30-second window. The phone needs no network, so it works in a basement, and it is free.

Its weakness is real: TOTP does not know what site you are on. If a rep lands on brandco-login.net and types her password and her code, an attacker relays both to the real site inside the 30-second window and is in. That is why passkeys matter.

Passkeys and WebAuthn as the direction of travel

WebAuthn is the browser standard behind passkeys. A passkey is a public/private key pair created by your device for one specific website. The private key never leaves the device, or the user's encrypted iCloud or Google keychain for synced passkeys. When you sign in, the site sends a challenge, the device asks you for Face ID or your PIN, signs the challenge, and returns the signature. Because the browser will only sign for the origin the key was created for, a phishing site gets nothing. That is what "phishing-resistant" means, and for a small team it is the largest practical security gain on offer.

The honest current state, as of July 2026. WebAuthn is the one place this book bends its no-dependency rule without apology: attestation parsing and the ceremony are exactly the kind of security-critical spec you adopt a maintained library for (SimpleWebAuthn is the common choice). The managed vendors are still catching up too — Supabase's own passkey support ships labeled experimental: "Passkey support is experimental. The API may change without notice. You must explicitly opt-in when creating the Supabase client." Its API exposes register/sign-in calls plus lower-level start/verify hooks, and is still evolving version to version.

Two limits matter for an ERP wherever you get WebAuthn: SSO users register their passkey with the IdP, not you, and passkeys are bound to the Relying Party ID, so changing your domain invalidates every one of them. And in most stacks a passkey today sits as a primary login method rather than a second factor that raises you to aal2.

The practical reading: ship TOTP now as the enforced second factor, and offer passkeys as a convenience sign-in for staff who want them, with the expectation that the API will change.

If phishing-resistant MFA is a hard contractual requirement today, that pushes you toward Clerk, WorkOS AuthKit or Auth0 as the identity layer, with your server accepting their tokens by verifying the IdP's JWKS — the same trust handshake your own sessions use. Clerk, Firebase Auth, Auth0, AWS Cognito and WorkOS are the five supported providers (checked July 2026), and each must sign its JWTs asymmetrically and publish a kid (key id) header, which tells the verifier which public key to check the signature against.

Note also that NIST forbids syncable authenticators at AAL3 because the private key is exportable; that does not affect you, because AAL2 is the right target for an ERP.

Recovery codes

People lose phones. Without recovery codes, a lost phone means you — the founder, at 11pm — manually resetting someone's MFA, which is itself an attack vector, because a caller claiming to be a locked-out employee is the classic social engineering script. Generate ten single-use codes at enrollment, store only their hashes, show them once, and require the user to confirm they saved them. Log every recovery-code use loudly.

Enforcement policy per role

RoleMFA requirementStep-up re-auth required for
Owner, AdminMandatory, no grace periodAdding users, changing roles, rotating API keys
Accounting / ControllerMandatoryPayment runs, bank detail changes, credit memos over threshold
Sales managerMandatoryPrice list changes, discount overrides
Sales repMandatory, 7-day grace at onboarding; device-bound factor on managed iPadsNothing routine
WarehouseOptional if on a supervised, single-purpose device on the warehouse networkNothing routine
Customer service, Read-onlyRecommended; mandatory if remoteNothing routine
External rep / agentMandatory, no exceptions — they are outside your controlExporting customer lists
Retailer portal userOptional, encouragedChanging ship-to addresses, adding portal users

The warehouse exception deserves a note. Scanner terminals and shared floor tablets get handled by the device, not the person: locked-down kiosk mode, on a network segment that only reaches your app, with a short PIN to identify who is picking. Forcing a TOTP prompt on someone wearing gloves at 6am produces one shared authenticator app taped to the wall, which is worse than no MFA at all.

Session management for business users

Two clocks govern every session and you need both.

Idle timeout ends the session after a period with no activity. It protects against the unlocked laptop in an open-plan office or a hotel lobby. OWASP puts common ranges at "2-5 minutes for high-value applications and 15-30 minutes for low risk applications".

Absolute timeout ends the session a fixed time after login regardless of activity. It bounds how long a stolen credential stays useful. For an application "intended to be used by an office worker for a full day", OWASP suggests an absolute timeout "between 4 and 8 hours".

Set them per role rather than globally. The policy below has survived contact with real apparel teams.

Role / contextIdle timeoutAbsolute timeoutRemember device?
Owner, Admin30 min8 hoursYes, 14 days, MFA still required at each new session
Accounting20 min8 hoursYes, 14 days
Sales manager / rep, desktop2 hours12 hoursYes, 30 days
Sales rep, managed iPadnone (app handles lock)7 days, extended to 14 during a declared show windowDevice is the trust anchor
Warehouse shared terminal10 min10 hours (covers a shift plus overtime)No
Customer service1 hour10 hoursYes, 14 days
External rep / agent30 min8 hoursNo
Retailer B2B portal1 hour24 hoursYes, 30 days

Your session table implements both clocks directly (Step 4's absolute_expiry and the idle window in Step 6); the managed consoles expose the same two as "Time-box user sessions" and "Inactivity timeout," which is the idle timeout, and "Single session per user" keeps only the most recent session — all three a few lines of your own SQL over the session table.

One behavior to understand, and one your design gets right by construction: enforcement happens when the session is next used, not by a background reaper. Your getSession (Step 6) refuses an expired row on the next request; a sweeper only tidies old rows. Managed systems behave the same way — expired rows are "progressively deleted from the database 24 hours after they expire". So a client holding a still-valid access token keeps working until that token expires. With the default one-hour token, your worst-case overshoot is one hour.

Per-role policy is just code, because the policy lives in your AUTH.mfaRequiredRoles list and your gate: a check in your request wrapper (the code older frameworks call the projects call middleware) that compares session age against the policy for the user's role and signs them out early. That is shown in step 6 of the implementation.

Refresh, and what "remember this device" should mean

With opaque sessions there is nothing for the client to refresh — the cookie points at a row your server slides forward on use (Step 6), so "remember this device" is simply a longer absolute_expiry on that row. (The managed alternative refreshes a token pair on every request and rate-limits it, commonly at 1,800/hour per IP with bursts to 30, comfortable for a building of tablets refreshing hourly behind one office IP, but a refresh-loop bug will hit it fast.

That per-IP limit has a trap the managed approach must handle and yours avoids: when auth calls originate from your server, the limiter sees one IP for the whole customer base. The fix there is IP address forwarding: enable it under Authentication → Rate Limits, then send the end-user address in an Sb-Forwarded-For header on requests made with a secret API key. Publishable and legacy keys are not supported for this.

"Remember this device" has exactly one safe meaning: skip the second factor on this device for N days. The user still types their password every single time. Implement it as a separate long-lived HttpOnly cookie holding a random device token, matched against a trusted_devices row recording the user, a hash of the token, the user agent, first- and last-seen timestamps, and an expiry. If a valid unexpired device token is present at login, let the user skip the TOTP prompt. Clear all rows for a user when their password or role changes.

Revoking on role change and offboarding

This is the part people skip and regret. Three triggers must kill access:

  • Role change. A demotion is a security event. If a sales manager becomes a sales rep, their existing JWT still carries role: sales_manager until it expires. Force the token to be reissued so the next one carries the new role.
  • Offboarding. When someone leaves, you deactivate the membership row, block their account, delete their trusted devices, revoke any API keys they created, and, this is the one that gets forgotten, invalidate any pending invitations they issued.
  • Credential change. Password change, email change, MFA factor removal.

This is exactly where owning the table wins. "Log this user out by id" is one UPDATE session SET revoked_at = now() WHERE user_id = $1 (Step 8) — instant, total, auditable. Managed token systems make it awkward: many have no admin "log out by id" at all, because the outstanding access token lives on the client until it expires. auth.admin.signOut() in the JavaScript client takes a jwt argument for the same reason. What you actually have are two supported levers.

-- Both levers are SQL, because both the sessions and the
-- memberships are your tables. Run inside one transaction.

-- LEVER 1: revoke every live session for the account. Takes
-- effect on that person's very next request. Reversible only
-- by signing in again (which is the point).
update session set revoked_at = now()
where user_id = :user_id and revoked_at is null;

-- LEVER 2: ban the account so a fresh sign-in is refused too.
update app_user set banned_at = now() where id = :user_id;
-- Lift it with:  update app_user set banned_at = null ...
-- The login handler (Step 6) checks banned_at before issuing.

Use both. Lever 1 stops the refresh token from working at all. Lever 2 makes sure that even if a token is somehow reissued, it carries a role that can read nothing. Neither one can claw back an access token that has already been issued — it stays valid until exp, which is exactly why you keep the access token lifetime at one hour rather than raising it. In the implementation section we wire this into a database trigger so it fires automatically whenever a membership row is deactivated or its role changes.

Build one habit here: after any password or email change, verify with your own eyes what your project does to other sessions. Sign in on two browsers, change the password in one, and reload the other. Do not take a blog post's word for it, including this one — the behavior is configurable and it has changed between releases.

The offline iPad, where a token must survive days

Chapter 6's PowerSync client keeps a local SQLite copy of the data the rep is allowed to see and syncs when a connection appears. SQLite is a complete database that lives in a single file on the device, with no server to run. Authentication has to bridge a gap the web never has: the device may be offline for hours or days, and it still needs a credential that PowerSync will accept when the connection returns.

PowerSync's own constraints are specific and you should design to them. The JWT must contain sub, aud, iat and exp. "The JWT must expire in 24 hours or less, and 60 minutes or less is recommended" — the gap between iat and exp must not exceed 86,400 seconds, and "JWTs older than 60 minutes are not accepted by PowerSync". Signing must use RSA, EdDSA or ECDSA verified through a JWKS endpoint; HS256 is for development only.

PowerSync's docs are explicit about why: "There is no way to revoke a JWT once issued without rotating the key."

So the token itself cannot survive days. What survives days is the refresh capability plus the local data. The correct architecture is:

OFFLINE                          BACK ONLINE
-------                          -----------
local SQLite: readable           refresh token -> new JWT (60m)
writes queue locally             JWT -> PowerSync connect
JWT expired: no sync             queued writes upload
app still fully usable           server re-checks permissions

Bounds:
  access token   60 min   (PowerSync max useful life)
  refresh token  bounded by session absolute timeout
  iPad session   7 days normal / 14 days show window
  local data     encrypted at rest by iOS + app lock

Read that as a contract. While offline the app needs no valid JWT at all, because it is reading local SQLite; it only needs one at the moment it reconnects. The risk you are bounding is "how long can a stolen iPad keep pulling fresh data", and the answer is "until the session's absolute timeout, or until you ban the account".

Bound it with four controls:

  1. The absolute session timeout on iPad sessions — 7 days, raised deliberately during a show and lowered again once the show ends.
  2. MDM: the iPad is supervised and remotely wipeable through Apple Business Manager plus your MDM tool.
  3. An app-level lock requiring Face ID after 15 minutes in the background, so a grabbed unlocked iPad is still one biometric away from your order book.
  4. Most important of all: the server re-validates every queued write on upload. When those queued orders arrive they pass through the same RLS policies as anything else, and writes from a revoked user are rejected.
Do not raise token lifetime to "fix" offline

The tempting shortcut is a 30-day JWT so the iPad never has trouble reconnecting. That hands anyone who extracts the token 30 days of unrevocable access to your customer list and pricing. PowerSync caps it at 24 hours for exactly this reason. Keep the token short, keep the refresh token as the long-lived thing, and keep the account bannable server-side.

Onboarding a customer

"Customer" here means a brand that buys your ERP, and the first hour of their life in your system sets the tone. The sequence that works:

Step 1: create the tenant and the first admin. Nobody can invite the first user because there is nobody to do the inviting. Solve it with a provisioning script you run: create the tenant row, create the first user with role owner, send them an invitation. Avoid self-serve tenant creation, because a rep signing up "to try it" ends up with a second orphan tenant containing three real orders.

Step 2: the owner invites their team. Owner enters email addresses and picks roles. The system sends invitations. This must be self-serve or you become the help desk forever.

Step 3: domain-verified auto-join, optionally. If the tenant has proven control of brandco.com, you can let anyone with a @brandco.com email join automatically at a preset default role. Prove domain control the same way every vendor does: give them a random string to publish as a DNS TXT record, then verify it.

Host:  _brandco-erp-verify.brandco.com
Type:  TXT
TTL:   300
Value: brandco-erp-verify=7c4f1a2e9b0d43a6b8e5c1f70d29a834

Check it from a terminal:
$ dig +short TXT _brandco-erp-verify.brandco.com
"brandco-erp-verify=7c4f1a2e9b0d43a6b8e5c1f70d29a834"

You generate the random value, store it against the tenant, and show the customer these exact record details. They add it at their DNS host. Your verification job runs dig or a DNS library, compares the value, and flips domain_verified_at when it matches. TTL 300 means five minutes of caching, so retries are quick.

Once verified, a new signup from that domain can auto-join. Keep auto-join at a low-privilege default role — read_only or customer_service — never at admin, because anyone who ever gets a mailbox at that domain, including a contractor, would inherit it.

Step 4: seat management. Decide early whether you charge per seat and whether the customer can add seats without talking to you. Track seats as active membership rows, not as auth users, because a user may belong to two tenants (an external agent who reps three brands). Show the owner a seat count and a limit, and block invitations that would exceed it with a clear message rather than a silent failure.

Invitation token design

The invitation token is a credential. Anyone holding it can become a user in someone else's tenant. Design it accordingly: unguessable, single-use, short-lived, bound to one email address, and stored hashed.

create extension if not exists pgcrypto;
create extension if not exists citext;

create table public.invitations (
  id             uuid primary key default gen_random_uuid(),
  tenant_id      uuid not null references public.tenants(id),
  email          citext not null,
  role           text not null
                 references public.roles(key),
  -- SHA-256 of the raw token; the raw token is never stored
  token_hash     bytea not null unique,
  invited_by     uuid not null references auth.users(id),
  expires_at     timestamptz not null
                 default now() + interval '7 days',
  accepted_at    timestamptz,
  accepted_by    uuid references auth.users(id),
  revoked_at     timestamptz,
  send_count     int not null default 0,
  last_sent_at   timestamptz,
  created_at     timestamptz not null default now()
);

-- One live invitation per email per tenant.
create unique index invitations_one_live
  on public.invitations (tenant_id, email)
  where accepted_at is null
    and revoked_at is null;

create index invitations_expiry
  on public.invitations (expires_at)
  where accepted_at is null;

-- RLS on, and deliberately NO policies: with no policy,
-- every client using the publishable key is denied. Only
-- server code holding the secret key can read this table.
alter table public.invitations enable row level security;

Walk through the choices:

  • bytea is the Postgres type for raw bytes.
  • token_hash stores SHA-256 of the token, never the token itself, so a leaked database backup does not hand over working invitations — exactly the reasoning behind hashing passwords.
  • citext is a case-insensitive text type, so Maria@ and maria@ resolve to the same invite.
  • expires_at defaults to seven days, long enough to survive a holiday, short enough that a forwarded email from March is dead.
  • accepted_at plus the unique partial index enforces single use and stops a tenant accumulating six live invites for the same person.
  • revoked_at lets an admin cancel.
  • send_count and last_sent_at let you rate-limit resends so the invite endpoint cannot be used to spam someone's inbox.

And enabling row level security with no policy at all is the strongest setting there is: nothing gets through except code holding the secret key.

import { randomBytes, createHash, timingSafeEqual }
  from 'node:crypto'

// PostgREST wants a bytea value as a hex literal string.
// Passing a raw Buffer silently stores the wrong bytes.
function toBytea(buf: Buffer) {
  return '\\x' + buf.toString('hex')
}

// 32 random bytes = 256 bits of entropy. base64url so it
// survives being pasted into a URL and an email client.
export function newInviteToken() {
  const raw = randomBytes(32).toString('base64url')
  const hash = createHash('sha256').update(raw).digest()
  return { raw, hash: toBytea(hash) }  // store hash, mail raw
}

export function hashToken(raw: string) {
  return toBytea(createHash('sha256').update(raw).digest())
}

// Used for the trusted-device cookie, where you fetch the
// row first and then compare. Constant time, so an attacker
// cannot narrow a value down by measuring how long it takes.
export function tokensMatch(a: Buffer, b: Buffer) {
  return a.length === b.length && timingSafeEqual(a, b)
}

randomBytes(32) uses the operating system's cryptographically secure random source. Do not use Math.random(), which is predictable and has been the root cause of real account-takeover bugs.

256 bits is far beyond the 64 bits of entropy OWASP requires for session identifiers, and it costs nothing. base64url encoding avoids the + and / characters that break in URLs.

The toBytea helper is the detail that bites people the moment a query builder or JSON layer sits between you and Postgres: a Buffer may serialize to an object, not bytes, so the row you insert never matches the row you look up. With postgres.js you pass the Buffer straight into a bytea column and it round-trips — one more thing the thin driver spares you. (If you ever front tables with an HTTP layer like PostgREST, it is the JavaScript client is actually talking to.) The raw token goes into the emailed link and is never written to your database or your logs.

// app/api/invitations/route.ts  (create an invitation)
import { newInviteToken } from '../auth/invite-token'
import { sendInviteEmail } from '../mail'
import { withTenant } from '@erp/db'

export async function createInvitationHandler(req, res) {
  const actor = await requireRole(req, ['owner', 'admin'])
  const { email, role } = await readJson(req)

  if (typeof email !== 'string' || !email.includes('@')) {
    return Response.json({ error: 'email required' },
                         { status: 400 })
  }
  if (!canGrant(actor.role, role)) {
    return Response.json(
      { error: 'cannot grant a role above your own' },
      { status: 403 })
  }

  const seats = await countActiveMembers(actor.tenantId)
  if (seats.used >= seats.limit) {
    return Response.json(
      { error: 'seat limit reached', seats }, { status: 409 })
  }

  const { raw, hash } = newInviteToken()   // raw mailed, hash stored

  // Runs inside the actor's tenant transaction — RLS scopes the
  // insert; no bypass credential is needed or wanted.
  const [inv] = await withTenant(actor, (tx) => tx`
    insert into invitations
      (tenant_id, email, role, token_hash, invited_by,
       send_count, last_sent_at)
    values (${actor.tenantId}, ${email.toLowerCase().trim()},
            ${role}, ${hash}, ${actor.userId}, 1, now())
    returning id, expires_at`)

  await sendInviteEmail({
    to: email,
    url: `${process.env.APP_URL}/invite/${raw}`,
    tenantName: actor.tenantName,
    inviterName: actor.displayName,
    expiresAt: inv.expires_at,
  })

  await audit('invitation.created', actor, {
    invitation_id: inv.id, email, role })

  return respond(res, 201, { id: inv.id })
}

Line by line:

  • we require the caller to be an owner or admin;
  • we sanity-check the email;
  • canGrant stops an admin from minting an owner, which is privilege escalation;
  • we check the seat limit before creating anything so the customer gets a clear 409 rather than a surprise invoice;
  • we generate the token and store only the hash;
  • we email the raw token in a URL;
  • and we write an audit record.

Notice there is no admin client and no bypass credential: the insert runs inside the actor's own tenant transaction, so RLS scopes it and the requireRole gate at the top is the only authority check. The less powerful the credential a handler holds, the fewer ways it can be misused.

// apps/admin/handlers/invitations-accept.ts
export async function acceptInvitationHandler(req, res) {
  const { token, password, fullName } = await readJson(req)

  // One generic error for every failure case, so the
  // endpoint cannot be used to probe for valid tokens.
  const bad = () => respond(res, 400,
    { error: 'This invitation is no longer valid.' })

  if (typeof token !== 'string' || token.length < 20) return bad()
  if (typeof password !== 'string' || password.length < 12) {
    return respond(res, 400,
      { error: 'Password must be at least 12 characters.' })
  }

  // No admin/bypass client: one SYSTEM-context transaction that
  // either commits whole or rolls back whole. Consume the
  // invitation by HASH, atomically, exactly once.
  const ok = await withTenant(SYSTEM, async (tx) => {
    const [inv] = await tx`
      update invitations set accepted_at = now()
      where token_hash = ${hashToken(token)}
        and accepted_at is null and revoked_at is null
        and expires_at > now()
      returning id, tenant_id, email, role`
    if (!inv) return false      // invalid / used / expired / raced

  // The invitee may already have an account - an outside
  // agent who reps three brands signs in once and belongs
  // to three tenants. Reuse the account, do not fail.
    let [u] = await tx`select id from app_user where email = ${inv.email}`
    if (!u) {
      [u] = await tx`
        insert into app_user (email, password_hash, email_verified_at)
        values (${inv.email}, ${await hashPassword(password)}, now())
        returning id`                // the invite proved the email
    }

    await tx`
      insert into memberships (tenant_id, user_id, role, status)
      values (${inv.tenant_id}, ${u.id}, ${inv.role}, 'active')`
    await auditTx(tx, 'invitation.accepted',
      { userId: u.id, tenantId: inv.tenant_id },
      { invitation_id: inv.id })
    return true
  })

  return ok ? respond(res, 200, { ok: true }) : bad()
}

The important details here are the ones that are easy to skip. Every failure returns the same message. If you distinguish "expired" from "not found", the endpoint answers a question an attacker wants answered, which tokens exist, and guessing becomes worthwhile. The password minimum is enforced on the server, not just in the form, because the form is not the only way to call this.

The account is created with email_confirm: true because possession of the invitation already proves control of the mailbox; making them confirm again is friction with no benefit. Note that tenant_id is deliberately not written into app_metadata here — the membership row is the single source of truth, and step 5's access token hook reads it fresh on every token.

If the membership insert fails and we created the auth user in this request, we delete it so we do not strand an account with no tenant; if the user already existed, we leave their account alone. The final update carries .is('accepted_at', null) and checks that a row actually came back, so two racing submissions cannot both succeed. That is what the "idempotency guard" comment means: running the request twice leaves the system in the same state as running it once.

When an employee leaves

Write this as a single function and a single button, because a checklist that a human runs will be run partially.

create or replace function public.offboard_member(
  p_tenant_id uuid,
  p_user_id   uuid,
  p_actor_id  uuid,
  p_reassign_to uuid default null
) returns void
language plpgsql
security definer
set search_path = public
as $$
begin
  -- SECURITY DEFINER runs as the function owner, so this
  -- check is the only thing standing between a caller and
  -- everyone else's account. Never remove it.
  if not exists (
    select 1 from memberships a
     where a.tenant_id = p_tenant_id
       and a.user_id   = p_actor_id
       and a.status    = 'active'
       and a.role in ('owner','admin')
  ) then
    raise exception 'not authorised to offboard members';
  end if;

  update memberships
     set status = 'inactive',
         deactivated_at = now(),
         deactivated_by = p_actor_id
   where tenant_id = p_tenant_id
     and user_id = p_user_id;

  update trusted_devices set revoked_at = now()
   where user_id = p_user_id and revoked_at is null;

  update api_keys set revoked_at = now(),
                      revoked_by = p_actor_id
   where tenant_id = p_tenant_id
     and created_by = p_user_id and revoked_at is null;

  update invitations set revoked_at = now()
   where tenant_id = p_tenant_id
     and invited_by = p_user_id
     and accepted_at is null and revoked_at is null;

  if p_reassign_to is not null then
    update customers set owner_user_id = p_reassign_to
     where tenant_id = p_tenant_id
       and owner_user_id = p_user_id;
  end if;

  insert into auth_events(tenant_id, actor_id, subject_id,
                          event, detail)
  values (p_tenant_id, p_actor_id, p_user_id,
          'member.offboarded',
          jsonb_build_object('reassigned_to', p_reassign_to));

  -- Postgres cannot make HTTP calls. Queue the ban for a
  -- worker (pg-boss, chapter 9) to send to the Admin API.
  insert into job_queue(kind, payload)
  values ('ban_user',
          jsonb_build_object('user_id', p_user_id));
end $$;

-- Lock the function down. Without these two lines any
-- signed-in user could call it over RPC.
revoke all on function public.offboard_member(
  uuid, uuid, uuid, uuid) from public, anon, authenticated;
grant execute on function public.offboard_member(
  uuid, uuid, uuid, uuid) to service_role;

One transaction checks that the caller is actually an owner or admin of that tenant, deactivates the membership, kills trusted devices so the "remember me" cookie stops working, revokes any API keys they created, cancels invitations they had outstanding, optionally reassigns their accounts to another rep so orders do not go dark, and writes the audit row. The queued job is the hand-off: a worker picks it up and calls the Admin API to ban the account, which is the part Postgres cannot do itself.

Note the deliberate absence of a delete — you never delete a user who has touched orders, because the ledger from chapter 1 references them forever. Deactivate, never delete.

And note the revoke/grant pair at the bottom: a security definer function runs with its owner's privileges, so leaving it callable by authenticated would let any signed-in user offboard anybody. A database function is reachable by anything that can call it, so "nobody would call it" is not a defense. Grants are.

Roles and permissions in practice

Roles are a compression scheme. Rather than assigning 200 individual permissions to each of 40 people, you define nine roles, assign permissions to roles, and assign one role per person per tenant. Here is a matrix that fits an apparel wholesale brand.

Four abbreviations recur in it:

  • ATS means "available-to-sell", the quantity of a style and size you can still promise a buyer after existing orders are subtracted.
  • PII means personally identifiable information, such as a buyer's direct phone number.
  • ASN means advance ship notice, the electronic packing list a retailer expects before your delivery arrives at their door.
  • GL means general ledger, the accounting book of record from chapter 1, so "GL mapping" is the rule that decides which ledger account a transaction lands in.
RoleCan seeCan changeCannotRow scope
OwnerEverything, including cost sheets, margin, payroll-adjacent data, billingEverything, including roles and billing—Whole tenant
AdminEverything except billingUsers, roles below owner, integrations, API keys, settingsGrant or remove owner; change billingWhole tenant
Sales managerAll customers, all orders, all reps' numbers, price lists, ATS, margin by stylePrice lists, discounts within policy, order approvals, rep-to-account assignmentEdit cost sheets, post payments, change GL mappingWhole tenant
Sales repOwn accounts, own orders, ATS, linesheets, own commissionCreate and edit own orders while status is draft or openSee other reps' accounts, see cost or margin, change prices beyond allowed discount bandOnly rows where they are the assigned rep
Customer serviceAll customers and orders, shipping status, returnsShip-to addresses, order notes, RA (return authorization) creation, split shipmentsChange prices, release credits, see costWhole tenant, read-heavy
WarehousePick tickets, allocations, inventory, ASN and carton dataPick, pack, ship confirm, inventory adjustments with reason codesSee prices, customers' credit status, or marginWhole tenant, inventory domain only
AccountingInvoices, payments, credits, chargebacks, aging, cost, marginPost payments, issue credit memos, resolve chargebacks, credit limitsCreate or edit sales orders; change ATSWhole tenant, finance domain
Read-onlyOrders, inventory, shipping — no cost, no margin, no customer contact PII beyond companyNothingAny write at allWhole tenant or scoped
External rep / agentOwn accounts only, own orders, ATS, linesheets, own commissionCreate and edit own draft ordersExport customer lists, see other agents, see cost, see any tenant-wide reportOwn accounts, plus tighter export limits than an internal rep

Two nuances that matter in this industry. First, cost is the crown jewel. FOB cost and landed cost tell a competitor or a departing rep exactly what your margin is. Cost visibility should be a separate permission, not implied by seniority — some sales managers get it, some do not, and that is a policy decision per customer.

Second, external agents are not employees. A multi-line showroom rep also sells three competing brands. They need order entry and ATS, and they should not be able to export a customer list to CSV. Make export a distinct permission.

How this composes with RLS from chapter 5

Chapter 5 established that every table carries tenant_id and that RLS policies filter on the tenant claim in the JWT. Roles sit inside tenant isolation as an inner layer. The composition is:

          request
             |
   [ tenant isolation ]   RLS: tenant_id = jwt tenant_id
             |            -- never bypassable, no exceptions
   [ role gate         ]  RLS: role in (allowed roles)
             |            -- what kind of user may touch this
   [ row scope         ]  RLS: owner_user_id = auth.uid()
             |            -- which rows within their tenant
   [ column policy     ]  views / column grants
             |            -- cost and margin hidden per role
   [ app-level UX      ]  hide buttons they cannot use
             |            -- convenience only, not security
          response

Read this top to bottom as layers of a sieve. Tenant isolation is absolute and is never relaxed for any role — even an owner cannot see another tenant. The role gate narrows by user type. Row scope narrows further, and this is where "a rep sees only their own accounts" lives. Column policy handles cost and margin, because hiding a column is different from hiding a row and RLS alone does not do it — you use a restricted view or column-level grants.

The app layer at the bottom only hides buttons. Treat that as a courtesy to the user and nothing more, because anyone can call the API directly and skip your interface entirely.

-- Helper functions read claims once, cheaply.
create or replace function public.jwt_tenant_id()
returns uuid language sql stable as $$
  select nullif(
    ((select auth.jwt()) -> 'app_metadata' ->> 'tenant_id'),
    '')::uuid
$$;

create or replace function public.jwt_role()
returns text language sql stable as $$
  select coalesce(
    ((select auth.jwt()) -> 'app_metadata' ->> 'role'),
    'none')
$$;

create or replace function public.jwt_aal()
returns text language sql stable as $$
  select coalesce((select auth.jwt()) ->> 'aal', 'aal1')
$$;

-- ORDERS: who may read and write which rows.
alter table public.orders enable row level security;

create policy orders_tenant_isolation
  on public.orders as restrictive
  for all to authenticated
  using (tenant_id = public.jwt_tenant_id());

create policy orders_read_scope
  on public.orders for select to authenticated
  using (
    case public.jwt_role()
      when 'owner'            then true
      when 'admin'            then true
      when 'sales_manager'    then true
      when 'customer_service' then true
      when 'accounting'       then true
      when 'warehouse'        then status in
             ('allocated','picking','packed','shipped')
      when 'read_only'        then true
      when 'sales_rep'        then rep_user_id = auth.uid()
      when 'external_agent'   then rep_user_id = auth.uid()
      else false
    end
  );

create policy orders_write_scope
  on public.orders for update to authenticated
  using (
    case public.jwt_role()
      when 'owner'          then true
      when 'admin'          then true
      when 'sales_manager'  then true
      when 'sales_rep'      then rep_user_id = auth.uid()
                                and status in ('draft','open')
      when 'external_agent' then rep_user_id = auth.uid()
                                and status = 'draft'
      else false
    end
  )
  -- WITH CHECK is applied to the row AFTER the update.
  -- Without it a rep could hand their order to someone
  -- else, or push it into a status they cannot touch.
  with check (
    case public.jwt_role()
      when 'owner'          then true
      when 'admin'          then true
      when 'sales_manager'  then true
      when 'sales_rep'      then rep_user_id = auth.uid()
                                and status in ('draft','open')
      when 'external_agent' then rep_user_id = auth.uid()
                                and status = 'draft'
      else false
    end
  );

-- Privileged finance tables additionally require MFA.
alter table public.payments enable row level security;

create policy payments_require_mfa
  on public.payments as restrictive
  for all to authenticated
  using (public.jwt_aal() = 'aal2');

The first three functions pull claims out of the verified JWT. Marking them stable lets Postgres call them once per statement rather than once per row, which matters on a 40,000-row order table. Wrapping auth.jwt() in (select ...) is the documented pattern that lets the planner cache the result.

The tenant policy is declared as restrictive, which means it is ANDed with everything else — no later permissive policy can widen past it. The read policy is a single case over the role claim, which reads like the matrix table above and is easy to review in a code review.

Note the two row-scoped roles at the bottom: sales_rep and external_agent only match rows where rep_user_id equals their own user id. That single line is "a rep sees only their own accounts", enforced by the database, applying equally to the web app, the iPad, a CSV export, and a mistake in a background job.

The write policy is deliberately narrower than the read policy: a rep can read an order in any status but can only edit it while it is draft or open. It carries both a using clause, which decides which existing rows may be targeted, and a with check clause, which decides whether the row is still allowed after the edit.

Leaving with check off is a common and expensive mistake: a rep could issue an update that sets rep_user_id to a colleague, and the database would allow it.

The final policy is the MFA gate — payments are unreachable unless the token says a second factor was verified, and because it is restrictive it cannot be bypassed by any other policy.

One gap to close in your own migration: the policies above cover reading and updating. Creating an order needs its own for insert ... with check (...) policy with the same role logic, or a rep will be able to read and edit orders and never create one.

Keep the role list short and boring

Every customer will ask for a custom role in month three. Resist until you have three customers asking for the same one. Nine roles you can hold in your head beat forty granular permissions you cannot audit. When you genuinely need finer control, add named permissions (can_view_cost, can_export) as boolean flags on the membership row and check those in policies, rather than multiplying roles.

Customer-facing access: the B2B portal

Most brands doing meaningful volume should give retailers logins. A portal that lets a buyer at a boutique see order status, download invoices, check ATS and reorder cuts a large share of your customer-service email. But the account model must be genuinely separate from staff, in four ways.

Four ways a portal account differs

Different identity space. A portal user belongs to a customer account, not to your tenant's staff directory. Model it as its own table with its own membership: portal_users linked to customers. Do not put a buyer at Nordstrom into your memberships table with a special role, because one mistaken policy and they are inside your staff surface.

Different default posture. Staff are default-deny with grants; portal users are default-deny with a very small allowlist: their own account's orders, their own invoices, published ATS and linesheets, their own shipping documents. Nothing else exists as far as their queries are concerned.

Different lifecycle. Retail buyers churn constantly and nobody tells you. Expect stale accounts. Auto-deactivate portal users after 12 months of no logins and email them a re-activation link. Let the retailer's own primary contact manage their colleagues' access, so you are not the help desk for someone else's staff turnover.

Different login method. Magic links or emailed codes earn their keep here. A buyer logs in four times a year. They will not remember a password, they will not install an authenticator app, and forcing either will kill adoption. Emailed one-time codes give a decent experience with no password to reset.

-- Portal users are scoped to a customer, not to staff.
create table public.portal_users (
  user_id     uuid primary key references auth.users(id),
  tenant_id   uuid not null references public.tenants(id),
  customer_id uuid not null references public.customers(id),
  role        text not null default 'buyer'
              check (role in ('buyer_admin','buyer','viewer')),
  status      text not null default 'active'
              check (status in ('active','inactive')),
  last_seen_at timestamptz,
  created_at  timestamptz not null default now(),
  unique (tenant_id, user_id)
);

alter table public.portal_users enable row level security;

create policy portal_users_read_self
  on public.portal_users for select to authenticated
  using (user_id = auth.uid());

create or replace function public.jwt_customer_id()
returns uuid language sql stable as $$
  select nullif(((select auth.jwt()) -> 'app_metadata'
                 ->> 'customer_id'), '')::uuid
$$;

-- Portal read access to orders: own customer only.
create policy orders_portal_read
  on public.orders for select to authenticated
  using (
    public.jwt_role() = 'portal'
    and tenant_id = public.jwt_tenant_id()
    and customer_id = public.jwt_customer_id()
    and status <> 'draft'
  );

The portal user carries two extra claims: role = 'portal' and a customer_id. The policy requires both to match, so a Nordstrom buyer sees Nordstrom orders and nothing else, and the status <> 'draft' clause stops them seeing an order your rep is still building.

Because the claims come from the signed JWT, and the JWT is minted by the access token hook in step 5 reading the portal_users table, the buyer cannot alter them. Note that portal deliberately does not appear in the staff roles table, so the staff policies above fall through to else false for a portal user.

Service accounts and API keys

Chapter 4's integrations need to authenticate too, and none of them can type a password or hold a phone. EDI (electronic data interchange, the fixed-format purchase-order and invoice files big retailers exchange), the 3PL's webhook (third-party logistics — the warehouse company that ships for you), the Shopify sync, the accounting export.

The mistake to avoid is issuing them a human user account: a service running as "maria@brandco.com" is invisible in audit logs, breaks when Maria leaves, and inherits every permission Maria has. Create explicit service accounts instead, with their own identity, their own scopes, and no ability to log in through the UI.

create table public.api_keys (
  id           uuid primary key default gen_random_uuid(),
  tenant_id    uuid not null references public.tenants(id),
  name         text not null,             -- 'ShipBob webhook'
  prefix       text not null,             -- 'ak_live_7fQ2'
  key_hash     bytea not null unique,     -- sha256(full key)
  scopes       text[] not null default '{}',
  created_by   uuid not null references auth.users(id),
  created_at   timestamptz not null default now(),
  expires_at   timestamptz,
  last_used_at timestamptz,
  last_used_ip inet,
  revoked_at   timestamptz,
  revoked_by   uuid references auth.users(id)
);

create index api_keys_lookup on public.api_keys (prefix)
  where revoked_at is null;

-- RLS on, no policies: server-side code only.
alter table public.api_keys enable row level security;

-- Scope vocabulary: verb:resource, no wildcards.
--   orders:read      orders:write
--   inventory:read   inventory:write
--   invoices:read    shipments:write
--   customers:read   webhooks:receive

The key you hand out looks like ak_live_7fQ2_9d3c1a8b....

  • The prefix column stores only the first segment, which lets you look the key up with an index and lets a human recognize it in a config file without exposing the secret.
  • key_hash stores SHA-256 of the whole key, so your database never contains a usable credential — same reasoning as the invitation token, and the same \x hex conversion applies when you write it.
  • scopes is an explicit array with no wildcard, so a 3PL webhook key that only needs to post shipment confirmations gets exactly shipments:write and nothing else.
  • last_used_at and last_used_ip let you find keys nobody uses and revoke them, and tell you fast whether a leaked key is being exercised from somewhere unexpected.
  • expires_at forces rotation to be a scheduled event rather than a crisis.

Rotation without downtime

Rotation means replacing a key while the integration keeps running. The pattern is always the same, and it applies to JWT signing keys, API keys and webhook secrets alike, so learn it once:

1. Issue key B alongside key A. Both valid.
2. Update the consumer's config to use B.
3. Watch api_keys.last_used_at for A. When it stops
   moving for a full cycle of that integration
   (24h for a nightly job, 1h for a webhook), proceed.
4. Revoke A. Set revoked_at, do not delete the row.
5. Log the rotation as an auth event.

Same shape for your JWT signing keys (if you mint sync/SSO tokens):
  publish standby key
  -> wait ~20 min for JWKS caches to clear
  -> rotate (standby becomes active)
  -> wait token-expiry + 15 min  (1h15m at a 1h token)
  -> revoke the previous key

Step 3 is the one people skip, and skipping it is how integrations break at 2am. The last_used_at column exists specifically so you can answer "is anything still using the old key?" with data instead of hope.

The signing-key timings are the standard JWKS ones: your JWKS endpoint is cached about 10 minutes at the edge and another 10 minutes inside consuming libraries, hence the 20-minute wait after creating a standby key; and before revoking a previously-used key you "wait for your configured access token expiry time plus 15 minutes", so one hour and fifteen minutes at the default token life. Revoke sooner than that and you sign out everyone still holding a token signed by the old key.

Two more rules. Never put an API key in the front-end bundle or in a git repository — use environment variables, and run a secret scanner such as gitleaks in CI (continuous integration: the automated checks that run on every push before code can merge).

Managed platforms formalize this split with two key types worth borrowing as a mental model: publishable keys (sb_publishable_...) run through the anon and authenticated Postgres roles and respect RLS; secret keys (sb_secret_...) use the service_role Postgres role, skip "any and all Row Level Security policies", and are refused with HTTP 401 when the request looks like it came from a browser. In your own stack the equivalent is stark: application requests connect as app_user (RLS applies), and exactly one command connects as the owner role (RLS bypassed) — the legacy anon and service_role JWT keys "will be deprecated by the end of 2026" (checked July 2026), so start new work on the new format.

The secret key bypasses every policy you wrote

Any code path holding sb_secret_... runs as service_role and RLS does not apply. That is why invitation acceptance, offboarding and impersonation in this chapter all do their own explicit permission check first. Treat every function that constructs an admin client as security-critical code and review it as such. The same warning covers security definer SQL functions: they run as their owner, so each one needs its own authorization check and its own revoke ... from authenticated.

Account security operations

Password reset done safely

The reset flow has four requirements:

  • The response must be identical whether or not the email exists, or you have built an account-enumeration tool.
  • The reset token must be single-use and short-lived.
  • Using it must invalidate existing sessions.
  • And the user must be told, at their old address, that a reset happened.
// Step 1: request a reset. Always returns the same thing.
export async function requestResetHandler(req, res) {
  const { email } = await readJson(req)

  // Mint a single-use token, store only its hash, mail the raw.
  // Do this whether or not the user exists — see below.
  const user = await findUserByEmail(email)
  if (user) {
    const { raw, hash } = newAuthToken()          // 32 random bytes
    await withTenant(SYSTEM, (tx) => tx`
      insert into auth_token (user_id, purpose, token_hash, expires_at)
      values (${user.id}, 'reset', ${hash}, now() + interval '1 hour')`)
    await sendResetEmail({ to: email,
      url: `${process.env.APP_URL}/account/new-password?t=${raw}` })
  }

  // Same response either way, so the endpoint cannot be used to
  // learn which addresses have accounts.
  respond(res, 200, {
    message: 'If that address has an account, we have sent a link.' })
}

Ignoring the result is the point. If you returned "no such user" for unknown addresses, anyone could paste in a list of emails and learn which ones are your customers' staff. Give the reset token the same shape as everything in auth_token: one hour, single use, and rate-limited to one request per address per 60 seconds.

The redirect target in that link must be an app-relative path you validate on the way out (the confirm route below rejects anything that is not a same-site path), or the route becomes an open redirect a phisher can point at your users. That validation is the self-hosted equivalent of a managed "redirect URL allowlist," and it lives in code you can read.

// Step 2: the /account/new-password route consumes the token
// and lets the holder set a password. No vendor in the loop.
import { hash as sha256 } from '../auth/hash'

export async function confirmResetHandler(req, res) {
  const url = new URL(req.url, process.env.APP_URL)
  const raw  = url.searchParams.get('t')
  const next = safeNext(url.searchParams.get('next'))  // same-site only
  if (!raw) return redirect(res, '/auth/error?e=missing')

  await withTenant(SYSTEM, async (tx) => {
    // Look up by HASH of the presented token, and consume it
    // atomically so a link works exactly once.
    const [t] = await tx`
      update auth_token set consumed_at = now()
      where purpose = 'reset' and token_hash = ${sha256(raw)}
        and consumed_at is null and expires_at > now()
      returning user_id`
    if (!t) return redirect(res, '/auth/error?e=invalid')
    // token good: render the set-password form for t.user_id,
    // and revoke their other sessions once the new password lands.
    return redirect(res, next)
  })
}

function safeNext(raw) {
  return raw && raw.startsWith('/') && !raw.startsWith('//') ? raw : '/'
}

This route is what your email template links to when you use {{ .TokenHash }} instead of {{ .ConfirmationURL }}. It swaps the hash for a real session and sets the cookies, then forwards to wherever the flow was heading. The same route serves signup confirmation, magic links, email change and reset, distinguished by the type parameter — one route, four flows.

The next check earns its place. Any route that redirects to a user-supplied URL is an open redirect, and an attacker will mail your customers a link that starts on your real domain and lands on theirs.

After the new password is set, email a notification to the address on file, and check with your own two-browser test what happened to the other sessions. Sending the "your password was changed" email is your job either way, and it is how a compromised account gets caught by its owner within minutes.

Email change verification

An attacker with a live session who can change the email address owns the account permanently, because every future reset goes to them. The defense is confirming both addresses. Do this by minting two auth_token rows — one mailed to the current address, one to the new — and completing the change only when both are consumed. (Managed consoles expose the same behavior as a "secure email change" toggle; here it is two token rows and an and.)

Then require re-authentication before the change: mail a one-time code to the address currently on file and require it back before you act — the same pattern as any other auth_token, used as a nonce argument to updateUser. A nonce is a value that works once and then stops working. Log the old and new values in your audit trail.

Lockout and rate limiting

Rate limiting protects three different things and they need different limits: credential guessing against one account; credential stuffing across many accounts from one source, which means replaying email-and-password pairs leaked from some other website; and email or SMS flooding that costs you money and reputation. The middle column below lists a managed vendor's defaults (Supabase's published reference) as sane starting numbers to copy into your own limiter.

EndpointA sane default to copyWhat to add yourself
Password sign-in (/token)Not separately limited; shares the token bucket with refreshPer-account: exponential backoff after 5 failures, cap at 15 min. Per-IP: 20 attempts / 10 min.
Token refresh1,800 / hour per IP, bursts to 30Alert if any single IP approaches it — usually a refresh loop bug. Turn on IP forwarding for server-rendered apps.
Token verification (/verify)360 / hour per IP, bursts to 30—
OTP / magic link30 OTPs / hour project-wide; 60-second window per userRaise the project limit once on custom SMTP; per-email cap of 5/hour
Signup confirmation, password reset60-second window per userSurface the wait in the UI so users stop hammering the button
MFA challenge / verify15 / hour per IP, with burstsLock the factor after 10 bad codes; require a recovery code
Built-in email sending2 emails per hour, project-wideConfigure custom SMTP before launch — this default will block real onboarding
Anonymous sign-ins30 / hour per IP, bursts to 30Disable entirely unless you need it

A token bucket is the usual shape for the IP-limited endpoints: each bucket holds ~30 requests, so an idle client tolerates a short spike, and sustained traffic above the rate is refused with HTTP 429.

Prefer exponential backoff over hard lockout in your own layer. A hard "locked for 30 minutes after 5 failures" rule is itself a denial-of-service tool: an attacker who knows your controller's email can lock them out during month-end close by failing five logins. Backoff with a cap slows an attacker to uselessness while a legitimate user who fixes their typo waits seconds.

Combine it with leaked-password checking so the passwords being guessed are not in the breach corpus to begin with, and with a CAPTCHA (Cloudflare Turnstile and hCaptcha are the common choices) triggered only after repeated failures rather than on every login.

You will need to see what a user sees. "It says my order is missing" is unanswerable without it. Casual impersonation is a backdoor. The version worth building is logged, consented, time-bounded and impossible to miss on screen.

// POST /api/support/impersonate
export async function POST(req: Request) {
  const actor = await requireRole(req, ['owner', 'admin'])
  const { targetUserId, reason, ticketRef } = await req.json()

  if (!reason || reason.length < 20) {
    return Response.json(
      { error: 'A written reason is required.' },
      { status: 400 })
  }

  // Consent must already exist and be unexpired.
  const consent = await getActiveConsent(targetUserId)
  if (!consent) {
    await notifyUser(targetUserId, 'impersonation_requested')
    return Response.json(
      { error: 'awaiting user consent' }, { status: 409 })
  }

  const session = await createImpersonationSession({
    actorId: actor.userId,
    targetUserId,
    tenantId: actor.tenantId,
    expiresAt: new Date(Date.now() + 30 * 60_000),
    readOnly: true,          // writes blocked by default
  })

  await audit('support.impersonation.started', actor, {
    target_user_id: targetUserId, reason, ticketRef,
    consent_id: consent.id, expires_at: session.expiresAt })

  await notifyUser(targetUserId,
    'impersonation_started', { actor: actor.displayName })

  return Response.json({ token: session.token })
}

Six controls in one handler:

  • A written reason of real length, so the audit log is useful six months later.
  • Explicit prior consent from the user, which they grant in their own settings and which expires.
  • A 30-minute hard expiry.
  • Read-only by default, so support can look without accidentally changing an order.
  • An audit event containing actor, target, reason and ticket reference.
  • And a notification to the user that it is happening.

In the UI, an impersonated session must show a permanent, loud banner, a colored bar across the top reading "Viewing as Maria Chen, read only — 24 minutes left, End session", because the worst impersonation incident is the one where the support person forgot they were in it.

The impersonation token should carry distinct claims: act (the real actor's id) alongside sub (the impersonated user). Then every audit record written during the session automatically shows both, and your RLS policies can refuse writes when act is present. Mint that token from your own signing key in your own service — an ordinary session row with an act claim naming the real operator, so the audit trail always shows who acted, not just as whom.

Breach response

Decide this before you need it, because you will decide badly at 3am. Write a one-page runbook covering the four scenarios that actually happen:

One user's credentials are compromised. Ban the account, force a password reset, revoke their trusted devices, review their audit trail for the last 30 days for exports and permission changes, notify the tenant's owner, then lift the ban once they are back on a clean device.

A secret key leaked (committed to git, pasted in Slack). Rotate the leaked credential — a new SESSION_SECRET or database password, deployed everywhere, the old one retired — then review database logs for access from unexpected IPs during the exposure window. Assume every row was readable and act accordingly.

Your JWT signing key is suspected compromised. Create a standby key, rotate, and revoke the previous key immediately rather than waiting the usual hour and fifteen minutes; accept that everyone gets signed out — which, since your sessions are a table, is one UPDATE session SET revoked_at = now(), no vendor permission required.

A customer's identity provider is compromised. Disable that SSO connection, which forces their users to your fallback method, and coordinate with their IT before re-enabling.

In all four cases: preserve logs before changing anything, write the timeline as you go, and tell the affected customer the same day. Depending on where your customer and their staff are, you may have statutory notification deadlines — chapter M covers the legal side. Your job here is to make sure the technical evidence exists.

Auditing authentication events

Chapter 5 covers the SOC 2 posture — SOC 2 being the security audit report that enterprise customers ask for before they will sign. This section is about the specific evidence auth must produce. The controls an auditor asks about — logical access, provisioning, deprovisioning, privileged access review — are all questions your auth logs must answer without you doing archaeology.

Log every one of these, as structured JSON, into a table you own:

EventMust captureWhy an auditor cares
auth.login.successuser, tenant, method, aal, ip, user agent, device idProves access is authenticated and traceable
auth.login.failureattempted email (hashed if you prefer), ip, reason classDetects brute force; shows monitoring exists
auth.mfa.enrolled / .removed / .recovery_useduser, factor type, actorMFA policy enforcement evidence
auth.logout and auth.session.revokeduser, session id, trigger (manual / role change / offboard)Deprovisioning is timely
member.invited / .accepted / .role_changed / .offboardedactor, subject, old role, new role, timestampAccess provisioning is authorized and reviewed
apikey.created / .rotated / .revokedactor, key prefix, scopesService account governance
support.impersonation.started / .endedactor, target, reason, consent, durationPrivileged access is controlled and consented
auth.password.changed / .reset / email.changeduser, actor, old and new emailAccount takeover detection
create table public.auth_events (
  id          bigint generated always as identity primary key,
  occurred_at timestamptz not null default now(),
  tenant_id   uuid,
  actor_id    uuid,        -- who did it
  subject_id  uuid,        -- who it was done to
  event       text not null,
  ip          inet,
  user_agent  text,
  session_id  uuid,
  detail      jsonb not null default '{}'::jsonb
);

create index auth_events_tenant_time
  on public.auth_events (tenant_id, occurred_at desc);
create index auth_events_subject_time
  on public.auth_events (subject_id, occurred_at desc);
create index auth_events_event_time
  on public.auth_events (event, occurred_at desc);

-- Append-only for clients: no inserts, updates or deletes.
-- Rows arrive from triggers and from the secret-key client.
revoke insert, update, delete on public.auth_events
  from authenticated, anon;

alter table public.auth_events enable row level security;

create policy auth_events_read
  on public.auth_events for select to authenticated
  using (tenant_id = public.jwt_tenant_id()
         and public.jwt_role() in ('owner','admin'));

Separating actor_id from subject_id is what lets you answer both "what did this admin do?" and "what was done to this user?" from the same table. Two column types need a word of explanation: inet is Postgres's type for an IP address, and jsonb is its binary JSON type, which stores a whole JSON object in one column and can be indexed.

The three indexes cover the three questions you will actually ask. Revoking insert, update and delete from the client roles makes the table append-only from the outside, in the same spirit as chapter 1's ledger — an audit log that can be edited is not evidence. RLS then limits reads to owners and admins within the tenant, so a rep cannot browse the login history of the whole company.

Two caveats to be honest about: the service_role key still bypasses all of this, and a database superuser can still edit rows, so if your auditor wants true immutability you ship a copy to append-only object storage as well.

Retention and what not to log

On retention: your logs live as long as your log storage keeps them — you set the number, a rotation policy on the box or a bucket lifecycle rule, rather than inheriting a plan tier's window. Ninety days is a common floor for auth events; whatever you pick, it is fine for debugging and useless for compliance. Keep your own auth_events table for 12 months hot in Postgres, then archive older rows to object storage as compressed JSONL (one JSON object per line, which is easy to append to and easy to grep) on a monthly job.

Twelve months hot covers a SOC 2 Type II observation window — Type II audits watch your controls over a stretch of time, commonly three to twelve months, and any "who changed this price in March" question. Do not log passwords, tokens, or full session identifiers — OWASP recommends logging "a salted-hash of the session ID instead of the session ID itself in order to allow for session-specific log correlation without exposing the session ID".

Log the failures as carefully as the successes. A spike in auth.login.failure for one account is an attack; a spike in permission denials for one user is either a misconfigured role or someone probing. Alert on rates, not on individual events, or you will mute the channel within a week.

A complete worked implementation: sessions you own

Everything above was principle. Here is the whole thing built from the standard library — email and password with mandatory TOTP for privileged roles, opaque cookie sessions in a database table, tenant claims your own server sets, and RLS doing the enforcement. There is no auth vendor in this design, and there is no auth framework either: the primitives you need — a slow password hash, random bytes, a constant-time compare, a TOTP verifier — all ship inside node:crypto. That is the whole reason this book can teach auth from zero without importing anyone's opinions: the dangerous parts are already written, correctly, by the people who wrote the runtime.

The one rule of rolling your own auth

"Never roll your own crypto" is real, and this design obeys it: every cryptographic operation below is a documented node:crypto call, not an invention. What you are building is the plumbing around those calls — a sessions table, a cookie, a login handler — which is ordinary application code that a managed platform would otherwise write for you at the cost of a dependency between your users and their login. If a requirement here ever exceeds what the standard library offers cleanly (WebAuthn attestation is the usual first one), that is the moment to add one audited library for that one job — deliberately, on the record, not by default.

Step 1: environment, and the absence of dependencies

# There is nothing to npm install for auth. node:crypto is built in.
# The only secrets auth needs live in the environment:

touch .gitignore
grep -qxF '.env' .gitignore || echo '.env' >> .gitignore

if [ -e .env ]; then
  echo ".env exists - edit it by hand."
else
  umask 077
  cat > .env <<'EOF'
# 32 random bytes, base64. Signs session cookies. Same value on
# every app instance, set once, NEVER regenerated at startup.
SESSION_SECRET=REPLACE_WITH_$(openssl rand -base64 32)
APP_URL=https://erp.brandco.com
DATABASE_URL=postgres://app_user:REPLACE@localhost:5432/erp
EOF
  chmod 600 .env
  echo "Created .env - fill in the real values now."
fi

Read what is not here: no client library, no publishable key, no service-role key, no NEXT_PUBLIC_ anything. The browser is never handed a credential of any kind — it holds one opaque cookie, and that is the entire client-side surface. SESSION_SECRET is the one value that must be identical across every app process and must survive restarts, because it signs the cookies one process issues and another must verify (chapter Q's shared-secret rule). It is set once in the box's environment; a process that generates its own at startup signs everybody out on every deploy.

Step 2: configuration, in files you can read

There is no dashboard. The settings a managed console would bury behind tabs are constants in one config module, checked into git, reviewed like any other code:

// apps/admin/auth/config.ts
export const AUTH = {
  // Password hashing (scrypt). Cost tuned to ~100ms on your box;
  // measure it, because too-fast is the whole vulnerability.
  scrypt: { N: 2 ** 15, r: 8, p: 1, keyLen: 32 },

  // Sessions.
  session: {
    cookieName: "erp_session",
    idleTimeout:     30 * 60 * 1000,        //  30 minutes
    absoluteTimeout:  8 * 60 * 60 * 1000,   //   8 hours (office)
    boothTimeout:    14 * 24 * 60 * 60 * 1000, // 14 days (offline iPad)
  },

  // Login throttling (chapter's lockout section).
  lockout: { attempts: 10, windowMs: 15 * 60 * 1000 },

  // Which roles must carry a second factor.
  mfaRequiredRoles: ["owner", "admin", "controller"],
} as const;

Two timeouts, not one, because the office controller and the show-floor iPad are different security problems (the chapter's session and offline sections). The MFA policy is a list a human can read and a reviewer can question, not a toggle in someone else's UI.

Step 3: email deliverability, or nothing else works

Password resets, magic links and verification mails are worthless if they land in spam, and self-hosting means the deliverability is yours too. You do not run a mail server — that way lies a blocked IP and a weekend lost — you relay through a transactional provider (Postmark, SES, Resend) over SMTP, and you publish three DNS records so receivers trust you: SPF (which servers may send as your domain), DKIM (a signature proving the mail was not altered), and DMARC (what to do when a message fails the first two). Chapter O sets these records; here, just know that the worker's mail seam (chapter 3) hands finished bytes to that relay and never blocks a login on it.

Step 4: the schema

Five tables carry the whole design. Every one is plain Postgres, under the same RLS and the same app_user role as everything else — auth is not a special citizen.

-- A person who can sign in. One row per human, global — a person
-- may belong to several tenants (the memberships live elsewhere).
create table app_user (
  id            uuid primary key default gen_random_uuid(),
  email         citext unique not null,     -- citext: case-insensitive
  password_hash text,                       -- null = SSO-only account
  email_verified_at timestamptz,
  created_at    timestamptz not null default now()
);

-- One row per active login. The cookie carries only this id
-- (signed); everything authoritative lives here, server-side.
create table session (
  id            uuid primary key default gen_random_uuid(),
  user_id       uuid not null references app_user(id) on delete cascade,
  tenant_id     uuid not null,              -- the workspace this login acts in
  created_at    timestamptz not null default now(),
  last_seen_at  timestamptz not null default now(),
  absolute_expiry timestamptz not null,
  revoked_at    timestamptz,                -- set on logout / offboarding
  user_agent    text,
  ip            inet
);
create index session_user_idx on session(user_id) where revoked_at is null;

-- Second factor. secret is encrypted at rest (Step 7).
create table user_totp (
  user_id       uuid primary key references app_user(id) on delete cascade,
  secret_enc    bytea not null,
  confirmed_at  timestamptz
);

-- Single-use tokens for reset / verify / invite. We store only a
-- hash, so a stolen database is not a stack of live links.
create table auth_token (
  id            uuid primary key default gen_random_uuid(),
  user_id       uuid references app_user(id) on delete cascade,
  purpose       text not null check (purpose in
                  ('reset','verify_email','invite')),
  token_hash    bytea not null,
  expires_at    timestamptz not null,
  consumed_at   timestamptz
);

-- Append-only trail of auth events (the chapter's audit section).
create table auth_event (
  id            bigint generated always as identity primary key,
  user_id       uuid references app_user(id),
  tenant_id     uuid,
  kind          text not null,   -- login_ok, login_fail, mfa_ok, revoke…
  ip            inet,
  at            timestamptz not null default now()
);

The shape is the doctrine: the session is server-side state, the cookie is a pointer to it, and "log this person out everywhere, now" is one UPDATE session SET revoked_at = now() — the revocation ceiling that JWT-only designs cannot offer.

Step 5: password hashing and the session cookie

Two functions carry the cryptography, and both are node:crypto calls in disguise. Hashing first — scrypt, deliberately slow, salt stored beside the hash:

// apps/admin/auth/password.ts
import { scrypt, randomBytes, timingSafeEqual } from "node:crypto";
import { promisify } from "node:util";
import { AUTH } from "./config";

const scryptAsync = promisify(scrypt);

export async function hashPassword(plain: string): Promise<string> {
  const salt = randomBytes(16);
  const key = (await scryptAsync(plain, salt, AUTH.scrypt.keyLen,
    { N: AUTH.scrypt.N, r: AUTH.scrypt.r, p: AUTH.scrypt.p })) as Buffer;
  // Store params WITH the hash, so raising cost never invalidates
  // stored passwords — old ones verify at their old cost.
  return `scrypt$${AUTH.scrypt.N}$${AUTH.scrypt.r}$${AUTH.scrypt.p}$` +
         `${salt.toString("base64")}$${key.toString("base64")}`;
}

export async function verifyPassword(plain: string, stored: string) {
  const [, N, r, p, saltB64, keyB64] = stored.split("$");
  const salt = Buffer.from(saltB64, "base64");
  const expected = Buffer.from(keyB64, "base64");
  const got = (await scryptAsync(plain, salt, expected.length,
    { N: +N, r: +r, p: +p })) as Buffer;
  // timingSafeEqual: a byte-by-byte compare that always takes the
  // same time, so an attacker learns nothing from how fast it fails.
  return got.length === expected.length && timingSafeEqual(got, expected);
}

Then the cookie. It carries the session id and a signature over it — nothing else, no claims a tamperer could rewrite, because the authoritative facts live in the session row the id points at:

// apps/admin/auth/cookie.ts
import { createHmac, timingSafeEqual } from "node:crypto";
const SECRET = Buffer.from(process.env.SESSION_SECRET!, "base64");

function sign(id: string) {
  const mac = createHmac("sha256", SECRET).update(id).digest("base64url");
  return `${id}.${mac}`;
}
export function readCookie(raw?: string): string | null {
  if (!raw) return null;
  const dot = raw.lastIndexOf(".");
  if (dot < 0) return null;
  const id = raw.slice(0, dot);
  const want = createHmac("sha256", SECRET).update(id).digest();
  const got  = Buffer.from(raw.slice(dot + 1), "base64url");
  return got.length === want.length && timingSafeEqual(got, want)
    ? id : null;   // bad signature = no session, no exception thrown
}
export function setCookie(res, id: string, maxAgeMs: number) {
  res.setHeader("Set-Cookie",
    `${"erp_session"}=${sign(id)}; HttpOnly; Secure; SameSite=Lax; ` +
    `Path=/; Max-Age=${Math.floor(maxAgeMs / 1000)}`);
}

HttpOnly keeps JavaScript (and therefore any XSS) from reading it; Secure keeps it off plain HTTP; SameSite=Lax stops other sites replaying it. The signature means a forged id is rejected before a database round-trip. This is the "classic session cookie" from earlier in the chapter, in thirty lines you own.

Step 6: the login handler and the per-request gate

Now the flow, as one route handler — the five-step choreography from chapter 10, specialized to login:

// apps/admin/handlers/login.ts
export async function loginHandler(req, res) {
  const { email, password } = parseLogin(await readJson(req));   // validate

  await withTenant(SYSTEM, async (tx) => {   // system ctx: pre-login
    if (await tooManyRecentFailures(tx, email, req.ip))          // lockout
      return respond(res, 429, { error: "try_again_later" });

    const user = (await tx`
      select * from app_user where email = ${email}`)[0];
    // Verify even when the user is missing, against a dummy hash, so
    // response time does not reveal which emails exist.
    const ok = user
      ? await verifyPassword(password, user.password_hash ?? DUMMY_HASH)
      : (await verifyPassword(password, DUMMY_HASH), false);
    if (!ok) {
      await recordEvent(tx, { email, kind: "login_fail", ip: req.ip });
      return respond(res, 401, { error: "invalid_credentials" });
    }

    // Password held. If the role needs a second factor, do NOT issue a
    // full session yet — issue a short "mfa pending" one and 200 with a
    // challenge. Only a passed TOTP (Step 7) mints the real session.
    if (await roleNeedsMfa(tx, user.id))
      return respond(res, 200, { mfa: await beginMfa(tx, user.id) });

    return issueSession(tx, res, user, req);   // the real thing
  });
}

async function issueSession(tx, res, user, req) {
  const [s] = await tx`
    insert into session (user_id, tenant_id, absolute_expiry,
                         user_agent, ip)
    values (${user.id}, ${defaultTenantFor(user)},
            now() + ${AUTH.session.absoluteTimeout} * interval '1 ms',
            ${req.headers["user-agent"]}, ${req.ip})
    returning id`;
  setCookie(res, s.id, AUTH.session.absoluteTimeout);
  await recordEvent(tx, { user_id: user.id, kind: "login_ok", ip: req.ip });
  respond(res, 200, { ok: true });
}

And the gate every other request passes — this is where the cookie becomes a bound tenant, the seam chapter 5 assumed existed:

// apps/admin/auth/session.ts
export async function getSession(req): Promise<Ctx | null> {
  const id = readCookie(cookieOf(req, AUTH.session.cookieName));
  if (!id) return null;
  return withTenant(SYSTEM, async (tx) => {
    const [s] = await tx`
      update session
      set last_seen_at = now()
      where id = ${id}
        and revoked_at is null
        and absolute_expiry > now()
        and last_seen_at > now() - ${AUTH.session.idleTimeout}
                                    * interval '1 ms'
      returning user_id, tenant_id`;      // one query: validate + slide
    return s ? { userId: s.user_id, tenantId: s.tenant_id } : null;
  });
}

One statement validates the session, enforces both timeouts, and slides the idle window — and returns nothing for a revoked or expired one, so a logout takes effect on the very next request. The tenantId it returns is exactly what withTenant binds into app.tenant_id, which is what the RLS policy reads. The whole circle closes here.

Step 7: TOTP, from the standard library

A time-based one-time code is a truncated HMAC of the current 30-second window under a shared secret — and HMAC is, again, node:crypto. Enrollment stores the secret encrypted (the envelope pattern from chapter 5); the challenge recomputes and compares in constant time, checking the adjacent window to tolerate clock skew:

// apps/admin/auth/totp.ts
import { createHmac, timingSafeEqual } from "node:crypto";

function codeFor(secret: Buffer, counter: number): string {
  const buf = Buffer.alloc(8);
  buf.writeBigInt64BE(BigInt(counter));
  const mac = createHmac("sha1", secret).update(buf).digest();
  const off = mac[mac.length - 1] & 0x0f;
  const bin = (mac.readUInt32BE(off) & 0x7fffffff) % 1_000_000;
  return bin.toString().padStart(6, "0");
}

export function verifyTotp(secret: Buffer, code: string): boolean {
  const now = Math.floor(Date.now() / 1000 / 30);
  for (const w of [now - 1, now, now + 1]) {          // ±1 window skew
    const want = Buffer.from(codeFor(secret, w));
    const got  = Buffer.from(code.padStart(6, "0"));
    if (got.length === want.length && timingSafeEqual(got, want))
      return true;
  }
  return false;
}

That is the entire second factor — RFC 6238, forty lines, no library. The enrollment QR code an authenticator app scans is just an otpauth:// URL; render it with any QR helper, or ship the URL as text. Recovery codes (the chapter's section) are ten random strings, each stored as a hash in auth_token with purpose recovery, consumed on use.

Step 8: revoke on role change and offboarding

Because a session is a row, not a self-contained token, the two operations that must be instant actually are:

-- Offboard: kill every active login for a person, everywhere.
update session set revoked_at = now()
where user_id = $1 and revoked_at is null;

-- Role downgrade: kill their sessions in the affected tenant, so the
-- next request re-resolves their (now lesser) permissions from scratch.
update session set revoked_at = now()
where user_id = $1 and tenant_id = $2 and revoked_at is null;

Wire the first to the offboarding path and the second to the membership-change verb, both in the same transaction as the change itself, and "their access ended when the record said it did" is a fact your audit trail can prove — not a promise about a token expiring within the hour. This is the concrete payoff of owning the session table, and it is the thing a JWT-only design has to build a revocation list to fake.

Step 9: verify it works

The proof is a handful of node:test cases over the functions above — no running auth service, no mocked vendor:

  • Hashing round-trips and rejects. verifyPassword(pw, await hashPassword(pw)) is true; a wrong password is false; and the same password hashed twice yields different strings (the salt is doing its job).
  • The cookie signature is unforgeable. readCookie(sign(id)) returns the id; flip one character of the signature and it returns null, without throwing.
  • Timeouts bite. Insert a session with last_seen_at in the past and assert getSession returns null; the same for one past its absolute expiry, and one whose revoked_at is set.
  • TOTP accepts the current window and rejects a stale one. Generate a code for now and accept it; generate one for now - 5 windows and reject it.
  • The cross-tenant test from chapter 5 still passes, because getSession feeds withTenant and the RLS fence does the rest — auth and isolation are one circuit, and this is where you see it close.

Nine steps, one dependency-free auth system, every dangerous operation a documented standard-library call. It is more code than pasting a vendor's SDK — a few hundred lines against a few dozen — and in exchange nothing sits between your users and their login, the revocation ceiling is a single UPDATE, and the day an auditor asks "show me exactly what happens when someone signs in," the answer is a file they can read rather than a vendor's documentation you have to trust.

Field notes & further reading

  • OWASP — Session Management Cheat Sheet: the reference this chapter's session table is built to satisfy — idle vs absolute timeouts, rotation, storage, and revocation.
  • RFC 7519 — JSON Web Token: what actually goes in a token's claims, and (§4.1) the registered names iss, aud, sub, exp the chapter's payload uses.
  • RFC 7517 — JSON Web Key and RFC 7518 (JWA): the JWKS format your rotation publishes and the ES256/RS256/HS256 algorithms it chooses among.
  • RFC 6238 — TOTP and RFC 4226 — HOTP: the exact algorithm behind Step 7's forty-line verifier; read them and the code stops being magic.
  • OWASP Session Management Cheat Sheet: cookie attributes, the __Host- prefix, session identifier length and entropy, the idle and absolute timeout ranges quoted above, and the salted-hash rule for logging session ids.
  • NIST SP 800-63B rev 4: the current word on password rules — no composition rules, no forced rotation, 15 characters for single-factor and a floor of 8 inside multi-factor — plus the AAL definitions and why PSTN/SMS is a restricted authenticator.
  • Google — Email sender guidelines: the requirements in force since 1 February 2024, including the 5,000-messages-per-day bulk threshold, SPF plus DKIM plus DMARC, From-domain alignment, and the 0.30% spam-rate ceiling that governs whether your password resets arrive.
  • Node — node:crypto: scrypt, HMAC, randomBytes and timingSafeEqual — every cryptographic primitive this chapter uses, already written and audited by the runtime's maintainers.
  • PowerSync — Custom authentication: the JWT contract for offline clients — required claims, the 86,400-second maximum lifetime, the 60-minute recommendation and rejection rule, the supported signature algorithms, and the JWKS key requirements.
Exercise

1. Build the invitation loop end to end. Create the invitations table from this chapter in a fresh local Postgres. Write the create endpoint and the accept endpoint. Then deliberately break it five ways and confirm each is rejected with the same generic message: accept an invitation twice, accept one you have manually expired by setting expires_at to yesterday, accept one you have revoked, accept with a token you invented, and accept a valid token twice from two terminals at the same moment. Confirm your auth_events table has a row for the successful path and that the raw token appears nowhere in your database or your server logs.

2. Prove the rep row-scope with two real tokens. Create two users with the sales_rep role in the same tenant, and three orders — one assigned to rep A, two to rep B. Sign in as each with curl, and confirm from the REST API (not the UI) that A sees exactly one order and B sees exactly two. Then have A attempt an update that sets rep_user_id to B on their own order, and confirm the with check clause refuses it. Change rep A's role to sales_manager in the database, confirm the trigger wrote an audit row, sign in again, and confirm A now sees all three. Finally set A's membership status to inactive, refresh the token, and confirm the claims contain role: none and every query returns zero rows.

3. Prove RLS is actually on. Run select relname, relrowsecurity from pg_class where relnamespace = 'public'::regnamespace order by 1 and confirm no table shows false. Then, with only the publishable key and no Authorization header, try to read memberships, api_keys and invitations over the REST API and confirm all three come back empty.

When all three exercises are done you should have: a working invite-and-accept flow whose tokens are unguessable, single-use and hashed at rest; a JWT that carries tenant, role and permission flags from your own tables; RLS policies that scope rows by rep, gate finance tables behind MFA, and refuse a row the user is not allowed to leave behind; and automatic audit plus account lockout whenever a role or status changes. That is a defensible access-control foundation for a real customer.