BOLA is not an authentication bug

•Michael•32 min read

The login worked, the token is valid, and the client still read someone else's statement.

Everyone hardens /login, requires a strong password, turns on MFA, rotates JWTs, and leaves GET /accounts/:id/statement trusting that whoever shows up with a valid token will only ever ask for their own id.

The number one API bug of the last few years is not injection, not weak passwords, not leaked tokens.

It's Broken Object Level Authorization, BOLA, the same old IDOR under a new name, first place in the OWASP API Security Top 10 in 2019 and again in 2023.

That second sentence usually reads like an opening statistic, it is, in fact, the subject of this whole piece.

Four years separate the two editions, and in between, the entire industry talked about BOLA: it became a chapter in every API security course, a free lab on PortSwigger, a mandatory item on vendor questionnaires, a selling point for half a dozen products, and none of that knocked the flaw off the top.

So the question that matters isn't "what is BOLA" (you probably already know), it's why a flaw that can be described in one sentence survives so much awareness of it.

Two answers, and the rest of this article makes the case for both.

The first: BOLA doesn't require bad code. It's born from the most natural way to write an endpoint: receive an id, look up the object, return it, the safe path takes one extra step, tying the lookup to the owner, and that step is easy to forget once in three hundred times, there is no moment where someone writes something obviously wrong; there's a moment where someone writes the obvious thing, and the obvious thing is insecure.

The second: no tool finds this for you. No scanner knows that statement 9930 should be invisible to user 42, because it has no way of knowing your domain's ownership rules, it sees two 200 responses and has no way to judge which one is wrong, there is no signature for "this object belongs to someone else," there is only business context, and business context is exactly what the tool doesn't have.

Put the two together: a flaw that depends on the developer remembering, multiplied by the number of endpoints, with no automated detector, that's not a ranking you beat with attention, it's a ranking you only beat by changing how the query itself gets written, so that forgetting results in denial.

Meanwhile, what gets written about API security almost always follows the same script: validate the JWT, don't use HS256, use RS256, expire the token in 15 minutes, turn on refresh token rotation, rate limit the login.

All of it correct, all of it about the front door, and the bug that leaks the most data in financial APIs lives inside the house, in a find(params[:id]) that sailed through code review because it looks like the most natural thing in the world, and it looks that way because it is.

I hunt this in financial services APIs, and the scary part isn't how rare it is, it's the opposite: it's in almost every endpoint that takes a client-supplied identifier, and it almost never shows up as the naive line from the textbook example.

This is the article I wish I'd read when I started testing APIs: why authentication isn't authorization, how to make forgetting fail closed instead of leaking, exactly where the check disappears, and how you prove the hole exists without relying on luck.

A note on scope: this is security engineering, not an attack recipe.

Every test described here assumes authorized scope, a pentest contract or a disclosure program's policy, and testing someone else's object without authorization is a crime, not research.

Image description

Fig. 1 - The front door gets audited, tested, and monitored; the second question usually doesn't exist in the code.


Authentication answers "who are you," authorization answers "are you allowed"

These are different questions, and the entire industry invests in the first one.

Authentication is the front door: token, password, MFA, session, and when it works, the server knows the request came from user 42.

Great, and that's where most systems stop asking.

Object-level authorization is the second question, asked on every single access: can user 42 see this specific resource? Not "can see statements," can see this statement, number 9931, which belongs to account 7, which belongs to user 13.

BOLA is what happens when the code answers the first question and assumes the second: the token is valid, so the server hands the data over, and nobody checked who owns the object.

It's worth pinning down the vocabulary, because these three neighboring flaws get confused constantly, and OWASP separated them on purpose:

AcronymWhat it breaksExample
BOLA (API1:2023)access to someone else's objectGET /statements/9930
BOPLA (API3:2023)access to a field that isn't yoursPATCH with "role": "admin"
BFLA (API5:2023)access to a function of another roleDELETE /admin/users/9

BOLA is horizontal movement: same role, someone else's object, BFLA is vertical: a role above your own, BOPLA sits in the middle, at the attribute level, and all three grow from the same root, and that root is what matters.

The most robust login in the world does not close this hole, because the hole is on the other side of the door.


Real problem #1: the id in the URL that nobody checks

I'll start with the textbook case, not because it's new, but because it's the canonical form, and every disguise in section 3 is a deformation of it, and if you already know this part, what matters is the end of the section, about the direction in which human error fails.

The textbook case is a line that looks harmless:

# app/controllers/statements_controller.rb
class StatementsController < ApplicationController
  before_action :authenticate_user!

  def show
    statement = Statement.find(params[:id])
    render json: statement
  end
end

The user is authenticated, the before_action ran, the token is valid, everything's green, and the controller fetches the statement by the id in the URL, without tying it to the logged-in user.

Legitimate request:

GET /statements/9931 HTTP/1.1
Authorization: Bearer <user 42's token>

The attack is swapping a number:

GET /statements/9930 HTTP/1.1
Authorization: Bearer <user 42's token>

Same token, same authenticated user, someone else's statement in the response, the server never asked who owns 9930, this is BOLA in its purest form, and in a financial system, the leaked object is a balance, a CPF (the Brazilian taxpayer ID), a transaction history, a Pix key, a home address.

The fix isn't validating the token harder, it's scoping the query by identity:

def show
  statement = current_user.statements.find(params[:id])
  render json: statement
end

The difference is the entire security of the endpoint, current_user.statements.find(9930) raises ActiveRecord::RecordNotFound because statement 9930 isn't in user 42's collection, and Rails translates that into a 404.

Authorization stopped being a check someone has to remember to write and became the very shape of the query.

Notice what changed in nature: the insecure version needs someone to add a verification line, and the secure version needs someone to remove the scope to break.

You inverted the sign of the human error: forgetting now results in denial, not leakage.

That inversion is the architectural point of this entire article, every time a protection depends on the developer remembering, it will fail on some one of the 300 endpoints, and every time it is the only available way to write the query, it only fails when someone deliberately steps off the path, and stepping off the path on purpose shows up in the diff.

The rule: never look up a raw client-supplied id, always look up starting from the owner.


Real problem #2: swapping 2 for 3 is too easy, so the industry hides the number

The most common reaction when a team discovers BOLA is wrong in one specific way: swap the sequential id for a UUID.

GET /statements/9930
GET /statements/e2b1c4a0-7f3d-4a11-9c2e-8f6b0d1a5e77

The reasoning is that nobody guesses a UUID, so nobody iterates, and it is true that it makes blind enumeration harder: out of 2^128 possibilities, brute force is off the table, but a UUID is obfuscation, not authorization, and the object stays accessible to anyone who has the identifier, and identifiers leak all the time:

  • In the body of another response from the same API, usually a listing that returns more fields than the screen uses.
  • In the Referer header, when the page with the id in the URL loads any third-party resource.
  • In application logs, load balancer logs, APM, corporate proxy history, Sentry.
  • In an email notification, a webhook to a partner, a URL shared over WhatsApp.
  • In an employee's mobile app, a screenshot pasted into a support ticket, an exported CSV.

The day that UUID shows up in any of these places, the endpoint goes back to leaking, because the ownership check never existed, you just made the bug harder to find, including for whoever is defending against it.

And there's a detail that almost always slips by: not every UUID is unpredictable, UUID v1 carries a timestamp and a MAC address, UUID v7, which became trendy for being sortable in a Postgres index, carries the creation time in milliseconds in its first 48 bits,1 and if your object gets created in response to an action the attacker triggers, they know the timestamp to millisecond precision, and the search space collapses.

Sortable in the index and unpredictable to the adversary are requirements that contradict each other, and most teams have only noticed the first one.

There's a variant of this same mistake that the framework actively encourages: the signed identifier, Rails's signed_id, GlobalID, the Active Storage blob URL, the "share" link that generates a token, and here the argument looks stronger than the UUID one, and it is: it isn't a value you guess, it's a value with an HMAC, genuinely unpredictable.

Unpredictable and still not authorization, a signed identifier is a capability URL, whoever has the value has the access, and the question "who's on the other end" never gets asked, it doesn't distinguish the owner from an intruder who received the link, so the whole leak list above applies just the same, with one problem of its own: these tokens usually carry a long expiration or none at all, and revoking one in practice means rotating the key and invalidating all of them at once.

A capability URL is a legitimate tool for the case where the capability is the model: an invitation, a public receipt, a short-lived temporary download, and it turns into a bug the day someone uses it as a shortcut to avoid writing the ownership check on a resource that has an owner.

UUIDs and signed tokens are good ideas for other reasons; as access control, they are worth zero.


Real problem #3: the disguises that get past code review

BOLA survives review because it almost never shows up as the naive example line, it shows up disguised, and each disguise has a different reason for fooling the reviewer.

In the body, not the URL. The id migrates into the JSON and the reviewer relaxes, because "the body is business data":

POST /transfers
Content-Type: application/json

{ "from_account": 7, "to_account": 20, "amount_cents": 5000 }

If the server trusts the from_account that came from the client instead of deriving it from the session, the user debits someone else's account, and an id in the body is exactly as dangerous as an id in the URL, and it's worse to audit, because it doesn't show up in the load balancer's access log.

In the nested object. The parent endpoint checks ownership and the child doesn't:

GET /accounts/7/transactions/8817

The code validates that account 7 is yours, then does Transaction.find(8817) without confirming that transaction 8817 belongs to account 7, the parent check gave a false sense of security, and it passes review twice over, because the reviewer sees an authorization happening in the first line and stops reading, and the safe form chains ownership all the way down to the leaf:

account = current_user.accounts.find(params[:account_id])
transaction = account.transactions.find(params[:id])

In the field that decides role. BOLA turns into BOPLA when the object is yours but the attribute shouldn't be:

PATCH /users/42
{ "name": "Maria", "role": "admin" }

Authenticated, authorized to edit your own profile, and role was on the list of accepted fields, mass assignment and BOLA are the same problem seen from different angles: the client wrote to an attribute it shouldn't control, and in Rails, permit is an allowlist, and that's exactly why it works, as long as nobody wrote params.require(:user).permit! on a Friday.

In the id inside the nested payload. This is the most profitable one in Rails, and it almost never shows up on a BOLA checklist, because the dangerous field isn't the resource's id, it's an id buried two levels down:

PATCH /users/42
{ "user": { "name": "Maria",
            "account_attributes": { "id": 99, "nickname": "main account" } } }

With accepts_nested_attributes_for :account, Rails interprets an id present in the nested attributes as "update this existing record," not "create a new one," and if record 99 doesn't belong to user 42, you just wrote to someone else's object through an endpoint that edits your own profile.

And permit doesn't save you here: account_attributes: [:id, :nickname] is on the allowlist because it needs to be, otherwise editing an existing record stops working, the allowlist is correct and the bug happens anyway, which is exactly the kind of thing that passes review, and what actually holds the line is resolving the record from the owner on the server, or a reject_if that confirms ownership before letting the id in.

In the bulk endpoint. This is the favorite one, because the check exists and still fails:

POST /statements/bulk
{ "ids": [9931, 9930, 9929] }

The code validates the first id, or validates inside an if that only covers the happy path, and where(id: ids) returns everything, and a bulk endpoint needs to compare quantity requested against quantity authorized and fail when they differ, not filter silently.

In GraphQL. The Node interface with node(id: ...) is a BOLA surface by construction: a generic resolver that fetches any object in the graph by global id, and if authorization lives in the field resolvers and not in the node loader, the attacker reaches the object through the side path, and the same holds for every traversed relation: me { account { transactions { ... } } } is authorized, transaction(id: 8817) { account { owner { document } } } might not be.

In reports and exports. GET /reports?account_id=7&format=csv is the same bug dressed up as BI, and report endpoints are usually written outside the resource-controller pattern, with hand-assembled SQL, and that's where the owner scope gets lost first.

In the object that doesn't look like an object. A file in a bucket with a predictable URL, a receipt PDF served by GET /files?path=..., an avatar, a ticket attachment, and if the download doesn't go through the same authorization layer as the resource it represents, you have BOLA sitting on top of an accidentally public S3 bucket.

In state, not ownership. The object is yours, but you can no longer act on it: cancel a transfer that already settled, reopen a closed ticket, edit a signed contract, and this is object-level authorization too, except the rule is the state machine, and no scanner sees this.

The pattern behind all of them: at some point the code trusted a value the client chose, to decide what the client can access.


Real problem #4: the check exists, and it's in the wrong place

A mature team doesn't forget authorization, a mature team puts authorization somewhere that doesn't cover every path.

The four most common wrong places:

In the frontend. The button disappears for anyone who isn't the owner, and the API stays wide open, this isn't access control, it's interface design, and the client is enemy territory: it reads the bundle, finds the route, and calls it directly.

In per-route middleware. Someone writes a rule that matches /accounts/* and thinks it covers everything, then /v2/accounts is born, or /internal/accounts, or a POST /graphql that reaches the same table without going through the protected route, and authorization tied to a URL pattern becomes debt the moment the API gets versioned.

In the controller, once per action. Better than the previous two, and still not enough: it covers the path the author remembered, the async job that reprocesses the same resource doesn't go through the controller, the console command doesn't, and the internal endpoint the data team built for their dashboard doesn't.

In the controller, but after the cache. The cruelest of the four, because here authorization exists, is correct, and runs, just too late to matter, a CDN, Rack::Cache, a fragment cache, a careless Cache-Control: public on an authenticated endpoint: any one of these makes user 42's response get served to user 13 without a single line of your code running, this isn't BOLA in the controller, it's BOLA in the layer sitting in front of it, and it never shows up in any test, because tests don't have a CDN, and the rule is short: a response that depends on identity is private, and if it's cached, identity is part of the cache key.

The right place is as close to the data as possible, in increasing order of guarantee:

  1. Scope in the query - current_user.statements.find(...), cheap, idiomatic, covers the application's own path.
  2. An explicit policy layer - Pundit, CanCan, a homegrown Authz module, makes the rule readable, testable, and reusable outside the controller.
  3. Row Level Security in the database - the last line, the one that holds even when the developer gets it wrong.

The policy layer solves the "every endpoint reimplements the rule" problem, and in Pundit, what matters isn't the single-object authorize, it's the collection's policy_scope:

class StatementPolicy < ApplicationPolicy
  class Scope < ApplicationPolicy::Scope
    def resolve
      scope.joins(:account).where(accounts: { user_id: user.id })
    end
  end

  def show?
    record.account.user_id == user.id
  end
end

And the hook that turns this into a guarantee, in ApplicationController:

after_action :verify_authorized, except: :index
after_action :verify_policy_scoped, only: :index

These two lines make the application fail in tests when someone writes a new action without authorizing it, this isn't documentation, it's a CI error, default-deny applied to the process, not just to the request.

And the bottom of the well, which is where I wish more teams would end up: RLS in Postgres.

ALTER TABLE statements ENABLE ROW LEVEL SECURITY;
ALTER TABLE statements FORCE  ROW LEVEL SECURITY;

CREATE POLICY statements_por_tenant ON statements
  USING (tenant_id = NULLIF(current_setting('app.tenant_id', true), '')::uuid);

Four details in this DDL that aren't just style:

  • FORCE ROW LEVEL SECURITY. Without it, the table owner bypasses the policy, and your Rails application almost certainly connects as the table owner, because it's the one that ran the migrations, and ENABLE on its own is a protection that doesn't protect against exactly who you need protecting from.
  • current_setting('app.tenant_id', true). The second argument makes the function return NULL instead of raising when the variable was never set in that session, tenant_id = NULL evaluates to NULL, which the policy treats as false: a connection with no tenant set sees no rows at all, it fails closed.
  • NULLIF(..., ''). This is the detail that only shows up in production, the true covers "the variable never existed in this session," which is the state of a freshly opened connection, but a custom GUC, once set, starts existing in the session with a reset value of empty string, so after the first SET LOCAL app.tenant_id on that connection, the commit doesn't reset the variable back to NULL, it resets it to '', from the second request onward, on that pooled connection, any query that runs before the next SET LOCAL gets '', and ''::uuid isn't false, it's invalid input syntax for type uuid: a 500 instead of a 404, intermittent, and proportional to how long connections live, and NULLIF brings that case back to NULL, which the policy already knows how to treat as denial.
  • SET LOCAL, not SET. The variable needs to be set per transaction, inside the transaction, so it doesn't leak between requests on the connection pool, SET LOCAL app.tenant_id = ... at the start of every request, and the pooled connection comes back clean on commit.

And now the honest part, which is where I see the most teams fool themselves: this policy does not solve the problem this article opened with.

It compares tenant_id, and if user 42 and user 13 are in the same tenant (two analysts at the same client company, two people on the same product, depending on how you modeled it), statement 9930 stays perfectly visible to both of them, the "last line of defense" promise is being kept against the cross-tenant leak from section 5, not against the horizontal access from section 1, which is the case that opened this article, and it's worth knowing which of the two you actually bought.

If you want both, the policy needs to go down to the ownership level, and at that point it stops being a column comparison and becomes a chain traversal:

CREATE POLICY statements_owned_by_user ON statements
  USING (
    tenant_id = NULLIF(current_setting('app.tenant_id', true), '')::uuid
    AND EXISTS (
      SELECT 1 FROM accounts a
       WHERE a.id      = statements.account_id
         AND a.user_id = NULLIF(current_setting('app.user_id', true), '')::uuid
    )
  );

This has a price, and the price is real: the EXISTS enters every execution plan against the table, and without an index on accounts (id, user_id) it hurts, and that's why the right call isn't "turn on ownership-level RLS everywhere," it's picking the two or three tables where a leak turns into a breach notification to Brazil's data protection authority (the ANPD), and paying the cost only there, threat model, not taste.

In both cases, RLS doesn't replace scoping the query; replacing it would mean paying that plan everywhere and losing readability, and it exists for the day someone wrote Statement.where(id: params[:ids]) in a report and nobody caught it, and its value is being exactly the one layer that doesn't depend on anyone having remembered.


Real problem #5: the object is yours, the tenant isn't

This is the BOLA that hurts most in fintech and in any B2B2C platform, and it doesn't show up in the textbook example because the textbook example only has one level of ownership.

In a multi-tenant architecture, ownership is a chain: a user belongs to a company, a company to a tenant, an account to a company, a transaction to an account.

Scoping by current_user covers the leaf and leaves the middle of the chain wide open, the classic case:

# looks safe, and it is
transaction = current_user.transactions.find(params[:id])

# and in the endpoint next door, written by someone else, six months later
company = Company.find(params[:company_id])
render json: company.transactions

The second endpoint checks that you're authenticated and never checks that company params[:company_id] belongs to your tenant, the attacker is a legitimate platform customer reading a competitor's operations on the same SaaS, and from the access log's point of view, it's perfectly normal traffic from a paying customer.

Three things that make a difference here:

  • The tenant never comes from the client. Not in a header, not in a query string, not in an editable claim, it's derived from the session on the server, and an X-Tenant-Id header the client can write is a BOLA wearing a feature's name.
  • The entire chain gets verified, not just the tip. If a resource has three levels of ownership, there are three chances to get it wrong, and the test needs to cover each hop in isolation.
  • A global identifier is the worst of both worlds. A sequential id shared across tenants turns any BOLA into a cross-tenant leak, if each tenant has its own numbering space, the same bug only leaks within the house, and that's not a defense, it's blast radius reduction, and blast radius is exactly what you negotiate when the incident happens.

A good question to bring to your next architecture review: how many endpoints in the system resolve the tenant from something the client sent? If you don't know the answer, it's greater than zero.


Real problem #6: 403 tells a story 404 doesn't

There's a debate that looks like bikeshedding and isn't: when the object isn't yours, what do you respond with?

ResponseWhat it revealsWhen to use it
404nothinghorizontal access, someone else's object
403the object existslack of permission within your own scope
401invalid credentialmissing or expired token

A 403 says "this exists and it's not yours," which already gives away the object's existence, and for an attacker profiling a target, that's half the work: with 403 versus 404 they can enumerate which ids exist, measure the size of the database, work out the daily growth rate, and know whether a specific person's account is on the platform.

404 hides even that: for anyone who isn't the owner, the resource simply doesn't exist, it's the same semantics as current_user.statements.find raising RecordNotFound, which is convenient, the idiomatic implementation already produces the right response.

403 still has a legitimate place: when the user knows the object exists because it's within their own scope, and the denial is about permission, not ownership, and an analyst who can't approve a transfer they can see right on their screen needs a 403, otherwise the interface is lying to them.

And the side channel isn't just the status code.

It's worth checking whether the denial response differs anywhere else:

  • Timing. A nonexistent object responds in 3 ms and someone else's object in 40 ms, because the second one got loaded before being denied, and the difference is an existence oracle.
  • Error body. {"error": "Statement not found"} versus {"error": "Not found"}, or a distinct error_code, or a trace_id that only shows up in one of the two cases.
  • Headers. Location, ETag, an X-Request-Id in a different format, a Content-Length that varies.
  • Counting and pagination. A listing's total that includes objects the listing doesn't return leaks the number, and sometimes the number is the data.
  • Rate limiting. Hitting a nonexistent object doesn't consume quota and hitting a real one does, yes, I have seen this.

The denial has to be indistinguishable, if it varies with the fact you're trying to hide, it isn't hiding it.


Real problem #7: a finding without proof is an opinion

A BOLA report needs to demonstrate the improper access, and the clean way to do that is the two-user test.

You need two test accounts in the application, A and B, both authorized under the engagement's scope, preferably with the same role, otherwise you're testing BFLA without realizing it, the procedure:

  1. Logged in as A, create a resource and note its identifier, say statement 9931.
  2. Logged in as A, request GET /statements/9931, it should work, A is the owner, and save the entire response, headers and timing included.
  3. Change only the token: send the same GET /statements/9931 with B's token.
  4. Compare byte for byte.

The decision table is short:

Response to B's tokenReading
200 with A's dataBOLA confirmed
200 with an empty or generic bodycheck field by field
404scoped by owner, probably safe
403explicit denial, leaks existence
401authentication problem, different test

The detail that separates a serious test from a guess: change only the token, keep everything else identical. Same URL, same method, same body, same headers, same field order in the JSON, and if the only variable is identity and the response changes from "leaked" to "denied," you've isolated the authorization flaw, and if you change two things at once, your report turns into a debate at the triage meeting.

Three refinements that raise the finding rate significantly:

Test with no token too. The third column of the test is the anonymous request, and it's uncomfortable how often it comes back 200 on an endpoint that "obviously" requires authentication, usually because the new route fell outside the before_action, or because there's an /internal path the gateway shouldn't be exposing.

Test both directions. A leaking read is bad, a leaking write is an incident, repeat the procedure with PATCH, DELETE, and with the id in the body, and an endpoint that denies the GET and accepts the PATCH exists more often than it should, because whoever wrote the rule thought about "seeing the data" and not "changing the data."

Automate the matrix, not the individual cases. What scales is generating the cartesian product of (endpoint × role × object owner) from the OpenAPI spec and checking the expected response for each cell, and Burp Autorize does this interactively while you browse; for regression, the right place is the team's own test suite:

# spec/requests/authorization_spec.rb
RSpec.describe "object-level authorization" do
  let(:owner)     { create(:user) }
  let(:intruder)  { create(:user) }
  let(:statement) { create(:statement, account: create(:account, user: owner)) }

  it "does not return another user's statement" do
    get "/statements/#{statement.id}", headers: auth_headers(intruder)

    expect(response).to have_http_status(:not_found)
    expect(response.body).not_to include(statement.account.number)
  end
end

This test costs ten minutes to write and stands guard forever, and one per resource family already changes a system's baseline.


Real problem #8: what can be automated, and what can't

Back to the opening thesis, now with what you can actually do about it.

No tool is going to hand you the list of your BOLAs, for the reason already stated: the ownership rule belongs to your domain, and a scanner has no way to know it, but "no tool solves this" isn't the same as "don't automate anything," and the distance between those two sentences is what separates a team that reduces its surface from a team that just worries.

What tooling can do, and it's worth turning on:

  • Static analysis of a local pattern. A lint rule that flags Model.find(params[:id]) outside a short list of approved exceptions, crude, noisy, and still worthwhile.
  • Policy coverage in CI. The verify_authorized checks from section 4, turning a missing authorization into a red build.
  • Surface diffing. Comparing the OpenAPI spec between releases and failing when a new endpoint shows up without a matching authorization test, and an endpoint with no inventory entry is an endpoint with no owner, and that's exactly where the flaw lives.

Notice what all three have in common: none of them looks for the vulnerability, all three look for the absence of the control. A query outside the pattern, an action with no policy, an endpoint with no test, you can't automate "this object belongs to someone else," which requires the domain, but you can automate "nobody here ever asked whose it is," which requires nothing beyond discipline of form, and that's exactly why this is the only automation that works.

The rest is architectural habit: scope every query by session identity, treat a client-supplied id as untrusted input even when it looks internal, and review every endpoint with a single question: if I swap this id for someone else's, what happens?

And this isn't theory, the list of major leaks caused by this flaw is embarrassingly mundane:

CaseYearWhat it was
First American Financial2019sequential document id, no authentication
USPS Informed Visibility2018API returned data for any account
Peloton2021any user's profile by id
Parler2021post by sequential id, API with no auth
Optus2022exposed endpoint with an incremental identifier

None of them needed an exploit, they needed a for loop and patience.


Bonus: three things called "authorization" in the same codebase

I'll close with a problem that isn't a flaw, it's an ambiguity, and because of that it doesn't fit any taxonomy, even though I've seen it cost half a day of incident response.

In a good chunk of the codebases I read, the word "authorization" is already taken, there's an AuthorizationService that validates the token, an authorize! method that only checks whether the session is active, an Authorization middleware that is, in fact, authentication, and in fintech there's a third meaning on top, the domain one: "authorizing a transaction" in the acquirer's sense, approving the card purchase.

Three concepts, one name, and then someone reads if authorized? in a controller, understands "ownership was checked," and moves on.

Separate the vocabulary while it only costs a rename: authenticate for identity, authorize for permission over an object, and the domain term (capture, approve, settle) for the transaction, and naming isn't pedantry when the wrong name makes the reviewer skip a check they think they've already seen.


Half an hour, one grep, and a number

No checklist, just an exercise that produces a number, because a number moves a meeting and a rhetorical question doesn't.

Run some variation of this against your codebase:

grep -rEn '\b[A-Z][A-Za-z0-9_]*\.(find|find_by|where)\(' app/ \
  | grep -E 'params\[' \
  | grep -vE 'current_user|current_account|policy_scope|authorize' \
  | tee /tmp/candidatos-bola.txt \
  | wc -l

Adjust it for your framework and your session variable names, what matters is the shape: a model queried directly, with a value coming from the client, without going through an identity scope.

A candidate isn't a bug, half of them will be legitimately public resources: a lookup table, a postal code search, a catalog, and the other half is the work.

What the number tells you:

LinesReading
zeroeither you scope everything, or the grep is wrong. Bet on the second and check it against an action you know is insecure
fewer than tenyou can review them one by one this week, and leave a regression test per resource family
dozensdon't review, change the pattern. verify_authorized in CI first, and the list becomes a backlog ordered by data sensitivity
hundredsthe problem isn't the list, it's that there's no secure path by default. RLS on the tables that would hurt, before any controller refactor

And most important: the grep only finds the canonical form. The five it will never catch, the ones you have to hunt by hand, are exactly the ones this article has spent the whole time describing:

  • the id in the POST body, because there's no visible params[:id]
  • the id inside nested _attributes, because permit is correct and the bug happens anyway
  • the leaf of a nested resource, because the parent got checked and the line looks authorized
  • the bulk endpoint, because the check exists and only covers one item
  • the state machine, because the object is yours and that action still isn't

If the grep came back with almost nothing and you felt relieved, these five are the reason not to.

Authentication is the door, object-level authorization is the lock on every drawer, most systems lock the door with great care and leave the drawers open.


References

Footnotes

  1. RFC 9562 (2024), which replaced RFC 4122 and standardized v6, v7, and v8, reserves 48 bits for a Unix timestamp in milliseconds in version 7 and leaves 74 bits of randomness, still a lot, but the prefix stops being a secret, and a known prefix is what turns scanning into something viable whenever a timing oracle exists. ↩

Loading comments...