The Password Database Was the Only Thing We Stored: PDF7's Case for Zero-Storage Authentication
PDF7 runs 25+ PDF tools — convert, compress, merge, split, edit — and none of them upload your file. Processing happens client-side, in your browser, in JavaScript. A document you drop into PDF7 does not transit our servers, and we do not have a copy of it after you close the tab. That is the entire product promise, and it is the reason legal teams, healthcare workers, financial advisors, and government users across 100 countries picked us over the alternatives.
We also, until June 2024, maintained a table of bcrypt-hashed passwords.
Both of those sentences were true at the same time, and the second one quietly cancelled the first. This post is the argument that ended it: why a hashed password database is still a stored credential, why hashing was not a defence we could rely on, what the week of deleting it actually involved, and what the numbers looked like between June and November 2024 while our user count went from 10,000 to 1,000,000.
Key Takeaways
- A hashed password database is still a database of credentials. PDF7 promised users we store nothing while running the one table an attacker most wants. bcrypt raises the cost of cracking a stolen hash; it does not change the fact that the hash was there to steal.
- Attack surface is binary, not proportional. Rate limiting, bot detection, and lockout policy all reduce how often the credential endpoint is hit. Only removing the endpoint changes whether it can be hit at all — which is why our credential stuffing figure went to zero and stayed at zero through 100x growth rather than degrading with volume.
- Deleting the password database took one week, and the deletion itself was the last step. Days 1–2 were threat modelling and configuration, days 3–4 integration and penetration testing, days 5–6 a feature-flagged rollout to 100% of traffic, day 7 validation and the secure destruction of the old credential store.
- Security operations spend fell 60% — measured on a different base than the 40% that was tagged "authentication." Those two numbers do not share a denominator and we are not going to draw them on one axis. The section below explains exactly what each one counts.
- Zero-storage did not fix everything. It moved our recovery root to email, made us dependent on identity providers we do not run, and left phishing and session theft untouched. Passkeys are how we are closing the largest remaining gap.
The One Thing We Stored
PDF7's architecture is a data minimization argument taken to its end point: if a document never reaches a server, no server can leak it, subpoena it, index it, or lose it in a backup. That is not a policy we enforce. It is a shape the system has.
A password database is the same argument running in reverse. It is a permanent, centralised, high-value collection of user secrets, sitting on infrastructure we operate, growing every time someone signs up. Every property that makes client-side document processing safe — nothing centralised, nothing retained, nothing worth attacking — was inverted in that one table.
The gap was not hypothetical, and it was not only philosophical. Users came to PDF7 specifically because we did not hold their documents. When an attacker took over one of those accounts, they landed inside a product that people had trusted with court filings and medical records. The account itself did not contain the documents. The trust violation did not care.
What finally made this a decision rather than a discomfort was writing the marketing copy. We could not write "we store nothing" without a footnote, and the footnote was the whole problem. A privacy claim that requires an asterisk is a privacy claim an attacker has already read.
What a Credential Stuffing Campaign Looks Like From the Defender's Side
Credential stuffing is the automated replay of username-and-password pairs stolen from other companies' breaches against your login endpoint, betting that some fraction of your users reused the pair. It is not an exploit of any bug in your system. It is an exploit of the fact that your system accepts passwords.
PDF7 absorbed thousands of these attempts a month before June 2024, and the volume rose with our visibility. Nothing about the traffic was sophisticated: distributed bot networks working through breach compilations, spread across enough source addresses to make IP reputation a blunt tool. Published industry estimates put the success rate of a stuffing campaign somewhere in the 0.5% to 2% range, which sounds small until you multiply it by an attacker's willingness to run billions of attempts across every site at once. On our side, it produced multiple successful account takeovers a month.
The structural problem is visible in the path the attack takes.
That distinction is the whole post in one image. Every control we had been buying and tuning worked on the volume of the attack. None of them worked on its viability.
Why Hashing Was Not an Answer
The standard objection to everything above is that PDF7 was not storing passwords, it was storing bcrypt hashes, and bcrypt is a deliberately slow one-way function designed to make offline cracking expensive. That objection is technically correct and strategically irrelevant, for three reasons.
Hashing sets a price, not a limit. An attacker who steals a bcrypt table cannot reverse it, but they can guess against it offline, with unlimited attempts, no rate limiting, no lockout, and no detection, for as long as they care to rent hardware. The work factor determines how much that costs. It does not determine whether it succeeds. Every weak or reused password in the table falls; the strong ones survive until the economics change.
Storage creates a target regardless of contents. A credential database is a honeypot in the literal sense: a concentrated store of things worth stealing, whose existence is public knowledge because every login form advertises it. Attackers do not need to know your hashing scheme to decide the table is worth going after. The decision to attack is made on the assumption that a table exists, and a login form is proof.
Hashing is a commitment to keep being right. bcrypt was appropriate when we chose it and the work factor was appropriate when we set it. Both are perishable. Owning a password database means owning an ongoing obligation to track algorithm deprecation, raise work factors as hardware improves, run migrations across the table, and enforce password policy — a maintenance stream with no end state and no upside, only the possibility of falling behind.
The clarifying question was not "is our hashing good enough." It was "what is the best outcome if this table leaks?" The best outcome for a well-hashed table is that most users are fine and a subset are not, plus a disclosure, plus a forced reset for a million people. The best outcome for a table that does not exist is nothing at all.
Zero-Storage, Defined Precisely Enough to Argue With
"Zero-storage authentication" is a term worth pinning down, because it is easy to say and easy to mean much less than it sounds like.
Zero-storage authentication means the authentication system holds no persistent user secret and no persistent personally identifiable information on its servers. There is no password field, no password hash, and no stored identifier that would be valuable to an attacker who compromised the authentication infrastructure. Authentication is established through ephemeral cryptographic tokens and verification events rather than by comparing a submitted secret against a stored one.
That definition rules out several things that are often described the same way:
- Passwords that are optional are not zero-storage. If the endpoint still accepts a password for any user, the table still exists and stage 3 of the attack path above is still live for everyone in it.
- Adding one social login is not zero-storage. It reduces how many users touch the password path without removing the path. PDF7 considered this and rejected it for exactly that reason: it would have improved a metric without changing an architecture.
- Encrypting the password column is not zero-storage. Encryption at rest moves the secret from the table to the key management system. It does not remove it.
The methods that satisfy the definition, and the three PDF7 runs today:
- One-tap login — the returning user authenticates with a single tap using a credential already held by their device or platform, with no secret typed and none stored by us.
- Google login — an OAuth 2.0 authorization flow that delegates identity verification to Google, so the credential is checked by an infrastructure we do not operate and never see.
- Email OTP — a time-limited one-time passcode delivered to the user's inbox, valid once, stored nowhere after it is consumed. This is the universal fallback for users with no social account, which matters across 100 countries more than it does in any single market.
The property that made this the right architecture for PDF7 specifically is not that it is more secure in the abstract. It is that it is the same argument as the product. Client-side processing and zero-storage authentication are one idea applied twice: do the work where the data already is, and keep nothing afterwards. A user does not have to trust either claim more than the other, because both are enforced by the shape of the system rather than by our conduct.
The Week We Deleted the Database
PDF7's engineering and security teams replaced the in-house password system with MojoAuth's MojoShield zero-storage authentication in one week in June 2024, with no downtime. The deletion was the last thing that happened, not the first.
Days 1–2 — threat model and configuration. Before any code, the security team mapped the threat model and attack surface explicitly: what an attacker gains at each point of compromise, before and after. That document is what made the rest of the week fast, because every subsequent decision had a stated criterion. In parallel: MojoShield configured, Google OAuth provider set up, email OTP configured with SPF and DKIM alignment so the passcodes would actually arrive, security monitoring dashboards built, and compliance documentation started.
Days 3–4 — integration and attacking it ourselves. Engineers replaced the password forms with passwordless interfaces via the JavaScript SDK, moved sessions onto validated JWTs, and set security headers — Content-Security-Policy, HSTS, X-Frame-Options. Then the security team spent the rest of the block trying to break it: session hijacking, CSRF, OAuth authorization-code interception, and deliberate attempts to evade the new bot detection. Penetration testing the new path before it took traffic was not optional for us, because the entire justification for the project was a security argument. Shipping it untested would have made the argument circular.
Days 5–6 — feature-flagged rollout. Deployment went out behind feature flags. New users got passwordless immediately; existing users received a communication campaign explaining what was changing and why, before the login screen changed under them. An unexplained change to a login page reads as phishing to exactly the security-aware users PDF7 attracts. Account linking flows moved existing accounts across without a forced reset. By the end of day 6, 100% of authentication traffic ran through the new system with zero security incidents, and the old password infrastructure was decommissioned.
Day 7 — validation, then deletion. Post-deployment validation, edge-case testing, log review, and the compliance rewrite: GDPR and CCPA documentation updated to describe an architecture that retains no credentials, which is a considerably shorter document than the one it replaced. Then the old password database was securely deleted.
That last step is worth separating out. Every intermediate state of this project still had the liability in it. A password table that no longer serves logins is still a password table, and a decommissioned system that still holds data is a breach waiting to be attributed to something nobody owns any more. The project was not finished when passwordless went live on day 6. It was finished when the rows were gone.
The Numbers, June to November 2024
PDF7 went live in June 2024. The table below compares the pre-migration baseline against November 2024, five months later. Active users grew 100x across that window, so the ratios carry the meaning, not the totals.
| Metric | Before (pre-June 2024) | After (Nov 2024) | Change |
|---|---|---|---|
| Active users | 10,000 | 1,000,000 | 100x |
| Credential stuffing attempts | Thousands/month | 0 | Attack vector removed |
| Account takeover incidents | Multiple/month | 0 | −100% |
| Security incidents, post-launch | — | 0 | Perfect record |
| Stored passwords | Hashed database | 0 | Liability removed |
| Password reset support load | Significant | 0 | Eliminated |
| Security operations budget | Baseline | −60% | See note below |
| Authentication uptime | Self-managed | 99.99% SLA | Enterprise-grade |
| Countries served | 100 | 100 | Unchanged |
Figures come from PDF7's own systems: authentication logs, edge and bot-detection logs, incident records, and the security operations budget line. The baseline column reflects the state of the in-house password system immediately before the June 2024 cutover; the after column was taken in November 2024. Nothing in this table is modelled. Where a pre-migration figure was tracked as a range or a category rather than a count, it appears here as a range or a category rather than being converted into a false precision.
Three of those rows need a reading note.
The zeros are structural, not a good quarter. Credential stuffing attempts are zero because a stuffing attempt aimed at an endpoint that accepts no password never becomes an attack in the first place; it is a request with nowhere to land. Password reset load is zero because there is no reset flow. Neither number is the output of a control that could be tuned badly, degrade under load, or be misconfigured during a deploy. That is the difference between a metric you defend and a metric you removed the possibility of.
"Multiple monthly" is what we can defend. Our narrative records also contain a broader estimate of compromised accounts derived from applying published 0.5%–2% stuffing success rates to our attempt volume. We are not publishing that as a measured figure, because it is an industry rate applied to our traffic rather than a count from our incident log. The incident log says multiple confirmed account takeovers per month. That is the number in the table.
Zero security incidents covers June to November 2024. It is a five-month record across a 100x increase in users, not a claim about the future. The reason we think it is durable is in the next two sections: the thing that would have needed to scale is gone.
The Budget Line That Fell 60%, and Why It Is Not the 40%
PDF7's security operations budget fell 60% after the migration. A separate figure, from before the migration, is that authentication-related security consumed 40%+ of that budget. Those two numbers get quoted next to each other, and they do not reconcile — because they are measured against different denominators.
The 40% is a composition figure: of the pre-migration security operations budget, more than two-fifths was explicitly tagged to authentication defence — bot detection tooling, rate limiting infrastructure, credential stuffing mitigation, password reset flows, and the monitoring around all of it.
The 60% is a change figure: the measured reduction in the total security operations line between the pre-migration baseline and November 2024. It is larger than the composition figure because removing the password system also removed spend that had never been tagged "authentication" in the budget — incident response hours for compromised accounts, the security engineering time that went into hash policy and password enforcement, and the scaling headcount the old architecture would have required as we grew. Those were categorised elsewhere before, and they went away too.
The reason to spell this out rather than pick one figure is that authentication cost savings are routinely presented as a single percentage with no stated base, which makes them impossible to compare against your own numbers. If you are trying to work out what this would be worth to you, the composition figure is the one to start from: find how much of your security spend is defending a credential endpoint, and treat everything beyond that as upside you will discover after the fact.
Why 100x Growth Cost Us Nothing in Security Architecture
PDF7 went from 10,000 to 1,000,000 active users between June and November 2024 without a single change to its authentication security architecture. That is the outcome we care about most, and it follows directly from the deletion rather than from anything we did afterwards.
Under the old system, scaling was going to be expensive in a specific way. More users meant a bigger credential table, which meant a more valuable target, which meant more stuffing volume aimed at it, which meant more bot detection tuning, more monitoring, more incident response capacity, and eventually security engineers hired specifically to defend the login endpoint. Every one of those costs was tied to the existence of the table. Growth multiplied them.
What replaced it is a managed multi-tenant authentication infrastructure with a 99.99% uptime SLA, where bot detection, adaptive rate limiting, and fraud prevention are operated by a team whose only job is that. PDF7 is a lean company. The realistic alternative was never "build enterprise-grade authentication security"; it was "run under-resourced authentication security and hope." Offloading it was not primarily a cost decision. It was an acknowledgement of who was going to be awake when the campaign ran at 3am.
The generalisable point: the reason 100x was free is that we removed the component whose cost scaled with users, not that we found a cheaper way to run it. If you are budgeting a security programme against a growth plan, the question worth asking about each control is whether it is defending something or removing something. The two behave completely differently at 100x.
What Zero-Storage Does Not Fix
PDF7's security team is aware that a post like this reads as a solved-problem story, so here are the parts that are not solved, each with what we are doing about it.
Email became the recovery root. With no password, the email inbox is the recovery path for a user who loses access to their device and their social account. That concentrates risk in a channel we do not control and did not harden. Our mitigation is that email OTP is a fallback rather than a primary method, and the codes are short-lived and single-use — but an attacker with inbox access is still an attacker with account access. This is a real trade, not an eliminated risk.
We depend on identity providers we do not run. Delegating to Google's OAuth infrastructure means inheriting Google's availability and Google's account security decisions for that share of our users. That is a favourable trade on security and an unfavourable one on control. Running multiple methods is the hedge: no single provider outage takes authentication down for everyone.
Phishing and session theft are untouched. Nothing about removing passwords stops a convincing fake login page or a stolen session token. Our answer is JWT sessions with configurable expiry and instant revocation, plus security headers — and, more substantially, passkeys. FIDO2 passkeys are cryptographically bound to the origin, which makes them phishing-resistant in a way that no OTP method is. That is the largest remaining gap in this architecture and the main reason passkeys sit at the top of the roadmap rather than further down it.
Coverage is not universal yet. Serving 100 countries with three methods leaves real gaps. WhatsApp authentication addresses the many markets — Indian, Southeast Asian, Latin American, Middle Eastern — where WhatsApp is not an app people also have but the default channel they already live in; Sign in with Apple, with its private email relay, suits the privacy-conscious iOS users who are close to PDF7's core audience; LinkedIn covers the professional and B2B segment; and SAML 2.0 and OIDC enterprise SSO is what B2B customers with central identity management require before they can adopt us at all. Each is a configuration on the same zero-storage foundation rather than a new security architecture, which is the point.
Cryptography has a shelf life. The same reasoning that made us distrust a fixed bcrypt work factor applies to the signature algorithms underneath token-based authentication. Post-quantum signatures are on the roadmap for the same reason the password table came off it: a security property you have to keep re-earning is a liability with a longer fuse, not an absent one.
What We Can Say Now That We Could Not Say Before
PDF7's compliance and marketing positions both changed on day 7, and the change was the same change in both cases: a claim stopped needing a footnote.
On compliance, GDPR and CCPA both push toward data minimization — collect and retain the minimum necessary. An architecture that retains no credentials satisfies that requirement structurally rather than procedurally, which shortens every conversation about it. The privacy documentation no longer describes a password database, its retention period, its hashing scheme, its breach notification path, or its deletion procedure, because there is nothing to describe. Audits ask what you store; "nothing" is a fast answer to verify.
On positioning, "we store nothing" became literally accurate. Before, PDF7's central claim was true about documents and false about credentials, and a security-literate user could work that out from the login form. Now the two halves of the product make the same promise. We would rather have a claim that survives inspection by the most skeptical person in the room, because those are the people who bring us legal filings and medical records.
The internal effect was smaller and more useful than either: the security team stopped spending its week on authentication. Bot detection tuning, stuffing mitigation, compromised-account investigation, forced resets, user notifications, and IP blocking were the work. That capacity went to document processing security and platform hardening — the parts of PDF7 that are actually ours to get right.
Frequently Asked Questions
What does zero-storage authentication actually mean?
Zero-storage authentication means the authentication system persists no user secret and no personally identifiable information on its servers — no password, no password hash, and no stored identifier worth stealing. PDF7 uses ephemeral cryptographic tokens and verification events instead of comparing a submitted secret against a stored one. The test PDF7 applies is whether an attacker who fully compromised the authentication infrastructure would obtain anything reusable. Passwords made optional, a single social login added alongside a password table, or an encrypted password column all fail that test.
If PDF7 stores no passwords, how does it recognise a returning user?
PDF7 recognises returning users through one of three verification events rather than through a stored secret. One-tap login uses a credential already held by the user's device or platform. Google login runs an OAuth 2.0 authorization flow where Google performs the verification. Email OTP sends a single-use, time-limited passcode to the user's inbox and discards it once consumed. In each case PDF7 receives a verification result and issues a session token; it never holds the thing being verified.
Can accounts still be taken over when there are no passwords to steal?
Account takeover via credential stuffing stopped entirely at PDF7 — from multiple confirmed incidents a month to zero between June and November 2024 — because stolen password pairs have no endpoint to be replayed against. Other takeover routes remain open. An attacker with access to a user's email inbox can complete an email OTP flow, and phishing or session-token theft are unaffected by removing passwords. PDF7's mitigations are short-lived single-use codes, revocable JWT sessions, and FIDO2 passkeys, which are origin-bound and therefore phishing-resistant in a way OTP is not.
What happened to PDF7's old password database?
PDF7 securely deleted the old password database on day 7 of the migration, after 100% of authentication traffic had moved to the new system on day 6 and post-deployment validation had passed. The sequencing was deliberate: a decommissioned password table that still holds rows carries the same breach liability as a live one while having no owner watching it. The compliance documentation was rewritten in the same step to describe an architecture that retains no credentials.
Does zero-storage authentication help with GDPR and CCPA compliance?
Zero-storage authentication satisfies the data minimization principle in GDPR and CCPA structurally rather than procedurally, which is what made it simpler for PDF7 to document. There is no credential retention period to define, no hashing scheme to justify, no password-database breach notification path to maintain, and no deletion procedure to prove, because there is no stored credential. This is a simplification of the compliance narrative, not a substitute for a compliance programme — PDF7 still maintains audit logs and the rest of its obligations.
What did zero-storage authentication cost PDF7, and what did it save?
PDF7's implementation took one week of engineering and security team time in June 2024, with no downtime. Security operations spend fell 60% against the pre-migration baseline, measured on the total security operations line in November 2024; separately, authentication defence had accounted for 40%+ of that budget before the migration. Those two figures use different denominators and should not be divided into each other. The larger saving is the one that never appeared as a line item: scaling from 10,000 to 1,000,000 users required no additional authentication security investment, because the component whose cost scaled with user count had been removed.
Conclusion
The mistake PDF7 made was not choosing bad password security. Our hashing was fine, our rate limiting worked, and our bot detection blocked most of what it saw. The mistake was treating the credential store as a component to secure when it was a component to remove.
Everything downstream followed from that reframe. The attack path terminated instead of narrowing. The budget fell by more than the tagged line predicted. The architecture absorbed 100x growth without a change, because the part that would have had to grow was gone. And a privacy claim that had needed a footnote for years stopped needing one.
If your product makes a data minimization promise of any kind, the useful exercise is to inventory what you actually store and ask which entries contradict the promise. For PDF7 the list had one item on it. Auditing that list took an afternoon; acting on it took a week; and the reason the result has held is that it is not a control that can be misconfigured — it is an absence.
About this post: written by the PDF7 Security Team. PDF7 is a privacy-first online PDF toolkit offering 25+ free tools with fully client-side document processing, serving users across 100 countries. All before-and-after figures are PDF7 first-party data — authentication logs, edge and bot-detection logs, incident records, product analytics, and the security operations budget — with the baseline taken immediately before the June 2024 cutover and the after column measured in November 2024. Where a pre-migration value was tracked as a category rather than a count, it is reported as a category. The 0.5%–2% credential stuffing success range is a published industry estimate used for planning and is not a PDF7 measurement. MojoAuth, the provider we moved to, has written up this engagement from its own side against the identical figures: see the PDF7 case study. Roadmap items describe priorities set during this period rather than features already shipped.