Why this matters
- VaultCommerce's first prototype stored orders in CSV files; a Black Friday traffic spike corrupted two files mid-write and lost twelve checkout records.
- Every backend service eventually shares data with other services, reports, and operators — files do not coordinate concurrent access.
- Interviewers use "why not just use files?" to test whether you understand concurrency, integrity, and query requirements before reaching for Postgres.
- Choosing the right storage layer early saves painful migrations when your catalog grows from hundreds to millions of SKUs.
What files get right
Flat files are simple. VaultCommerce's founders exported products.csv and orders.csv from a spreadsheet and loaded them in Java:
List<Product> products = csvReader.read("products.csv", Product.class);
For a solo developer and a read-only demo, that works. No server install, no schema migrations, easy to inspect in Excel.
File storage strengths
- Zero setup — Create a file, write bytes, done.
- Human-readable — Open in a text editor or spreadsheet for debugging.
- Portable — Copy the file to another machine; no network dependency.
The problems appear the moment two users checkout at the same time.
Where files break down
Production pain points
- Concurrency — Two threads appending to
orders.csvwithout locking produce interleaved or truncated rows. - Partial writes — A crash during
FileOutputStream.write()leaves half a record on disk with no rollback. - Query cost — Finding all orders for customer
user-8842means scanning every line in a multi-gigabyte file. - Integrity — Nothing stops you from inserting an order with
product_id = 999when product 999 does not exist.
VaultCommerce learned this on launch day when parallel checkouts overwrote the same file offset. The on-call engineer spent six hours reconciling payments against a corrupted ledger.
What a database adds
A relational database like Postgres wraps your data in a server process that handles locking, crash recovery, and structured queries:
SELECT o.id, o.total_cents, p.name
FROM orders o
JOIN order_items oi ON oi.order_id = o.id
JOIN products p ON p.id = oi.product_id
WHERE o.customer_id = 'user-8842'
ORDER BY o.created_at DESC;
The engine returns matching rows in milliseconds because it maintains indexes and a query planner — not because it reads faster disks.
Database guarantees over files
- Concurrent access — Row-level locks let thousands of checkouts proceed without corrupting each other.
- Atomic writes — A transaction either commits entirely or rolls back; no half-written orders.
- Declarative queries — Express what you need; the optimizer decides how to fetch it.
- Constraints — Foreign keys reject orphan order lines before they reach production data.
When files still make sense
Not every byte belongs in a database. VaultCommerce keeps product images in S3, static config in YAML, and one-off analytics exports as Parquet files on disk. Files excel at large immutable blobs and batch pipelines where no concurrent writer competes for the same path.
The migration moment
The signal to move off files is not team size — it is concurrent writers plus the need for reliable queries. VaultCommerce switched to Postgres when the second microservice needed to read order status while the checkout service was still writing it.
# application.yml — VaultCommerce order service
spring:
datasource:
url: jdbc:postgresql://localhost:5432/vaultcommerce
username: vault_app
One connection pool, one source of truth, and every service queries the same consistent snapshot.
Quick recall
Everything you need if you only revisit this box.
- Files work for prototypes and immutable blobs; they fail under concurrent writes and lack query efficiency.
- Databases provide locking, crash recovery, indexes, and declarative SQL over structured data.
- VaultCommerce moved to Postgres after CSV corruption during parallel checkouts.
- Use object storage for images and files for analytics exports; use a database for shared mutable state.
- The trigger to migrate is concurrent writers plus integrity and query requirements, not headcount.
Test yourself
Answer these before moving on — recall is what makes it stick.