Friday, 2 October 2026

Scalability Starts in the Code Before It Reaches the Infrastructure

Scalability Starts in the Code Before It Reaches the Infrastructure

When developers talk about scalability, the conversation often jumps quickly to infrastructure.

Load balancers.

More servers.

Redis.

Read replicas.

Kubernetes.

Queues.

CDNs.

Those things matter.

But I think one of the biggest mistakes we make is assuming that scalability begins when infrastructure becomes more sophisticated.

It usually begins much earlier.

It begins in the code.

Infrastructure can multiply capacity. It cannot permanently compensate for inefficient architecture.

If every request performs unnecessary work, adding servers simply lets you perform more unnecessary work at the same time.

A scaling problem can start with one innocent request

Imagine an endpoint that returns a customer's dashboard.

At first, the application has 100 users.

The request does something like:

Load customer
Load subscriptions
Load invoices
Load payments
Calculate totals
Load notifications
Call external API
Render response

Everything feels fine.

Then traffic grows.

Now that same request is executed thousands of times.

Maybe every dashboard visit runs 30 database queries.

Maybe the external API is called every time even though the result changes once per hour.

Maybe historical totals are recalculated from thousands of rows for every request.

The application may still work.

But its cost grows directly with traffic.

Eventually someone says:

“We need more servers.”

Maybe you do.

But first I would ask:

Why does one request cost this much?

The cheapest scaling improvement is often doing less work

Suppose one request takes:

Database queries: 40
External calls: 3
CPU-heavy calculation: 1
Response time: 800ms

Now imagine reducing that to:

Database queries: 8
External calls: 0
Cached calculation: 1
Response time: 150ms

You have changed the scaling characteristics of the system before touching the infrastructure.

That is why performance and architecture are closely related.

Good scaling often means reducing the amount of work required per unit of traffic.

Query design matters early

Databases are usually one of the first places bad scaling assumptions become visible.

A feature may work perfectly during development with 500 records.

Then production contains 20 million.

This query:

SELECT *
FROM orders
WHERE customer_id = ?
ORDER BY created_at DESC;

may be fine.

Until the right index does not exist.

Or until the application loads all rows when it only needs twenty.

Or until every row triggers another ORM query.

The famous N+1 problem is a good example.

This:

Load 100 orders
    ↓
Load customer for each order

can quietly turn one logical operation into 101 database queries.

No amount of elegant controller code makes that scalable.

Understanding data access is part of designing the feature.

Data structures matter too

Scalability is not only about databases.

Sometimes code itself grows badly with input size.

Imagine checking whether every user appears in another list.

A simple implementation might repeatedly scan the entire list.

With ten users, nobody cares.

With ten million comparisons, suddenly the algorithm matters.

You do not need to turn every business application into a computer science exercise.

But developers should understand one basic question:

What happens to this code when the amount of data becomes 10x or 100x larger?

That question alone catches a surprising number of problems.

Blocking work should be questioned

Another common scaling issue is doing everything inside the user's request.

Imagine checkout performs:

Save order
Charge payment
Generate PDF invoice
Resize uploaded images
Send confirmation email
Notify analytics platform
Call CRM
Send webhook
Return response

Some of those operations may genuinely belong in the request.

Others may not.

If the user does not need the work completed before receiving the response, it may belong in a background job.

The request path might become:

Validate
Save order
Charge payment
Queue follow-up work
Return response

Now the user waits for less work.

And the system can process asynchronous workloads separately.

Queues are infrastructure, but the decision to separate synchronous and asynchronous responsibility is an architectural decision.

Stateless code scales more easily

Horizontal scaling becomes much simpler when any application instance can handle any request.

But that depends heavily on how state is managed.

Suppose server A stores temporary user state in process memory.

The next request reaches server B.

Now the state is missing.

Teams often discover this only after adding a second server.

The infrastructure did not create the problem.

It exposed an assumption already present in the code.

The same thing happens with:

  • local filesystem sessions,
  • static mutable state,
  • in-memory caches treated as authoritative,
  • background jobs depending on request-local state,
  • application instances owning information they should not own.

Designing state deliberately makes future scaling much easier.

Caching cannot fix unclear ownership

Caching is another area where scalability discussions quickly become infrastructure discussions.

“Add Redis.”

Sometimes that helps.

But before caching something, I want to understand:

What is the source of truth?
How stale can this value become?
What invalidates it?
Who owns invalidation?
What happens if the cache disappears?

Without those answers, the cache can become a second database with worse consistency rules.

Good caching begins with understanding the data.

Clear boundaries help systems scale selectively

Suppose an application contains:

Accounts
Orders
Billing
Reporting
Notifications

Now Reporting becomes extremely expensive.

If the architecture has reasonable boundaries, you may be able to scale Reporting differently.

Maybe its jobs get dedicated workers.

Maybe its database reads use replicas.

Maybe it eventually becomes an independent service.

But if everything shares implementation details and data freely, scaling one capability independently becomes difficult.

This is one reason modular architecture matters even before microservices enter the conversation.

Good boundaries give you more scaling options later.

Measure before redesigning

There is also an opposite mistake:

designing complicated architecture because you imagine huge traffic in the future.

That usually creates complexity before there is evidence that the complexity is needed.

I prefer something closer to:

Build simply
    ↓
Measure
    ↓
Identify pressure
    ↓
Optimize the expensive path
    ↓
Scale the necessary component

Observability matters here.

Without measurements, performance discussions become guesses.

You need to know:

  • slow endpoints,
  • expensive queries,
  • queue latency,
  • cache hit rates,
  • dependency latency,
  • error rates,
  • resource usage.

Then you can respond to real pressure.

Scaling is a chain

I tend to think about scalability in layers:

Efficient code
    ↓
Efficient data access
    ↓
Clear state ownership
    ↓
Async where appropriate
    ↓
Caching where justified
    ↓
Horizontal scaling
    ↓
Specialized infrastructure

You might eventually need every layer.

But skipping directly to the bottom rarely fixes weaknesses at the top.

Frameworks cannot make bad assumptions scalable

Laravel, Symfony and other mature PHP frameworks already provide powerful scaling tools.

Queues.

Caching.

Database connections.

Events.

Workers.

Middleware.

Scheduling.

But the framework cannot decide whether your request performs unnecessary work.

It cannot automatically decide which business operation should be asynchronous.

It cannot determine whether your data ownership makes sense.

It cannot know whether your cache invalidation strategy is correct.

Those are engineering decisions.

Scalability is really about growth without collapse

A scalable system is not simply one that can run on many servers.

It is one whose design allows capacity to grow without complexity growing uncontrollably with it.

Sometimes that requires more infrastructure.

Sometimes it requires fewer queries.

Sometimes it means moving work out of the request.

Sometimes it means changing an algorithm.

Sometimes it means fixing a boundary.

And sometimes the best scaling improvement is deleting work the system never needed to perform in the first place.

So before asking:

“How do we scale this application?”

I think it is worth asking something simpler:

“What exactly does the application do for one request, and how much of that work is actually necessary?”

That is often where scalability really starts.

No comments:

Post a Comment