Saturday, 10 October 2026

Why Horizontal Scaling Exposes Bad State Management

Why Horizontal Scaling Exposes Bad State Management

A PHP application can run perfectly well on one server for years.

Then traffic grows.

You add a second application instance behind a load balancer.

And suddenly strange things start happening.

Users get logged out.

Uploads disappear.

Background jobs cannot find temporary files.

One request sees state created by another server.

A deployment works on one node but fails on another.

The problem often gets blamed on “horizontal scaling.”

But horizontal scaling usually did not create the problem.

It exposed an assumption that was already there.

The application was treating local state as if it belonged to the whole system.

One server hides a lot

Imagine an application running like this:

User
 ↓
Server A
 ↓
PHP Application

Server A stores:

  • sessions locally,
  • temporary files locally,
  • cached values in process memory,
  • generated reports on its filesystem.

Everything appears to work because every request lands on the same machine.

Now introduce a load balancer:

             ┌── Server A
User → LB ───┤
             └── Server B

Request 1 goes to Server A.

Request 2 goes to Server B.

Now the assumptions change.

If Server A owns information Server B needs, the application is no longer behaving like one system.

Sessions are the classic example

Suppose PHP sessions are stored on the local filesystem.

The user logs in through Server A.

Their session is written there.

The next request reaches Server B.

Server B cannot find the session.

From the user's perspective:

“I was randomly logged out.”

The quick fix might be sticky sessions:

User 1 → always Server A
User 2 → always Server B

That can be useful in some situations.

But it may also hide the underlying architectural question:

Should this state belong to one application instance?

Often, the better answer is a shared session store.

For example:

Server A ─┐
          ├── Shared Session Store
Server B ─┘

Now either server can handle the next request.

Local files create the same problem

Consider a user uploading a document.

Server A receives it and saves:

/tmp/uploads/document.pdf

The next request reaches Server B and tries to process it.

The file is missing.

Again, horizontal scaling exposed a state-ownership problem.

If the file belongs to the application rather than one process, it probably needs storage that all relevant instances can access.

That might be:

  • object storage,
  • shared storage,
  • a database-backed workflow,
  • or another deliberate persistence mechanism.

The specific solution depends on the application.

The important part is recognizing ownership.

In-memory caches can quietly become authoritative

An in-memory cache is useful.

But it becomes dangerous when the application starts treating it as the source of truth.

Suppose Server A contains:

$productCache[$id] = $product;

Server B has its own completely separate memory.

Now:

Server A → price = $20
Server B → price = $25

Both servers are functioning exactly as designed.

The architecture is not.

This does not mean local caches are bad.

They can be excellent for data that is safe to recreate and tolerate temporary inconsistency.

The issue is assuming local memory represents globally consistent application state.

Static state becomes more dangerous in long-running processes

Horizontal scaling is not only about multiple machines.

Modern PHP applications may also have long-running workers.

Imagine:

final class CurrentTenant
{
    public static ?string $id = null;
}

One execution sets:

Tenant A

The next execution happens inside the same process.

If cleanup fails, it may inherit Tenant A's state.

Add multiple persistent workers and now every process has its own mutable copy of “current” state.

This creates two different problems:

Across processes:
state is inconsistent

Within a reused process:
state can become stale

Both come from the same architectural mistake:

execution-specific state has been given the wrong lifetime.

Ask what owns the state

A useful question for almost every value in an application is:

How long should this exist, and who owns it?

Some examples:

Application configuration
→ application lifetime

Current request
→ execution lifetime

Current tenant
→ execution lifetime

Database row
→ persistent business state

Local cache entry
→ disposable optimization

Temporary calculation
→ operation lifetime

Problems begin when those boundaries blur.

For example:

Current tenant
stored as application-wide static state

or:

Business-critical file
stored only on one web server

or:

Authoritative inventory count
stored only in local memory

Scaling exposes these mismatches quickly.

Queues expose similar assumptions

Imagine an HTTP request creates an order and dispatches a background job.

The job assumes it can access:

CurrentUser
CurrentTenant
request-local files
temporary memory
active database transaction

But the job may execute:

  • seconds later,
  • on another worker,
  • on another server,
  • after the original request is gone.

A queue forces the application to answer:

What information does this operation actually need?

That information should normally travel through an explicit job payload or be retrieved from an authoritative source.

The original request's ambient state should not be assumed to still exist.

Stateless does not mean “no state”

This distinction matters.

People often say scalable web applications should be stateless.

That does not mean the application has no state.

Business applications are full of state.

Customers.

Orders.

Sessions.

Payments.

Documents.

Feature configuration.

The better interpretation is:

Any application instance should not depend on unique local mutable state in order to process the next request correctly.

State still exists.

It simply has an explicit owner.

Horizontal scaling should become boring

Ideally, adding another application instance should look like:

1 instance
   ↓
2 instances
   ↓
5 instances
   ↓
10 instances

without changing application correctness.

The extra nodes should primarily provide capacity.

If adding Server B forces you to redesign authentication, storage, job handling, caching and tenant context all at once, those boundaries were probably implicit before scaling.

This is also why execution lifetime matters in EvolvePHP

One of the runtime principles behind EvolvePHP is separating:

Application
Execution
Transient

An application-level service can live for a long time.

The current user should not.

The current tenant should not.

The active transaction should not.

An HTTP request, queue message or scheduled task should have an explicit execution boundary.

And when that execution ends, its state should end with it.

This matters whether the application runs on:

PHP-FPM
long-running workers
multiple containers
queue consumers
future persistent runtimes

The deployment model may change.

The ownership rules should remain understandable.

Scale exposes architecture

Horizontal scaling is often described as:

“Just add more servers.”

Infrastructure may make adding those servers easy.

Application architecture determines whether doing so is safe.

Before scaling horizontally, I would ask:

  • Where are sessions stored?
  • Where are uploads stored?
  • What state exists only in process memory?
  • Which caches are authoritative?
  • Can jobs run on any worker?
  • Does execution-specific context leak into application lifetime?
  • Can any instance handle the next request?

If those answers are clear, horizontal scaling becomes much less mysterious.

Because the hardest part of adding another server is often not the server.

It is discovering how much your application secretly depended on the first one.

No comments:

Post a Comment