Wednesday, 30 September 2026

Observability Is Not Logging

Observability Is Not Logging

A lot of systems say they have observability because they have logs.

They do not.

Logs are useful, sometimes essential, but they are only one part of understanding what a system is actually doing.

A system can produce thousands of log lines every minute and still be nearly impossible to diagnose when something goes wrong.

The real question is:

Can you understand the internal state of the system from the signals it produces?

That is what observability is really about.

Logs tell you what happened

A typical application might log:

User logged in
Order created
Payment failed
Queue job started
Queue job completed
Database timeout

Useful.

But when production is failing, those messages often leave important questions unanswered.

For example:

Why did checkout become slow?

Which dependency caused it?

Was the slowdown isolated to one tenant?

Did a retry make the problem worse?

Was the request blocked on the database?

Did the queue backlog first?

Which deployment introduced the change?

Logs may contain pieces of the answer.

Observability should help connect them.

The three signals are not enough by themselves

People often describe observability as:

Logs
Metrics
Traces

That is a useful model, but simply collecting all three does not automatically make a system observable.

You can have millions of metrics and still not know what happened.

You can have distributed traces that nobody can interpret.

You can have perfect logs with no useful context.

The important part is the relationship between the signals.

A request might look like:

HTTP Request
    ↓
Authentication
    ↓
Order Service
    ↓
Database
    ↓
Payment API
    ↓
Queue

If something takes 4.7 seconds, I want to know where those 4.7 seconds went.

That is where tracing becomes powerful.

Context is what turns data into evidence

Imagine seeing this log:

Payment provider timeout

Now compare it with:

execution_id: 9f73...
tenant: acme
order: ORD-18291
provider: stripe
attempt: 2
duration: 4.8s
trace_id: 7ab1...

The second message is not just more detailed.

It connects that event to a specific execution.

Now other telemetry from the same request, queue job or worker task can be correlated.

That is much more useful.

Observability becomes harder in long-running systems

Traditional PHP has an interesting advantage.

The process usually dies after the request.

That naturally destroys request-specific state.

Long-running workers change that.

Imagine a worker processing:

Execution 1 → Tenant A
Execution 2 → Tenant B
Execution 3 → Tenant C

If telemetry context from Execution 1 leaks into Execution 2, your monitoring system may confidently report the wrong tenant, user or trace.

That is worse than missing telemetry.

It is misleading telemetry.

In persistent runtimes, observability needs lifecycle discipline too.

Context should begin with the execution.

And it should end with the execution.

Errors are not the only thing worth observing

A mature system should help answer questions before everything fails.

For example:

Are requests becoming slower?

Is memory usage gradually increasing?

Are retries rising?

Is one dependency getting slower?

Are queue jobs taking longer?

Are workers being recycled more often?

Are database connections failing?

Are cleanup failures occurring?

These are often early signals.

By the time users report the outage, the system may have been warning you for hours.

Good observability helps turn those warnings into something engineers can act on.

Architecture should be observable too

This is an area I think deserves more attention.

Most observability platforms focus heavily on runtime performance.

That is important.

But architecture has runtime consequences too.

Imagine being able to see:

Module dependency violations

Unexpected cross-module calls

Execution-scoped services retained too long

Cleanup failures

Quarantined workers

High retry rates on one capability

Persistent memory growth

Unexpected data ownership crossings

Those signals tell you more than whether CPU usage is high.

They tell you when architectural assumptions are starting to break down.

This influences how I think about EvolvePHP

Observability is one of the areas I want EvolvePHP to treat as part of application architecture, not as something bolted on afterward.

That means thinking about:

  • execution identifiers,
  • structured observations,
  • trace context,
  • metrics,
  • lifecycle events,
  • module boundaries,
  • cleanup outcomes,
  • persistent-runtime diagnostics.

The idea is not that EvolvePHP should replace tools like OpenTelemetry or existing monitoring platforms.

Quite the opposite.

The framework should make it easier to produce useful evidence that those tools can consume.

A framework understands things that an external monitoring tool may not.

It knows when an execution starts.

It knows which module is running.

It knows when cleanup succeeds.

It knows when a process has become unsafe to reuse.

Those are valuable signals.

Observability should reduce uncertainty

When production breaks at 2 AM, nobody wants more dashboards.

They want answers.

What changed?

Where did the failure start?

Who was affected?

Was the problem local or systemic?

Can the process continue safely?

Did retrying make things worse?

What should we fix first?

That is the standard I think observability should be judged by.

Not how many logs the system produces.

Not how many charts appear on a dashboard.

But how quickly the telemetry turns uncertainty into understanding.

Logging records events.

Observability helps explain the system.

No comments:

Post a Comment