← All insights Engineering

Why Optimizing One System Can Break Another

Performance work changes behaviour. Other systems may have been relying on the old behaviour.

Far Boundary · · 3 min read

You make a system faster and a different system starts failing. Items appear late, data goes missing, and players see states that should not exist. The cause is usually not the optimisation itself. It is that something else depended on a behaviour you changed without knowing it was a dependency.

Behavioural contracts nobody wrote down

Every system promises things that its neighbours quietly rely on. Examples: an event fires before a value is updated, results arrive in a particular order, a lookup always reflects the latest change, a loop finishes within the same frame. These are contracts. They are rarely documented, and an optimisation can break them.

How optimisations change behaviour

Timing
Batching, deferring or spreading work across frames means results arrive later or in a different order.
Caching
Stored results may be stale when something else changes the source.
Reduced update rates
Systems that sampled the value often now see it less often.
Reduced replication
Clients that relied on seeing something no longer receive it, or see it later.
Removed work
A step that looked unnecessary turned out to set something another system read.
Shared data changes
Switching a structure type can change ordering or iteration behaviour.

Where it hurts most

Networking assumptions
The server and client agree on a sequence. Change when messages are sent and the sequence breaks, producing desyncs or duplicated actions.
Data consistency
Saved data, in-memory state and what the player sees must agree. Delayed or batched writes can leave them out of step, especially around a player leaving.
Race conditions
Faster code can expose races that slower code hid by chance.
Interacting systems
Combat, inventory, effects and UI all touching the same values.

Why isolated benchmarks are not enough

A benchmark measures one thing in a quiet environment. It will not show that another system now reads stale data. It also measures speed, not correctness. A system can be faster and wrong.

Validate an optimisation safely

  1. State the target and the metric. What must improve, by roughly how much, measured how?
  2. Write down the contract. What does this system promise to others: timing, ordering, data freshness?
  3. Search for dependents. Find every place that reads its results, listens to its events or touches its state.
  4. Add checks first. Before changing anything, add tests or assertions that capture the current behaviour.
  5. Change one thing at a time so you know which change caused which effect.
  6. Test with real conditions: many players, joining and leaving, latency, long sessions.
  7. Compare behaviour, not just speed. Run old and new side by side where possible and check that outputs match.
  8. Roll out gradually, to a subset of servers if you can, and watch error and behaviour metrics.
  9. Keep an easy way back.

When behaviour must change

Sometimes the optimisation needs a different contract, such as updates arriving every few frames. Then change the dependents deliberately, with their own tests, instead of hoping nothing minds.

A useful habit

Record the contracts you discover. A short note beside each system listing what it guarantees makes the next optimisation safer.