We spent six weeks with broken lobbies earlier this year. Not broken in a way that crashed sessions or kicked players. Broken in a quieter way: matches that were technically valid but felt deeply uneven. Players with similar historical rankings getting paired into lopsided exchanges that neither side found satisfying. The telemetry was fine. The ELO numbers said things should work. The sessions said otherwise.
The core problem turned out to be that we were reading skill at the wrong time and in the wrong way.
The Static Score Problem
Standard matchmaking systems do something reasonable and usually insufficient. They assign each player a rating, update that rating after each match, and use current ratings to build balanced lobbies. ELO is the classic version. TrueSkill is a more sophisticated variant. Both share the same fundamental architecture: you measure past performance, infer current skill from that history, and match accordingly.
The problem with this for our specific game is that a player's skill at any given session is not their average historical performance. It can be lower, if they are warming up or coming off a long break. It can be dramatically higher, if they are deep in a flow state and playing some of the best game of their life. Historical ratings capture neither of these states. They capture a statistical average over a history that may not represent what that player is doing right now.
For a casual game, this gap is acceptable. For a competitive shooter where matchmaking quality is directly tied to whether the player is having a good time, the gap matters a lot. A player who is genuinely playing above their rating will stomp through a lobby matched to their history. A player who is cold will get torn apart by opponents matched to who they normally are. Both outcomes produce sessions that feel wrong, and the rating system has no way to know they happened.
What Live Signal Reading Actually Requires
Replacing historical-only matching with something that incorporates live performance data sounds straightforward. It is not. We made three different attempts at this before landing on an approach that held together architecturally.
The first attempt was naive. We tracked kill-death ratio over the first three rounds and used that as a real-time adjustment factor applied to the historical rating. The result was worse than the baseline. Early-round variance is too high in competitive shooters. A player might go zero-for-three in round one because of map positioning decisions that have nothing to do with their mechanical skill. Reading that as a downward signal and adjusting the match configuration mid-session creates chaos. The lobby adjusted incorrectly, twice in eight seconds at one point during a test session, and players reported feeling like the game was actively working against them.
The second attempt moved to tracking longer-window performance signals: first-strike rate, utility usage timing, rotation response to audio cues. These are better indicators of actual skill level than kills in the first few rounds. But collecting them with enough reliability to make a matching decision introduces latency that conflicts with the responsiveness requirements of the backend. We were hitting 200 to 250 millisecond delays on the signal pipeline during contested match states. That is not catastrophic, but it is enough to create frame-level artifacts in the adaptive system sitting next to the matchmaking layer, and those artifacts were visible in the gameplay.
The third attempt, which is what we are running now, separates the signal collection from the matching decision into two different timescales.
The Two-Timescale Approach
The architecture we landed on treats the live skill signal as a confidence modifier on the historical rating rather than a replacement for it. The historical rating remains the primary factor. The live signal narrows or widens the acceptable match bracket based on what is actually observed in the current session.
Signal collection runs at its natural speed, which is slower than the game tick and uncoupled from it. We are not trying to read skill at 64Hz. We are reading it at roughly 0.5Hz across a rolling 90-second window, aggregating the behavioral indicators into a single confidence vector per player. That confidence vector says: this player's current session looks like their historical rating, or it looks like it is running hot, or it looks like they are not at full performance yet.
The matching decision uses that confidence vector to adjust the bracket tolerance. A player running hot gets matched into a slightly tighter bracket where the historical ratings of opponents are closer to the ceiling of what the hot player can handle. A player running cold gets matched into a slightly wider bracket with more tolerance on the downside. Neither adjustment is large. The historical rating is still the anchor. But the live confidence vector stops the system from making obviously wrong decisions when a player is clearly not performing at their average.
The key architectural constraint is that the live signal cannot override the historical rating by more than a calibrated delta. We set that delta based on playtest data across sessions. Allowing too large a delta reintroduces the variance problem we saw with the first approach. The live signal is real but noisy, and the historical rating is less precise but stable. You want the stable signal to dominate, with the live signal providing a correction layer, not a replacement.
What We Got Wrong on Lobby Assembly
Fixing the skill signal architecture did not fix everything. We still had an assembly problem.
Lobby assembly in a 5v5 format means you are not just matching individual skill levels. You are building two teams where the aggregate skill on each side is balanced and the internal skill variance within each team is compatible with team-based play. A team where four players are at one performance tier and one player is significantly above or below that tier will not function well as a team, even if the opposing team has the same aggregate rating.
We were assembling lobbies by solving for aggregate skill delta between teams. That produced lobbies that looked balanced in the numbers but felt unbalanced in the game because intra-team variance was uncontrolled. We have since added a variance constraint to the assembly algorithm. Both teams need to fall within a maximum intra-team skill range, not just match each other in aggregate. That single constraint cut our post-match "unfair" feedback reports substantially across the later playtest sessions.
The Honest State of Where We Are
The two-timescale architecture is solid. The assembly variance constraint works. Six weeks of broken lobbies taught us more about what matchmaking actually requires than six months of reading about it would have.
We are not claiming the system is finished. The confidence vector calibration needs more data at the tails. Players who are genuinely outlier performers in either direction are still harder to match well than players in the middle of the distribution, and they always will be. The assembly algorithm trades wait time against match quality in ways we are still tuning.
What we can say is that the system no longer produces the kind of sessions we were seeing in those six weeks. The mechanism is structurally correct. The remaining work is calibration, and calibration requires players, which is what the early access program is for.