What We Learned From Watching 50 Competitive Players
All News

What We Learned From Watching 50 Competitive Players

9 min read

In the summer of 2025, we ran a research program that I still think about more than anything else we have done as a studio. We recruited fifty competitive FPS players across a wide skill range, from mid-ranked to players who had placed in regional tournaments, and we sat next to each of them for three-hour sessions while they played. We did not use them as playtesters for our own game. We watched them play whatever they normally played. We took notes. We occasionally asked questions. We tried not to interfere.

The goal was to build a ground-level picture of how competitive players actually engage with opponent AI: what they notice, what they complain about, what they adapt to, and what they give up on. Leaderboard data tells you who is winning. It does not tell you what winning players are thinking.

How Skill Looks From a Distance vs. Up Close

The first thing that surprised us was how different skill looked from a spectator position compared to rank position. Players we had recruited from mid-tier brackets because their ratings suggested moderate skill were, in many cases, doing sophisticated things that their ratings did not capture. They were tracking opponent positions through sound cues with high accuracy, holding angles in ways that suggested genuine map knowledge, and making conscious decisions about resource management that showed clear strategic thinking.

At the same time, some of the players from higher brackets were visibly coasting on mechanical precision alone. Their aim was excellent. Their game sense, in the sense of reading the game state and making predictive decisions, was not that much sharper than the mid-bracket players. Their ratings reflected the aim gap more than any strategic difference.

This matters for AI design because most skill-based opponent scaling systems are calibrated against rating data. If ratings are mostly measuring mechanical precision rather than game sense, then an AI tuned to replicate a rating-tier opponent will be miscalibrated along the dimensions that actually determine how players experience difficulty. We spent a significant amount of time after this program revisiting how we define skill tiers in our own system.

The Mental Model Problem

The second major finding was about mental models of AI opponents. Almost every player we watched had an explicit mental model of how the AI in their current game operated, and they were using that model to make decisions. What surprised us was how often the model was wrong in specific ways and how much that wrongness shaped their play.

A representative case: one player we observed was convinced that the AI in her game varied its aggression based on how many kills she had accumulated in a round. She had built a playstyle around this belief, playing conservatively early to accumulate a kill buffer before pushing aggressively. When we analyzed the match logs afterward, there was no such correlation in the AI behavior. The aggression variation was responding to positional factors, not kill count. Her mental model was plausible but incorrect, and it was shaping her strategy in ways that sometimes helped her and sometimes actively hurt her for reasons she could not diagnose.

This was not an isolated case. We saw this pattern across almost all the players we observed. The mental models were diverse, often plausible, and frequently wrong. The gap between what players believed about the AI and what the AI was actually doing was a constant source of frustration that players attributed to randomness, difficulty spikes, or unfairness rather than to a misread of the system.

What Players Actually Notice

We tracked what players verbalized during sessions, either to themselves or to us when we asked questions. The things they noticed and talked about fell into a small number of categories that we did not expect going in.

Positioning was mentioned far more often than aim. Players would comment on an opponent appearing in an unexpected location, holding an unusual angle, or moving in a way that suggested awareness of where the player was. They talked about this as though it reflected intelligence. Poor aim from an AI went largely unremarked unless it was severe enough to feel obviously broken.

Timing was the second most mentioned category. Players commented on opponents that seemed to "know" when to push versus hold, when to retreat and when to commit. This temporal intelligence read as much more sophisticated than raw reaction speed. Several players described good AI timing as feeling like playing against someone who had practiced the map for years.

Direct adaptation was mentioned by a minority of players, and interestingly, the players who mentioned it were not always correct that they were being adapted to. Some players who were not being adapted to described the AI as reading them. Some players who were being adapted to did not comment on it at all. Adaptation is harder to consciously detect than positioning or timing.

Frustration Patterns

We paid particular attention to moments of player frustration because those moments are where AI design failures are most visible. The frustration patterns we observed fell into three types.

The first was illegibility frustration. Players would lose a round and be genuinely unable to explain why. They would replay the last few seconds in their head, verbalize their reasoning, and arrive at nothing. This was the most common type and also the most corrosive to session satisfaction. Players who understood why they lost were willing to accept the loss. Players who could not understand it reported the game as unfair at much higher rates, even when their actual performance metrics were similar to the legibility-frustrated group.

The second was pattern frustration. Players had developed a counter to a specific AI behavior and that counter stopped working without an observable change in the AI's approach. In most cases, when we looked at the session data, the AI had not changed. The player's execution of the counter had become slightly inconsistent. But from the player's perspective, a strategy that used to work had mysteriously failed, which produced a specific kind of anger.

The third type, which appeared mostly at higher skill levels, was ceiling frustration. Players felt that no matter what they did, they could not push the AI past a certain difficulty. They were winning every round by a margin that felt too wide. The AI was not giving them a real challenge. This was the inverse of the illegibility problem but produced a similar outcome: players disengaged from the session.

What This Changed in Our Approach

We came away from this program with several concrete changes to how we approach AI design. We added explicit legibility requirements to our counter-strategy system: every adaptation the AI makes must be observable by a paying-attention player within a short window. We revised our skill-bracket definitions to include a game-sense dimension alongside the mechanical precision metric we had been using. And we built a ceiling-avoidance mechanism into the difficulty scaling to prevent the third frustration type from occurring in extended sessions.

More broadly, the program made us much more cautious about designing from behavioral data alone. The behavioral logs from those fifty sessions showed patterns that we could have analyzed and acted on without ever sitting next to the players. But sitting next to them gave us context that no log could capture. Why players made certain decisions, what they were trying to do when they made unusual choices, how they talked about the game when something went wrong. That context is what turned the behavioral observations into actual design knowledge.

We plan to run a second program of this kind once we have an internal playable that we can watch players engage with directly. The questions will be different. But the format will be the same.