Most playtesting focuses on the middle. You recruit players near the median skill level, you observe how they engage with the systems you have built, you patch the problems that show up most frequently. That approach works well for finding the sharp edges in the average experience. It does not tell you much about what happens at the extremes.
We ran six weeks of closed sessions specifically designed to test the extremes. We wanted to know what a very high-skill player experienced when facing our adaptive AI, and we wanted to know what a player at the lower end of competitive skill experienced with the same systems. The problems we found were not the ones we anticipated.
The Setup
We recruited two groups. The high-skill group were competitive FPS players with ranked history in games that use similar mechanical frameworks. These are people who understand movement, economy, and spatial awareness at a practiced level. They know what good competitive game feel looks like from inside a session, and they notice immediately when something is off.
The lower-skill group were players who engage with competitive shooters but sit below the median performance bracket in the games they play. They care about improvement, they play regularly, but they have not broken through to a level where they would describe themselves as competitive. This group is often underrepresented in closed playtest sessions because studios recruit toward skill, not toward breadth.
Each group ran two-hour sessions twice per week across the six weeks. We did not mix the groups. We observed each group independently and compared findings after the sessions. The game build they played was the same build with the same parameters throughout, which was important: we wanted to see how the same system behaved across the skill range without confounding the results by adjusting the build between groups.
What Broke for High-Skill Players
The high-skill players had no problem with the AI mechanics. Aim, movement, and reaction timing were all credible to them. The problem they surfaced was in the adaptation model's ceiling behavior.
A high-skill player typically has a deep repertoire. They do not run the same strategy twice in a row. They read the game state and make decisions based on what is happening now, not from habit. Our adaptation model had been tuned to pattern-match on behavioral repetition. Against a player with a shallow or habitual repertoire, this works. Against a high-skill player who is actively varying their approach, the model has nothing strong to pattern-match on and the AI's counter-behavior becomes generic rather than specific.
High-skill playtesters described this as the AI feeling "flat" in extended sessions. The first few rounds felt fresh. After round six or seven, when the adaptation layer had no strong signal to work with, the AI defaulted to aggressive forward pressure regardless of what the player was doing. Aggressive forward pressure is a sensible default behavior. It is not what you want from a system that is supposed to be reading your game and responding to it.
The fix for this is not to make the adaptation model more aggressive. It is to recognize when the model has no confident pattern to act on and shift to a different mode of play. We are working on a secondary behavioral mode for low-confidence states that does not just fall back to generic pressure. The adaptation layer should be honest about what it does and does not know, and its behavior should reflect that honestly.
What Broke for Lower-Skill Players
The lower-skill group surface completely different problems. The adaptation model worked correctly for them in the sense that it identified their patterns quickly. Lower-skill players tend to have more consistent behaviors that pattern-match more clearly. The AI found those patterns and countered them effectively.
The problem was that the adaptation pressure felt punishing rather than instructive. High-skill players who get countered can identify what happened and adjust. Lower-skill players often could not make that identification. The AI was doing something specific in response to their specific behavior, but the players could not see the connection. They experienced it as the game being hard in an opaque way rather than as a signal that they were doing something the AI had learned to expect.
This creates a game feel problem that is separate from the technical question of whether the adaptation is working. The adaptation was working correctly. But a system that adapts faster than the player can read what it is responding to produces frustration, not engagement. The player needs to be able to see what the AI has learned about them, at least partially, in order for the adaptation to feel like a fair challenge rather than an arbitrary difficulty spike.
We added a behavioral tell layer to the AI in response to this feedback. The AI now signals, through positioning and timing decisions that are legible without explanation, what behavior it has learned to expect. A player who has been rushing a specific corridor will notice, if they pay attention, that the AI pre-positions against that corridor. The signal is not a tutorial message. It is behavioral. But it gives the player something to read and respond to, which changes the experience from opaque pressure to a recognizable contest.
The Cross-Skill Insight
The most useful finding from running both groups simultaneously was the confirmation that what "working correctly" means differs completely between skill brackets.
For high-skill players, a correctly working adaptive AI needs to handle repertoire depth. It needs to do something meaningful when facing a player who varies their approach. For lower-skill players, a correctly working adaptive AI needs to be legible. Its adaptation needs to be visible enough that the player can connect the cause and effect.
These are not contradictory requirements, but they require different things from the system. A single parameter set that optimizes for one group will underperform for the other. We are now building bracket-aware behavioral mode selection into the system: the AI reads the skill level of its current opponent and selects from a set of behavioral modes calibrated to that bracket. High-skill opponents face an AI that handles ambiguity better. Lower-skill opponents face an AI that is more communicative in its behavior.
We want to be clear: this is not making the AI easier for lower-skill players. It is making it more readable. The difficulty level scales to the bracket. The behavioral communication style also scales to the bracket, but in the direction of clarity, not in the direction of pulling punches.
What Six Weeks Actually Looks Like
A note on the process itself: six weeks of closed playtesting on a system that is still in development is exhausting and generative in roughly equal measure. We had three sessions where the build was in a state we were not confident in and we ran them anyway, because the feedback from an unstable build is still useful feedback. Players can tell you what feels wrong even when they cannot tell you why. Several of the most useful findings came from sessions where the build was rough.
The thing that these sessions consistently demonstrate is that you cannot model the player experience from inside the studio. We knew the adaptation system worked technically. We had reviewed the behavioral logs and confirmed the model was updating correctly. None of that predicted what the high-skill group would say about ceiling behavior, and none of it predicted the legibility problem the lower-skill group surfaced. You have to run the sessions.