Skip to main content

Lesson 26: Playtesting & Telemetry

  • Module 13: Balance & Testing
  • Lesson 26 of 27
  • โฑ๏ธ About 1 h 30 min (instruction + lab)

You can't see your own game with fresh eyes, but your players can. In this lesson you plan playtests, instrument a game so it records what players actually do, and analyze that data carefully enough that it tells you the truth instead of what you hoped to hear.

๐ŸŽฏ Learning Objectives

By the end of this lesson, you will be able to:

  • Plan a playtest session with a clear question, a protocol and respect for testers' privacy.
  • Build telemetry with three channels: semantic events, a heatmap in seconds and a time-sampled path.
  • Save and load event logs safely with the csv and json modules.
  • Compute completion rates, per-task success rates, funnels and SUS scores that never divide by zero.
  • Compare two versions of a design with a permutation test and effect size, and report "no clear difference" when that is the truth.

Project: Instrumented Playtest, a three-task test level that logs every attempt, saves a CSV and prints an honest report.

In This Lesson

๐Ÿ”ฌ Why Playtest, and What to Ask

A playtest is a small experiment. You start with a question ("can new players find the second pad?"), watch real people play, collect evidence, and change the game based on what you learn. Then you test again. You already met the qualitative half in Game Dev I's Capstone Part 2: Polish & Playtest: sit quietly, don't explain, and watch where players hesitate. This lesson adds the quantitative half: numbers the game records for you.

A five-step loop of numbered nodes: Build (prototype, vertical slice, define goals), Test (select testers, create protocol, run session), Observe (heatmaps, path tracking, think-aloud), Analyze (quantitative, qualitative, synthesis) and Iterate (prioritize, implement, validate), which leads back to Build.
The playtest loop: build, test, observe, analyze, iterate. Each pass should answer a specific question.
Kind of testTypical questionBest evidence
UsabilityCan players figure out what to do?Observation, failed attempts, time to first success
Difficulty and balanceIs it too hard, too easy, or lopsided?Deaths, completion, the performance window from Difficulty & DDA
Fun and engagementDo players want to keep going?Interviews, voluntary replays, where they quit
A/B comparisonIs version B better than version A?The same metric measured on two groups

Numbers tell you what and where; watching and asking tell you why. A heatmap can show that everyone lingers by the decoy pad, but only a tester thinking aloud will tell you "I thought the lighter green one was the goal."

๐Ÿ“‹ Planning a Session

Write a one-page protocol before anyone plays, so every tester gets the same session:

  1. The question. One or two, specific: "Do players find pad 2 within 30 seconds?"
  2. Who. Testers who match your audience, and ideally people who have never seen the game.
  3. The script. What you say at the start (the same words each time), what task you give, and that you won't help. Ask them to think aloud.
  4. What gets recorded. Telemetry events, your notes with timestamps, and a short questionnaire afterward.
  5. What counts as done. Finishing the tasks, or a time limit.

๐Ÿ”’ Consent and privacy

Tell testers what the game records and why, and ask before you start. Identify them only by a code such as T01, keep names in a separate place (or not at all), record gameplay events rather than anything personal, and store the files locally. If you ever collect data from players of a released game, the rules where you and they live apply; learn them before you ship analytics.

โœ… Growth Mindset: Watching Someone Struggle With Your Game Is Hard

The first time a tester can't find something you think is obvious, you'll want to jump in and explain. Don't; and don't take it personally. Every confusing moment you watch is one that thousands of future players won't have to suffer. Designers who seem to have great instincts mostly have a long history of being surprised by testers. You're building that history now.

๐Ÿ“ก Designing Telemetry

Telemetry is the game writing its own lab notebook. The lab records three channels, because each answers a different question and none can be rebuilt from the others:

  • Events (what happened, and when): task_start, task_attempt, task_success, task_fail, session_end. Only meaningful moments, never "a frame passed".
  • Heatmap (where time was spent): seconds spent in each 40 ร— 40 px cell. Adding dt rather than 1 per frame makes a 144 Hz tester's map comparable to a 60 Hz tester's.
  • Path (in what order): the position every 0.2 s of game time, again independent of frame rate.
@dataclass
class Event:
    t: float                   # seconds since the session started
    session: str
    tester: str
    kind: str                  # task_start, task_attempt, task_success, task_fail, session_end
    data: dict = field(default_factory=dict)


class Telemetry:
    """Three channels: events (what happened), heat (where time was spent), path (in what order)."""
    def __init__(self, session, tester):
        self.session = session
        self.tester = tester
        self.clock = 0.0
        self.events = []
        self.heat = Counter()          # (cx, cy) -> seconds spent in that cell
        self.path = []                 # (t, x, y) every SAMPLE_EVERY seconds
        self._since_sample = SAMPLE_EVERY

    def tick(self, pos, dt):
        self.clock += dt
        self.heat[(int(pos.x // CELL), int(pos.y // CELL))] += dt
        self._since_sample += dt
        if self._since_sample >= SAMPLE_EVERY:
            self._since_sample = 0.0
            self.path.append((round(self.clock, 2), round(pos.x), round(pos.y)))

    def log(self, kind, **data):
        self.events.append(Event(round(self.clock, 3), self.session, self.tester, kind, data))

Log every attempt, not just successes

If the game only logs successes, a tester who pressed the button nine times on the wrong pad looks exactly like one who got it first time. Log the attempt first, then the outcome, with a reason for failures:

t.log("task_attempt", task=self.task, x=round(self.pos.x), y=round(self.pos.y))
if PADS[self.task].collidepoint(self.pos):
    t.log("task_success", task=self.task)
else:
    on_other = DECOY.collidepoint(self.pos) or any(r.collidepoint(self.pos) for r in PADS.values())
    t.log("task_fail", task=self.task, reason="wrong pad" if on_other else "not on a pad")

Now "task 2: 1 success in 6 attempts, 4 of them on the wrong pad" points straight at the decoy, before you even look at the heatmap.

๐Ÿ’พ Saving Data Safely

CSV looks simple enough to write by hand: join the values with commas. It breaks the moment a value contains a comma, a quote or a newline, and a JSON dictionary always contains commas. This complete program shows the difference:

import csv
import io
import json

row = {"t": 12.5, "tester": "T03", "kind": "task_fail",
       "data": json.dumps({"task": 2, "reason": "wrong pad"})}

# By hand: the commas inside the JSON become extra columns.
by_hand = ",".join(str(v) for v in row.values())
print(len(by_hand.split(",")), "columns by hand:", by_hand)

# With the csv module: fields that contain commas or quotes are quoted for you.
buffer = io.StringIO()
writer = csv.DictWriter(buffer, fieldnames=list(row))
writer.writeheader()
writer.writerow(row)
buffer.seek(0)
back = next(csv.DictReader(buffer))
print(len(back), "columns with csv:", json.loads(back["data"]))
5 columns by hand: 12.5,T03,task_fail,{"task": 2, "reason": "wrong pad"}
4 columns with csv: {'task': 2, 'reason': 'wrong pad'}

The lab writes a fixed set of columns and stores each event's variable details as one JSON string in the data column. Open files with newline="" as the csv docs require, and put logs in a folder next to the script (Path(__file__).parent / "playtest_logs") so they never land somewhere surprising.

def write_events(path, events):
    """Write events with csv.DictWriter: commas, quotes and newlines in the data are escaped for us."""
    path.parent.mkdir(parents=True, exist_ok=True)
    with open(path, "w", newline="", encoding="utf-8") as f:
        writer = csv.DictWriter(f, fieldnames=FIELDS)
        writer.writeheader()
        for e in events:
            writer.writerow({"t": e.t, "session": e.session, "tester": e.tester, "kind": e.kind,
                             "data": json.dumps(e.data, sort_keys=True)})

The lab then reads the file back and analyzes what was saved, not what is in memory. That small habit catches format bugs on day one instead of in the middle of your analysis.

๐Ÿ“Š Metrics That Don't Lie

Analysis code runs on messy data: a session with no testers yet, a task nobody attempted, a tester who skipped a step. Every metric has to handle those cases on purpose.

def completion_rate(events, n_tasks):
    """Share of testers who finished every task, or None when there are no testers."""
    groups = by_tester(events)
    if not groups:
        return None
    done = sum(1 for evs in groups.values()
               if len({e.data.get("task") for e in evs if e.kind == "task_success"}) >= n_tasks)
    return done / len(groups)
  • Return None, not 0, when there is nothing to measure. "0% completion" and "no data" are very different findings.
  • Per-task success rate is successes รท attempts, and None when there were no attempts.
  • A funnel counts testers who completed task 1, then tasks 1 and 2, and so on. Each stage is a subset of the one before, so it can never go above 100% or rise. If your funnel rises, it is counting the wrong thing.
  • Hotspots are the cells with the most seconds: heat.most_common(3). A hotspot away from any goal usually means confusion.

A standard questionnaire: SUS

The System Usability Scale (SUS) is a widely used ten-question survey. Each answer is 1 (strongly disagree) to 5 (strongly agree). Odd-numbered statements are positive ("I thought it was easy to use") and even-numbered ones negative ("I found it unnecessarily complex"), so they are scored in opposite directions and the total is scaled to 0โ€“100:

def sus_score(responses):
    """System Usability Scale: 10 answers from 1 to 5; odd items positive, even items negative."""
    if len(responses) != 10 or any(r not in (1, 2, 3, 4, 5) for r in responses):
        raise ValueError("SUS needs exactly 10 answers, each from 1 to 5")
    total = sum((r - 1) if i % 2 == 0 else (5 - r) for i, r in enumerate(responses))
    return total * 2.5

Python counts from 0, so index 0 is question 1 (positive). A tester who answers 3 to everything scores exactly 50; the best possible answers (5, 1, 5, 1, ...) score 100. A SUS score is not a percentage, so compare it with your own earlier builds rather than reading it as a grade.

โš–๏ธ Comparing Two Versions Honestly

You try two tutorial layouts on six testers each. Layout B's average time is 1.8 seconds faster. Is B better? Maybe, or maybe you just happened to give B slightly faster testers. With small groups, random variation alone produces differences like that all the time.

A permutation test asks the question directly: if the layout made no difference, how often would shuffling the testers into two random groups produce a gap at least this big? It needs only the standard library:

import random
import statistics


def permutation_p(a, b, rounds=5000, seed=1):
    rng = random.Random(seed)
    observed = abs(statistics.mean(a) - statistics.mean(b))
    pooled = list(a) + list(b)
    hits = 0
    for _ in range(rounds):
        rng.shuffle(pooled)
        diff = abs(statistics.mean(pooled[:len(a)]) - statistics.mean(pooled[len(a):]))
        hits += diff >= observed - 1e-12
    return (hits + 1) / (rounds + 1)


layout_a = [48, 55, 61, 52, 70, 58]      # seconds to finish the tutorial (practice data)
layout_b = [51, 49, 57, 60, 54, 62]
print(f"A {statistics.mean(layout_a):.1f} s, B {statistics.mean(layout_b):.1f} s, "
      f"p = {permutation_p(layout_a, layout_b):.3f}")

layout_c = [34, 38, 31, 40, 36, 33]
print(f"A {statistics.mean(layout_a):.1f} s, C {statistics.mean(layout_c):.1f} s, "
      f"p = {permutation_p(layout_a, layout_c):.3f}")
A 57.3 s, B 55.5 s, p = 0.668
A 57.3 s, C 35.3 s, p = 0.003

The p-value is that "how often by chance" fraction. For A versus B, shuffled groups produced a gap at least that big about two times in three: there is no evidence that B is better. For A versus C, it almost never happened by chance, so C really does look faster for testers like these.

  • Report the size of the difference too. Cohen's d (difference in means รท pooled standard deviation) says whether a difference is big enough to matter. With enough testers, even a trivial difference gets a small p-value.
  • "No clear difference" is a valid result. The lab's compare() only names a faster version when p is below 0.05, and otherwise says to test more players. Declaring a "winner" from any gap at all is how teams ship changes that did nothing.
  • Decide the metric before you test. If you check twenty metrics, one of them will look "significant" by chance.

If you prefer a library, SciPy provides Welch's t-test (pip install scipy, then scipy.stats.ttest_ind(a, b, equal_var=False)). The permutation test above has the advantage of making no assumptions about the shape of your data and needing nothing beyond Python itself.

๐Ÿ‹๏ธ Practice Exercise: Instrumented Playtest

Objective: finish the telemetry and analysis of a three-task test level so it logs every attempt, saves a CSV that survives any data, and reports metrics and A/B comparisons honestly.

Time: about 45 minutes, plus 10 minutes to test a classmate. Starter file: playtest_starter.py (your instructor has it). The level, drawing and input work. The CSV is joined by hand and never read back, failed presses aren't logged, one metric can divide by zero, the funnel can rise, the heatmap counts frames, and the A/B verdict declares a winner from any difference. The numbered to-do comments in it match these steps.

  1. Play the starter once and read what it prints when you close it. (โ‰ˆ 3 min)
  2. To-do 1 and 2: write events with csv.DictWriter (data as JSON) and read them back with csv.DictReader. (โ‰ˆ 10 min)
  3. To-do 3: log a task_attempt on every press and a task_fail with a reason when it misses. (โ‰ˆ 6 min)
  4. To-do 4: make the per-task success rate None when there were no attempts. (โ‰ˆ 3 min)
  5. To-do 5: make the funnel count testers who completed all tasks up to k. (โ‰ˆ 6 min)
  6. To-do 6: add dt seconds to the heatmap instead of 1 per frame. (โ‰ˆ 2 min)
  7. To-do 7: use permutation_p in compare() and only name a faster layout when p < 0.05. (โ‰ˆ 7 min)
  8. Swap machines with a classmate. Give them the one-sentence task, stay silent, and take notes. Then compare your notes with the report. (โ‰ˆ 10 min)

You are done when:

  • the saved CSV opens in a spreadsheet with exactly five columns;
  • the report lists attempts and successes per task, with None or no rate when nothing was attempted;
  • pressing SPACE on the decoy produces a task_fail with reason wrong pad;
  • the built-in A/B example prints "no clear difference";
  • you can name one change to the level based on your classmate's session.
๐Ÿ’ก Hint

For to-do 2, csv.DictReader gives you strings: convert t with float() and data with json.loads(). For to-do 5, the condition is "for every task number i from 1 to k, this tester has a task_success for i": that's an all(...) around an any(...).

โœ… Example Solution

If your instructor hands you the lab file, you will see a few extra lines marked lab runtime near the top and and frame_budget() in the loop. They let the instructor's checker run the program automatically; when you run it yourself they do nothing.

"""Instrumented Playtest: Advanced Lesson 26 practice exercise (solution).

A tiny three-task test level wired for telemetry. The tester walks to the
glowing pad and presses SPACE to activate it; one decoy pad looks almost the
same. Every press is logged as a task_attempt (with task_success or
task_fail), time spent per map cell builds a heatmap, and a path is sampled
every 0.2 s. When the window closes, the events are written with the csv
module, read back, and analyzed with empty-safe metrics.

Keys: WASD or arrows move, SPACE activates, H toggles the heatmap, P the path.
"""
import csv
import json
import random
import statistics
import time
from collections import Counter
from dataclasses import dataclass, field
from pathlib import Path

import pygame


WIDTH, HEIGHT = 800, 520
CELL = 40                      # heatmap cell size, px
SAMPLE_EVERY = 0.2             # seconds between path samples
SPEED = 220                    # px/s
LOG_DIR = Path(__file__).parent / "playtest_logs"
FIELDS = ["t", "session", "tester", "kind", "data"]
TEXT = (235, 235, 235)


@dataclass
class Event:
    t: float                   # seconds since the session started
    session: str
    tester: str
    kind: str                  # task_start, task_attempt, task_success, task_fail, session_end
    data: dict = field(default_factory=dict)


class Telemetry:
    """Three channels: events (what happened), heat (where time was spent), path (in what order)."""
    def __init__(self, session, tester):
        self.session = session
        self.tester = tester
        self.clock = 0.0
        self.events = []
        self.heat = Counter()          # (cx, cy) -> seconds spent in that cell
        self.path = []                 # (t, x, y) every SAMPLE_EVERY seconds
        self._since_sample = SAMPLE_EVERY

    def tick(self, pos, dt):
        self.clock += dt
        self.heat[(int(pos.x // CELL), int(pos.y // CELL))] += dt
        self._since_sample += dt
        if self._since_sample >= SAMPLE_EVERY:
            self._since_sample = 0.0
            self.path.append((round(self.clock, 2), round(pos.x), round(pos.y)))

    def log(self, kind, **data):
        self.events.append(Event(round(self.clock, 3), self.session, self.tester, kind, data))


def write_events(path, events):
    """Write events with csv.DictWriter: commas, quotes and newlines in the data are escaped for us."""
    path.parent.mkdir(parents=True, exist_ok=True)
    with open(path, "w", newline="", encoding="utf-8") as f:
        writer = csv.DictWriter(f, fieldnames=FIELDS)
        writer.writeheader()
        for e in events:
            writer.writerow({"t": e.t, "session": e.session, "tester": e.tester, "kind": e.kind,
                             "data": json.dumps(e.data, sort_keys=True)})


def read_events(path):
    with open(path, newline="", encoding="utf-8") as f:
        return [Event(float(r["t"]), r["session"], r["tester"], r["kind"], json.loads(r["data"]))
                for r in csv.DictReader(f)]


# --- analysis: every function is safe with no data --------------------------------------
def by_tester(events):
    groups = {}
    for e in events:
        groups.setdefault(e.tester, []).append(e)
    return groups


def completion_rate(events, n_tasks):
    """Share of testers who finished every task, or None when there are no testers."""
    groups = by_tester(events)
    if not groups:
        return None
    done = sum(1 for evs in groups.values()
               if len({e.data.get("task") for e in evs if e.kind == "task_success"}) >= n_tasks)
    return done / len(groups)


def task_stats(events):
    """task -> (successes, attempts, success rate or None)."""
    stats = {}
    for e in events:
        if e.kind in ("task_attempt", "task_success"):
            s, a = stats.get(e.data["task"], (0, 0))
            stats[e.data["task"]] = (s + (e.kind == "task_success"), a + (e.kind == "task_attempt"))
    return {task: (s, a, s / a if a else None) for task, (s, a) in stats.items()}


def funnel(events, n_tasks):
    """How many testers completed task 1, then 1 and 2, ... Never above 100% and never rising."""
    groups = by_tester(events)
    rows = []
    for k in range(1, n_tasks + 1):
        count = sum(1 for evs in groups.values()
                    if all(any(e.kind == "task_success" and e.data.get("task") == i for e in evs)
                           for i in range(1, k + 1)))
        rows.append((k, count, 100 * count / len(groups) if groups else None))
    return rows


def hotspots(heat, top=3):
    return heat.most_common(top)


def sus_score(responses):
    """System Usability Scale: 10 answers from 1 to 5; odd items positive, even items negative."""
    if len(responses) != 10 or any(r not in (1, 2, 3, 4, 5) for r in responses):
        raise ValueError("SUS needs exactly 10 answers, each from 1 to 5")
    total = sum((r - 1) if i % 2 == 0 else (5 - r) for i, r in enumerate(responses))
    return total * 2.5


def cohens_d(a, b):
    """Difference in means divided by the pooled standard deviation."""
    if len(a) < 2 or len(b) < 2:
        return None
    pooled = (((len(a) - 1) * statistics.variance(a) + (len(b) - 1) * statistics.variance(b))
              / (len(a) + len(b) - 2)) ** 0.5
    return (statistics.mean(a) - statistics.mean(b)) / pooled if pooled else None


def permutation_p(a, b, rounds=5000, seed=1):
    """Two-sided permutation test on the difference in means (standard library only).

    Shuffle the group labels many times and count how often a difference at least
    as large as the real one appears by chance."""
    if not a or not b:
        return None
    rng = random.Random(seed)
    observed = abs(statistics.mean(a) - statistics.mean(b))
    pooled = list(a) + list(b)
    hits = 0
    for _ in range(rounds):
        rng.shuffle(pooled)
        diff = abs(statistics.mean(pooled[:len(a)]) - statistics.mean(pooled[len(a):]))
        hits += diff >= observed - 1e-12
    return (hits + 1) / (rounds + 1)


def compare(name_a, a, name_b, b, alpha=0.05):
    """An honest one-line verdict: a difference is only claimed when the test supports it."""
    if len(a) < 2 or len(b) < 2:
        return "Not enough data to compare."
    p = permutation_p(a, b)
    d = cohens_d(a, b)
    ma, mb = statistics.mean(a), statistics.mean(b)
    summary = f"{name_a} mean {ma:.1f} vs {name_b} mean {mb:.1f}, d = {d:.2f}, p = {p:.3f}"
    if p < alpha:
        faster = name_a if ma < mb else name_b
        return f"{summary}: {faster} is faster in this sample."
    return f"{summary}: no clear difference; test more players before deciding."


# --- the test level ------------------------------------------------------------------------
PADS = {1: pygame.FRect(660, 60, 70, 70), 2: pygame.FRect(80, 380, 70, 70), 3: pygame.FRect(370, 220, 70, 70)}
DECOY = pygame.FRect(170, 380, 70, 70)       # looks like pad 2: a deliberate usability trap


class PlaytestSession:
    def __init__(self, tester="T01"):
        self.telemetry = Telemetry(session=time.strftime("%Y%m%d-%H%M%S"), tester=tester)
        self.pos = pygame.Vector2(100, 100)
        self.task = 1
        self.telemetry.log("task_start", task=1)

    @property
    def finished(self):
        return self.task > len(PADS)

    def update(self, move, dt):
        if move.length_squared() > 0:
            self.pos += move.normalize() * SPEED * dt
            self.pos.x = max(10.0, min(WIDTH - 10.0, self.pos.x))
            self.pos.y = max(10.0, min(HEIGHT - 10.0, self.pos.y))
        self.telemetry.tick(self.pos, dt)

    def action(self):
        """SPACE: every press is an attempt; it either succeeds or fails with a reason."""
        if self.finished:
            return
        t = self.telemetry
        t.log("task_attempt", task=self.task, x=round(self.pos.x), y=round(self.pos.y))
        if PADS[self.task].collidepoint(self.pos):
            t.log("task_success", task=self.task)
            self.task += 1
            if not self.finished:
                t.log("task_start", task=self.task)
        else:
            on_other = DECOY.collidepoint(self.pos) or any(r.collidepoint(self.pos) for r in PADS.values())
            t.log("task_fail", task=self.task, reason="wrong pad" if on_other else "not on a pad")

    def draw(self, screen, font, show_heat, show_path):
        screen.fill((24, 26, 34))
        if show_heat and self.telemetry.heat:
            most = max(self.telemetry.heat.values())
            overlay = pygame.Surface((WIDTH, HEIGHT), pygame.SRCALPHA)
            for (cx, cy), secs in self.telemetry.heat.items():
                alpha = max(0, min(255, int(40 + 180 * secs / most)))
                overlay.fill((230, 70, 60, alpha), (cx * CELL, cy * CELL, CELL, CELL))
            screen.blit(overlay, (0, 0))
        if show_path and len(self.telemetry.path) > 1:
            pygame.draw.lines(screen, (90, 190, 240), False, [(x, y) for _, x, y in self.telemetry.path], 2)
        pygame.draw.rect(screen, (70, 120, 90), DECOY, border_radius=8)
        for n, rect in PADS.items():
            color = (120, 230, 140) if n == self.task else (60, 80, 70) if n > self.task else (40, 60, 50)
            pygame.draw.rect(screen, color, rect, border_radius=8)
            screen.blit(font.render(str(n), True, TEXT), rect.move(28, 24).topleft)
        pygame.draw.circle(screen, (250, 250, 250), self.pos, 10)
        status = "All tasks done - close the window" if self.finished else f"Task {self.task}: stand on the green pad, press SPACE"
        screen.blit(font.render(status, True, TEXT), (10, 8))
        screen.blit(font.render(f"events {len(self.telemetry.events)}   H heatmap   P path", True, TEXT), (10, 30))


def report(events, heat):
    lines = []
    rate = completion_rate(events, len(PADS))
    lines.append("Completion: " + ("no testers" if rate is None else f"{rate:.0%}"))
    for task, (s, a, r) in sorted(task_stats(events).items()):
        lines.append(f"Task {task}: {s}/{a} attempts succeeded" + ("" if r is None else f" ({r:.0%})"))
    lines.append("Hotspots (cell, seconds): " + ", ".join(f"{c}:{s:.1f}" for c, s in hotspots(heat)))
    return lines


def main():
    pygame.init()
    screen = pygame.display.set_mode((WIDTH, HEIGHT))
    pygame.display.set_caption("Instrumented Playtest")
    clock = pygame.time.Clock()
    font = pygame.font.Font(None, 24)               # created once
    session = PlaytestSession()
    held = set()
    show_heat, show_path = True, True

    running = True
    while running:
        dt = min(clock.tick(60) / 1000, 0.05)
        for event in pygame.event.get():
            if event.type == pygame.QUIT:
                running = False
            elif event.type == pygame.KEYDOWN:
                held.add(event.key)
                if event.key == pygame.K_SPACE:
                    session.action()
                elif event.key == pygame.K_h:
                    show_heat = not show_heat
                elif event.key == pygame.K_p:
                    show_path = not show_path
            elif event.type == pygame.KEYUP:
                held.discard(event.key)

        move = pygame.Vector2(
            (pygame.K_RIGHT in held or pygame.K_d in held) - (pygame.K_LEFT in held or pygame.K_a in held),
            (pygame.K_DOWN in held or pygame.K_s in held) - (pygame.K_UP in held or pygame.K_w in held))
        session.update(move, dt)
        session.draw(screen, font, show_heat, show_path)
        pygame.display.flip()

    pygame.quit()
    t = session.telemetry
    t.log("session_end", finished=session.finished)
    path = LOG_DIR / f"session_{t.session}_{t.tester}.csv"
    write_events(path, t.events)
    events = read_events(path)                      # analyze what was saved, not what is in memory
    print(f"Saved {len(events)} events to {path.name}")
    for line in report(events, t.heat):
        print(line)
    # Example A/B data (made up for practice): seconds to finish the tutorial, two layouts.
    print(compare("Layout A", [48, 55, 61, 52, 70, 58], "Layout B", [51, 49, 57, 60, 54, 62]))
    print("Playtest closed cleanly.")


if __name__ == "__main__":
    main()

๐Ÿ““ Learning Journal

Take five minutes to write in your learning journal (a notebook or a plain text file works). Jot down:

  • Key concepts you learned today
  • Techniques that clicked (and the ones that haven't, yet)
  • Questions or confusion to bring to the next session
  • Ideas to try in your own game
  • Progress and feelings: how did this lesson go for you?

โœ๏ธ This lesson's prompts:

  1. What did your classmate do that you did not expect? Did the telemetry show it, or only your notes?
  2. Write the one question you most want a playtest of your capstone to answer, and the one metric that would answer it.
  3. How did it feel to report "no clear difference"? Why is that still a useful result?

๐Ÿ“ Summary

Playtesting is a loop of small experiments. You plan each session around a question, with a script and consent. The game records three channels: meaningful events (including every attempt and its outcome), a heatmap of seconds and a path sampled by time. You save with the csv module so no value can break the file, read back what you saved, and compute metrics that handle missing data by returning None. When comparing two versions, a permutation test and an effect size tell you whether a difference is real and whether it matters, and "no clear difference" is an honest, useful answer.

๐ŸŽ“ Key Takeaways

  • Start every playtest with a specific question; watch silently and ask testers to think aloud.
  • Log attempts and outcomes, not just successes; time-based heatmaps and paths work at any frame rate.
  • Use csv.DictWriter/DictReader with newline=""; never join CSV fields by hand.
  • Metrics return None for "no data"; funnels count testers who completed every earlier step.
  • SUS: odd items score r โˆ’ 1, even items 5 โˆ’ r, times 2.5.
  • Claim a difference only when the test supports it, and report its size.

๐Ÿ”ญ Looking Ahead

Everything in this course now comes together. In Advanced Capstone: Build, Playtest, Ship you plan and build a polished game that uses at least two Advanced systems, run a peer playtest with these methods, and ship it.

โ“ Common Questions

How many testers do I need?

For finding usability problems, a handful of sessions per round, repeated after each fix, is a common practice. For comparing two versions with numbers, you usually need many more, which is exactly why the lab's six-per-group example shows no clear difference.

Why seconds in the heatmap instead of visit counts?

Counting one per frame means a tester on a 144 Hz monitor "visits" every cell more than twice as much as one at 60 Hz. Adding dt measures time, which is what you actually care about.

Should I use JSON Lines instead of CSV?

JSON Lines (one JSON object per line) is a good choice when events have very different fields. CSV opens directly in spreadsheets. The lab combines them: fixed CSV columns plus a JSON data column.

Why is the permutation test seeded?

So the same data always gives the same p-value, which makes the tests and your reports repeatable. With 5000 shuffles, a different seed changes the p-value only slightly.

What's a "good" p-value?

0.05 is a common convention, not a law. What matters more is deciding your question and threshold before testing, reporting the effect size, and not treating 0.049 and 0.051 as opposite truths.

๐ŸŽฏ Quick Quiz

Question 1: Why does the lab log a task_attempt on every press instead of only logging successes?

Question 2: Why use csv.DictWriter instead of ",".join(...)?

Question 3: Six testers per layout: A averages 57.3 s, B averages 55.5 s, and the permutation test gives p = 0.668. What is the honest conclusion?

Question 4: A tester answers 3 to all ten SUS questions. What SUS score is that?

Question 5: Why does the funnel count testers who completed tasks 1 through k, rather than everyone who completed task k?

๐ŸŒŸ Going Further

  • Many sessions: write a small script that reads every CSV in playtest_logs/, concatenates the events and prints the funnel across all testers.
  • Heatmap image: save the heatmap overlay with pygame.image.save at the end of a session, so you can put several side by side.
  • SUS form: after the last task, show the ten SUS statements one by one and record answers 1โ€“5 with the number keys.
  • Instrument your RTS: log path_failed events in Tiny RTS and see which map tiles cause them.
  • Read the docs: csv, statistics and random (shuffle and seeded Random).
  • Coming up in Game Dev III: Advanced: the capstone studio, where a peer playtest with these tools is one of your milestones.