Lesson 16: Adaptive Audio & DSP
You will make sound effects out of pure numbers, process them with echo, filtering and overdrive, and build a soundtrack that tightens as danger gets closer. Audio that reacts to play is one of the strongest ways to make a game feel alive, and with numpy it takes surprisingly little code.
๐ฏ Learning Objectives
By the end of this lesson, you will be able to:
- Convert between numpy arrays and pygame-ce Sounds, matching the mixer's sample rate and channel count.
- Synthesize sound effects with numpy: tones, pitch sweeps, seeded noise and envelopes.
- Apply echo, a low-pass filter and soft clipping that work on mono and stereo arrays, and clip safely to 16-bit.
- Build layered adaptive music on reserved Channels that stay in sync and fade by time, not by frame.
Project: an Adaptive Soundtrack whose bass and drums fade in as a hunter gets closer, with an echoing sound effect on the space bar.
In This Lesson
๐ข Sound Is Just Numbers
A speaker cone moves in and out, and digital audio is simply a list of where the cone should be, measured tens of thousands of times a second. Each measurement is a sample; the number per second is the sample rate (44,100 is common). pygame-ce's mixer stores samples as 16-bit integers from โ32,768 to 32,767, and a stereo sound has two columns, left and right.
In Audio Mixing & Spatial Sound you played sounds from files and balanced them on buses. Here you will make the samples yourself, as a numpy array, and hand them to pygame.sndarray:
import numpy as np
import pygame
pygame.mixer.pre_init(44100, -16, 2) # ask for 44.1 kHz, 16-bit, stereo
pygame.init()
rate, size, channels = pygame.mixer.get_init() # what you actually got
print("rate", rate, "size", size, "channels", channels)
t = np.arange(int(0.5 * rate)) / rate # half a second of time stamps
wave = np.sin(2 * np.pi * 440 * t) * 0.3 # an A note at 30% volume, floats -1..1
pcm = (wave * 32767).astype(np.int16) # to 16-bit integers
if channels > 1:
pcm = np.column_stack([pcm] * channels) # shape (n, 2): same signal left and right
beep = pygame.sndarray.make_sound(pcm)
print("shape", pcm.shape, "length", round(beep.get_length(), 2), "s")
beep.play()
pygame.time.wait(600) # a tiny script: wait for the beep, then quit
pygame.quit()
- Ask, then check.
pre_init()is a request.pygame.mixer.get_init()returns the rate, sample format and channel count you really got, and your arrays must match them. - Shape must match the channels. A stereo mixer needs a 2-D array of shape
(n, 2). Passing a 1-D array raisesValueError: Array must be 2-dimensional for stereo mixer.np.column_stackcopies a mono signal into both columns. - Work in floats. Do all the math on floats from โ1 to 1, and convert to
int16only at the very end. pygame.time.wait()is fine in a ten-line test script. In a game it would freeze the loop; sounds play on their own while your loop keeps running.
๐๏ธ Synthesizing Sound Effects
Retro games built their whole soundscape from a few simple ingredients, and you can too. Each is a numpy one-liner:
def sweep(f_start, f_end, seconds, rate):
"""A tone whose pitch slides from f_start to f_end Hz (float samples, -1..1)."""
freq = np.linspace(f_start, f_end, int(seconds * rate))
phase = 2 * np.pi * np.cumsum(freq) / rate # add up the phase sample by sample
return np.sin(phase)
def noise(seconds, rate, seed=0):
"""White noise from a seeded generator, so the same seed gives the same sound."""
return np.random.default_rng(seed).uniform(-1.0, 1.0, int(seconds * rate))
def decay(samples, rate, speed):
"""Multiply by a fading envelope: loud at the start, exp(-speed * t) after."""
t = np.arange(len(samples)) / rate
shape = np.exp(-speed * t)
return samples * (shape if samples.ndim == 1 else shape[:, None])
def laser(rate):
return decay(sweep(1400, 180, 0.35, rate), rate, 7.0) * 0.8
def explosion(rate):
t = np.arange(int(0.9 * rate)) / rate
rumble = np.sin(2 * np.pi * 55 * t) * 0.5
return decay(noise(0.9, rate, seed=3) * 0.7 + rumble, rate, 5.0)
- Why
cumsumfor a sweep? A sine's phase is how far around the circle it has turned. At a fixed pitch the phase grows steadily, sosin(2ฯยทfยทt)works. When the pitch changes, each sample must advance the phase by that moment's frequency, so the phase is a running total:cumsum(freq) / rate. Plugging a changingfstraight intosin(2ฯยทfยทt)produces the wrong pitch. - Seeded noise.
np.random.default_rng(seed)is numpy's generator object, the same one-generator-per-system habit you use withrandom.Random(seed). The explosion sounds identical every run, which makes it testable. - Envelopes shape everything. The same noise is a hiss, a snare or an explosion depending on how fast it fades. Try different
speedvalues.
๐๏ธ Effects That Work in Stereo
DSP (digital signal processing) is math on sample arrays. Three effects cover a lot of game audio. Each one below works on a mono array of shape (n,) and a stereo array of shape (n, 2), because game code handles both.
Echo, with its tail
def echo(samples, rate, delay=0.25, feedback=0.45, repeats=4):
"""Add `repeats` quieter copies, each `delay` seconds later. The result is LONGER, so the
tail is kept. Works for (n,) and (n, channels) arrays."""
step = int(delay * rate)
n = len(samples)
out = np.zeros((n + step * repeats,) + samples.shape[1:], dtype=np.float64)
for k in range(repeats + 1):
out[k * step:k * step + n] += samples * feedback ** k
return out
An echo is copies of the sound, each later and quieter. The last copy starts delay ร repeats seconds after the first, so the output must be that much longer. Cutting it back to the original length chops off exactly the echoes you wanted to hear. samples.shape[1:] is () for mono and (2,) for stereo, so one line builds the right shape for both.
Low-pass: a moving average
def low_pass(samples, taps=20):
"""Moving-average filter (a simple FIR low-pass), run on each channel separately."""
kernel = np.ones(taps) / taps
if samples.ndim == 1:
return np.convolve(samples, kernel, mode="same")
return np.stack([np.convolve(samples[:, c], kernel, mode="same")
for c in range(samples.shape[1])], axis=1)
Averaging each sample with its neighbors smooths out fast wiggles (high frequencies) and keeps slow ones (low frequencies): the muffled sound of a noise heard through a wall or underwater. np.convolve only accepts 1-D arrays, so a stereo array is filtered one channel at a time. A handy property to check it with: a tone whose period is exactly taps samples is canceled completely, because each average covers one full cycle. At 44,100 Hz, 20 taps cancels 2,205 Hz (44,100 รท 20), while a 100 Hz tone passes almost unchanged.
Clipping: soft and hard
def soft_clip(samples, drive=1.0):
"""Overdrive: tanh squashes peaks smoothly, and never goes past -1..1."""
return np.tanh(samples * drive)
def to_int16(samples):
"""Float -1..1 to int16. Clipping FIRST stops loud peaks wrapping around to the other sign."""
return (np.clip(samples, -1.0, 1.0) * 32767).astype(np.int16)
Mixing and echo can push samples past 1.0: the explosion above peaks at about 1.15. Converting that straight to int16 is out of range, and numpy doesn't promise any particular result. On the machine used for this lesson, 1.2 ร 32,767 became โ26,216: a full-volume click of the opposite sign. np.clip first turns it into ordinary, much gentler clipping. np.tanh goes further and rounds peaks off smoothly, which is the "overdrive" sound of a distorted guitar; turning up drive gives more grit.
โ Growth Mindset: Train Your Ears Like You Trained Your Eyes
Audio bugs are harder to see than graphics bugs, and "it sounds weird" can feel impossible to debug. It isn't; you just need a way to look. Draw the waveform (the warm-up lab synth_dsp_solution.py does), print np.abs(samples).max() to catch clipping before you hear it, and compare lengths before and after an effect. Every sound designer started out unable to name what was wrong with a sound. Each clipped click you track down yet again is your ear getting better.
๐ผ Adaptive Music With Layers
Many games score a scene with vertical layers: several tracks written to play together, such as a calm pad, a bass line and drums. All of them play all the time; the game only changes their volumes. Exploring, you hear the pad. A threat appears and the bass fades in. Combat starts and the drums join. Because the layers never stop, they never fall out of step with each other.
Try it: drag the intensity and watch each layer move toward its target at a steady speed. Press Start sound to hear three synthesized loops built with the same formulas as the practice exercise.
Channels that belong to the music
pygame.mixer.music streams one track at a time, so it can't play three layers at once. Use three Sound objects on three Channels instead, and reserve those channels so sound effects never steal them:
pygame.mixer.set_reserved(len(layers)) # channels 0-2: music only
for i, (layer, make) in enumerate(zip(layers, (pad_layer, bass_layer, drum_layer))):
layer.channel = pygame.mixer.Channel(i)
layer.channel.set_volume(0.0)
layer.channel.play(to_sound(make(RATE)), loops=-1) # all start now, all loop
set_reserved(3)keeps channels 0 to 2 out offind_channel(), the call your sound effects use, so a burst of effects can't cut off the bass.loops=-1repeats forever. Every layer is exactly one bar long (2 seconds at 120 BPM), so they loop together. Starting them one after another in the same frame usually puts them on the same mixer buffer; if the mixer happens to run between twoplay()calls, they start one buffer apart (pygame-ce's default buffer is 512 samples, about 12 ms at 44.1 kHz).- All three start at volume 0. "Bringing in the drums" is then only a volume change.
Fades measured in seconds
FADE_PER_SECOND = 0.8 # volume units per second (0 -> 1 in 1.25 s)
def approach(current, target, max_step):
"""Move current toward target by at most max_step (never overshoots)."""
if current < target:
return min(target, current + max_step)
return max(target, current - max_step)
# every frame:
for layer in layers:
layer.volume = approach(layer.volume, targets[layer.name], FADE_PER_SECOND * dt)
layer.channel.set_volume(layer.volume)
Adding a fixed 0.01 to a volume every frame would make fades twice as fast at 120 FPS as at 60. Multiplying the fade rate by dt makes a fade take the same time on every machine, the same rule you apply to movement. Channel.fadeout() and play(fade_ms=...) exist too, but they fade all the way out to silence and stop, or fade in from silence as playback starts. Layers need to glide to any volume and back while they keep playing, so this lesson moves the volume itself.
๐๏ธ Practice Exercise: Adaptive Soundtrack
Objective: make a three-layer soundtrack that fades its bass and drums in as a patrolling hunter gets closer to you, with an echoing zap on the space bar.
Time: about 40 minutes. Starter file: adaptive_music_starter.py (your instructor has it). The layers are synthesized and the scene, movement and HUD are done. Its numbered to-do comments match the steps below.
- Run the starter and move with the arrows or WASD. The pad plays once and stops, and nothing reacts to the hunter. (โ 3 min)
- Make every layer loop forever by passing
loops=-1toplay(). (โ 2 min) - Write
intensity_from_distance(): 1.0 insidenear, 0.0 beyondfar, a straight line between. (โ 8 min) - Write
layer_targets(): bass 0.7 above intensity 0.3, drums 0.8 above 0.65, the pad always at 0.55. (โ 7 min) - Write
approach()so volumes glide toward their targets by at mostFADE_PER_SECOND * dtper frame, never overshooting. (โ 10 min) - Play: walk toward the hunter and away again, and press SPACE near and far. Change
FADE_PER_SECONDand feel the difference. (โ 10 min)
You are done when:
- the bass fades in as you approach the hunter's ring and the drums join when you are close, then both fade out as you leave;
- the HUD bars glide rather than jump, and a full fade takes about the same time at any frame rate;
- SPACE plays the zap with its echo tail and never cuts off the music;
- closing the window prints
Music layers started: 3, the number of zaps and the final volumes.
๐ก Hint
Test approach() by hand in the Python shell before running the game: approach(0.0, 1.0, 0.3) should give 0.3 and approach(0.9, 1.0, 0.3) should give 1.0, not 1.2. If the music cuts out when you zap, check that set_reserved() runs before the layers start and that the zap uses find_channel(), not Channel(0).
โ Example Solution
The lab file your instructor runs also contains a short block marked lab runtime and and frame_budget() in the loop, so the checker can run it for a fixed number of frames. They are left out here and do nothing when you run it yourself.
"""Adaptive Soundtrack: Advanced Lesson 16 practice exercise (solution).
Three music layers (pad, bass, drums) are synthesized with numpy, started
together on three reserved Channels so they loop in sync, and faded in and
out by how close the hunter is to you. Fades move at a fixed speed per
SECOND, so they take the same time at any frame rate.
arrows / WASD move SPACE zap (a sound effect with echo) Esc quit
"""
from dataclasses import dataclass
import numpy as np
import pygame
SIZE = W, H = 960, 540
RATE = 44100 # replaced by the mixer's real rate
BPM = 120
BAR_SECONDS = 4 * 60 / BPM # one bar of four beats = 2.0 s: every layer is exactly this long
FADE_PER_SECOND = 0.8 # volume units per second (0 -> 1 in 1.25 s)
SPEED = 240 # player speed, pixels per second
# ----------------------------------------------------------------------------- synthesis
def beat_times(rate):
return np.arange(int(BAR_SECONDS * rate)) / rate
def pad_layer(rate):
"""A soft A-minor chord that swells over the bar."""
t = beat_times(rate)
chord = sum(np.sin(2 * np.pi * f * t) for f in (220.0, 261.63, 329.63)) / 3
swell = 0.6 + 0.4 * np.sin(np.pi * t / BAR_SECONDS)
return chord * swell * 0.5
def bass_layer(rate):
"""A plucked low note on every beat."""
t = beat_times(rate)
beat = 60 / BPM
since_beat = t % beat
return np.sin(2 * np.pi * 55.0 * t) * np.exp(-since_beat * 6.0) * 0.7
def drum_layer(rate):
"""A kick on beats 1 and 3, a noise hi-hat on every half beat."""
t = beat_times(rate)
half = 30 / BPM
kick_t = t % (2 * 60 / BPM)
kick = np.sin(2 * np.pi * 60.0 * kick_t * np.exp(-kick_t * 8.0)) * np.exp(-kick_t * 9.0)
hat = np.random.default_rng(5).uniform(-1, 1, len(t)) * np.exp(-(t % half) * 40.0) * 0.25
return kick * 0.8 + hat
def zap(rate):
"""A short falling blip, with a 3-repeat echo tail added."""
n = int(0.18 * rate)
freq = np.linspace(900, 300, n)
blip = np.sin(2 * np.pi * np.cumsum(freq) / rate) * np.exp(-np.arange(n) / rate * 18.0) * 0.6
step = int(0.12 * rate)
out = np.zeros(n + 3 * step)
for k in range(4):
out[k * step:k * step + n] += blip * 0.5 ** k
return out
def to_sound(samples):
"""Float -1..1 to a pygame Sound shaped for the mixer (clipped first)."""
pcm = (np.clip(samples, -1.0, 1.0) * 32767).astype(np.int16)
channels = pygame.mixer.get_init()[2]
if channels > 1:
pcm = np.column_stack([pcm] * channels)
return pygame.sndarray.make_sound(np.ascontiguousarray(pcm))
# ----------------------------------------------------------------------------- adaptive logic
def intensity_from_distance(distance, near=90.0, far=420.0):
"""1.0 when the hunter is within `near` pixels, 0.0 beyond `far`, linear in between."""
return max(0.0, min(1.0, (far - distance) / (far - near)))
def layer_targets(intensity):
"""Target volume for each layer. The pad is always there; bass and drums join as it heats up."""
return {
"pad": 0.55,
"bass": 0.7 if intensity > 0.3 else 0.0,
"drums": 0.8 if intensity > 0.65 else 0.0,
}
def approach(current, target, max_step):
"""Move current toward target by at most max_step (never overshoots)."""
if current < target:
return min(target, current + max_step)
return max(target, current - max_step)
@dataclass
class Layer:
name: str
channel: object = None # a pygame.mixer.Channel, or None without audio
volume: float = 0.0
def main():
global RATE
pygame.init()
screen = pygame.display.set_mode(SIZE)
pygame.display.set_caption("Adaptive Soundtrack")
clock = pygame.time.Clock()
font = pygame.font.Font(None, 28)
layers = [Layer("pad"), Layer("bass"), Layer("drums")]
zap_sound = None
try:
pygame.mixer.init()
RATE = pygame.mixer.get_init()[0]
pygame.mixer.set_reserved(len(layers)) # channels 0-2: music only
for i, (layer, make) in enumerate(zip(layers, (pad_layer, bass_layer, drum_layer))):
layer.channel = pygame.mixer.Channel(i)
layer.channel.set_volume(0.0)
layer.channel.play(to_sound(make(RATE)), loops=-1) # all start now, all loop
zap_sound = to_sound(zap(RATE))
except pygame.error:
pass # no audio device: logic still runs
started = sum(1 for layer in layers if layer.channel is not None)
player = pygame.Vector2(160, 270)
hunter = pygame.Vector2(800, 270)
hunter_vel = pygame.Vector2(0, 150) # patrols up and down
held = set()
dirs = {pygame.K_LEFT: (-1, 0), pygame.K_a: (-1, 0), pygame.K_RIGHT: (1, 0), pygame.K_d: (1, 0),
pygame.K_UP: (0, -1), pygame.K_w: (0, -1), pygame.K_DOWN: (0, 1), pygame.K_s: (0, 1)}
intensity = 0.0
zaps = 0
running = True
while running:
dt = clock.tick(60) / 1000
for event in pygame.event.get():
if event.type == pygame.QUIT:
running = False
elif event.type == pygame.KEYDOWN:
if event.key == pygame.K_ESCAPE:
running = False
elif event.key == pygame.K_SPACE:
zaps += 1
channel = pygame.mixer.find_channel() if zap_sound else None
if channel is not None: # None = every free channel is busy
channel.play(zap_sound)
elif event.key in dirs:
held.add(event.key)
elif event.type == pygame.KEYUP:
held.discard(event.key)
move = pygame.Vector2()
for key in held:
move += dirs[key]
if move.length_squared() > 0:
player += move.normalize() * SPEED * dt
player.x = max(20, min(W - 20, player.x))
player.y = max(20, min(H - 20, player.y))
hunter += hunter_vel * dt
if not 60 <= hunter.y <= H - 60:
hunter_vel.y = -hunter_vel.y
hunter.y = max(60, min(H - 60, hunter.y))
intensity = intensity_from_distance(player.distance_to(hunter))
targets = layer_targets(intensity)
for layer in layers:
layer.volume = approach(layer.volume, targets[layer.name], FADE_PER_SECOND * dt)
if layer.channel is not None:
layer.channel.set_volume(layer.volume)
screen.fill((14, 16, 26))
pygame.draw.circle(screen, (60, 40, 50), hunter, 420, 1) # "far" ring
pygame.draw.circle(screen, (230, 70, 80), hunter, 18)
pygame.draw.circle(screen, (90, 220, 140), player, 14)
screen.blit(font.render(f"intensity {intensity:.2f}", True, (240, 240, 240)), (20, 16))
for i, layer in enumerate(layers):
y = 56 + i * 34
screen.blit(font.render(layer.name, True, (220, 220, 230)), (20, y))
pygame.draw.rect(screen, (50, 54, 76), (100, y, 200, 20))
pygame.draw.rect(screen, (120, 200, 255), (100, y, int(200 * layer.volume), 20))
pygame.display.flip()
print(f"Music layers started: {started}")
print(f"Zaps: {zaps}")
print("Final volumes: " + ", ".join(f"{layer.name}={layer.volume:.2f}" for layer in layers))
pygame.quit()
if __name__ == "__main__":
main()
๐ Learning Journal
Take five minutes to write in your learning journal (a notebook or a plain text file works). Jot down:
- Key concepts you learned today
- Techniques that clicked (and the ones that haven't, yet)
- Questions or confusion to bring to the next session
- Ideas to try in your own game
- Progress and feelings: how did this lesson go for you?
โ๏ธ This lesson's prompts:
- Describe a moment in a game you love where the music changed with what you were doing. How do you think it was built: layers, a new track, or something else?
- Which game state in your capstone idea would make a good "intensity" value, and which three layers would you write for it?
- What surprised you about making sounds from numbers instead of files?
๐ Summary
You treated sound as what it is inside the computer: an array of samples at a known rate, shaped to match the mixer's channels. You synthesized sweeps, noise and envelopes with numpy, then processed them with an echo that keeps its tail, a moving-average low-pass that filters each channel separately, and clipping that happens before the conversion to 16-bit so nothing wraps around. Finally you built vertical layers: loops of equal length started together on reserved channels, with volumes that glide toward targets at a rate per second driven by the game's intensity.
๐ Key Takeaways
- Read
pygame.mixer.get_init()and shape arrays to match:(n,)for mono,(n, 2)for stereo. - Do audio math in floats from โ1 to 1; clip, then convert to
int16last. - Pitch sweeps need a running phase (
cumsum); seeded generators make sounds repeatable. - Effects must handle both shapes; an echo makes the sound longer, and
np.convolveruns per channel. - Adaptive layers loop together on reserved channels and change only their volumes, fading by
rate ร dt.
๐ญ Looking Ahead
In Sockets, TCP & UDP, the course turns to multiplayer: how two programs talk over a network, what latency and bandwidth mean for a game, and your first echo server.
โ Common Questions
Why do I get "Array must be 2-dimensional for stereo mixer"?
The mixer is stereo and you passed a 1-D (mono) array. Stack it into two columns with np.column_stack([pcm, pcm]), or build it with the channel count from pygame.mixer.get_init() as to_sound() does.
My synthesized sound clicks at the start or end. Why?
A sound that starts or stops mid-wave jumps from some value to silence in one sample, which you hear as a click. Fade in over a few milliseconds and fade out at the end (an envelope), or end the sound where the wave crosses zero.
Can I apply these effects to sounds loaded from files?
Yes. pygame.sndarray.array(sound) gives you the samples as an int16 array; divide by 32,768 to get floats, apply the effect, then build a new Sound with to_sound(). Do it once when loading, not every time the sound plays.
Should I generate music in real time?
Not sample by sample in Python: the mixer needs tens of thousands of samples a second and a Python loop can't keep up reliably. Generate or load whole loops ahead of time, as this lesson does, and adapt by changing volumes, which is cheap.
What if the player's machine has no audio device?
pygame.mixer.init() raises pygame.error. The exercise catches it and keeps running with the layer logic but no sound, which is also what lets the lab checker run it on a server.
๐ฏ Quick Quiz
Question 1: The mixer is stereo and make_sound(pcm) fails because pcm has shape (22050,). What is the fix?
Question 2: Why does sweep() compute its phase with np.cumsum(freq)?
Question 3: Why is the output of echo() longer than its input?
Question 4: Why do all three music layers start at once and keep looping at volume 0 instead of starting when needed?
Question 5: The fade uses approach(volume, target, FADE_PER_SECOND * dt). What does multiplying by dt guarantee?
๐ Going Further
- Stingers: when the intensity crosses 0.65, play a short one-shot "alarm" sound on a free channel, but only once per crossing.
- Horizontal re-sequencing: instead of layers, switch between whole loops (explore, combat) at the next bar line by counting bars from the moment the music started.
- A better low-pass: a one-pole filter,
y[i] = y[i-1] + a * (x[i] - y[i-1]), gives a smoother, adjustable cutoff. Write it in a loop over a sound once at load time, and compare it with the moving average. - Docs: pygame-ce sndarray and mixer (
set_reserved,Channel); numpy convolve.