Lesson 8: Utility AI
A hungry critter that sees food and an enemy at the same time has to weigh them, and the answer should change as it gets hungrier or more hurt. In this lesson you build utility AI, where every option earns a score and the best score wins, and you learn the three rules that keep those scores honest: clamp them, apply personality once, and add hysteresis.
🎯 Learning Objectives
By the end of this lesson, you will be able to:
- Build a
decide()loop that gates actions by requirements and picks the highest score. - Compare linear, quadratic, exponential, sigmoid and gaussian response curves and choose one for a need.
- Combine considerations so every score stays in [0, 1], and explain how a negative score inverts a personality trait.
- Apply personality multipliers in exactly one place and measure how aggression changes a fight's outcome.
- Build hysteresis and measure how it reduces action switching.
Project: Utility Critter, a creature that eats, fights, flees and explores, with live score bars.
In This Lesson
⚖️ Scores Instead of Priorities
Deciding what to have for lunch isn't a fixed priority list. If you're starving, the nearest sandwich wins; if you're only a bit hungry, you might walk further for something better; if you're late for a meeting, you skip lunch. Each option gets a mental score from several considerations (hunger, distance, time), and you pick the best.
That's utility AI. It's the natural next step after the Behavior Trees lesson: a tree's Selector always tries Escape before Gather because of where the branches sit, while a utility agent re-ranks its options every tick from the current situation. The whole decision is one loop:
def decide(agent, world, current=None, hysteresis=HYSTERESIS):
"""Return (best_action, final_scores). The current action gets a small bonus."""
finals = {}
for name, base in base_scores(agent, world).items():
if base is None:
finals[name] = None # requirement failed: can't be chosen
continue
finals[name] = base * personality_multiplier(agent, name)
best, best_value = "explore", -1.0
for name, score in finals.items():
if score is None:
continue
value = score + (hysteresis if name == current else 0.0)
if value > best_value:
best, best_value = name, value
return best, finals
There are two layers. Requirements are yes/no gates: you can't eat without food, so eat gets None and is skipped entirely. Scores rank everything that is possible. Keeping the gates separate means a zero score still means "possible, but not appealing", which is useful when you debug.
| Behavior tree | Utility AI | |
|---|---|---|
| Choice made by | Child order, first success wins | Highest score this tick |
| Great for | Clear rules ("always flee from fire") | Trade-offs that shift with needs |
| Tuning knobs | Tree structure | Curves, weights, traits |
| Classic failure | A branch that can't be interrupted | Flickering between near-equal scores |
They also combine well: a behavior tree can have a leaf that runs a utility decision, or a utility option can run a small tree.
📈 Response Curves
Every consideration starts as a number in [0, 1]: hunger 0.7, health 0.3, "enemy closeness" 0.9. A response curve turns that input into how much the agent cares. The curve's shape is a design decision:
- Linear: caring grows steadily.
- Quadratic / exponential: barely cares until the value is high, then cares a lot. Good for hunger: a critter shouldn't abandon everything for a snack.
- Sigmoid: a soft switch around a center value. Good for "am I healthy enough to fight?".
- Gaussian: peaks at a sweet spot and falls off on both sides, such as a preferred distance from the player.
Here are the five curves as Python, with a table of outputs you can compare against the figure:
import math
def clamp01(x):
return max(0.0, min(1.0, x))
def linear(x):
return clamp01(x)
def quadratic(x):
return clamp01(x) ** 2
def exponential(x, power=3):
return clamp01(x) ** power
def sigmoid(x, k=10, center=0.5):
return 1 / (1 + math.exp(-k * (clamp01(x) - center)))
def gaussian(x, center=0.5, width=0.2):
return math.exp(-((clamp01(x) - center) ** 2) / (2 * width * width))
for curve in (linear, quadratic, exponential, sigmoid, gaussian):
row = " ".join(f"{curve(x):.2f}" for x in (0.0, 0.3, 0.5, 0.7, 1.0))
print(f"{curve.__name__:12}{row}")
linear 0.00 0.30 0.50 0.70 1.00
quadratic 0.00 0.09 0.25 0.49 1.00
exponential 0.00 0.03 0.12 0.34 1.00
sigmoid 0.01 0.12 0.50 0.88 0.99
gaussian 0.04 0.61 1.00 0.61 0.04
Every curve clamps its input first. A need that drifts to 1.02 because of a rounding step must not produce a score above 1, and a quadratic of a slightly negative input would otherwise turn it positive.
✖️ Combining Scores Safely
An action usually has several considerations. The exercise's fight score is "healthy enough" × "good odds" × "enemy close enough":
odds = clamp01(1 - 0.5 * enemy["strength"]) # stronger enemy, worse odds
scores["fight"] = sigmoid(agent["health"]) * odds * clamp01(danger * 2)
Multiplying clamped considerations has two useful properties. The product stays in [0, 1], and any consideration at 0 vetoes the action: no amount of health makes fighting appealing when there's nobody close enough to fight.
The negative-score trap
A tempting formula, and one you will find in many tutorials, subtracts a threat term: combat - threat. That can go below zero, and a negative score breaks every multiplier applied after it. Watch what a personality trait does to it:
import math
def clamp01(x):
return max(0.0, min(1.0, x))
def sigmoid(x, k=10, center=0.5):
return 1 / (1 + math.exp(-k * (clamp01(x) - center)))
def old_fight_score(health, enemy_strength, aggression):
"""The tempting way: subtract a threat, then multiply by the trait."""
combat = sigmoid(health)
threat = enemy_strength * 0.5
return (combat - threat) * aggression
def new_fight_score(health, enemy_strength, aggression):
"""This lesson's way: clamp every consideration, multiply, apply the trait once."""
odds = clamp01(1 - 0.5 * enemy_strength)
return sigmoid(health) * odds * (0.5 + clamp01(aggression))
for aggression in (0.3, 0.9):
old = old_fight_score(0.3, 1.0, aggression)
new = new_fight_score(0.3, 1.0, aggression)
print(f"aggression {aggression}: old score {old:+.3f} new score {new:+.3f}")
aggression 0.3: old score -0.114 new score +0.048
aggression 0.9: old score -0.343 new score +0.083
With the subtract-then-multiply formula, the more aggressive critter gets the lower fight score, the exact opposite of the trait's meaning, so you could never watch aggression make a critter braver. With every consideration clamped to [0, 1], a bigger trait can only raise the score, and the exercise shows the difference in an actual fight.
✅ Growth Mindset: Print the Numbers
Utility bugs hide because the agent still does something plausible. When a critter behaves oddly, you don't need a better theory yet; you need the numbers. Print or draw every score each tick (the exercise's bars do this) and look for the one that's negative, above 1, or not changing when it should. The fix is usually one clamp away.
🎭 Personality, Applied Once
Personality makes fifty critters from one scoring function feel like fifty individuals. Each trait is a number in [0, 1], mapped to a multiplier from 0.5 to 1.5 (a trait of 0.5 gives 1.0, which is neutral), and applied to exactly one action:
TRAIT_OF = {"eat": None, "fight": "aggression", "flee": "caution", "explore": "curiosity"}
def personality_multiplier(agent, action):
"""Trait 0..1 becomes a multiplier 0.5..1.5 (a trait of 0.5 is neutral). Applied ONCE."""
trait = TRAIT_OF[action]
return 1.0 if trait is None else 0.5 + clamp01(agent["traits"][trait])
"Applied once" matters. A common slip is a score_explore that already multiplies by curiosity, followed by a personality step that multiplies by curiosity again, so the trait's effect is squared: tiny for timid critters and oversized for curious ones, and nobody could tell from reading either function. Keep considerations about the situation in the scoring functions, and traits about the character in one table.
With scores kept in range, aggression really does change outcomes. In the lab's test, a critter fights an enemy it can only just beat: with aggression 0.2 it breaks off at about 56% health and runs, and with aggression 0.8 it stays in and wins.
🧲 Hysteresis: Stop the Flicker
When two options score almost the same, tiny changes flip the winner back and forth every frame. The critter takes a step toward the enemy, a step away, a step toward. Players read that as a broken AI. The fix is hysteresis, the same idea as the gap between NEAR and SAFE in the Behavior Trees lesson: the rule for switching to an action is stricter than the rule for staying with it. Here the current action gets a small bonus, so a rival must beat it by a clear margin:
HYSTERESIS = 0.1 # the current action's bonus: a rival must win by more
value = score + (hysteresis if name == current else 0.0)
The difference is dramatic. In the lab's 12-second duel test, a timid critter switched actions 4 times with a bonus of 0.1 and 200 times with no bonus, and without hysteresis even the aggressive critter never finished its fight, because it kept dithering. Keep the bonus small: too big, and the critter commits to stale choices, such as eating while an enemy closes in.
Here is a browser version of the exercise. Add food and an enemy, change aggression, and switch hysteresis off to watch the switch counter race:
💡 Why this matters
Utility AI shines for characters with needs: survival creatures, colony sims, companions who should feel like they have moods. The score bars double as a design tool: a designer can look at them and say "fleeing should win sooner here" without reading code.
🏋️ Practice Exercise: Utility Critter
Objective: finish a critter that scores eat, fight, flee and explore every frame, so its choices change with hunger, health and personality without flickering.
Time: about 50 minutes. Starter file: critter_starter.py (your instructor has it). The curves, eat and explore scores, the world, movement and score bars are done. The numbered to-do comments (1 to 3) in the file follow steps 2 to 4 below, in order.
- Run the starter. Press F: the critter walks to food once it's hungry enough. Press E: fight and flee say "not possible". (≈ 3 min)
- Score fight and flee in
base_scores(), clamping every consideration. (≈ 12 min) - Write
personality_multiplier(): 1.0 for actions without a trait, otherwise 0.5 + the trait. (≈ 5 min) - Add the hysteresis bonus for the current action in
decide(). (≈ 5 min) - Test aggression: press E, watch the fight, then restart and press A three times before pressing E. Compare how each fight ends. (≈ 10 min)
- Press H to turn hysteresis off during a fight and watch the switch counter. (≈ 5 min)
You are done when:
- every bar stays between 0 and its personality multiplier, and never goes negative;
- at the default aggression the critter breaks off a fight and flees when it's hurt, and at aggression 0.8 it fights on and wins;
- with hysteresis on, the chosen action changes only a few times per fight; with it off, the switch counter climbs fast;
- closing the window prints the final action, the switch count and the hysteresis state.
💡 Hint
Fight and flee only make sense while the enemy is within DANGER_RANGE, so score them inside if danger > 0: and leave them None otherwise. If a braver critter flees sooner, some consideration went negative: wrap it in clamp01(). If curiosity feels far too strong, check that no scoring function reads a trait.
✅ Example Solution
The lab file has a few extra lines marked lab runtime near the top and and frame_budget() in the loop, so the instructor's checker can run it automatically. They do nothing when you run it yourself, and they are left out here.
"""Utility Critter: Advanced Lesson 8 practice exercise (solution).
A critter scores eat / fight / flee / explore every frame and does the best
one. F drops food, E drops an enemy, A/Z change aggression, C/X change caution,
H toggles hysteresis, SPACE pauses. The bars on the right are the live scores.
Close the window to quit.
"""
import math
import random
import pygame
PLAY_W, HEIGHT, PANEL_W = 700, 480, 300
SPEED = 120.0 # px/s
FLEE_SPEED = 150.0 # px/s
DANGER_RANGE = 300.0 # px: enemies farther than this are no threat
HYSTERESIS = 0.1 # the current action's bonus: a rival must win by more
ACTIONS = ["eat", "fight", "flee", "explore"]
TRAIT_OF = {"eat": None, "fight": "aggression", "flee": "caution", "explore": "curiosity"}
BAR_COLORS = {"eat": (250, 204, 21), "fight": (239, 68, 68), "flee": (249, 115, 22), "explore": (168, 85, 247)}
# ---------------------------------------------------------------- response curves
def clamp01(x):
return max(0.0, min(1.0, x))
def linear(x):
return clamp01(x)
def exponential(x, power=3):
return clamp01(x) ** power
def sigmoid(x, k=10, center=0.5):
return 1 / (1 + math.exp(-k * (clamp01(x) - center)))
def closeness(distance, max_range):
"""1.0 when touching, 0.0 at max_range or farther."""
return 1 - clamp01(distance / max_range)
# ---------------------------------------------------------------- scoring
def base_scores(agent, world):
"""Each action's score in [0, 1], or None when its requirement fails.
Considerations are multiplied, and every one is clamped to [0, 1] first,
so a score can never go negative (a negative score would flip the meaning
of every multiplier applied to it).
"""
scores = {name: None for name in ACTIONS}
food = nearest(agent["pos"], world["food"])
enemy = nearest(agent["pos"], world["enemies"])
if food is not None:
near_food = 0.5 + 0.5 * closeness(agent["pos"].distance_to(food), 600)
scores["eat"] = exponential(agent["hunger"]) * near_food
if enemy is not None:
danger = closeness(agent["pos"].distance_to(enemy["pos"]), DANGER_RANGE)
if danger > 0:
odds = clamp01(1 - 0.5 * enemy["strength"]) # stronger enemy, worse odds
scores["fight"] = sigmoid(agent["health"]) * odds * clamp01(danger * 2)
scores["flee"] = linear(1 - agent["health"])
scores["explore"] = 0.6 * linear(agent["boredom"])
return scores
def personality_multiplier(agent, action):
"""Trait 0..1 becomes a multiplier 0.5..1.5 (a trait of 0.5 is neutral). Applied ONCE."""
trait = TRAIT_OF[action]
return 1.0 if trait is None else 0.5 + clamp01(agent["traits"][trait])
def decide(agent, world, current=None, hysteresis=HYSTERESIS):
"""Return (best_action, final_scores). The current action gets a small bonus."""
finals = {}
for name, base in base_scores(agent, world).items():
if base is None:
finals[name] = None
continue
finals[name] = base * personality_multiplier(agent, name)
best, best_value = "explore", -1.0
for name, score in finals.items():
if score is None:
continue
value = score + (hysteresis if name == current else 0.0)
if value > best_value:
best, best_value = name, value
return best, finals
# ---------------------------------------------------------------- world
def nearest(pos, things):
if not things:
return None
key = (lambda t: pos.distance_to(t)) if isinstance(things[0], pygame.Vector2) else \
(lambda t: pos.distance_to(t["pos"]))
return min(things, key=key)
def new_agent():
return {"pos": pygame.Vector2(PLAY_W / 2, HEIGHT / 2), "hunger": 0.3, "health": 1.0, "boredom": 0.5,
"traits": {"aggression": 0.5, "caution": 0.5, "curiosity": 0.5},
"wander": pygame.Vector2(PLAY_W / 2, HEIGHT / 2)}
def move_toward(agent, target, speed, dt):
offset = target - agent["pos"]
agent["pos"] += offset.clamp_magnitude(speed * dt)
return offset.length()
def act(agent, world, action, dt, rng):
"""Carry out one frame of the chosen action, then let the needs drift."""
if action == "eat":
food = nearest(agent["pos"], world["food"])
if food is not None and move_toward(agent, food, SPEED, dt) < 8:
world["food"].remove(food)
agent["hunger"] = 0.0
elif action == "fight":
enemy = nearest(agent["pos"], world["enemies"])
if enemy is not None and move_toward(agent, enemy["pos"], SPEED, dt) < 22:
enemy["hp"] -= 0.35 * dt
agent["health"] -= 0.3 * enemy["strength"] * dt
if enemy["hp"] <= 0:
world["enemies"].remove(enemy)
elif action == "flee":
enemy = nearest(agent["pos"], world["enemies"])
if enemy is not None:
away = agent["pos"] - enemy["pos"]
if away.length_squared() == 0:
away = pygame.Vector2(1, 0)
agent["pos"] += away.normalize() * FLEE_SPEED * dt
else:
if move_toward(agent, agent["wander"], SPEED * 0.6, dt) < 8:
agent["wander"] = pygame.Vector2(rng.uniform(30, PLAY_W - 30), rng.uniform(30, HEIGHT - 30))
agent["boredom"] -= 0.25 * dt
agent["hunger"] = clamp01(agent["hunger"] + 0.04 * dt)
if action != "explore":
agent["boredom"] += 0.05 * dt
agent["boredom"] = clamp01(agent["boredom"])
enemy = nearest(agent["pos"], world["enemies"])
if enemy is None or agent["pos"].distance_to(enemy["pos"]) >= DANGER_RANGE:
agent["health"] += 0.05 * dt # only heals when no enemy is in range
agent["health"] = clamp01(agent["health"])
agent["pos"].x = max(10, min(PLAY_W - 10, agent["pos"].x))
agent["pos"].y = max(10, min(HEIGHT - 10, agent["pos"].y))
def draw_panel(surface, font, agent, finals, action, hysteresis_on, switches):
x0 = PLAY_W + 12
pygame.draw.rect(surface, (15, 23, 42), (PLAY_W, 0, PANEL_W, HEIGHT))
surface.blit(font.render("Scores (after personality)", True, (203, 213, 225)), (x0, 12))
for i, name in enumerate(ACTIONS):
y = 44 + i * 40
score = finals[name]
surface.blit(font.render(name, True, (226, 232, 240)), (x0, y))
if score is None:
surface.blit(font.render("not possible", True, (100, 116, 139)), (x0 + 70, y))
continue
color = BAR_COLORS[name] if name == action else tuple(c // 2 for c in BAR_COLORS[name])
pygame.draw.rect(surface, color, (x0 + 70, y, max(2, int(140 * score)), 18))
surface.blit(font.render(f"{score:.2f}", True, (226, 232, 240)), (x0 + 220, y))
lines = [f"doing: {action}",
f"hunger {agent['hunger']:.2f} health {agent['health']:.2f}",
f"boredom {agent['boredom']:.2f}",
f"aggression {agent['traits']['aggression']:.1f} caution {agent['traits']['caution']:.1f}",
f"hysteresis {'ON' if hysteresis_on else 'off'} switches {switches}",
"F food E enemy A/Z C/X H SPACE"]
for i, text in enumerate(lines):
surface.blit(font.render(text, True, (203, 213, 225)), (x0, 230 + i * 26))
def main():
pygame.init()
screen = pygame.display.set_mode((PLAY_W + PANEL_W, HEIGHT))
pygame.display.set_caption("Utility Critter")
clock = pygame.time.Clock()
font = pygame.font.Font(None, 22)
rng = random.Random(11)
agent = new_agent()
world = {"food": [], "enemies": []}
action, switches = "explore", 0
hysteresis_on, paused = True, False
running = True
while running:
dt = min(clock.tick(60) / 1000, 0.05)
for event in pygame.event.get():
if event.type == pygame.QUIT:
running = False
elif event.type == pygame.KEYDOWN:
traits = agent["traits"]
if event.key == pygame.K_f:
world["food"].append(pygame.Vector2(rng.uniform(30, PLAY_W - 30), rng.uniform(30, HEIGHT - 30)))
elif event.key == pygame.K_e:
spot = agent["pos"] + pygame.Vector2(160, 0).rotate(rng.uniform(0, 360))
spot.x = max(20, min(PLAY_W - 20, spot.x))
spot.y = max(20, min(HEIGHT - 20, spot.y))
world["enemies"].append({"pos": spot, "hp": 1.2, "strength": 0.5})
elif event.key in (pygame.K_a, pygame.K_z):
step = 0.1 if event.key == pygame.K_a else -0.1
traits["aggression"] = round(clamp01(traits["aggression"] + step), 1)
elif event.key in (pygame.K_c, pygame.K_x):
step = 0.1 if event.key == pygame.K_c else -0.1
traits["caution"] = round(clamp01(traits["caution"] + step), 1)
elif event.key == pygame.K_h:
hysteresis_on = not hysteresis_on
elif event.key == pygame.K_SPACE:
paused = not paused
choice, finals = decide(agent, world, action, HYSTERESIS if hysteresis_on else 0.0)
if not paused:
if choice != action:
switches += 1
action = choice
act(agent, world, action, dt, rng)
screen.fill((30, 41, 59))
for food in world["food"]:
pygame.draw.circle(screen, (250, 204, 21), food, 7)
for enemy in world["enemies"]:
pygame.draw.circle(screen, (70, 40, 40), enemy["pos"], DANGER_RANGE, 1)
pygame.draw.circle(screen, (239, 68, 68), enemy["pos"], 11)
pygame.draw.circle(screen, (96, 165, 250), agent["pos"], 12)
draw_panel(screen, font, agent, finals, action, hysteresis_on, switches)
pygame.display.flip()
pygame.quit()
print(f"final action: {action}; switches: {switches}; hysteresis {'on' if hysteresis_on else 'off'}")
print(f"aggression {agent['traits']['aggression']:.1f}, health {agent['health']:.2f}")
if __name__ == "__main__":
main()
📓 Learning Journal
Take five minutes to write in your learning journal (a notebook or a plain text file works). Jot down:
- Key concepts you learned today
- Techniques that clicked (and the ones that haven't, yet)
- Questions or confusion to bring to the next session
- Ideas to try in your own game
- Progress and feelings: how did this lesson go for you?
✍️ This lesson's prompts:
- Pick one need from your own life (sleep, hunger, boredom). Which response curve fits how it changes your choices, and why?
- Describe a character in your capstone who would be better with utility AI than with a behavior tree, or the other way around.
- What did the switch counter teach you about hysteresis that the explanation alone didn't?
📝 Summary
Utility AI turns decisions into arithmetic. Requirements gate out impossible actions; every possible action gets a score built from considerations in [0, 1], shaped by response curves and combined by multiplication; the highest score wins. Three habits keep it trustworthy: clamp every consideration so no score goes negative (a negative score turns a trait upside down), apply each personality trait once in a single table, and give the current action a small hysteresis bonus so near-ties don't make the agent flicker. Drawing the scores as bars makes all of this visible.
🎓 Key Takeaways
- Two layers: requirements say what's possible, scores say what's best.
- Response curves shape how much a need matters; the shape is a design choice.
- Multiply clamped considerations: the score stays in [0, 1] and any zero vetoes.
- Negative scores invert multipliers; clamp and they can't.
- Apply personality once, in one table, as a 0.5–1.5 multiplier.
- A small hysteresis bonus turns hundreds of switches into a handful.
🔭 Looking Ahead
Your agents can move, find paths and decide. Next they need somewhere worth exploring: in Dungeons & Caves (BSP, Cellular Automata) you generate the levels they'll roam.
❓ Common Questions
Should I add considerations or multiply them?
Multiplying makes every consideration a potential veto, which is usually what you want ("no enemy in range" should kill the fight option). Adding lets a strong consideration make up for a weak one. Some designs use a weighted average; if you do, still clamp each input first.
Multiplying many considerations makes scores tiny. Is that a problem?
It can be: an action with five considerations is penalized just for having more of them. A common fix is a compensation factor that raises the product toward the average as the count grows. Keep it in mind if you add many considerations; with two or three it rarely matters.
Should the agent always pick the top score?
Not always. Picking randomly among the top few, weighted by score, makes a crowd of identical agents look less robotic. Use a random.Random instance, and keep hysteresis so the chosen action still sticks.
How often should decide() run?
Every frame is fine for a handful of agents. With many, run decisions a few times per second and stagger them across agents; the actions themselves still update every frame with dt.
Where do GOAP, fuzzy logic and decision trees fit?
They are different decision techniques, and this lesson doesn't teach them. Goal-oriented action planning (GOAP) searches for a sequence of actions that reaches a goal; the Going Further section has a pointer.
🎯 Quick Quiz
Question 1: In decide(), what happens to an action whose requirement fails (its base score is None)?
Question 2: A subtract-then-multiply fight score is −0.38 before personality. Critter X has aggression 0.3 and critter Y has 0.9, applied as plain multipliers. What happens?
Question 3: Why does this lesson multiply clamped considerations instead of adding them?
Question 4: With HYSTERESIS = 0.1, the critter is eating (score 0.40) and explore now scores 0.45. What does it do this tick?
Question 5: The lesson's sigmoid has steepness 10 and center 0.5. About what does it return for a health of 0.7?
🌟 Going Further
- Commitment time: besides the score bonus, require an action to run for at least 0.5 seconds (a timer in seconds, decreased by
dt) before it can be replaced, unless a rival beats it by 0.3. - Weighted random choice: pick among the top two actions with probability proportional to score, using a
random.Randominstance. Keep hysteresis on and compare how natural it looks. - A herd: spawn ten critters with random traits in [0, 1] and watch how differently they handle the same enemy.
- Data-driven curves: move each action's curve and trait into a JSON file so designers can tune them without editing Python.
- Read more: Dave Mark's Behavioral Mathematics for Game AI covers response curves in depth, and Jeff Orkin's GOAP papers (from the game F.E.A.R.) introduce planning, a different approach to the same problem.