AlgebraKart AI — Current Architecture (as-built)
This document describes what the AI subsystem in game/AlgebraKart/ai/ actually does today, traced from source, not what it was designed to do. It is the baseline for the LSTM / time-based-evolution work described in ai-next-architecture.md.
Lineage: the design is a C++ port of Samuel Arzt's Deep Learning Cars (Unity, March 2017) — an evolutionary-algorithm-trained feed-forward network with no backpropagation. Weights are the genome; selection/crossover/mutation are the only learning signal.
1. Component map
| File | Role |
|---|---|
ai/genotype.{h,cpp} | One member of the population: a flat std::vector<float> of parameters + evaluation and fitness scalars. |
ai/neural_layer.{h,cpp} | A dense layer: weights[neuronCount+1][outputCount], one activation function. |
ai/neural_network.{h,cpp} | An array of NeuralLayer* plus a shared topology int array. |
ai/agent.{h,cpp} | Pairs a Genotype with a freshly-built NeuralNetwork (ffn) and a NetworkActor (the in-world vehicle/character). |
ai/agent_controller.{h,cpp} | Per-frame brain loop: reads sensors + flags, computes reward, runs the FFN, decides controls. |
ai/sensor.{h,cpp} | One raycast probe in a fixed local direction. |
ai/flag.{h,cpp} | A boolean detector fed into the network as 0/1. |
ai/agent_movement.{h,cpp} | Converts control outputs into NetworkActor steering/throttle. |
ai/agent_fsm.{h,cpp}, ai/fsm.{h,cpp} | State machine scaffolding (largely unused by the control path). |
ai/genetic_algorithm.{h,cpp} | Selection, recombination, mutation, termination. |
ai/evolution_manager.{h,cpp} | Static singleton owning the population, agents, controllers, actors; statistics file writing. |
GenomeVisualization.{h,cpp} | A 3D + UI network viewer. Compiled but never instantiated anywhere. |
Ownership at runtime is all through EvolutionManager's statics:
EvolutionManager (static singleton)
├── GeneticAlgorithm* geneticAlgorithm — owns currentPopulation_ / prevPopulation_ (Genotypes)
├── int* ffnTopology — shared, non-owned by the networks that point at it
├── vector<shared_ptr<Agent>> agents — index-aligned
├── vector<shared_ptr<AgentController>> agentControllers — index-aligned
└── vector<shared_ptr<NetworkActor>> networkActors — index-aligned
The three vectors are parallel arrays keyed by agent index; nothing enforces that they stay the same length, and several call sites index one from the size of another.
2. Network topology
Built in EvolutionManager::buildNeuralNetwork() (ai/evolution_manager.cpp:149), with NUM_NEURAL_LAYERS = 4 (Constants.h:9):
ffnTopology = [8, 7, 5, 3, 3] // 5 entries, 4 "layers"
layer 0: 8 → 7
layer 1: 7 → 5
layer 2: 5 → 3
layer 3: 3 → 3
weightCount = Σ (topology[i] + 1) * topology[i+1] = 9·7 + 8·5 + 6·3 + 4·3 = 133.
So every genotype is a 133-float vector, and that is the entire genome.
Documented input map (evolution_manager.cpp:163)
| # | Input | Source |
|---|---|---|
| 0 | sensor: left | raycast (1,0,0) in vehicle space |
| 1 | sensor: up | raycast (0,1,0) |
| 2 | sensor: front | raycast (0,0,1) |
| 3 | sensor: right | raycast (-1,0,0) |
| 4 | on-track flag | numWheelsContact > 0 |
| 5 | un-flipped flag | timeSinceNoWheelContact <= 5s |
| 6 | velocity-consistent flag | smoothedSpeedKm > 1.0 |
| 7 | (unused — comment lists a "distance meter" that is never wired in) | — |
The comment block lists 8 semantic inputs numbered 1–8 but the input layer is 8 wide and AgentController::update() only fills sensors.size() + flags.size() = 4 + 3 = 7 slots. Input 7 (ffnInput[7]) is read uninitialised (agent_controller.cpp:139, the array is new double[7] — actually one element shorter than the input layer, so the network would read past the end if it read its inputs at all; see §5.1).
Output map
| # | Output | Consumer |
|---|---|---|
| 0 | steering | AgentMovement::setInputs |
| 1 | acceleration | AgentMovement::setInputs |
| 2 | action state | AgentMovement::setInputs |
Activation function: MathHelper::softSignFunction for every layer, assigned in the Agent constructor (agent.cpp:37). NeuralLayer defaults to sigmoid but that default is always overwritten.
3. The evolution loop
AlgebraKart::InitEvolutionManager() (AlgebraKart.cpp:774)
└─ EvolutionManager::startEvolution()
├─ buildNeuralNetwork() → ffnTopology, weightCount = 133
├─ new GeneticAlgorithm(133, POPULATION_SIZE = 8)
├─ wire events:
│ allAgentsDied → evalFinished
│ fitnessCalculationFinished → writeStatisticsToFile, checkForTrackFinished
│ terminationCriterion → checkGenerationTermination
│ algorithmTerminated → onGATermination
└─ geneticAlgorithm->start()
├─ buildPopulation() → 8 Genotypes, 133 random params each
├─ initializePopulation() → setRandomParameters(-1, 1)
└─ evaluation(currentPopulation) → EvolutionManager::startEvaluation()
└─ for each genotype: new Agent + new AgentController, agentsAliveCount++
Per frame, AlgebraKart::HandleAI() (AlgebraKart.cpp:748) walks the agent list and calls controller->update(timeStep) for every controller whose actor is alive.
A generation is supposed to end like this:
AgentController::die() → Agent::kill() → agentDied event → EvolutionManager::onAgentDied()
└─ when agentsAliveCount hits 0 → allAgentsDied() → evalFinished()
└─ GeneticAlgorithm::evaluationFinished()
├─ fitnessCalculationMethod() fitness = evaluation / meanEvaluation
├─ sort by fitness descending
├─ checkTermination() generationCount >= RestartAfter (100)
├─ remainderStochasticSampling() → intermediate population
├─ randomRecombination() → new population (best 2 carried over verbatim)
├─ mutateAllButBestTwo()
└─ evaluation(newPopulation) → rebuild agents, generationCount++
In practice this never fires. Both die() call sites in AgentController::update() are commented out (agent_controller.cpp:162 for the no-wheel-contact timeout, agent_controller.cpp:396 for the checkpoint timeout) and win() is disabled (agent_controller.cpp:252). The only live die() path is AlgebraKart.cpp:1807, which is itself inside a commented-out block. So agentsAliveCount never reaches zero, allAgentsDied never fires, and the population stays on generation 1 forever.
There is no wall-clock or step budget on a generation — the only termination signal is "every agent died".
4. Reward / fitness
AgentController::update() accumulates a scalar reward into genotype->evaluation every frame (agent_controller.cpp:205):
o1 = (leftDist - 90) / 310 // want large
o2 = (upDist - 20) / 65 // want small (touching ground)
o3 = (frontDist - 90) / 310 // want large
o4 = (rightDist - 90) / 310 // want large
o5 = distanceCovered > 0 ? 1 - 1/distanceCovered : 0
reward += ((o1*0.2 + (1 - o2*0.2) + o3*0.2 + o4*0.2 + o5*0.2) * dt * 0.02)
Then fitness = evaluation / mean(evaluation) across the population (genetic_algorithm.cpp:162).
Properties of this reward worth naming:
- It is unbounded in time: an agent that simply sits still with clear sensors accrues
reward at roughly the same rate as one that drives. The
1 - (o2*0.2)term is positive for every plausibleo2, so the floor of the per-frame reward is ≈0.8·dt·0.02for doing nothing. Survival, not skill, dominates. o5saturates almost immediately:1 - 1/distanceCoveredis ≥ 0.99 once the vehicle has moved 100 units, so distance stops discriminating very early.- Raw sensor distances (0.01…180) are pushed through hardcoded min/max constants
(
90/400,20/85) that do not match the sensor's actual range (sensorStretch = 180), soo1/o3/o4are frequently negative or > 1. - There is no term at all for the two other skills the game wants to train — algebra answer correctness and beat quality.
5. Defects that make the current network non-functional
These are not stylistic notes. Each one independently breaks the learning loop, and together they mean the neural network currently has no causal influence on agent behaviour whatsoever. They are the reason "improve the agents' ability to learn" cannot start with adding LSTM on top.
5.1 The forward pass never reads its inputs, and never applies the bias
NeuralLayer::processInputs() (ai/neural_layer.cpp:61):
double *biasedInputs = new double[neuronCount + 1];
for (int i = 0; i < neuronCount; i++)
biasedInputs[i] = 0.0; // ← should be inputs[i]
biasedInputs[neuronCount] = 1.0; // bias
for (int j = 0; j < outputCount; j++)
for (int i = 0; i < neuronCount; i++) // ← stops before the bias row
sums[j] += biasedInputs[i] * weights[i][j];
The inputs parameter is never copied into biasedInputs, so every term of the sum is 0 * w. The bias row is written but the accumulation loop bounds exclude it. Therefore sums[j] == 0 for all j, and every layer returns activation(0) — 0.0 for softSign, 0.5 for sigmoid — regardless of sensor readings and regardless of the weights.
5.2 Layers are not chained
NeuralNetwork::processInputs() (ai/neural_network.cpp:40) passes the original inputs to every layer rather than feeding each layer the previous layer's output:
for (int i = 0; i < numLayers; i++) {
double *layerOutput = layers[i]->processInputs(inputs); // ← always `inputs`
...
outputs = layerOutput;
}
Only the last layer's result is returned. The network is effectively a single 3→3 layer with three dead layers attached — and once §5.1 is fixed, this becomes a heap out-of-bounds read, because layer 1 expects 7 inputs and would be handed the 8-element (actually 7-element, see §5.6) input array.
5.3 The genotype→weights mapping collapses to one number
Agent::Agent() (ai/agent.cpp:49):
for (k = layers) for (i = neurons) for (j = outputs) {
std::vector<float> parameters = genotype->getParameterCopy();
for (int p = 0; p < genotype->getParameterCount(); p++)
ffn->layers[k]->weights[i][j] = parameters[p]; // ← overwritten 133 times
}
The innermost loop assigns all 133 parameters to the same weight slot, so each weight ends up holding parameters[132]. Every weight in the network is identical, and 132 of the 133 genes are inert. Crossover and mutation on genes 0–131 cannot change behaviour.
(It is also O(layers · n · m · 133) with a vector copy per weight — but that is the smaller problem.)
5.4 Crossover produces null offspring
GeneticAlgorithm::completeCrossover() takes its two output genotypes by value (genetic_algorithm.h:69, genetic_algorithm.cpp:373):
static void completeCrossover(shared_ptr<Genotype> p1, shared_ptr<Genotype> p2,
float swapChance,
shared_ptr<Genotype> offspring1, // ← by value
shared_ptr<Genotype> offspring2);
The assignments at the end of the function mutate local copies. The caller's offspring1 / offspring2 in randomRecombination() stay null, so every recombined child pushed into newPopulation is a null shared_ptr. If a generation ever did complete, EvolutionManager::startEvaluation() would construct Agents from those nulls and dereference them in the constructor.
5.5 Remainder stochastic sampling shadows its loop index
genetic_algorithm.cpp:203:
for (int i = 0; i < currentPopulation.size(); i++) {
if (currentPopulation[i]->fitness < 1) break;
for (int i = 0; i < (int) currentPopulation[i]->fitness; i++) { // ← inner i shadows outer i
... currentPopulation[i]->getParameterCopy() ...
}
}
The inner loop redefines i, so both the bound and the indexed element refer to the inner counter. The intended "insert floor(fitness) copies of genotype i" becomes a self-referential walk from index 0, which selects the wrong parents and can index out of range. The break on fitness < 1 also means selection stops at the first below-average individual rather than skipping it — with POPULATION_SIZE = 8 and a reward that clusters tightly (§4), the intermediate population is routinely smaller than 2, which trips the badPop guard and freezes evolution outright.
5.6 Input buffer is one element shorter than the input layer
agent_controller.cpp:139 allocates new double[sensors.size() + flags.size()] = 7 doubles, but the input layer is 8 wide. Genotype::setRandomParameters() (genotype.cpp:59) also ignores its minValue/maxValue arguments and always draws from Random(0,1), so initial genomes are strictly non-negative — half the weight space is unreachable at initialisation and only reachable later via mutation.
5.7 The FFN output is multiplied by zero
Even with everything above fixed, AgentController::update() (agent_controller.cpp:373) does:
float ffnW = 0.0f; // weight on the neural network's opinion
float senW = 1.0f; // weight on the hand-written heuristic
controlInputs[k] = controlInputsFFN[k]*ffnW + controlInputsSensors[k]*senW;
controlInputs[1] = 17.0f; // acceleration hardcoded, overriding both
So the vehicles are driven entirely by the hand-coded if/else sensor heuristic in lines 266–354, with throttle pinned at a constant. And that heuristic is itself inert on steering: sensorSteer is initialised to 0 and never assigned (only the unused sensorSteerTarget is), and controlInputsSensors[2] is never initialised at all.
5.7b The movement layer discards steering a second time
Even the heuristic's steering never reaches the car. AgentMovement::applyInput() (ai/agent_movement.cpp:46) repeats the same pattern one layer down:
float w1 = 0.0f; // weight on the agent's steer
float w2 = 1.0f; // weight on the vehicle's waypoint steer
float wyPtSteer = networkActor_->vehicle_->getDesiredSteer();
float steerVar = Random(0.0f, 0.4f);
// float steerControl = ((steerInputNorm_ * deltaTime)*w1) + (wyPtSteer*deltaTime*w2);
float steerControl = (wyPtSteer * deltaTime * w2) * steerVar;
steerInputNorm_ — the value the whole pipeline exists to compute — is dropped, and steering comes from the vehicle's own waypoint follower multiplied by a random number in [0, 0.4]. This matters because it is independent of §5.7: fixing ffnW alone would still have changed nothing about how the cars drive.
Then Vehicle::FixedUpdate() (network/Vehicle.cpp:409) seeds newSteering from the analog steerLevel but immediately overrides it with ±1.0 whenever the LEFT/RIGHT buttons are set — and applyInput() always sets a button when it sets a level. So the final steering signal reaching the physics is full-lock bang-bang, in the direction of a randomised waypoint follower.
setInputs() also normalises the action channel by a running min/max (ai/agent_movement.cpp:163) that is (action_ - action_) / (action_ - action_) on the first call — a 0/0 NaN in actionNorm_ for that frame.
5.8 Copy semantics on the network types are unsound
NeuralNetwork::deepCopy() (neural_network.cpp:80) stores the address of a stack-local NeuralLayer into the new network and returns by value — a dangling pointer plus a leaked layer. getTopologyCopy() heap-allocates and returns *copy, leaking the original and returning a shallow copy that shares layers with a destroyed object. Neither NeuralNetwork nor NeuralLayer follows the rule of three/five while owning raw pointers. Nothing currently calls these two methods, which is the only reason they have not crashed anything.
5.9 Sensors update at 0.5 Hz
Sensor::update() only performs a raycast when lastRaycast > RAYCAST_TIME_WAIT (sensor.h:17 = 2.0 seconds). Between casts, output and hit hold the previous result. At 60 fps, that means the "sensory data" driving the controller is up to 2 seconds stale — at 100 km/h the vehicle has travelled ~55 m since the reading. hit is also latched: setHit(true) is called on a hit but never reset to false on a miss, so once a sensor hits anything it reports "hit" forever.
5.10 Restart discards everything learned
checkGenerationTermination → defaultTermination returns true at generation 100, which calls onGATermination() → restartAlgorithm() → startEvolution(), which constructs a brand-new GeneticAlgorithm with a fresh random population. There is no checkpointing of elite genomes to disk on this path (checkForTrackFinished()'s save block is commented out, evolution_manager.cpp:356), so 100 generations of selection are thrown away.
6. What the agents actually do today
Putting §5 together, the behaviour loop that is really running is:
raycast (every 2s, latched hit flag)
→ hand-written if/else in agent_controller.cpp:266-354
→ steering always 0, throttle always 17.0
→ AgentMovement, which discards steering again and substitutes
vehicle->getDesiredSteer() * Random(0, 0.4)
→ Vehicle, which overrides the analog level with full lock
→ NetworkActor
and, in parallel, a genetic algorithm that:
generates 8 random 133-float vectors
→ maps all 133 to a single weight value per network
→ runs a network whose output is discarded
→ accumulates a reward that mostly measures elapsed time
→ and never completes a generation because nothing ever dies
The GenomeVisualization component that would let you see any of this is fully written (GenomeVisualization.cpp, 630 lines: 3D neuron spheres, weight-coloured connection geometry, a UI stats panel) but is never registered, never attached to a node, and never updated — no file outside itself references it. Its hardcoded topology (9,10,8,3, GenomeVisualization.cpp:31) also does not match the real one (8,7,5,3,3), and UpdateNeuronColors() fills activations with random numbers because the network exposes no way to read them.
It also could not have drawn anything even if it had been wired up. All three of the resources it loads are absent from this project's data:
| Requested | Data/ | CoreData/ |
|---|---|---|
Models/Sphere.mdl | missing | missing |
Materials/UnlitSolid.xml | missing | missing |
Materials/UnlitVertexColor.xml | missing | missing |
CreateNeuron() does not null-check cache->GetResource<Model>(...), so every neuron would have been a StaticModel with a null model and null material — silently invisible.
7. Other domains that exist but are not in the learning loop
The game already has the two non-driving skills the roadmap wants to train, but neither is connected to the GA:
- Algebra.
AlgebraKart.h:621-647: the server picks a shared equation, broadcasts it, clients answer withSubmitEquationAnswer(choiceIndex)(bound to F2–F6), the server scores it inProcessEquationAnswer()and rotates the equation every 45 s. Agents have no input describing the equation and no output that can answer it. - Beat making.
beat/MusicVoiceAgent.{h,cpp}: three algorithmic voices (Aria / Terra / Click) driven byMarkovVoiceGeneratoroff the sharedSequencer's step event. The note choice is a fixed Markov walk — no evolved parameters, no fitness, no feedback from the population.
8. Summary of the baseline
| Property | Current value |
|---|---|
| Model class | Feed-forward, evolved weights, no backprop |
| Topology | 8 → 7 → 5 → 3 → 3 (133 weights) |
| Memory / recurrence | None |
| Population size | 8 (POPULATION_SIZE, evolution_manager.h:100) |
| Generation boundary | "all agents dead" — never reached in practice |
| Generation cap | 100, then full random restart with no checkpoint |
| Selection | Remainder stochastic sampling (index-shadowing bug) |
| Recombination | Uniform crossover, p_swap = 0.6 (returns nulls) |
| Mutation | p = 0.3 per gene, ±2.0 amplitude, elite-2 preserved |
| Sensor rate | 0.5 Hz (2 s raycast interval), latched hit flag |
| Network influence on behaviour | Zero (ffnW = 0.0, and the forward pass is constant anyway) |
| Skills trained | Driving only, and only via a hand-written heuristic |
| Visualisation | Written, not wired up |
The next document, ai-next-architecture.md, describes the target design: a repaired forward pass, an LSTM recurrent core whose weights are still evolved, a redesigned sensory vector, wall-clock 15-minute epochs with checkpointing, per-skill fitness for driving / algebra / beats, and the 3D "mind view" bound to the M key.