← All design docs

Turiya Demo: "Maya Space" — Collaborative Learning System

Status: partially implemented. MusicVoiceAgent, MarkovVoiceGenerator, and VoiceVisualization exist and work locally (§2/§3/§3.1/§5) — Aria/Terra/Click are algorithmic-only in this pass, per the code's own note, human seat-claim (§2, §6, §8) isn't wired up yet. Multiplayer networking (§4 BPM sync, §6 CSP messages/session state) and the persistence/replay extension (§7) are still design-only — MSG_CSP_BPM_CHANGE doesn't exist in network/csp/CSP_messages.h yet. Engine: Turiya (Urho3D fork + SLikeNet/libdatachannel WebRTC transport + Emscripten browser client). Builds on the existing beat/ synth stack and network/ CSP stack in algebrakart/v1/Source/Game/AlgebraKart/.

1. One-liner

Four voices share one room and one clock. A human sits down at a drum machine; three algorithmic agents each hold a synthesized voice (high lead, low bass, metronome click). Every sound each voice makes grows a 3D wireframe structure out of the ground beneath it — a tube whose spine is straight and whose ring radii are shaped by that voice's waveform amplitude and by a Mandelbrot escape-time field. The whole room shares one BPM, and the human can push it up or down live.

2. Session & roles

One session = one Scene, one shared Sequencer clock (beat/Sequencer.h), four voice slots:

SlotOccupantRoleRegister
0Human (main player)Drum machine — steps a 16-step pattern across 3 sample channelsPercussive / broadband
1Agent "Aria"High-pitch synth lead~880–3520 Hz
2Agent "Terra"Low-pitch synth bass~40–160 Hz
3Agent "Click"Metronome-type instrument, plays strictly on the beatShort transient, fixed pitch

This maps directly onto Recorder::CreatePattern(Beat* channel1_, Beat* channel2_, Beat* channel3_, ...) and world1.ms_pattern in db/schema.sql, which already models "three rows per captured beat tick, one per channel" — that schema was written for exactly this 3-non-drum-voice shape, we're just finally filling it.

Slots 1–3 are human-fillable from v1, not an AI-only permanent fixture: each is a seat, not a species. At session start, any unclaimed slot among 1–3 runs its Markov-chain pattern generator (§3.1) as the default occupant; a second/third/fourth human joining the session claims a slot the same way slot 0 is claimed, and that voice switches from algorithmic mode to direct human input (§6, "Actors"). Both occupancy modes reuse the same NetworkActor/ClientObj plumbing and the same Synthesizer/register — only where the next note comes from changes. Treat "agent" here as "seat occupant, algorithmically-driven by default," not "AI system" — the genetic/neural Agent framework already in ai/agent.h is a plausible future backing for the pattern generator (evolve melodic genotypes instead of a hand-tuned Markov table), but is overkill for the demo — see §11 for how that framework's actual per-frame decision structure compares to the Markov approach used here.

3. Sound design

All four voices are Synthesizer instances (beat/Synthesizer.h) pushing samples into a BufferedSoundStream, played through SoundSource3D so they are spatialized at each voice's pedestal (§6). Reuse osc1_/osc2_ + filter_/accumulator_ already on Synthesizer; don't add a second synth engine.

  • Drum machine (human) — Sequencer in "player-driven" mode: the human toggles steps across 3 sample channels (kick/snare/hat-equivalent) via Sampler, same as the existing Beat/Sampler step data. This is the one voice that plays pre-recorded samples rather than a live oscillator — keep that distinction, it's the percussive anchor the other three lock to.
  • Aria (high) — two detuned oscillators (osc1_/osc2_) an octave+ apart in the 880–3520 Hz range, short envelope, note choice driven by the shared Markov pattern generator (§3.1) walking a fixed scale.
  • Terra (low) — single sine/saw oscillator at 40–160 Hz, longer envelope, plays on beat subdivisions less frequently than Aria (e.g. once per bar or every other beat) so the register separation stays audible; same Markov generator, sparser transition table.
  • Click (metronome) — very short, fixed-pitch transient (think woodblock/rimshot), fires exactly once per beatTimeStep_ tick, always — this voice's whole job is to be the audible, visible heartbeat of the room, so it's the one voice that does not use the Markov generator (its "pattern" is fixed: every tick, no variation).

3.1 Pattern generation: Markov chain seeded by BPM

Aria and Terra each own a small transition matrix over scale-degree states (Aria: full scale, dense transitions; Terra: root/fifth/octave-heavy, sparse transitions). The RNG driving state transitions is reseeded from bpm_, not free-running:

seed = hash(bpm_, bar_, voiceId)

recomputed once per bar. Consequences worth designing around:

  • Deterministic across clients: any two clients that agree on bpm_ and bar_ (already kept in sync via MSG_CSP_STATE, §4/§6) independently regenerate the identical note sequence for an algorithmic voice, with zero per-note network traffic. This extends the "don't send geometry, send tick parameters" argument in §5.5 one level further: for algorithmic voices, not even note-on events need to cross the network — only the shared clock does.
  • Tempo changes audibly reshape the pattern, not just its speed — a BPM push changes seed, so Aria/Terra's melodic material shifts at the next bar boundary alongside the tempo. This is a feature (the human's BPM control becomes a compositional lever, not just a playback-speed knob), but flag it in playtesting in case it reads as "the melody randomly changed" rather than "I changed that."
  • Human-controlled mode (§2, §6): when a human claims slot 1 or 2, the Markov generator for that voice stops being consulted — input maps directly to scale-degree triggers (same scale/transition-matrix states, just picked by a button press instead of a weighted random draw), so the register and note vocabulary stay identical between algorithmic and human control. Click (slot 3) has no meaningful "human mode" beyond manually confirming/tapping the beat — low priority, can stay algorithmic-only if scope is tight.

4. Tempo & BPM

Sequencer already initializes bpm_ = 90 (beat/Sequencer.cpp:174) — that's not a coincidence we need to introduce, it's the existing default, confirmed correct.

  • Sequencer becomes session-scoped, not per-voice: one shared instance computes beatTimeStep_ from bpm_/beatsPerBar_, and all four Synthesizers (drum machine included) read ticks from it. Today Sequencer owns one Synthesizer — extend it to drive N registered voices off the same beat clock instead of adding N sequencers.
  • Only the human's slot (slot 0) may change bpm_. Suggested range 40–220, step ±1 on tap/hold, ±5 on a modifier. Reject/clamp out-of-range requests server-side, not just in UI.
  • Sync model: the server (or session host, in P2P topology) owns the authoritative bpm_/beat_/bar_/currTime_. A BPM change is a small reliable message broadcast over the existing CSP channel (network/csp/CSP_messages.h) — add MSG_CSP_BPM_CHANGE alongside the existing MSG_CSP_INPUT (32) / MSG_CSP_STATE (33), e.g. 34, carrying {newBpm, effectiveAtBeat}. Clients apply it at the next bar boundary, not instantly, so playback doesn't audibly hiccup mid-bar. Regular MSG_CSP_STATE snapshots continue to reconcile beat_/bar_ drift the same way position/velocity drift is reconciled today.
  • Display: a shared BPM readout (UI text, not per-player) visible to all four participants, ticking visibly on each beat (flash/pulse) so latency in the numeric readout doesn't matter as much as the visual pulse being locally smooth.

5. Visualization: the Mandelbrot waveform spike

This is the centerpiece and the least like anything already in the codebase — GenomeVisualization.cpp gives us the CustomGeometry pattern to build on (CreateComponent<CustomGeometry>() → BeginGeometry(0, TRIANGLE_STRIP) → DefineVertex(...)), but nothing in-repo currently does audio-driven or fractal-driven mesh generation — this is genuinely new work, not a reskin.

5.1 Shape

Per voice, one mesh: a straight 3D spine (a line segment anchored at that voice's ground pedestal, extending upward or outward — pick one consistent axis per pedestal, e.g. radially outward from room center so all four structures are legible from a central vantage point) with a ring of vertices around the spine at each sample point, forming a tube. The ring radius at each point is what "draws" the waveform — a quiet moment is a narrow neck, a loud transient is a wide bulge, and the fractal term makes the bulge's silhouette spiky/organic instead of a smooth lathe.

pedestal ── spine (straight 3D vector, fixed direction) ──────────▶
            ring0   ring1   ring2   ring3   ...   ringN (newest)
             ○        ◯        ◎        ⊛              ← radius = f(amplitude, mandelbrot)

5.2 Data feeding the mesh

Each voice already produces what's needed, mostly via existing types:

  • Time-domain amplitude — peak or RMS of the samples a Synthesizer pushed to its BufferedSoundStream this tick.
  • Frequency-domain data — AudioSpectrumAnalyzer (beat/AudioSpectrumAnalyzer.h, muFFT-backed) already computes an FFT spectrumView from a sample buffer. Reuse it per-voice instead of building a second FFT path; feed it the same buffer segment the amplitude value above was measured from.

Bundle both into a small per-tick struct, e.g.:

struct VoiceVizTick {
    unsigned voiceId;       // 0..3
    float amplitude;        // normalized 0..1
    float bandEnergy[8];    // 8 coarse spectrum bins, normalized 0..1
    unsigned tickSeed;      // deterministic RNG seed for this tick (see 5.5)
};

5.3 Mandelbrot modulation

For a ring at normalized position t along the spine (0 = oldest visible segment, 1 = newest), map (t, amplitude) into the complex plane and run the standard iteration:

c = center + zoom * complex(2*t - 1, amplitude * yRange)
z0 = 0
z_{n+1} = z_n^2 + c, until |z| > 2 or n == maxIter
escape = n / maxIter   // 0..1, 0 = inside the set (fully "spiky"/dense), 1 = far outside (smooth)

center and zoom are fixed per voice (Aria/Terra/Click/drum each own a different, fixed window into the Mandelbrot set, chosen for visual variety — this is a tuning constant, not something computed at runtime).

Ring radius:

radius(t) = baseRadius
          * amplitude              // waveform amplitude drives overall size
          * (1 + spikeGain * (1 - escape))   // mandelbrot term drives spikiness

For per-vertex (not just per-ring) organic detail, perturb each of the segmentCount vertices around a ring by re-running the same escape-time function with the angular position folded in:

c_vertex = c(t) + microZoom * complex(cos(angle), sin(angle)) * bandEnergy[angle_to_band(angle)]

i.e. each of the 8 bandEnergy bins governs a wedge of the ring's circumference, so a voice with strong high-frequency content gets visibly spikier geometry on that side of the tube, while amplitude alone governs overall girth. This is the concrete mechanism behind "mixed with the waveform amplitude or other frequency domains" — amplitude scales the tube, frequency bins steer where the fractal spikes stick out.

5.4 Growth over time

The tube grows forward, it isn't redrawn from scratch:

  • Each VoiceVizTick (one per Sequencer beat subdivision the voice actually emits sound on — silent ticks emit no new ring, so the tube's ring density itself shows activity) appends one new ring at the far end of the spine and advances the spine tip forward by a fixed segment length.
  • Keep a bounded ring-count window per voice (e.g. 256 rings) as a circular buffer — oldest ring is dropped as a new one is appended, so the mesh doesn't grow unbounded over a long session. Rebuild the CustomGeometry from the live window each time a ring is added/dropped (cheap: TRIANGLE_STRIP between consecutive rings, same call pattern as GenomeVisualization::conn.geometry->BeginGeometry(0, TRIANGLE_STRIP)).
  • Fade older rings toward the trailing edge of the window (vertex color alpha, or shrink toward baseRadius*amplitudeFloor) so the structure reads as "growing forward out of the ground" rather than "an infinite static pipe."
  • On a fresh session / voice reset, the spine retracts to zero length and regrows — useful as the "round start" beat.

5.5 Networking the visualization (don't send geometry)

Send the compact VoiceVizTick struct over the network, not vertices — each client regenerates identical geometry locally from the same deterministic Mandelbrot function and the same tick data, the same way client-side prediction already works for actor state in network/csp/. tickSeed exists only if any non-deterministic jitter is added for texture; if the mesh generation stays a pure function of (t, amplitude, bandEnergy), drop the seed field entirely. This keeps visualization sync cheap regardless of segmentCount/ring density, and keeps the four structures perfectly consistent across all viewers without a spectator ever needing the authoritative audio stream.

5.6 Placement

Four pedestals arranged around the room (circle or square), one per voice, each spine oriented up-axis (straight up from its pedestal) — confirmed by an in-engine prototype comparing up-axis against radial-outward; up-axis reads more clearly as "growing" and avoids the four tubes visually colliding mid-room. Drum machine pedestal is where the human's own avatar/rig stands; Aria/Terra/Click pedestals hold a simple idle-animated marker (no full character rig needed for the demo) unless claimed by a human (§2, §8), in which case that pedestal's marker is replaced by the claiming player's rig.

6. Networking & session architecture

Reuse, don't replace:

  • Transport: SLikeNet fork over WebRTC (native server via libdatachannel
  • Civetweb, browser client via Emscripten) — already proven end-to-end in algebrakart-engine/webrtc-echo-demo. No new transport work needed.
  • Actors: each of the 4 slots is a NetworkActor/ClientObj. Slot 0 is always driven by real Controls input (extend the existing NTWK_CTRL_* bitmask in network/NetworkActor.h with e.g. NTWK_CTRL_BPM_UP / NTWK_CTRL_BPM_DOWN and step-toggle controls for the drum pattern). Slots 1–3 are mode-switchable per seat: unclaimed, they're server-simulated by the Markov generator (§3.1) and merely replicated, same as any AI-driven NetworkActor today; claimed by a joining human, they switch to real Controls input (scale-degree trigger bitmask, same idea as NTWK_CTRL_* but mapped to Aria/Terra's transition-matrix states instead of movement/drum steps) and the server stops consulting the Markov generator for that voice until the seat is released. Seat claim/release is a small reliable message (join session already has an equivalent handshake for slot 0; extend it to carry a slot index instead of assuming slot 0).
  • Session state sync: extend MSG_CSP_STATE payloads (or add a sibling message) to include {bpm_, beat_, bar_} and each voice's latest VoiceVizTick, piggybacking on the snapshot cadence that already exists rather than inventing a new tick rate.
  • New CSP message: MSG_CSP_BPM_CHANGE (see §4) for the low-frequency, reliable BPM-change event, separate from the high-frequency unreliable per-tick viz data.

7. Persistence & replay (stretch, but cheap given what exists)

Recorder/world1.ms_pattern/world1.ms_time_code already capture 3-channel beat data to Postgres via ODBC (beat/Recorder.cpp, db/schema.sql). For this demo:

  • Reuse as-is for the drum machine channel data.
  • If replay of the full room (not just the drum pattern) is wanted, extend ms_pattern with a voice_id column and store each voice's VoiceVizTick alongside its sample index — since the mesh is a pure function of tick data (§5.5), persisting ticks is sufficient to deterministically regenerate the entire visual session on playback. ReplaySystem/ReplayUI (ReplaySystem.cpp, ReplayUI.cpp) already exist as the playback shell; this demo would be their first real multi-voice content.
  • Out of scope for a first pass: treat as a natural v2, don't block the demo on schema migration.

8. Controls & UX

  • Human (slot 0): existing movement/interact controls to walk up to the drum machine and "sit," then step-sequencer input (toggle steps per channel) + BPM up/down.
  • Human (slots 1–2, optional): walk up to the Aria/Terra pedestal and "claim" it the same way slot 0 is claimed; while held, scale-degree triggers replace the Markov generator for that voice (§3.1, §6). Releasing the seat hands the voice back to the algorithmic generator.
  • Spectator (any unclaimed additional player): free camera around the room to view all four growing structures.
  • Shared BPM readout, pulsing on-beat, visible to everyone in the session regardless of slot.

9. Implementation plan

  1. Shared clock: lift Sequencer from "owns one Synthesizer" to "drives N registered voices"; confirm bpm_ default and clamp range; wire human BPM controls through NTWK_CTRL_* → MSG_CSP_BPM_CHANGE.
  2. Three algorithmic voices: stand up Aria/Terra/Click as Synthesizer instances; implement the Markov transition tables + bpm_/ bar_-seeded RNG (§3.1) for Aria/Terra, fixed on-tick trigger for Click; confirm they sound distinct (register + envelope) before touching visuals.
  3. Seat claim/release: extend the join handshake to carry a slot index (0–3) instead of assuming slot 0; wire scale-degree Controls input for a claimed Aria/Terra seat as an alternate note source to the Markov generator, gated on claim state. Can land after step 2 lands and sounds right algorithmically — human-control is an input-source swap, not a new sound path.
  4. Per-voice viz tick pipeline: wire AudioSpectrumAnalyzer + peak/RMS amplitude into a VoiceVizTick emitted per beat subdivision per voice.
  5. Mandelbrot mesh generator: standalone, testable function (t, amplitude, bandEnergy[8]) -> ring vertices; unit-test the escape-time math in isolation before wiring to CustomGeometry. Land with placeholder constants (ring-window size, maxIter) and a debug overlay showing mesh rebuild cost per tick — real values come from the profiling pass in step 9, not from guessing now.
  6. Growing-tube component: new VoiceVisualization component (sibling to GenomeVisualization, same CustomGeometry idiom) that consumes VoiceVizTick events and maintains the bounded ring window.
  7. Network wiring: extend CSP messages/snapshot payload; verify all four structures render identically on two clients from the same tick stream.
  8. Room + pedestals: level dressing, camera, shared BPM readout UI.
  9. Perf profiling pass (WASM/browser client target): sweep ring-window size and maxIter against actual per-tick mesh-rebuild cost (step 5's debug overlay) across the four simultaneous voice meshes; lock final constants from measured data, not the placeholder values.
  10. (stretch) Recorder/replay extension per §7.

10. Open questions

  • Resolved — pattern generator: Markov chain seeded by bpm_/bar_, see §3.1.
  • Resolved — human-fillable slots: yes, slots 1–3 are human-fillable from v1 via seat claim/release, see §2, §6, §8, and implementation step 3.
  • Resolved — tuning constants: ring-window size and maxIter are not fixed in this doc; they're set empirically via the dedicated profiling pass (implementation step 9), against the real per-tick mesh-rebuild cost on the WASM/browser client target, once steps 5–6 exist to measure.
  • Resolved — spine orientation: up-axis, confirmed by in-engine prototype comparing it against radial-outward, see §5.6.

Algebra Kart's racing bots make a comparable kind of real-time, per-frame decision — worth comparing directly since it's the other live "AI agent" system in this codebase, and it's tempting to assume the two share machinery. They don't, on purpose.

What the steering pipeline does. Every frame, each bot blends two signal sources: a small feed-forward neural network (8 raycast/state inputs, trained across generations by the genetic algorithm in ai/) and a hand-tuned reactive heuristic (steer away from a blocked side, slow if boxed in), weighted network × 0.3 + heuristic × 1.0 — the heuristic leads, the network refines. That blended signal then either corrects a waypoint-following baseline (a valid track waypoint exists this frame) or takes full steering authority (fallback — no waypoint found), before being eased into final steer/throttle. Full diagram: Architecture section, monkeymayastudios.com.

Why Aria/Terra/Click don't use this. §2 already flagged ai/agent.h's GA/NN framework as a plausible future backing for the voice pattern generator, and called it overkill for this demo. The actual implementation (MusicVoiceAgent/MarkovVoiceGenerator) confirms that call was right for now: voices need note choice that's deterministic and network-cheap — every client independently regenerating the identical sequence from (bpm, bar, voiceId) with zero per-note traffic (§3.1) — not an evolved policy network, whose output isn't naturally cheap to keep identical across clients the way a reseeded Markov chain is.

Where they're the same shape. Both systems are "an agent blends multiple signal sources into one decision, every tick": sensors + network + heuristic for a kart; shared-clock-seeded RNG for a voice. If Maya Space voices ever do move toward evolved behavior, the steering pipeline's blend, then fall back structure is the closer template to reach for than starting from scratch — recent room/session state standing in for raycasts, an evolved network standing in for the kart's network, and the Markov chain demoted to playing the same reactive-fallback role the heuristic plays today.