Topics: Multi-Agent Systems • Information Theory • Semantic Drift • KL-Divergence • State Space Models
An empirical research framework designed to analyze multi-agent Large Language Model (LLM) conversations as observable, continuous-state dynamical systems.
Rather than relying on artificial heuristics or context-truncation proxies, this lab measures intrinsic non-Markovian dynamics directly from next-token logit distributions, quantifying how much causal weight the deep conversational history exerts on the present output versus the immediate past.
Topics: Multi-Agent Systems • Information Theory • Semantic Drift • KL-Divergence • State Space Models
Natural language reasoning is an open, stochastic process rather than a closed, deterministic Boolean system. When two LLM agents converse without human intervention, they participate in an accelerated process of synthetic social alignment, navigating a high-dimensional semantic space to negotiate shared context.
To evaluate how artificial memory operates during these interactions, we reject artificial perturbations (such as synthetic noise or forced text deletion). Instead, we query the underlying probability distributions of the models to measure explicit information-theoretic metrics as the conversation unfolds.
To measure whether a conversation is genuinely non-Markovian (dependent on deep history) or has degraded into a simple 1-step chain, we evaluate the Kullback-Leibler (KL) Divergence between two probability distributions at every turn
-
$P(x_t \mid x_{\lt t})$ [Non-Markovian Distribution]: The model's probability distribution for the next token given the full conversation history. -
$Q(x_t \mid x_{t-1})$ [Markovian Distribution]: The model's probability distribution for the exact same token given only the immediate preceding turn.
-
High
$D_{KL}$ : The deep context ($t-2, t-3, \dots, 0$ ) is actively constraining current output. The system is operating in a strongly non-Markovian regime. -
Low
$D_{KL} \approx 0$ : The deep context provides no additional predictive utility over$x_{t-1}$ . The conversation has collapsed into a 1-step Markov chain (e.g., small talk, repetitive loops, or complete memory loss).
Using high-dimensional embedding vectors mapped via Principal Component Analysis (PCA), each model response is represented as a coordinate in semantic vector space (
-
Consensus (Convergence): Measured via Inter-Agent Distance Decay:
$$D_t = | A_t - B_t |$$ As models establish a shared context space and align their logic,$D_t$ monotonically decreases. -
Semantic Drift (Random Walk): Measured via Mean Squared Displacement (MSD) relative to the initial prompt (
$A_0$ ):$$MSD_t = | A_t - A_0 |^2$$ Linear or exponential growth in$MSD_t$ signals that the model has lost its original context anchor. -
Incoherent Looping (Limit Cycles): Detected via Step Velocity (
$V_t = | A_t - A_{t-1} |$ ) and spatial Autocorrelation to identify periodic orbits or frozen conversational states.
This framework is built to compare two fundamentally distinct AI memory architectures:
| Property | Transformer (e.g., Llama, GPT-4) | State Space Model (e.g., Mamba) |
|---|---|---|
| Memory Mechanism | Uncompressed KV-Cache ( |
Compressed Recurrent Hidden State ( |
| Theoretical |
Maintains high |
Forced state compression causes intrinsic, continuous |
| Failure Mode | Memory tax (high compute/memory footprint). | Recall tax (accidental compression/erasure of critical historical tokens). |
llm-conversation-dynamics/
├── .devcontainer/
│ └── devcontainer.json # Automated GitHub Codespaces configuration
├── .github/
│ └── copilot-instructions.md # Grounding instructions for GitHub Copilot
├── src/
│ ├── __init__.py
│ ├── agent_loop.py # Core multi-agent autonomous conversation engine
│ ├── non_markovian_kl.py # Logit & KL Divergence extraction engine
│ └── plot_dynamics.py # 3-Panel visualization dashboard (PCA, D_KL, MSD)
├── .gitignore
├── requirements.txt
└── README.md
- Click Code
$\rightarrow$ Codespaces$\rightarrow$ Create codespace on main. - Add your API keys (
OPENAI_API_KEY,ANTHROPIC_API_KEY) to your repository's Codespace Secrets. - Run the baseline multi-agent loop:
python src/agent_loop.py
- Extract explicit non-Markovian dynamics (
$D_{KL}$ ):python src/non_markovian_kl.py
- Generate the analysis dashboard:
python src/plot_dynamics.py
All experimental runs export structured JSON datasets containing token sequences, raw embedding arrays, logprobs, and computed