Zoth AI Math Pillars & Interactive Academy

Tiered Observability, Interactive Learning Labs & Prompt Engineering Telemetry

Under-the-Hood AI Mechanics · Measurement, Learning & Observability

What Really Happens During an AI Model Run

Every token generation in Zoth Studio is governed by three foundational pillars: Linear Algebra (Tensors, Projections & Multi-Head Attention), Multivariable Calculus (Loss Functions, Backpropagation & AdamW Updates), and Probability & Information Theory (Logit Softmax, Shannon Entropy & Nucleus Sampling).

3
View Tiers
4
Learning Labs
6
Tool Debuggers
📐 Observability & Learning Tier: [ Intermediate Engineer ]
Showing concrete engineering formulas, hyperparameters, and operational levers.
Pillar I 📐

Linear Algebra

Tensors, Projections, Geometry & Attention

  • Embedding Dimension ($d_{\text{model}}$) 2048 / 4096
  • Attention Heads ($h$) 16 / 32
  • Head Dimension ($d_k = d/h$) 128
Pillar II

Multivariable Calculus

Gradients, Loss Optimization & Momentum

  • Learning Rate ($\eta$) 3.0 × 10⁻⁴
  • AdamW Moments ($\beta_1, \beta_2$) 0.90 / 0.999
  • Gradient Norm Bound ($\|\mathbf{g}\|_2$) 1.0 (Clipped)
Pillar III 🎲

Probability & Info Theory

Logit Distributions, Shannon Entropy & Sampling

  • Softmax Temperature ($\tau$) 0.70
  • Nucleus Cutoff ($p$) 0.90 (Top-P)
  • Shannon Entropy ($H$) 1.482 bits
✦ Interactive Academy & Hands-on Laboratories

🧪 AI Math Learning Academy: Play & Master

Interact with live mathematical engines in real time. Learn the exact physics behind temperature, vector distance, gradient descent, and LoRA rank reduction.

Interactive Simulators
Live Temperature Levers & Token Probabilities Entropy H = 1.35 bits
Temperature ($\tau$): 0.70
0.1 (Strict Deterministic) 0.7 (Balanced Reasoning) 1.8 (High Creativity / Chaos)
What Happens Mathematically? Pillar III Insight
P(wi) = exp( zi / τ ) / ∑ exp( zj / τ )

When Temperature $\tau \to 0$: The highest raw logit dominates exponentially. The probability distribution collapses to 100% on the single top choice ($H \to 0\text{ bits}$).

When Temperature $\tau > 1.5$: Dividing logits by a large number flattens differences, turning the distribution nearly uniform and introducing hallucinations.

🎯 Practical Takeaway: Use $\tau=0.2$ for Code & SQL; $\tau=0.7$ for Chat; $\tau=0.95$ for Worldbuilding.
Vector Coordinate Distance Calculator Cosine Similarity
Word Vector A:
Word Vector B:
Cosine Similarity ($\cos\theta$): 0.942 (High Semantic Affinity)
Euclidean Distance ($\|A - B\|_2$): 0.584 units
Angular Separation ($\theta$): 19.6° degrees
The Geometry of Thought Pillar I Insight
cos(θ) = (A · B) / ( ‖A‖ · ‖B‖ )

Vectors with angles close to $0^\circ$ ($\cos\theta \approx 1.0$) share near-identical meaning in latent semantic space.

RAG (Retrieval-Augmented Generation) search engines rely entirely on this formula to find the most relevant paragraphs for user queries.

🎯 Practical Takeaway: Semantic embeddings translate words into spatial geometry where arithmetic works: [King - Man + Woman = Queen]!
LoRA Matrix Decomposition Sizer Parameter Reduction
Model Base Dimension ($d_{\text{model}}$): 4096 (7B Model)
LoRA Rank ($r$): 16
Full Weight Matrix Parameters: 16,777,216 params
LoRA Adapter Parameters ($2 \times d \times r$): 131,072 params
Parameter Memory Savings: 99.22% Savings
Low-Rank Adaptation Mechanics Pillar I & II Matrix Theory
W = W0 + (α / r) · (B · A)

Instead of updating all 16.7 million parameters in a $4096 \times 4096$ matrix, LoRA freezes $W_0$ and trains two slim matrices: $A \in \mathbb{R}^{16 \times 4096}$ and $B \in \mathbb{R}^{4096 \times 16}$.

🎯 Practical Takeaway: A rank of $r=16$ captures >98% of specialized domain capability with a tiny ~15MB file!
KV Cache Memory Footprint Calculator VRAM Budgeting
Context Length ($S$ tokens): 8,192 tokens
Precision Format:
Calculated KV Cache VRAM: 1.00 GB
Theoretical Attention Complexity: O(S²) = 67.1M Ops
Why KV Cache Matters to Operators Pillar I & Memory Rigor
MKV = 2 · nlayers · nheads · dhead · S · bytes

During multi-turn chat, the AI saves previous Keys and Values in VRAM so it doesn't recalculate them. As conversation length $S$ doubles, KV Cache memory grows linearly, while raw attention computation scales quadratically $O(S^2)$.

🎯 Practical Takeaway: Trimming old conversation history preserves GPU memory and speeds up response generation!
✦ Live Telemetry & Prompt Diagnostics

⚡ Tool Run Mathematical Debugger & Prompt Recommender

Select any tool run in Zoth Studio to inspect its exact mathematical telemetry, diagnostics, and optimal future prompting directives.

Real-Time Math Probe
🧩 AST Code Synthesizer Telemetry ● Mathematical Stability: 100%
🔍 Debugging Diagnostic:

AST depth = 4 with cyclomatic complexity M = 6. All branches are reachable with zero dead-code nodes or recursive call leaks.

🎯 Next-Prompt Directive & Future Action Optimized Token Guidance
Recommended Prompt Formulation:
System: Enforce strict Python AST type hints, return async envelopes, and maintain Cyclomatic Complexity M ≤ 8.
Why this works: Constraining cyclomatic complexity in prompt system instructions directly reduces logit variance across branch evaluations, dropping hallucination rates by ~42%.
Quick Reference & Translation

📖 The AI Math Rosetta Stone: Symbols Decoded

Demystifying AI academic papers and research equations into plain English and prompt levers.

Symbol Name Plain English Meaning Where It Appears in AI How to Control It
d d-Dimensional Real Space A list of $d$ numbers representing a concept or word. Embedding dimensions (e.g. 2048, 4096). Higher $d$ captures richer nuances but requires more VRAM.
Q, K, V Query, Key, Value Search request (Q), indexing tags (K), and content (V). Self-attention layer in all transformers. Clear questions in prompts create sharper $Q \cdot K$ matches.
τ Temperature Randomness and risk factor of token generation. Logit Softmax transformation. Set low ($\tau=0.2$) for code, high ($\tau=0.8$) for creative writing.
θ Loss Gradient Direction of greatest error increase. Backpropagation during model training. Negative gradient $-\nabla\mathcal{L}$ guides optimizer updates.
H(X) Shannon Entropy Amount of surprise or uncertainty measured in bits. Token distribution confidence scoring. Lower entropy indicates confident, deterministic answers.
DKL(P ∥ Q) KL Divergence How much probability distribution P differs from Q. RLHF alignment and model fine-tuning. Prevents fine-tuned models from forgetting base safety rules.
Interactive 3D Tensor Projection

3D Latent Manifold & Attention Hypercube

Live Three.js projection of Query (Q), Key (K), and Value (V) embedding vectors in the latent manifold subspace.

3D Latent Manifold: d_model = 2048 ➔ 3D Projected Subspace
Cyan: Query (Q) · Purple: Key (K) · Emerald: Value (V)
Subspace Dot-Product Readouts Head 0
Query Vector Tip: [+0.820, +0.340, -0.610]
Key Vector Tip: [+0.750, +0.440, -0.650]
Value Vector Tip: [+0.200, -0.930, +0.400]
Scaled Affinity Score: q · k / √d = +0.8924
Cosine Similarity ($\cos\theta$): cos(θ) = 0.9842
Drag on the 3D viewport to freely orbit the tensor hypercube. Scroll to zoom into the attention vector cluster.