How the vector knows
4,831 company-FYs × 500 tickers → 154 feats cat([x·m,m]) → 20 towers 24-d → transformer 4L4H fusion CLS 128→64-d L2 where cosine = company similarity. Same MTNN cockpit as hoops (12,966 seasons → 48-d) now on SEC EDGAR + yfinance. Purity@10 0.7057 lift 6.32× over random 0.1117, cross-ticker 0.4013, 500 tickers 2015-2024. v4 17 towers → 64-d concat, v6 20 towers → 154 feats → 64-d transformer.
Evaluation — sector coherence gate • engineering QA not predictive
Loading assets/eval_scoreboard.json… 4831 points • same-ticker neighbors allowed unless cross-ticker metric • silhouette -0.0034 vs perm -0.0204 sector overlap heavy but above chance. IC>0 gate proves forward-looking, not just label coherence.
Forward returns vs analyst estimates — isotonic calibration (like hoops preseason Vegas over/unders)
Preseason props analogy → equities forward
Hoops measures beating Vegas preseason over/unders & player props via historical training data. Equities analog: forward returns vs analyst estimates — embedding should encode forward 1M/3M/6M/12M signal above chance, then calibrated isotonic to remove systematic bias. IC 6M 0.007 >0 gate proves forward knowledge not just sector label. Triple-barrier hit 10% before -7% 63d 0.2189 vs random 0.10.
Loading assets/forward_calibration_isotonic.json… (162 thresholds, bias 5.76%→0.00% <1%, IC 0.878→0.881 preserved, y_min -1 y_max 1 clip)
Architecture diagram — truthful boxes mask→towers→fusion→embed→heads
Input 154 feats cat([x·m,m]) where m∈{0,1} mask = 1 if measured that FY, 0 if never-measured (banks missing inventory, early FY missing DEF14A). 20 towers 20×96h GELU LN →24-d + skip ×2 → 20×24=480 +12 FY macro emb=492 → fusion gated 192h Attn+gate or transformer CLS 4L4H 128-d 4H FF512 drop0.12 → CLS 128→512→64 L2 unit sphere → heads archetype 8 / sector 11 / profile 14 / next 14 / skills 12 / valuation / market. ~300K gated ~580K transformer. ONNX WASM ExecuTorch mobile.
154 cat(x·m,m) → 20×(2d_in→160→32) → 480+12=492 → 192h gated / 4L4H CLS 128 → 64 L2 ||v̂||=1 cos=v̂·ŵ → heads 8/11/14/14/12
Training cockpit — what ships now and what trains next
Current deployed — equities_mtnn_v4_concat_d64_b1_lowmem
Truthful deployed model from assets/manifest.json. No fabricated recall/CQS — provenance honestly notes sector-coherence as the only measured geometry QA. Rows 4,051 (3439 tickers incl history) vs live 4,831 (500 S&P). full_history 7,370 tickers • 7,310 active • 60 defunct • 59 pre-1980 • 21 sixties kept in manifest for Explorer search. Wiki embeddings 16-d fused (continuity 0.8 unit-norm).
Hill-climb 01→05 + loss weights
Loss weights (real from ARCHITECTURE.md): archetype 0.25, sector 0.15, profile 0.12, next 0.10, skills 0.20, valuation 0.12, market 0.12, others 0.05-0.08. NCE temp 0.08 same-ticker FY+1 + feature-dropout views drop_p 0.12, same-sector hard-neg boost 0.2-0.3. Hybrid SupCon archetype optional. AdamW bias/LN no decay, OneCycle 10% warmup linear, clip 1.0, batch 512-1024, 60 epochs ~2 min Alienware.
CQS equities mirror hoops: 0.35 recall@10 same-ticker-next-FY +0.25 archetype purity@20 +0.10 sector acc +0.20 next_R2 +0.10 market dir bonus. Gate CQS≥baseline+0.005 recall within 0.02. Baseline target equity synth continuity0.72 recall 0.90+ purity 0.60+ CQS0.65+. Checkpoint upgraded to composite proxy 0.5*recall+0.5*purity fixing early-epoch restore bug.
v6 transformer — exact training command
python pipeline/train_mtnn.py --epochs 60 --dim 64 --tower-width 24 --fusion transformer --n-layers 4 --n-heads 4 --tower-hidden 96 --tower-blocks 2 --drop-p 0.12 --nce-player-weight 0.65 --nce-arch-weight 0.35 --lr 0.001 --lr-schedule onecycle --warmup-pct 0.1 --weight-decay 0.0002 # v6_real_154 variant (20 towers): # python pipeline/train_mtnn.py --epochs 60 --dim 64 --tower-width 24 --tower-hidden 96 --fusion transformer --n-layers 4 --n-heads 4 --features 154 --towers 20 --conditioning gpr_epu_comm
Preprocessing: per-FY z-score winsor ±4 like hoops per-season, cat([x·m,m]) mask where wiki_embeddings always present but distress/altman/piotroski/beneish conditioned on GPR/EPU/comm YoY: x(1+GPR*0.5) etc. FY embedding 12-d learned for macro regime (RATE_10Y, VIX, CREDIT_SPREAD, GDP). Optim: AdamW no decay bias/LN. Scheduler OneCycle warmup 10% linear. Loss: InfoNCE 0.08 + archetype 8-way + sector 11-way + profile 14-d + next 14-d + skills 12 + valuation/market.
How MTNN actually works for companies
154 feats cat([x·m,m]) → 20 towers 96h GELU LN →24-d (+masked flag) + skip ×2 → 20×24=480 concat + 12-d FY macro emb =492 → gated Attn+gate 192h / transformer CLS 4L4H 128-d →128→64 L2 unit sphere. Every box is real dim from
feature_manifest_v6_real.json. v4 17 towers 122-138 feats, v6 20 towers 154 feats incl distress_altman 12, piotroski 10, beneish_quality 10.Company A FY + Company B FY → fuse in 64-d weighted avg (0.65 player-ish ticker continuity +0.35 archetype) → nearest real among 4,831 company-FYs by cosine argmin. Powers equity “Chimera Twin” like hoops Chimera daily puzzle donor A+B → target.
income 11 (REV..EBITDA_MARGIN) • balance 10 (ASSETS..INV_CAP) • cashflow 7 (OCF..CAPEX_REV) • growth 9 • profitability 9 ROE/ROA/ROIC+margins • leverage_liquidity 7 • efficiency 5 • per_share 5 EPS/BVPS • market_price 10 ret1/3/6/12M vol30/90/252 beta mom • valuation 8 PE/PB/EV… • management_neo 14 NEO/CEO/board • ownership 6 INST/INSIDER • disclosure_text 6 MDA/RISK • sector_context 3 • macro_regime 4 • form 6 earn surprise • bbref_bridge 2 AltmanZ/Piotroski proxy • distress_altman 12 • piotroski 10 • beneish_quality 10
v̂ = v/||v||₂, ||v̂||=1, cos = v̂·ŵ. 64-d normalized powers k-NN sector purity@10 0.7057 vs random 0.1117, cross-ticker 0.4013 removes same-ticker trivial inflation. 11 GICS sectors, 8 archetypes on sphere.
Data flow — truthful boxes measured from ARCHITECTURE.md
Input 154 feats cat([x·m,m]) • 20 towers 96h GELU LN →24-d + skip ×2 • SE-style gated 192h or transformer CLS 4L4H 128-d 4H FF512 drop0.12 →CLS 128→512→64 L2 unit sphere → heads archetype 8 / sector 11 / profile 14 / next 14 / skills 12 / valuation / market. Tap any tower/tile.
Loading tower activations, family share per head… 20 towers from feature_manifest_v6_real.json.
Tap input, tower, or head to see exact slice indices, feats, and masked handling. No ticker in X. FY 12-d appended post-tower.
Guess & grade — heads decode 64-d
Same MLP heads as rightmost column. Tap a row to lock trace back through network. Heads share embedding but have independent MLPs 64→128→out.
Archetype — 8-way k-means business model
Sector — 11-way GICS
Financial Craft — 12×(64→32→1) percentile
Next-FY profile — 14-d forecast + profile current
Probe any company-FY
Loading architecture… 20 towers 96→24 ×2 LN GELU skip, 12-d FY macro emb, 480+12=492→192→64 L2 ~300K. Per-FY z-score winsor ±4, cat([x·m,m]) masked where disclosure/management missing pre-2018. No ticker leak in X.
Select a company to see how 154 SEC/market features become a 64-d analog fingerprint. Same ticker adjacency drives recall but cross-ticker purity proves shape, not label.
Compare off — toggle to diff two company-FYs and see Δ in tower activations and cosine.
What drove this prediction — tower attribution
Embedding map — 64-d rendered in 3D PCA • 4,831 points
Slow auto-rotate • drag to orbit • scroll to zoom • green=selected — 4,831 points L2-normalized 64-d → PCA 3. Reuses assets/real_pca.json (x,y,z) like hoops. Archetype colours match Chimera equation.
Loading nearby companies… cosine in 64-d • FY continuity 0.8
From 10-K to fingerprint
01 Gather → 02 Vectors (equities mirror hoops 01→02)
Pull CompanyFacts XBRL US-GAAP + DEF 14A proxy + yfinance 2015-2024. 4,831 company-FYs after filter (500 S&P tickers, gaps). Per-FY z-score like hoops per-season so 2018 and 2024 both µ=0 of their FY. Mask m creates cat([x·m,m]) so model knows never-measured vs zero (banks missing inventory, early FY missing disclosure). SEC conditioning scalars GPR/EPU/commodity YoY.
- income, balance, cashflow, growth 9, profitability 9 (ROE/ROA/ROIC) → 11/10/7/9/9 =46 feats
- leverage 7, efficiency 5, per_share 5, market_price 10 ret/vol/beta/mom, valuation 8 → 35
- management_neo 14 NEO count CEO age/tenure/founder/comp/board, ownership 6, disclosure_text 6 MDA sent/FOG → 26
- sector_context 3, macro_regime 4 RATE/VIX/CREDIT/GDP, form 6 earn surprise, bbref_bridge 2 Altman/Piotroski proxy → 15
- distress_altman 12 Altman X1-X5 Z/Z' distance to default Δ, piotroski 10 F-SCORE, beneish_quality 10 M-SCORE + flag =32 → total 154
03 Train → 04 Export → 05 Deploy (transformer upgrade)
Residual towers 20×96→24 ×2 LN GELU + skip =480, +12 FY =492, gated 492→192→64 L2 or transformer tokens 20×24→128 proj + FY 12→128 + CLS 128 =22 tokens → Transformer 128-d 4L 4H FF512 drop0.12 → CLS 128→512→64 L2. ~300K params gated, ~580K transformer. Heads MLP decode independent but share emb. ONNX WASM + PCA 3 for web. FY embedding guards macro drift like hoops Procrustes drift.json chained RᵀR=I root 1996-97.
Arch spec — manifest.json + feature_manifest_v6_real.json
Loading assets/manifest.json…
Loading assets/feature_manifest_v6_real.json…