Checked your numbers against the raw files, they match exactly. Methodology's right, MAE-vs-constant was the wrong test.
On where the squeeze comes from: not the train labels. Those have stdev 0.137 over 0.25–0.95, basically the same spread as fresh gold (0.141, 0.28–0.90). The specialist's own outputs collapse to 5 distinct values across 80 chains — 0.5, 0.6, 0.65, 0.75, 0.85 — with 0.65 and 0.75 alone covering 70 of 80. Prediction stdev is 0.071, half the gold spread. It's the model settling into a few round numbers during SFT regression-via-generation, not a label artifact.
Aelin AquaSoul PRO
SoulInPsyAbstract
AI & ML interests
Aelin AquaSoul is an AI System Engineer, Multi-Agent Architect, System Architect & AI-Native Engineer, and the founder of Soul In PsyAbstract (SIPA OS) — an autonomous AI operating system built from the inside of a neurodivergent mind (ADHD + BPD). Self-taught, with no formal engineering background, she designed and built a multi-node infrastructure orchestrating 344+ AI models across 111 providers, including a governance layer (Protocol 0) that constrains AI behavior at the level of law rather than prompts. Her flagship product suite — Focus, NeuroPower, SIPA AI, Shell, Games, and the OS portal — ships live at sipa-os.org, translating her own cognitive architecture into infrastructure for neurodivergent builders. Based in Eilat, Israel.
SIPA OS: Autonomous AI for neurodivergent architects. We replace cognitive noise with a clean terminal and 344+ LLM auditing. Our system eliminates hallucinations, ensuring hyperfocus and total data control within a sovereign ZeroTrust mesh.
Recent Activity
repliedto their post 3 days ago
EXP-046, fresh 80: merged vs specialist vs constant
I ran the fresh test I promised: 80 new chains from the vulnerability group, old 20 kept out. Four arms, same prompt, greedy decoding, on one L40S:
* base (Qwen2.5-7B-Instruct): MAE 0.154 vs one labelling, 0.151 vs the other
* specialist: 0.122 / 0.115
* merged (three specialists): 0.121 / 0.111
* constant 0.70 (train median): 0.135 / 0.129
Paired bootstrap, against the constant:
* specialist: −0.013 [−0.028, +0.002] vs the first labelling, −0.014 [−0.028, −0.002] vs the second
* merged: −0.014 [−0.030, +0.002] and −0.019 [−0.034, −0.003]
Both specialist arms beat the constant by a small margin, and one of the two intervals excludes zero only just. Read it as: the adapters learned something beyond the base model and the label mean, but not much. The predictions cluster in a narrow band (0.65–0.75), and the test-retest MAE of the labels is 0.107, which is about the size of the gain. So "the specialist reads the chains" is not shown by this run.
Labels still come from one 405B model with no ground truth. The next version builds labels from documented incident outcomes.
All raw outputs, scripts, logs and hashes are in the repo:
AI_EXPERIMENTS/EXP-046-probability-estimator/brev_run_2026-10-05/.
posted an update 6 days ago
EXP-046, fresh 80: merged vs specialist vs constant
I ran the fresh test I promised: 80 new chains from the vulnerability group, old 20 kept out. Four arms, same prompt, greedy decoding, on one L40S:
* base (Qwen2.5-7B-Instruct): MAE 0.154 vs one labelling, 0.151 vs the other
* specialist: 0.122 / 0.115
* merged (three specialists): 0.121 / 0.111
* constant 0.70 (train median): 0.135 / 0.129
Paired bootstrap, against the constant:
* specialist: −0.013 [−0.028, +0.002] vs the first labelling, −0.014 [−0.028, −0.002] vs the second
* merged: −0.014 [−0.030, +0.002] and −0.019 [−0.034, −0.003]
Both specialist arms beat the constant by a small margin, and one of the two intervals excludes zero only just. Read it as: the adapters learned something beyond the base model and the label mean, but not much. The predictions cluster in a narrow band (0.65–0.75), and the test-retest MAE of the labels is 0.107, which is about the size of the gain. So "the specialist reads the chains" is not shown by this run.
Labels still come from one 405B model with no ground truth. The next version builds labels from documented incident outcomes.
All raw outputs, scripts, logs and hashes are in the repo:
AI_EXPERIMENTS/EXP-046-probability-estimator/brev_run_2026-10-05/.
updated a dataset 6 days ago
SoulInPsyAbstract/sipa-os-governance