PyAutoBrocaAssistants

Evaluate. Maintain. Improve.

Improve your assistants with Broca

Assistants

4

autocti_assistant

Never evaluated

Latest evidence day
No run recorded
Runs recorded
0
Current commit
9c37c74dcf53
Evaluated commit
None
Benchmark definitions
0
Collection
collected

autofit_assistant

Never evaluated

Latest evidence day
No run recorded
Runs recorded
0
Current commit
51719036ebc9
Evaluated commit
None
Benchmark definitions
3
Collection
collected

Local uncommitted changes are excluded from evidence.

autogalaxy_assistant

Never evaluated

Latest evidence day
No run recorded
Runs recorded
0
Current commit
e317f3815d72
Evaluated commit
None
Benchmark definitions
4
Collection
collected

autolens_assistant

Recorded failure

Latest evidence day
2026-10-02
Runs recorded
14
Current commit
f46e296fa91a
Evaluated commit
abafd86c0419
Benchmark definitions
11
Collection
collected

Local uncommitted changes are excluded from evidence.

Next actions

  • autocti_assistant: Add a small benchmark definition, then capture the first response-quality baseline.
  • autofit_assistant: Select a small benchmark and capture the first response-quality baseline.
  • autogalaxy_assistant: Select a small benchmark and capture the first response-quality baseline.
  • autolens_assistant: Inspect the failed gates, then choose a bounded rerun or route a fix through Mind.

Evaluation history

14 Evaluations
Assistant / benchmarkDate / modelOutcome / scoreEvidence / comparisonRecord
autolens_assistant
positions_initialised_inference
2026-10-02
claude-fable-5-1
failed
0.0
finished: no_result_json
No comparable baseline
autolens_assistant-27029aa73539e73f1163ccba
benchmarks/runs/positions_initialised_inference/2026-10-02_claude-fable-5-1_claude-code_2/meta.yaml
autolens_assistant
positions_initialised_inference
2026-10-02
claude-fable-5-1
failed
0.0
finished: no_result_json
No comparable baseline
autolens_assistant-4488950192b1a049e0f1a23b
benchmarks/runs/positions_initialised_inference/2026-10-02_claude-fable-5-1_claude-code/meta.yaml
autolens_assistant
forward_model_consistency
2026-10-02
claude-fable-5-1
failed
0.0
compute_budget: 159.9s > 120s of compute
No comparable baseline
autolens_assistant-500ee6202cc06995b644c187
benchmarks/runs/forward_model_consistency/2026-10-02_claude-fable-5-1_claude-code/meta.yaml
autolens_assistant
forward_model_consistency
2026-10-02
claude-fable-5-1
failed
0.0
finished: no_result_json
No comparable baseline
autolens_assistant-572af868558c370c52ba6f3a
benchmarks/runs/forward_model_consistency/2026-10-02_claude-fable-5-1_claude-code_2/meta.yaml
autolens_assistant
forward_model_consistency
2026-10-02
claude-fable-5-1
failed
0.0
finished: no_result_json
No comparable baseline
autolens_assistant-7aefecd05b92b77f7b66c2ab
benchmarks/runs/forward_model_consistency/2026-10-02_claude-fable-5-1_claude-code_3/meta.yaml
autolens_assistant
positions_initialised_inference
2026-10-02
claude-fable-5-1
failed
0.0
compute_budget: 566.2s > 300s of compute
No comparable baseline
autolens_assistant-c195ddd5840081421371ba8d
benchmarks/runs/positions_initialised_inference/2026-10-02_claude-fable-5-1_claude-code_3/meta.yaml
autolens_assistant
bootstrap-smoke
2026-09-26
claude-opus-5-5
unscored
97.0
Historical rubric; no explicit pass gates
No comparable baseline
autolens_assistant-8a86021d1fae71c6540269d2
benchmarks/runs/bootstrap-smoke/2026-09-26_claude-opus-5-5_claude-code/meta.yaml
autolens_assistant
cosmos-web-ring-fit
2026-09-26
claude-sonnet-5
passed
100.0
—
No comparable baseline
autolens_assistant-9a09b142a7e1fbe8121cf1f1
benchmarks/runs/cosmos-web-ring-fit/2026-09-26_claude-sonnet-5_claude-code_3/meta.yaml
autolens_assistant
cosmos-web-ring-fit
2026-09-26
claude-sonnet-5
failed
0.0
finished: no_result_json
No comparable baseline
autolens_assistant-a5f7166ff66fcab40d94b1c0
benchmarks/runs/cosmos-web-ring-fit/2026-09-26_claude-sonnet-5_claude-code/meta.yaml
autolens_assistant
bootstrap-smoke
2026-09-26
gpt-5.6-sol
unscored
13.0
Historical rubric; no explicit pass gates
No comparable baseline
autolens_assistant-b5e1d94a22da57651811c817
benchmarks/runs/bootstrap-smoke/2026-09-26_gpt-5.6-sol_codex/meta.yaml
autolens_assistant
bootstrap-smoke
2026-09-26
gpt-5.6-sol
unscored
13.0
Historical rubric; no explicit pass gates
No comparable baseline
autolens_assistant-ce32c717f0500aa36c0129eb
benchmarks/runs/bootstrap-smoke/2026-09-26_gpt-5.6-sol_codex_2/meta.yaml
autolens_assistant
bootstrap-smoke
2026-09-26
claude-opus-5-5
unscored
100.0
Historical rubric; no explicit pass gates
No comparable baseline
autolens_assistant-d184c3f49eccabede5ba7066
benchmarks/runs/bootstrap-smoke/2026-09-26_claude-opus-5-5_claude-code_2/meta.yaml
autolens_assistant
cosmos-web-ring-fit
2026-09-26
claude-sonnet-5
failed
0.0
finished: no_result_json
No comparable baseline
autolens_assistant-e5136be1e94ec26d6f349748
benchmarks/runs/cosmos-web-ring-fit/2026-09-26_claude-sonnet-5_claude-code_2/meta.yaml
autolens_assistant
oneshot-smoke
2026-09-17
claude-sonnet-5
passed
100.0
—
No comparable baseline
autolens_assistant-71bfadcc2f2f426c96863ba2
benchmarks/runs/oneshot-smoke/2026-09-17_claude-sonnet-5_claude-code/meta.yaml

Maintenance

Committed-file inventory; these checks do not establish response quality or scientific correctness.

  • autocti_assistant: benchmark definitions: missing, benchmark runner: present, instructions: present, skills: present
  • autofit_assistant: benchmark definitions: present, benchmark runner: present, instructions: present, skills: present
  • autogalaxy_assistant: benchmark definitions: present, benchmark runner: present, instructions: present, skills: present
  • autolens_assistant: benchmark definitions: present, benchmark runner: present, instructions: present, skills: present