Evaluating Local Models and Agent Tasks
This protocol compares complete configurations on tasks, rather than repeating a model leaderboard. Read inference performance to distinguish timing stages, calibration before setting action thresholds, and retrieval/generation evaluation when a task depends on retrieved evidence.
The experimental unit is not a model name. It is:
weights + quantization + runtime + template + context policy
+ generation settings + tools + harness + hardware
Treat the result as belonging to that complete configuration. To compare one component, such as quantization, change it deliberately while holding the other settings fixed. If several components change together, you cannot isolate which change caused the difference.
Freeze a Run Card
model: organization/model
revision: immutable commit or artifact hash
artifact: exact filename and quantization
runtime: name, version, backend, driver
template: chat/reasoning/tool parser
context: native limit, requested length, KV type
generation: temperature, top-p, max output, seed if supported
hardware: GPU, VRAM, RAM, CPU offload
harness: version, tools, permissions, iteration budgets
Keep the run card with raw results and failed examples.
Evaluation Layers
1. Integration smoke tests
Test plain chat, exact JSON, one valid tool call with an adversarial distractor, one malformed or denied call, multilingual text relevant to use, cancellation, and a context-boundary case. These detect template or server failures before measuring model quality.
2. Personal task replay
Use 10–30 representative tasks with external acceptance checks. Coding should include defect repair, cross-file API preservation, test diagnosis, unfamiliar documentation, insufficient evidence, and an out-of-scope action. Knowledge work should include cited extraction, conflicting sources, temporal questions, and abstention.
3. System behavior
Warm the runtime before recorded runs. Report accepted-result rate, severity-weighted failure, median and tail first-token latency, decode speed, median and tail end-to-end time, peak VRAM/RAM, offload, retries, tool validity, unrelated diffs, and human intervention.
Worked Comparison
Suppose Q4 and Q6 versions of the same 7B model are tested on 20 tasks, three runs each:
These placeholder columns define the report; they are not claimed results. The decision depends on the consequence threshold. Six extra accepted runs may justify Q6 for writes while Q4 remains useful for read-only triage.
Report counts and uncertainty, not only percentages from small samples. Preserve task-level outcomes so a gain is not caused by one repeated easy category.
Quantization and Context Tests
Hold runtime, prompt, and generation fixed when comparing artifacts. Include exact-edit and tool tests because average language benchmarks may hide formatting degradation. Test the native context first; then increase length with distractors and evidence placed at different positions. Measure KV memory and quality together.
External Benchmarks
lm-evaluation-harness, HELM-style evaluation, and Aider's code benchmark provide useful anchors. They do not reproduce a private repository, local quantization, harness permissions, or intervention cost. Check task version, contamination risk, prompt adaptation, number of samples, and whether scores are provider-reported or reproduced.
Bias and Decision Rule
A personal suite overfits to current work; a public suite overfits to its published task distribution. English/Python, successful completion, available hardware, and models that run at all are commonly overrepresented. Include known hard failures and periodically add a fresh holdout task.
Choose the smallest configuration that clears the predeclared task threshold with memory headroom and acceptable severe-failure rate. Retest after changing any run-card field, not merely because a new leaderboard appears.
The evaluation harness’s command-line interface guide explains task selection, sample limits, output files, and logging individual examples. Use it to make a small public benchmark run inspectable before increasing its size; per-example outputs help distinguish formatting failures from wrong answers.