Skip to main content
Phase 2 evaluation system

Evidence that keeps its limitations visible.

LLMSlim’s offline, deterministic evaluation separates measured results from derived session totals and work that is not applicable to the released local engine.

Release gatePASS

Offline runner · canonical JSON · independent labels

MeasuredDerivedNot applicable
442tests passing0 failures
92.94%branch coverage90% release gate
22evaluation categories24 curated samples
4languagesEnglish · Hindi · Chinese · Japanese
375generated schemas18 measured catalogs
0provenance violationspass security regression

Small corpus. Explicitly scoped.

The 24 synthetic, independently labelled samples are a transparent regression set—not population-level proof. The default suite exercises extractive compression; provider-dependent rewrite and hybrid results are not reported as local measurements.

Start with the released Python API
Instruction retention93.8%16 applicable labels · measured
Entity retention85.1%19 applicable labels · measured
Lexical Jaccard89.2%24 samples · proxy, not semantic equivalence
Median latency0.026 msp95 1.889 ms · local wall-clock

The schema tax is visible.

Measured synthetic tool contracts show the context cost of resending a whole catalog. This is measurement, not schema compression, selection, or lazy loading.

MeasuredDerived
Catalog tools
Schema complexity

Contracts retain names, parameter types, required fields, enums, nested shape, constraints, and meaningful defaults in the research comparison.

Measured full catalog12,739 tokenstiktoken · complex · 64 tools
Tokens per tool200
16 turnsDerived203,824 tokens
32 turnsDerived407,648 tokens

Why send 64 contracts when the model may need one?

Future research: contract-safe minification, relevance evaluation, selective loading, and lazy schema hydration.Read methodology

4 / 5 checks passed.

The Markdown structural sample is invalid after compression. It is retained as a visible limitation; JSON, XML, YAML, and code checks are not presented as a blanket preservation claim.

Known limitation

Priority cannot be earned by imperative text.

RAG, tool, and assistant context cannot be elevated to protected priority solely by wording. This mitigates compression-induced instruction elevation; it does not claim complete prompt-injection prevention.

0 observed boundary violations

Results exist as checked-in artifacts.

Dataset composition, mode, iteration metadata, classification, and environment details are emitted with the canonical output. Run the repository suite on your own hardware before making performance decisions.

Open the repository

Claims carry their own class.

Measured

Observed from the deterministic suite: retention, ratios, structure, latency, schema catalog tokens, and security regressions.

Derived

Explicit arithmetic from measured catalog tokens, such as 16- and 32-turn totals assuming a full catalog is resent each turn.

Not applicable

Provider-dependent rewrite/hybrid outcomes, dollar savings, and tool-selection quality are not fabricated for the default local suite.