01 / RESEARCH / EVALUATION
LLM Skills & Metacognition
Per-skill evaluation of operational Modus Ponens capability.INDEPENDENT WORK / 01
QUESTION
Can a reusable reasoning procedure be measured directly, and how does direct skill performance relate to downstream tasks that require that procedure plus additional requirements?
METHOD
- 01Screened and deep-read 9 approved papers before design.
- 02Built 60 direct Modus Ponens probes and 60 downstream tasks per model.
- 03Used deterministic scoring and preserved raw responses for review.
VERIFIABLE EVIDENCE
- 120 deterministic evaluation items: 60 direct probes and 60 downstream tasks per model.
- GPT-5.6 Luna: 60/60 skill, 60/60 task. DeepSeek V4 Flash: 58/60 skill, 42/60 task. Qwen 3.7 Max: 60/60 skill, 50/60 task.
- Observed downstream failures were attributed separately instead of being treated as proof that the target skill was absent.
INTERPRETATION BOUNDARY
This evaluates an operational reasoning procedure across equivalent representations; it does not claim to isolate a latent internal skill module.
TECHNICAL IMPLEMENTATION
PythonJSONLdeterministic scoringevaluation design