LAURENCE FANG
← RETURN TO EVIDENCE

01 / RESEARCH / EVALUATION

LLM Skills & Metacognition

Per-skill evaluation of operational Modus Ponens capability.INDEPENDENT WORK / 01

QUESTION

Can a reusable reasoning procedure be measured directly, and how does direct skill performance relate to downstream tasks that require that procedure plus additional requirements?

METHOD

  1. 01Screened and deep-read 9 approved papers before design.
  2. 02Built 60 direct Modus Ponens probes and 60 downstream tasks per model.
  3. 03Used deterministic scoring and preserved raw responses for review.

VERIFIABLE EVIDENCE

  • 120 deterministic evaluation items: 60 direct probes and 60 downstream tasks per model.
  • GPT-5.6 Luna: 60/60 skill, 60/60 task. DeepSeek V4 Flash: 58/60 skill, 42/60 task. Qwen 3.7 Max: 60/60 skill, 50/60 task.
  • Observed downstream failures were attributed separately instead of being treated as proof that the target skill was absent.

INTERPRETATION BOUNDARY

This evaluates an operational reasoning procedure across equivalent representations; it does not claim to isolate a latent internal skill module.

TECHNICAL IMPLEMENTATION

PythonJSONLdeterministic scoringevaluation design