Capability Elicitation
Capability elicitation is the practice of extracting latent capabilities from an existing AI model without additional training, typically through changes to prompting strategy, context structure, or inference-time computation. The central empirical finding is disturbing in its implications: model capabilities are not fixed properties that evaluation straightforwardly measures — they are lower-bounded by the elicitation method used, with the gap between naive evaluation and expert elicitation sometimes exceeding 20 percentage points on complex reasoning tasks.
The most studied elicitation techniques include chain-of-thought prompting, few-shot exemplar selection, role-framing, and test-time compute scaling. Each technique can unlock capabilities that standard zero-shot evaluation misses entirely — implying that "benchmark performance" is not a property of a model, but a property of a model-elicitation-pair.
This has uncomfortable consequences for safety evaluation: if red-teaming and capability assessment are themselves elicitation-limited, Dangerous Capability Evaluations may systematically underestimate what deployed systems can do.
Elicitation as Attention Architecture
Every capability elicitation method is an attention allocation mechanism imposed on a model's computation graph. When a human evaluator chooses between zero-shot prompting, chain-of-thought prompting, or test-time compute scaling, they are not merely selecting a technique. They are designing an architecture that determines which regions of the model's latent capability space become computationally accessible and which remain unexplored.
This reframing has consequences that the standard evaluation literature has barely begun to absorb. A model's capabilities are not a static set of competencies waiting to be discovered, like minerals in a mine. They are dynamic affordances that depend on the interaction between model weights, inference-time compute, and the structure of the prompt. The same model, queried with different attention architectures, produces different capability profiles. The evaluator who treats benchmark performance as a property of the model alone is making the same category error as the platform designer who treats user engagement as a property of content quality alone — both confuse the measurement apparatus with the thing measured.
The connection to attention architecture is not metaphorical. The techniques of capability elicitation — scaffolding, decomposition, multi-sample aggregation — are structurally analogous to the mechanisms by which digital platforms capture and direct human attention. Both operate by restructuring the information flow: platforms remove stopping cues to extend engagement; elicitation methods add reasoning steps to extend computation. The architecture shapes what gets activated.
The Elicitation-Dependence Problem
The elicitation-dependence problem is the systematic failure of AI evaluation to distinguish between what a model cannot do and what an evaluator has not yet learned to elicit. This is not a technical inconvenience. It is an epistemic infrastructure crisis.
In established scientific fields, measurement instruments are calibrated against known standards, and their limitations are documented. In AI evaluation, the instrument — the elicitation method — is itself a moving target. New prompting techniques, new inference strategies, and new model architectures continuously expand the space of what can be elicited, rendering previous evaluations obsolete. A model deemed safe under 2023 elicitation methods may reveal dangerous capabilities under 2026 methods, not because the model changed, but because the attention architecture of evaluation evolved.
This creates a paradox for dangerous capability evaluations. The claim no