ML//model//GPT//prompt engineering//automatic prompt engineering
Automatic prompt engineering (APE) is the search for good instructions to a language model by an algorithm instead of by hand, scoring candidate prompts on a set of examples and keeping the best, and it is used when a prompt will run thousands of times and a few points of accuracy are worth an afternoon of compute. The model's weights stay frozen: what is optimized is the text in front of the input.
Automatic prompt engineering (APE) is the search for good instructions to a language model by an algorithm instead of by hand, scoring candidate prompts on a set of examples and keeping the best, and it is used when a prompt will run thousands of times and a few points of accuracy are worth an afternoon of compute. The model's weights stay frozen: what is optimized is the text in front of the input.
The loop has three parts. A generator proposes candidate instructions (often a model asked to write them, or to rewrite the best ones so far), an evaluator runs each candidate over a small benchmark built for the task and scores the outputs, and a selector keeps the winners for the next round. The original APE paper (Zhou and colleagues, 2022) did exactly this and found instructions as good as human-written ones on most of its tasks; one of its finds was a better zero-shot chain of thought trigger than the famous let's think step by step.
Static APE searches once, offline, for one prompt that works well across the inputs of a task, for instance classifying maintenance tickets into a closed list of categories. The search can be a genetic algorithm (select the best-scoring prompts, recombine them, mutate them, score again), beam search over rewrites, or the compilers of DSPy, which also choose the few-shot examples.
Dynamic APE builds the prompt in place for each input: it reads the document first (a contract, a fault log), picks the instructions and context that fit its kind, and only then asks. It adapts better and costs a planning call per request.
The score is a fitness function, and it decides everything. It should measure what the task really needs (exact-match or F1 for extraction, a checked answer for reasoning) and can subtract cost and latency per call; a fitness computed on twenty examples will happily overfit them, so the winner is confirmed on held-out ones.
The prompt is a parameter, the benchmark is the specification.
Automating the search moves the human work from writing strings to building the evaluation set and the metric, and a weak metric gets optimized as faithfully as a good one.
Each evaluation is a model call, so a search over 50 candidates on 200 examples is 10,000 calls; sampling the examples, scoring cheap candidates on a few items first and stopping early are what keep the bill down. When the gains plateau, the next step is changing the weights (fine-tuning) rather than the words (prompt engineering).