Elicit: Impact of Prior Turns on LLM Accuracy

How Prior Turns Affect Large Language Model Accuracy on Information Extraction Tasks

Executive summary

Across recent empirical work, prior context can both help and hurt extraction-style accuracy, but the dominant pattern is that accuracy degrades when the model must integrate constraints or evidence spread across many prior turns or long contexts, motivating explicit re-grounding, memory management, and position-bias mitigations.

In multi-turn chat, models can get “lost” as dialogue history grows, and memory-related abilities “decline significantly in extended conversations,” with attention often concentrating on the most recent instruction rather than earlier constraints, producing instruction-following failures and inconsistency over time.

In few-shot in-context learning (ICL), prior “turns” take the form of demonstrations, and performance can depend strongly on which examples appear, how many appear, and in what order they appear, with example ordering alone producing “accuracy fluctuating by ≈ 12 percentage points” and search methods improving accuracy by “10.5 percentage points over random selection.”

In long-context settings, evidence position inside the context window is a first-order driver of accuracy, with a consistent “lost-in-the-middle” effect where accuracy is best when evidence appears near the beginning or end and worse when evidence is in the middle, including worst-case drops of “more than 20%” in multi-document QA and broad degradations at “25% ∼ 75% depth.”

Mechanistically, these effects are repeatedly tied to attention allocation patterns (recency in chat and U-shaped attention in long contexts), uncertainty spikes and ambiguity accumulation, demonstration interference or spurious correlations in ICL, and failures to maintain or update a stable world state across turns, which can yield contradictions in “40.50%” of context-sensitive queries.

Mitigation evidence supports interventions that (i) reconstruct or distill history when uncertainty spikes (ERGO), (ii) encourage abstention and reduce premature commitments in multi-turn rollouts (RLAAR), (iii) use structured state/memory (notes or knowledge graphs), (iv) repeat or re-state task instructions near the final query, (v) select and format demonstrations carefully (including contrastive negatives), and (vi) correct position bias using attention calibration, positional encoding re-scaling, or training that places crucial evidence throughout the context window.

Scope and definitions

This review uses “prior turns” broadly to cover three empirically studied regimes: (1) prior conversational turns in multi-turn chat that accumulate dialogue history and constraints, (2) prior in-context demonstrations in few-shot prompting that function as training-like turns inside the prompt, and (3) earlier content in long context windows where the position of evidence (beginning, middle, end) affects retrieval and correctness via “lost-in-the-middle” patterns.

This review uses “information extraction” broadly to include structured prediction tasks that require faithful selection and formatting of information, including relation extraction where “VANILLA prompting is difficult to handle the overlapping relations,” document extraction where removing “layout-aware demonstrations” drops F1, and QA-as-extraction settings such as multi-document QA where performance drops when gold evidence appears mid-sequence.

Because the evidence base operationalizes extraction in different ways (structured outputs, relation/event extraction, multi-doc QA, and context-sensitive conversational queries), the report emphasizes directly measured accuracy/F1-style deltas and intervention effects rather than a single unified metric definition.

Effect 1 Multi turn history

Multi-turn chat introduces prior-turn effects through accumulated context, incremental constraints, and the need to preserve and update task state over time, and recent work documents that models can get “lost” in conversation with “substantial performance degradation in multi-turn conversations compared to single-turn interactions.”

A common observed pattern is that memory-related abilities decline as conversations extend, including findings that “memory-related abilities (i.e., IE, IM, MR) decline significantly in extended conversations,” which directly implies that extraction-like tasks depending on instruction and entity/state memory can become less reliable with more turns.

How prior turns hurt

One empirically supported degradation mode is constraint accumulation across turns, where evaluation shows that “model performance deteriorates as additional layers of conditions or constraints are applied across conversation turns, even when accurate retrieved contexts are provided,” indicating that simply retrieving relevant prior content does not guarantee that the model correctly integrates it for accurate extraction or decision-making.

A complementary attention-level explanation is that models can over-weight the latest turn, as “the attention heatmap reveals that the model mostly concentrates on the most recent instruction rather than distributing its focus across all relevant turns,” which can cause earlier extraction rules, schema constraints, or exceptions to be ignored when generating final structured outputs.

Another degradation mode is “middle-of-conversation” weakness, where “memory degradation is more severe in the middle of a conversation than at the beginning or end,” consistent with a non-uniform attention allocation to dialogue history and with “attention to the middle part of the conversation” being “markedly lower” than to the beginning and end.

A further failure mode is premature commitment and error propagation, because in multi-turn settings “models frequently answer the question prematurely and subsequently fail to recover,” and “pollute the context with [their] own guesses,” making it “progressively harder” to correctly integrate later clarifications that would otherwise correct extraction mistakes.

Finally, multi-turn contexts can expose state-update limitations, where conversational models can “fail to update world states consistently,” producing “contradictions in 40.50% of context-sensitive queries,” which can translate to extraction errors whenever extracted facts or entity attributes must be revised across turns.

How prior turns help

Prior conversational turns can help when they provide relevant, personalized information, but the benefit is conditional because “irrelevant conversation history can be harmful,” highlighting that useful history must be filtered by relevance rather than included wholesale.

Conversational drift dynamics also indicate that not all history changes are equally harmful, because “abrupt prompt-to-prompt semantic drift sharply increases the hazard of inconsistency, whereas cumulative drift is counterintuitively protective,” suggesting that stable conversational trajectories can support longer-horizon task state so long as the shifts are not sudden shocks.

Quantitative effects

Quantitative results show that multi-turn degradation can be large and model-dependent, including reports that performance drops can be “as large as 73% for certain models” in multi-turn interactions on multiple-choice tasks, illustrating that prior-turn effects can be severe even for tasks that resemble extraction-as-selection.

At the same time, targeted interventions can recover large fractions of lost accuracy, such as ERGO improving “average performance by 56.6% compared to standard multi-turn baselines” and reducing multi-turn “unreliability” by “35.3%,” which supports the view that degradation is not inevitable and can be mitigated with adaptive context management.

Other intervention results include mitigation of “LiC performance decay (62.6% to 75.1%)” and improved “calibrated abstention rates (33.5% to 73.4%),” which operationalizes a concrete reliability improvement when prior turns make the model uncertain or under-specified.

Structured context distillation also shows measurable gains, as “the context distilled by our method enables the model to achieve a 14.1% higher score on the final question-answering task compared to the baseline,” suggesting that summarizing or distilling prior turns into a smaller state representation can improve final-step accuracy.

Medical multi-turn QA examples illustrate that performance can change across turns, with one system improving “from 43.1% at turn 2 to 50.6% at turn 10” and another rising “from 52.2% to 58.2%,” which indicates that not all multi-turn exposure produces monotonic declines and that task-specific protocols can yield modest gains with more interaction.

Finally, long-term memory modules can substantially improve accuracy, where “integrating the long-term memory produces the strongest results: 0.93 total accuracy,” compared to weaker setups, supporting memory-augmented architectures as a path to stable extraction across long dialogues.

Effect 2 Few shot demonstrations

Few-shot ICL can be understood as a structured form of “prior turns” inside a single prompt, where earlier demonstrations provide task patterns, label distributions, formatting cues, and implicit heuristics that shape extraction outputs in later test instances.

The empirical evidence shows that demonstration choice, number, ordering, and formatting each independently influence correctness, including substantial performance swings from example ordering and large drops when key demonstration attributes (such as layout cues in document extraction) are removed.

Demonstration selection

Demonstration selection quality can materially affect performance, including approaches that add contrastive negatives, where “contrastive in-context learning” uses both correct and incorrect examples to expose “typical mistakes,” which is aligned with extraction use cases where false positives and boundary errors are frequent.

Selection strategies that enforce balance and coverage can also improve structured extraction, as “I 2 CL with balance and coverage delivers significantly improvement compared to TableIE with random selection,” supporting the idea that representative, diverse examples are more instructive than randomly selected ones.

Quantitatively, selection strategies can improve ICL performance over random baselines by “4.59% on text-davinci-003 and 6.74% on gpt-4,” which indicates that retrieval or curation of demonstrations can deliver meaningful gains without changing the base model.

Number of demonstrations

The number of demonstrations can help, with reports that “the effect of each IE task increases as the number of shots increases,” which supports the practical heuristic that adding relevant shots can increase extraction quality up to a point.

However, adding demonstrations can also hurt depending on their label type and interference patterns, as evidence shows that adding more positive demonstrations can cause “counterintuitive degradation” while adding more negative demonstrations can yield “improvement,” which is consistent with interference or spurious correlations among examples.

Task-specific k-sensitivity is also documented in relation extraction settings, where “setting k to 5 delivers the best results across two datasets,” implying that there is often an intermediate optimal demonstration count rather than a monotone relationship.

Ordering and position

Example ordering can swing performance widely even when the demonstration set is fixed, with evidence that “example ordering plays a crucial role in performance” and that accuracy can fluctuate by “≈ 12 percentage points” as ordering changes.

Search and scoring methods can exploit this sensitivity, as OptiSeq “improves accuracy by 10.5 percentage points over random selection,” implying that ordering is not just noise but a manipulable factor with large practical returns for structured generation tasks.

Mechanistically, ordering sensitivity is consistent with the claim that LLMs “treat the entire prompt-content and order-as a holistic sequence,” which implies that demonstration order changes the effective conditioning context and thus the most probable structured output sequence.

Formatting and schemas

Demonstration formatting is a critical component for extraction pipelines that require machine-readable outputs, and document extraction work shows that removing “layout-aware demonstrations” can drop performance “around 10 F1 score on CORD,” indicating that demonstration content can encode task-relevant structural cues beyond pure text semantics.

Formatting demonstrations can be explicitly constructed “to guide LLMs to predict labels in a desired format for easy extraction,” which connects demonstration design directly to downstream parseability and extraction fidelity.

Tabular or structured prompting formats can also support extraction, as TableIE is designed to “generate organized and concise outputs” by “incorporating explicit structured table information into ICL,” which indicates a mechanism where explicit schema constraints reduce ambiguity in structured outputs.

Contrastive formatting can further improve error avoidance, as one approach introduces a “flag instruction behind the response to differentiate positive and negative samples,” thereby giving the model an explicit channel for distinguishing correct vs incorrect patterns in the prompt context.

More generally, small prompt-format changes can have large performance effects, including evidence that changing “a single example delimiter” can lead to “18.3% -29.4%” performance differences, and that delimiter choice can increase “attention scores for the target key” by “25%,” highlighting that extraction reliability can be sensitive to low-level prompt tokenization and attention routing effects.

Quantitative effects

Quantitative examples include a relation extraction improvement of “0.54 points” on ACE05-E relative to another method, which illustrates that even in structured IE benchmarks, careful ICL design can yield measurable gains.

Saturation effects also appear in few-shot extraction-like settings, where gains beyond a certain shot count can become “very marginal (only 0.0016),” suggesting diminishing returns and motivating adaptive demonstration selection rather than simply increasing k indefinitely.

Effect 3 Long context position

Long-context settings introduce prior-context effects via evidence placement inside the context window, and empirical work documents a robust “lost-in-the-middle” phenomenon where “accuracy drops significantly for information near the center of the context window.”

Controlled evidence placement experiments show that performance “significantly decreases when x gold is placed within the middle of the input prompt” compared to beginning or end placements, providing direct causal evidence that position alone can change accuracy in multi-document reasoning and extraction-like retrieval settings.

Positional degradation patterns

A common aggregate pattern is that models degrade when the answer is located in the middle range, with evidence that “most models exhibited a degradation in performance at the 25% ∼ 75% depth” and “best performance” when the answer is at “0% depth or 100% depth,” consistent with a U-shaped position-accuracy curve.

The position penalty can be large enough to drop below closed-book baselines, as GPT-3.5-Turbo multi-document QA “can drop by more than 20%” in the worst case, and performance in large-document settings can be “lower than performance without any input documents,” which indicates that prior context can actively harm correctness when the model fails to retrieve mid-context evidence.

Long-context evaluation difficulty is also highlighted by benchmark-level drops, where “long-context LLMs show a 30%∼60% performance drop” on LONGMEMEVAL variants, reinforcing that long contexts can induce large failures in memory and extraction fidelity even when models are marketed as long-context capable.

Quantitative mitigation effects

Mitigation methods can meaningfully reduce mid-context penalties, including attention calibration that, “when the gold document is placed mid-sequence,” offers improvements of “6-15 pp,” and shows a performance curve “almost entirely above” the baseline curve in most cases.

The same line of work reports that found-in-the-middle calibration can improve RAG performance “by up to 15 percentage points,” implying that position-bias correction can benefit downstream generation quality when retrieval results are placed in long prompts.

Positional encoding interventions can also help, as Ms-PoE “reducing the gap accuracy by approximately 2% to 4%,” and “re-scaling the indices of positional encoding” can yield “average accuracy gain of up to 3.8” on ZeroSCROLLS across multiple LLMs, supporting a plug-and-play approach to reducing lost-in-the-middle without full retraining.

Training-based mitigations explicitly aim to remove the bias, as some work introduces training “to explicitly teach the model that the crucial information can be intensively present throughout the context, not just at the beginning and end,” indicating that data distribution over positions can be an actionable training lever for extraction accuracy.

Evaluation-time reordering can also help, because one “possible solution” is to run completion multiple times while “randomly permuting documents” and selecting the best response, which treats position sensitivity as a nuisance factor to average out at inference time.

Counterevidence and boundary conditions

Not all tasks or models show position effects, as one study reports “no evidence of a context position effect for simple Q&A fact retrieval in modern Gemini 2.5 Flash” even near million-token scale, suggesting that position bias can be substantially reduced in some modern systems or for simpler retrieval regimes.

At the same time, the evidence base indicates that position bias and other long-context stressors can coexist, including findings that lost-in-distance and lost-in-the-middle have “independent, compounding effects,” and that “the sheer length of the input alone can hurt LLM performance, independent of retrieval quality and without any distraction.”

Mechanisms

Across the three regimes, a recurring mechanism class is attention misallocation, including recency overweighting in multi-turn chat where attention “mostly concentrates on the most recent instruction,” and U-shaped attention biases in long-context settings where tokens at beginning and end receive higher attention “regardless of their relevance.”

Long-context studies provide mechanistic hypotheses for primacy, proposing that the primacy effect arises from causal masking that biases attention toward earlier tokens, and describing “attention sinks” where early tokens “disproportionately attract most of the attention weight” despite limited semantic content.

Architectural analyses also connect positional bias to representation geometry, with findings that the causal mask introduces “positional hidden states” that correlate with absolute positions and convey positional information to the model, providing a concrete internal pathway for position-dependent errors in extraction-style retrieval.

Multi-turn degradation is additionally linked to uncertainty and ambiguity accumulation, where ERGO computes Shannon entropy “to detect spikes in uncertainty that indicate breakdown in comprehension,” implying that prior context can push the model into unstable internal states that degrade extraction fidelity unless actively corrected.

Another cross-cutting failure mechanism is error compounding through premature answers, where models “answer the question prematurely” and then “pollute the context with [their] own guesses,” which can bias later extractions and make corrected outputs less likely given the polluted context distribution.

For few-shot ICL, the mechanism evidence emphasizes interference and spurious correlations, where adding more demonstrations of a given type can cause “counterintuitive degradation,” indicating that the model can latch onto superficial patterns across examples rather than the intended extraction rule.

ICL work also identifies selection and evaluation distortions such as “Template Bias,” which can bias evaluations of information gain and thereby mislead demonstration selection policies that aim to optimize extraction performance.

Finally, conversation-state failures provide a mechanistic account of contradictions in multi-turn contexts, where models fail to update world states consistently and produce contradictions in “40.50%” of context-sensitive queries, which is directly relevant whenever extraction requires consistent entity attributes, temporal states, or evolving records across turns.

Mitigations

The mitigation literature suggests that improving extraction accuracy under prior turns requires actively managing what the model conditions on, when it conditions on it, and how that information is structured, rather than simply increasing context length or appending more demonstrations.

Multi turn mitigations

Uncertainty-triggered reconstruction is supported by evidence that “when such spikes occur, ERGO triggers entropy-guided prompt reconstruction,” which is explicitly framed as mitigating “accumulated ambiguity” and restoring coherence in multi-turn settings.

Policy-optimization approaches also support multi-turn robustness, including “a fully dynamic, multi-turn rollout during policy optimization,” and a reward signal that makes abstention valuable when “the current information is insufficient,” which aims to prevent premature extraction commitments that would later contaminate the context.

Structured memory approaches reduce reliance on raw history, including maintaining “structured notes of current plans and intermediate conclusions rather than full histories,” and transforming dialogue history into an OWL-compliant knowledge graph to “mitigate context decay.”

Instruction refocusing strategies are also empirically supported, as “repeating the task before the last query proves to be an effective mitigation strategy,” which addresses recency overweighting by ensuring the final prompt segment contains the task constraints the model is most likely to attend to.

Lightweight interventions can also shift drift trajectories, with evidence that drift “stabilizes at finite levels” and can be “shifted downward by lightweight interventions such as goal reminders,” and that when targeted interventions are introduced, “the equilibrium shifts to lower divergence levels.

Few shot mitigations

Contrastive demonstrations are supported by approaches that use both positive and negative examples and introduce explicit flags to differentiate them, aiming to guide the model away from typical extraction mistakes by showing counterexamples in-context.

Structure-aware retrieval and coverage criteria are supported by methods that “only select and annotate a few samples” based on “internal triple semantics,” and by DPP-based selection where the example set is chosen with the “highest probability,” reflecting a diversity-aware approach to building more informative extraction prompts.

Order-robustness is also treated as a mitigation target, including “order-agnostic inference” methods such as Batch-ICL, which is motivated by large ordering sensitivity effects and aims to reduce extraction variance across prompt permutations.

Adaptive control of the number of demonstrations is supported by AICL, where the likely reason for gains is its ability to “adapt the number of examples to use,” thereby preventing degradation from “non-relevant (not useful) examples.”

Long context mitigations

Attention calibration provides a direct mitigation of positional bias, as found-in-the-middle “mitigate this positional bias through a calibration mechanism” that helps the model “attend to contexts faithfully according to their relevance” even when they are in the middle.

Positional encoding re-scaling is supported by Ms-PoE, which “re-scales the indices of positional encoding” to enhance performance across LLMs and reduce lost-in-the-middle gaps, indicating that position bias can be reduced with post-hoc adjustments to positional representations.

Training for position-agnostic recall is supported by methods that explicitly teach that crucial information can occur throughout the context rather than only at the edges, implying that distributional training changes can reduce extraction failures under mid-context evidence placement.

Inference-time ensembling via permutation is supported as a practical workaround, where repeatedly permuting documents and selecting the best response is proposed as a way to counteract position bias without modifying the model.

Practical guidance for practitioners

Designing robust extraction pipelines in LLM systems requires treating prior context as a potentially adversarial variable: the same information can produce different outputs depending on whether it is embedded in long dialogue history, encoded as demonstrations, or placed in the middle of a long retrieval context.

The recommendations below prioritize interventions with direct empirical support in the provided evidence base, and they emphasize controlling relevance, structure, and position to reduce predictable failure modes.

First, in multi-turn extraction workflows, avoid unbounded history accumulation and instead distill history into a structured state, because structured notes “rather than full histories” and knowledge-graph state representations are explicitly proposed to reduce context decay and keep reasoning grounded in distilled information.

Second, add explicit re-grounding near the final query, because attention can concentrate on the most recent instruction and repeating the task before the last query is an effective mitigation, which is directly aligned with production extraction flows where final-turn correctness matters most.

Third, incorporate uncertainty-aware resets or reconstructions when long dialogues or incremental constraints produce confusion, because ERGO uses entropy spikes as a signal of breakdown and triggers reconstruction to restore coherence, which provides a concrete operational pattern for when to summarize, reset, or rebuild prompts.

Fourth, in few-shot extraction, treat demonstration selection and ordering as tunable parameters rather than fixed defaults, because selection strategies can improve performance by 4.59%–6.74% over random and ordering can shift accuracy by about 12 percentage points, with search methods adding 10.5 points over random ordering/selection baselines.

Fifth, include contrastive negatives and explicit formatting constraints when feasible, because contrastive ICL explicitly uses incorrect examples to expose typical mistakes and uses flag instructions to differentiate positive and negative patterns, and formatting demonstrations are used to guide outputs into easily extractable formats.

Sixth, in long-context extraction or RAG settings, assume a lost-in-the-middle risk unless validated otherwise, because multi-doc QA and controlled-depth experiments show systematic degradation at mid-depth and >20% worst-case drops, and apply mitigation such as attention calibration with 6–15 pp mid-sequence improvements when the gold document is in the middle.

Finally, when model choice allows, validate whether your target model exhibits position bias for your task class, because some evidence reports no position effect for simple fact retrieval in Gemini 2.5 Flash, while other benchmarks report large drops for long-context systems, implying that model and task regime matters for position-sensitivity risk management.

The table below summarizes actionable levers, the dominant failure they address, and representative quantitative effects reported in the evidence.

Regime Actionable lever Supported rationale Representative quantitative effect
Multi-turn chat Repeat task near final query Recency-weighted attention can over-focus on the latest instruction, and repeating the task is effective.Han 2025 +1 Multi-turn performance drops can be as large as 73% in some settings, motivating mitigations.Hankache et al. 2025
Multi-turn chat Uncertainty-triggered reconstruction Entropy spikes indicate comprehension breakdown and trigger reconstruction to restore coherence.Khalid et al. 2025 ERGO improves average performance by 56.6% and reduces unreliability by 35.3%.Khalid et al. 2025
Few-shot ICL Curate demos and avoid interference Adding more positive demos can degrade accuracy due to spurious correlations, so selection matters.Chen et al. 2023 Selection improves performance by 4.59% on text-davinci-003 and 6.74% on GPT-4 vs random.Li et al. 2024
Few-shot ICL Optimize or robustify ordering Ordering can cause ≈12 pp accuracy swings and can be optimized by search.Bhope et al. 2025 OptiSeq improves accuracy by 10.5 pp over random selection.Bhope et al. 2025
Long context Attention calibration Calibrates relevance attention to reduce U-shaped positional bias.Hsieh et al. 2024 Mid-sequence improvements of 6–15 pp and up to 15 pp RAG improvements.Hsieh et al. 2024
Long context Positional encoding re-scaling Re-scaling positional indices reduces lost-in-the-middle gap and improves accuracy across models.Zhang et al. 2024 Reduces gap by ~2–4% and yields up to 3.8 average accuracy gain on ZeroSCROLLS.Zhang et al. 2024

Open questions and limitations

A central limitation of the current evidence is that degradation mechanisms differ by regime and task, spanning lost-in-conversation effects, demo interference, and positional attention bias, which makes it difficult to map a single root cause to all extraction failures under prior context.

For multi-turn chat, there is strong evidence for attention imbalance and state-update failures, but uncertainty remains about which mechanism dominates for a specific extraction deployment, especially when tasks combine constraint accumulation, entity updates, and tool-augmented retrieval in the same conversation.

The drift literature suggests that drift can stabilize and can be shifted by goal reminders, and that abrupt semantic drift increases inconsistency hazard, but the operational translation from drift metrics to concrete extraction error predictors remains unresolved in the provided evidence.

For long-context models, evidence indicates large drops on benchmarks and strong positional bias mechanisms, but there is also counterevidence of minimal position effects in at least some modern systems for simple fact retrieval, leaving open which architectural or training changes account for this difference across models and task regimes.

The long-context literature also indicates that length, position, and distance can be separable stressors, since lost-in-distance and lost-in-the-middle have independent, compounding effects, and length alone can hurt performance even without distraction, which complicates evaluation design for extraction pipelines that rely on retrieved contexts of varying length.

For few-shot ICL, open issues include predicting whether adding demonstrations will help or hurt in a given extraction prompt, because demonstration interference can cause degradation and ordering sensitivity can produce large swings even with fixed demo sets.

Prompt decomposition as an adjacent lever

Although not directly about multi-turn dialogue history, prompt decomposition provides evidence that structuring intermediate steps can improve extraction accuracy, with claims that prompt design matters because “different prompts towards same tasks can cause large variations in the model predictions.”

In relation extraction, summarize-and-ask (SumAsk) is proposed to “transform RE inputs to the effective question answering (QA) format,” motivated by the difficulty of completing multiple reasoning processes in one step and by evidence that “summarization consistently improves the overall performance.”

Quantitatively, reported results on NYT show large micro-F1 increases from VANILLA to SUMASK across models, including values consistent with ChatGPT improving from 39.5 to 65.6 and GPT-J improving from 20.4 to 55.7 in the provided table, which supports structured decomposition as a practical companion to prior-turn management when extraction requires multi-step reasoning over overlapping relations.