Scholarly status: This paper establishes CAMDLE as a conceptual framework and research agenda. It synthesises established literature, articulates a proposed design, describes a prototype apparatus, and identifies hypotheses for empirical study. Evaluchat has not yet established that CAMDLE improves language proficiency, writing quality, critical thinking, or academic integrity outcomes. Claims are separated into established literature, proposed synthesis, product description, and open hypotheses.
Constrained AI-Mediated Dialogic Language Education (CAMDLE)
Evaluchat research and discussion paper · 2 August 2026
Executive summary
Contemporary generative AI can produce fluent academic prose from a short request, sharply reducing the effort required to obtain a polished textual product. This capability sharpens a longstanding problem in writing education: when the intended learning outcome includes planning, argumentation, language development, revision, or metacognitive control, the final text is an incomplete and increasingly ambiguous signal of the process that produced it.
This problem gives contemporary force to Bloom’s (1984) 2-Sigma challenge: how can education approximate the benefits of expert one-to-one tutoring at scale? GenAI makes scalable, responsive assistance technically plausible, but it also sharpens an automation-versus-agency paradox. The same system that can extend access to expert-like support can also automate the planning, linguistic, and evaluative activity that writing instruction is meant to develop.
CAMDLE proposes dialogic constraint as a theoretical response to this paradox: place a conversational AI writing partner inside a process-based environment, but make substantial drafting assistance conditional on the learner first contributing ideas, evidence, questions, and language through dialogue. The constraint is not intended to prove who typed each word. It functions instead as a pedagogical condition for preserving learner agency: the learner must negotiate meaning before more powerful assistance becomes available, while the system retains the responsiveness and scalability that make GenAI educationally significant.
In this sense, CAMDLE reframes the 2-Sigma challenge for an era of generative systems. It does not claim to produce two-standard-deviation gains. Instead, it offers a design hypothesis: scalable dialogic assistance, proportionate drafting support, and teacher-facing process evidence may bring some benefits of tutoring into writing education without treating automation as a substitute for learner judgment.
Evaluchat is a prototype implementation of this idea. It combines a dialogue panel, a drafting canvas, conditional scaffolding, revision history, and teacher-facing process signals. The platform is an apparatus for research, not evidence that the proposed mechanism works.
The central research question is threshold calibration: what counts as sufficient dialogic contribution to unlock drafting support, and how does that threshold vary by task type, proficiency level, language background, and learner strategy?
1. The problem: product-based writing under ubiquitous GenAI
Writing assessment traditionally treats the submitted text as a proxy for planning, reasoning, language proficiency, and authorship. That proxy was always imperfect, but generative systems have changed the cost of producing fluent text. Recent reviews describe benefits in fluency, cohesion, organisation, vocabulary, and feedback, alongside concerns about over-reliance, hallucination, bias, diminished metacognitive engagement, and unequal access (Aljuaid, 2024; Urzúa et al., 2025; Sanz-Tejeda et al., 2026).
The psychometric issue can be stated as construct-irrelevant variance (Messick, 1989). When a learner submits text substantially produced or transformed by a generative system, the text may reflect the model’s language production, the learner’s prompting and selection skill, access to the system, and the conditions of use — as well as the learner’s own planning, reasoning, or language proficiency. If the intended inference concerns independent cognitive processing or authorship, those influences contaminate the score rather than simply adding noise. The same AI use could be construct-relevant for an assessment of AI-mediated composing, but it is construct-irrelevant to an inference about unaided composing unless that use is explicitly part of the construct.
The implication is not that final texts are worthless. It is that a final text alone no longer supports the same inferences: it may describe product quality under particular assistance conditions, but it is a weak and ambiguous basis for inferring the cognitive process that produced it. An assessment system should make relevant interaction evidence visible and ask the teacher to interpret it in context.
Boundary: Process signals are mechanical observations. Keystrokes, paste events, focus changes, and interaction counts may prompt a conversation, but they do not prove who authored a sentence or whether learning occurred.
2. Evaluchat as a research apparatus
Evaluchat is a browser-based prototype with two related workspaces:
- Dialogue panel: the learner discusses the task with a conversational language model, develops ideas, asks for feedback, and negotiates wording.
- Drafting canvas: the learner writes, edits, accepts, rejects, and revises a document.
The model is constrained from generating a complete assignment or substantial canvas prose from a single low-effort request. After the learner has supplied enough relevant contribution, limited co-generation may unlock. This is proportional scaffolding: assistance is conditional on conceptual and linguistic work visible in the dialogue.
The prototype records process context such as dialogue turns, canvas changes, paste events, session pacing, and scaffold unlocks. Two evidentiary layers should be distinguished: keystroke telemetry describes observable interface events (insertions, pastes, focus, timing) without directly observing attention, intention, or planning; cognitive process evidence is the more contentful record — explanations, claims and evidence in dialogue, responses to feedback, and transformations of AI suggestions — which can support arguments about reasoning when read in context, but remains indirect and requires human interpretation.
Evaluchat isolates one defined interaction pathway. It cannot detect all off-device or mediated behaviour, including retyping, dictation, paper notes, or assistance from a second device. Signals are context for human judgment, not an “integrity score.”
3. The CAMDLE construct
Constrained AI-Mediated Dialogic Language Education (CAMDLE) is the proposed name for a learning design in which a generative agent is prevented from independently producing the whole assignment, and drafting support is released conditionally after learner contribution across successive dialogic iterations.
- Dialogic contribution precedes substantial generation. Learners articulate, question, select, or defend ideas before higher-powered assistance becomes available.
- The constraint is proportional, not absolute. Learners may use assistance; the design regulates when and how much assistance is available.
- The learner remains the executive controller. The learner evaluates, accepts, rejects, and revises AI suggestions.
- Process is evidence, not verdict. Dialogue and revision traces support teacher interpretation without pretending to prove authorship.
- The threshold is an empirical variable. It must be calibrated against learning, equity, usability, and circumvention outcomes.
CAMDLE is a proposed synthesis. The component theories are established research traditions; their combination, implementation, and predicted effects are not established by the existence of those traditions.
4. The theoretical synthesis
Sociocultural theory and scaffolding
Vygotsky’s account of mediated development provides a rationale for treating assistance as contingent on what a learner can currently do with support (Vygotsky, 1978). Operationally, the ZPD should be treated as a task- and time-specific difference between independent performance and performance with contingent assistance — not a latent quantity readable from message length or an unlock threshold. When the “more knowledgeable other” is an LLM, the term describes a functional role in a particular interaction, not a claim that the system is a teacher or globally more knowledgeable than the learner.
Extended mind and cognitive offloading
The extended mind thesis treats external artefacts as potential components of cognitive activity when actively integrated into problem solving (Clark & Chalmers, 1998). Cognitive offloading can reduce burden, but its educational effect depends on what is offloaded and what the learner continues to control (Risko & Gilbert, 2016). Cognitive Load Theory sharpens the boundary: AI can reduce extraneous operational load (syntax, formatting, routine retrieval) while becoming substitutive when it takes over executive functions such as warranting an argument, adopting a stance, or deciding what a revision should accomplish.
CAMDLE distinguishes strategic offloading of operational burdens from substitutive offloading of conceptual work. This is a hypothesis about interaction and learning, not a fact that can be inferred from one metric.
Self-regulated learning
Self-regulated learning models describe cycles of forethought, performance, and self-reflection (Zimmerman, 2000). The interface makes planning, composing, evaluating, and revising available for study; it does not automatically produce self-regulation.
Cognitive process writing
Process models describe composing as recursive planning, translating, reviewing, and knowledge transformation (Flower & Hayes, 1981; Bereiter & Scardamalia, 1987). The relevant outcome is not “more chat” or “more text,” but better explanations, decisions, and substantive revision.
Formative assessment
Formative assessment uses evidence during learning to improve the learner’s next move (Black & Wiliam, 1998). Generative dialogue is only formative when feedback is understood, evaluated, and acted upon. The teacher therefore remains in the interpretive loop.
A related design problem is Bloom’s (1984) 2-sigma challenge: one-to-one tutoring with mastery-learning techniques produced large achievement gains relative to ordinary classroom instruction, motivating the search for scalable approximations. Several Bloom alterable variables map onto CAMDLE’s intended pattern—tutorial-style dialogue, feedback-corrective cycles, time on task, and required participation—without equating an engagement unlock with mastery learning or claiming measured 2-sigma outcomes. Bloom motivates why a scalable dialogic apparatus is worth building; threshold calibration remains the primary empirical question.
5. Evidence map
| Literature stream | Supports | Does not establish |
|---|---|---|
| Process writing | Recursive planning, translating, and reviewing | That chat logs proxy all composing cognition |
| Formative assessment | Feedback can improve the next learning move | That AI feedback equals expert teacher feedback |
| Bloom 2-sigma / tutoring | One-to-one tutoring plus mastery techniques can produce large gains; scalable approximations are a long-standing design problem | That GenAI dialogue produces 2-sigma gains, or that an engagement unlock equals mastery learning |
| Self-regulated learning | Planning, monitoring, and reflection matter | That interface activity represents self-regulation |
| Cognitive offloading | External tools can reduce task burden | That offloading improves durable learning |
| GenAI writing reviews | AI can support fluency, organisation, feedback, and language development | That unconstrained assistance produces independent proficiency |
| L2 / EAP writing | AI raises questions of voice, critical literacy, and equity | That machine fluency is a fair human benchmark |
| Process-based AI assessment | Interaction traces can be analysed as candidate evidence | That any trace is a validated learning measure |
Recent work on critical GAI literacy, L2 writing, hybrid feedback, and process assessment makes this programme tractable while setting a high bar. Those studies support investigating process evidence; they do not validate Evaluchat’s particular threshold or telemetry model.
6. Research propositions
- Threshold calibration: there is a range of contribution levels at which additional AI drafting support improves progress without reducing explanation, evaluation, or revision.
- Conditional assistance: compared with unconstrained assistance, conditional assistance increases sessions containing learner-generated explanations, questions, or revisions.
- Learning transfer: a better supported text does not demonstrate learning; transfer requires independent or delayed writing, explanation, or revision tasks.
- Input and output: exposure, transcription, and learner-generated production may have different effects and should remain an open comparison.
- Equity: a single lexical or turn-volume threshold may disadvantage different proficiency levels, L1 backgrounds, disabilities, or communication styles.
- Teacher interpretation: process evidence is most defensible when read with the transcript, draft, assignment context, and student explanation.
7. Proposed research programme
Primary study: threshold calibration
Begin with a mixed-methods, design-based study in an EAP or L2 academic writing context. Compare pre-registered threshold policies, including an unconstrained-assistance condition where appropriate. Measure blinded baseline and post writing, independent transfer writing, language measures, dialogue and revision traces, student explanation of decisions, motivation, agency, workload, circumvention, and teacher workload.
The primary endpoint must be specified before deployment. Candidate endpoints include independent writing improvement, quality of learner explanations, or an evidence-centred process-and-outcome measure. A higher final essay score alone is insufficient.
Secondary studies
- Compare genuine learner production, reading/listening exposure, and transcription or relay patterns.
- Test differential unlock rates, frustration, and learning outcomes across L2 proficiency and L1 groups.
- Study whether process evidence improves teacher feedback and student-teacher conversations.
- Examine mechanisms involving planning, vocabulary, evaluation, revision, and metacognitive explanation.
- Document when constraint creates productive effort and when it creates avoidable circumvention.
Design-based research is appropriate because threshold, interface, teacher practice, and measures must be refined together in authentic settings (The Design-Based Research Collective, 2003). It does not remove the need for comparison groups, transparent measures, ethics review, or pre-specified analyses.
8. Ethics, limitations, and governance
A South African deployment involving minors should address POPIA, institutional approval, parental or guardian consent, student assent, data minimisation, retention, access control, and de-identification. The platform should collect only signals needed for the research question and should not expose students to hidden or punitive profiling.
- Process signals are circumstantial and incomplete.
- A constraint can create frustration, inequity, or strategic gaming.
- Dialogue volume is not equivalent to language quality or learning.
- AI can hallucinate, reproduce bias, or provide poor feedback.
- L2 writers should not be judged against monolingual machine fluency.
- Novelty effects and commercial interest may affect early findings.
- Results from one age group, genre, language, or institution may not generalise.
The responsible claim is modest: CAMDLE offers a testable way to structure AI-mediated writing and expose more interaction for human interpretation. Whether it improves learning, for whom, and under which thresholds remains to be established.
9. Implications for educators and institutions
For educators, the immediate implication is not to adopt a numerical engagement score. It is to make the intended learning process explicit: define what students must explain, decide, and revise; teach interrogation of AI output; assess independent transfer, not only supported products; read dialogue and revision evidence in context; explain what is collected and why; and give students a way to challenge or explain process evidence.
For institutions, process-based AI assessment should complement — not replace — human judgment, assessment design, accessibility support, and clear policy. Tools should not turn uncertain telemetry into automated misconduct findings.
10. Collaboration invitation
Evaluchat is available as a prototype apparatus for educators and researchers who want to investigate constrained AI-mediated writing. Possible collaborations include a feasibility pilot, a co-designed threshold-calibration study, an independent postgraduate project, or a methodological critique of process evidence.
Academic collaborators should retain ownership of the research question, protocol, analysis, and publication decisions. Evaluchat can contribute prototype access and implementation support while disclosing its commercial interest in positive findings.
Download the white paper (PDF) hello@evaluchat.com
Research questions in this paper are invitations to investigate, not answers presented in advance.
CAMDLE research status
11. Conclusion
CAMDLE is a proposal about how to design AI-mediated writing so that the learner’s contribution remains visible and consequential. Its most valuable claim is not that a gate automatically creates learning. It is that the gate creates a manipulable research variable: the amount and kind of contribution required before drafting assistance becomes available.
That variable can be tested. A credible research programme should be willing to find that some thresholds do not work, that effects differ across learners, that human feedback is necessary, or that the constraint creates more frustration than learning. The white paper’s purpose is to make those findings possible — not to pre-announce them.
References
Aljuaid, H. (2024). The impact of artificial intelligence tools on academic writing instruction in higher education: A systematic review. DOI.
Bereiter, C., & Scardamalia, M. (1987). The psychology of written composition.
Black, P., & Wiliam, D. (1998). Assessment and classroom learning. Assessment in Education, 5(1), 7–74.
Bloom, B. S. (1984). The 2 sigma problem: The search for methods of group instruction as effective as one-to-one tutoring. Educational Researcher, 13(6), 4–16. DOI.
Clark, A., & Chalmers, D. (1998). The extended mind. Analysis, 58(1), 7–19. DOI.
The Design-Based Research Collective. (2003). Design-based research: An emerging paradigm for educational inquiry. Educational Researcher, 32(1), 5–8.
Flower, L., & Hayes, J. R. (1981). A cognitive process theory of writing. College Composition and Communication, 32(4), 365–387.
Goulart, L., Matte, M. L., Mendoza, A., et al. (2024). AI or student writing? Journal of Second Language Writing, 66, 101160. DOI.
Messick, S. (1989). Validity. In R. L. Linn (Ed.), Educational measurement (3rd ed., pp. 13–103). Macmillan.
Risko, E. F., & Gilbert, S. J. (2016). Cognitive offloading. Trends in Cognitive Sciences, 20(9), 676–688. DOI.
Sanz-Tejeda, A., et al. (2026). The impact of generative AI on academic reading and writing. Frontiers in Education. DOI.
Urzúa, C. A. C., et al. (2025). Effects of AI-assisted feedback via generative chat. Education Sciences, 15(10), 1396. DOI.
Vygotsky, L. S. (1978). Mind in society.
Zimmerman, B. J. (2000). Attaining self-regulation. In Handbook of self-regulation.
Critical GAI Literacy in doctoral academic writing (2024).
Evidence-centered Assessment for Writing with Generative AI (2024).
Assessing students’ DRIVE (2025).
The role of generative AI and hybrid feedback in improving L2 writing skills (2025).