Study portfolio
SL-001Stage 02: Protocol draftingPlanned intervention study

Scientific reasoning instruction with AI assistance

Research questionCan structured AI-assisted instruction improve how graduate learners generate, test, and revise scientific hypotheses?

SL-001 public status

Current stage
Protocol drafting
Stage 02 of 05
Recruitment
Not started
No participants are being enrolled
Results
None
This case file contains no study findings
Last reviewed
28 August 2026
Public planning record

PLANNED STUDY CASE FILE

A Learning Question With a Testable Decision

This case file treats scientific reasoning as a teachable practice rather than a fluency contest. The proposed intervention would prompt learners to state mechanisms, generate discriminating predictions, identify rival explanations, and revise confidence after critique. The study would test learning and transfer, not merely whether learners produce longer or more polished answers.

Learning context
A research-methods seminar, laboratory rotation, or doctoral training module in which learners already have enough domain knowledge to judge evidence but vary in formal practice with hypothesis construction.
Decision this study should support
Determine whether a structured AI-assisted lesson warrants a larger multisite effectiveness study and which instructional components should be retained, revised, or removed.

Protocol objectives

  1. 01

    Estimate the contrast in rubric-scored scientific reasoning after structured AI-assisted instruction.

  2. 02

    Test whether any improvement transfers to a new topic without access to the AI tutor.

  3. 03

    Measure whether learner confidence becomes better calibrated to answer quality.

  4. 04

    Identify failure modes, including deference to plausible but weak AI critiques.

DESIGN CANDIDATE

What Would Be Tested

Every element remains provisional until the protocol is registered. Unknowns are shown as unknowns instead of being filled with unsupported precision.

Design candidate
Counterbalanced crossover with matched lesson units and a delayed transfer assessment
Unit of assignment
Individual learner, subject to program constraints and contamination assessment
Conditions
Structured AI tutor and instructor-authored active-learning materials
Masking
Outcome raters blinded to condition where response formatting permits
Sample size
Not set. A priori precision or power analysis is required after the primary outcome and clustering structure are fixed
Study duration
Two instructional units plus an independently authored delayed transfer task

Learning sequence

01

Elicit

Record the learner's initial hypothesis, mechanism, predictions, confidence, and evidence before assistance.

02

Challenge

Present structured prompts for rival explanations, falsifiers, hidden assumptions, and measurement limits.

03

Revise

Require a traceable revision that states what changed, what did not change, and why.

04

Transfer

Assess the same reasoning moves on a new scientific problem without the original scaffolding.

MEASUREMENT

Outcomes Defined Before Observation

Candidate roles may change during protocol review. Any primary outcome will be fixed before data collection or access to relevant outcome data.

OutcomeRoleOperational definitionTiming
Scientific reasoning scoreCandidate primaryBlinded analytic-rubric score covering mechanism, discriminating predictions, alternatives, testability, and evidence alignment.Immediate post-instruction
Delayed transfer scoreCandidate primaryRubric score on a novel domain-adjacent problem completed without access to the study tutor.Delay to be specified
Calibration errorCandidate secondaryAbsolute difference between normalized confidence and observed rubric performance.Baseline, post-instruction, and transfer
Unsupported uptakeSafety measureCount and severity of revisions that adopt an AI suggestion without adequate evidence or valid reasoning.During assisted lesson
Instructional timeImplementationActive time required to complete the lesson, with idle-time rules prespecified.Each lesson

ANALYSIS DISCIPLINE

A Plan That Can Report Unfavorable Results

Final estimands, models, exclusions, missing-data rules, multiplicity decisions, and stopping conditions will be specified in the registered protocol where applicable.

  1. 01

    Specify one primary contrast before outcome data are inspected and distinguish confirmatory from exploratory analyses.

  2. 02

    Use a model that accounts for repeated observations and any cohort, course, or instructor clustering supported by the final design.

  3. 03

    Report effect estimates with uncertainty intervals, score distributions, missingness, and inter-rater reliability.

  4. 04

    Test sensitivity to prior domain knowledge, prior AI use, task order, rubric specification, and exclusion decisions.

  5. 05

    Report null, inconclusive, and unfavorable results with the same outcome definitions used for favorable results.

RESEARCH INTEGRITY

Implementation, Access, and AI Disclosure

The record must be detailed enough to audit what learners experienced, what the AI system could do, and where qualified humans remained responsible.

Implementation record

  • Lesson objectives, instructor guidance, allowed resources, and time allocation
  • Tutor system instructions, model version, settings, retrieval sources, and tool availability
  • Learner interaction trace with privacy-preserving identifiers where review permits
  • Deviation log covering interruptions, instructor assistance, and technical failures
  • Rater training materials, adjudication rules, and rubric version

Equity and access

  • Measure prior access to paid AI systems and familiarity with prompt-based tools.
  • Review tasks for linguistic complexity unrelated to the target scientific skill.
  • Offer an accessible interface, keyboard operation, sufficient time, and documented accommodation procedures.
  • Report participation and outcomes by prespecified groups only when sample size, privacy, and interpretability permit.

AI-system disclosure

  • Provider, model name, model version, and access dates
  • System instructions, lesson prompts, tools, retrieval corpus, and relevant settings
  • Rules for feedback, refusal, uncertainty statements, and prohibited answer disclosure
  • Human-authored content and checkpoints that remain outside model control
  • Known model changes or service incidents during the study window

VALIDITY REGISTER

Main Risks and Planned Responses

These responses reduce specific risks. They do not eliminate uncertainty or guarantee that the final design will support a causal claim.

R01

Task contamination

Use newly authored transfer tasks, record search exposure where feasible, and keep scorers separate from lesson design.

R02

Rater expectancy

Blind condition labels, standardize response formatting, train raters, and report reliability before adjudication.

R03

Novelty and motivation

Measure prior AI use and learner experience, and avoid interpreting engagement as learning.

R04

Instructor variation

Standardize essential lesson components and record permissible instructor adaptation.

R05

Short-term performance only

Include delayed transfer and avoid claims about durable learning without follow-up evidence.

STUDY GATES

Status Changes Require an Exit Record

A stage label is a public claim. The record moves forward only when its stated exit condition is documented.

01

Concept

Complete

Question, intended inference, learning context, and initial validity risks are recorded.

02

Protocol drafting

Current

Primary outcomes, allocation, power or precision rationale, and analysis plan are fixed.

03

Review

Not started

Scientific, ethics, privacy, accessibility, and security determinations are documented.

04

Registration and execution

Not started

A time-stamped protocol is public before enrollment or outcome inspection.

05

Reporting

Not started

Results, uncertainty, deviations, validity limits, artifacts, and corrections are released.

EVIDENCE CONTEXT

Real Sources That Inform the Protocol

These external sources provide context for design decisions. They are not Santaros Labs outputs, endorsements, or evidence that this proposed study has been completed.

01

Scientific Reports · 2025

AI tutoring outperforms in-class active learning: an RCT introducing a novel research-based design in an authentic educational setting

Provides a real higher-education comparison and shows why instructional design, time, and learning outcomes should be measured separately.

Open source
02

Institute of Education Sciences · Living standard

Standards for Excellence in Education Research

Informs falsifiable hypotheses, outcome relevance, open science, implementation records, generalizability, and reporting of uncertainty.

Open source
03

UNESCO · 2023

Guidance for generative AI in education and research

Frames human-centered pedagogical design, privacy, capacity, and institutional responsibility.

Open source

OPEN SCIENCE PLAN

Planned Public Open

  • Preregistered protocol
  • Lesson and tutor specification
  • Outcome rubric and rater manual
  • Synthetic demonstration dataset
  • Analysis code and environment record
  • De-identified data or a governed access statement

Availability will depend on consent, ethics review, licensing, privacy, security, and institutional requirements. A restriction will be explained rather than presented as open access.

METHODS COLLABORATION

Improve This Study Before Registration

We welcome educators, learning scientists, domain researchers, statisticians, research software specialists, and governance reviewers who can strengthen the protocol.

Start a research inquiry
A research inquiry does not imply study enrollment, institutional approval, funding, or authorship.