Inspect
Evaluate dataset documentation, license, provenance, missingness, and analysis suitability before code is written.
Research questionCan learners produce a more reproducible scientific analysis when AI assistance is embedded inside explicit provenance instruction?
SL-002 public status
PLANNED STUDY CASE FILE
This case file connects computational training with research integrity. Learners would not be scored only on obtaining the expected answer. They would also be evaluated on whether an independent analyst can reconstruct data selection, transformations, code execution, model use, judgment calls, and the computing environment from the submitted record.
Estimate the contrast in blinded reproduction success between instructional conditions.
Measure which provenance fields most often prevent or enable successful reproduction.
Test transfer to a new dataset and analysis environment.
Quantify hidden manual decisions and unsupported analytic changes introduced during AI use.
DESIGN CANDIDATE
Every element remains provisional until the protocol is registered. Unknowns are shown as unknowns instead of being filled with unsupported precision.
Evaluate dataset documentation, license, provenance, missingness, and analysis suitability before code is written.
State the question, intended estimand or analytic target, exclusions, transformations, and acceptance tolerances.
Run analysis while capturing code, environment, AI interactions, sources, and material human decisions.
Transfer the package to a blinded analyst and record every ambiguity, failure, and undocumented dependency.
MEASUREMENT
Candidate roles may change during protocol review. Any primary outcome will be fixed before data collection or access to relevant outcome data.
| Outcome | Role | Operational definition | Timing |
|---|---|---|---|
| Reproduction success | Candidate primary | Whether a blinded analyst obtains the target outputs within prespecified numerical and qualitative tolerances. | After package submission |
| Provenance completeness | Candidate co-primary | Proportion of required provenance fields that are accurate, specific, and sufficient for reuse. | At package audit |
| Unresolved decision count | Candidate secondary | Number of material analytic choices that cannot be reconstructed from the submitted record. | During reproduction |
| Transfer performance | Candidate secondary | Reproducibility score for a new dataset completed without the original teaching examples. | Delayed task |
| Time and support burden | Implementation | Learner time, instructor assistance, and reproducer time under prespecified logging rules. | Throughout each task |
ANALYSIS DISCIPLINE
Final estimands, models, exclusions, missing-data rules, multiplicity decisions, and stopping conditions will be specified in the registered protocol where applicable.
Prespecify the reproduction tolerance and adjudication rule before packages are evaluated.
Estimate condition contrasts with models appropriate to the assignment unit, repeated tasks, and course clustering.
Report failure categories, not only an aggregate reproduction rate.
Conduct sensitivity analyses using stricter and looser tolerances declared in advance.
Separate pedagogical effectiveness, technical reliability, and time cost in interpretation.
RESEARCH INTEGRITY
The record must be detailed enough to audit what learners experienced, what the AI system could do, and where qualified humans remained responsible.
VALIDITY REGISTER
These responses reduce specific risks. They do not eliminate uncertainty or guarantee that the final design will support a causal claim.
Define scientifically acceptable output tolerances and include analyses with more than one valid implementation.
Measure baseline skills, block or stratify assignment if justified, and report heterogeneous effects cautiously.
Version environments, archive container definitions, and timestamp all externally hosted dependencies.
Log instructional support and distinguish independent completion from assisted completion.
Describe the tested configuration precisely and avoid generalizing beyond evaluated systems and tasks.
STUDY GATES
A stage label is a public claim. The record moves forward only when its stated exit condition is documented.
Learning problem, intended decision, candidate tasks, and reproducibility definition are recorded.
Assignment unit, primary tolerance, outcomes, and analysis plan are fixed.
Scientific, education, privacy, accessibility, and infrastructure reviews are documented.
Protocol, tasks, environments, and scoring rules are time-stamped before use.
Reproduction outcomes, failure taxonomy, costs, deviations, and artifacts are released.
EVIDENCE CONTEXT
These external sources provide context for design decisions. They are not Santaros Labs outputs, endorsements, or evidence that this proposed study has been completed.
Scientific Data · 2016
Supports rich metadata, persistent identification, provenance, interoperability, and reuse across data, tools, and workflows.
Open sourceNational Institute of Standards and Technology · 2021
Provides a real scientific-computing context for reusable datasets, models, and executable research infrastructure.
Open sourceInstitute of Education Sciences · 2022
Informs decisions about what education data can be shared and what documentation is needed for responsible reuse.
Open sourceOPEN SCIENCE PLAN
Availability will depend on consent, ethics review, licensing, privacy, security, and institutional requirements. A restriction will be explained rather than presented as open access.
Next portfolio record
SL-003Teaching evidence synthesis under scientific disagreementMETHODS COLLABORATION
We welcome educators, learning scientists, domain researchers, statisticians, research software specialists, and governance reviewers who can strengthen the protocol.
A research inquiry does not imply study enrollment, institutional approval, funding, or authorship.