Research protocol · trial begins December 2026
What we are about to measure, and what we have not settled.
- What
- A two-arm non-inferiority trial. Licensed teachers supported by human mentors, against unlicensed facilitators supported by an AI mentor, with the four-to-eight person teaching team held constant in both arms.
- When
- December 2026, at the online Japanese-language school we open for Japanese-speaking children living outside Japan. Registered before the first lesson is recorded, not after.
- What we are asking for
- A research lead and a measurement specialist. Five design values are still open — allocation, margin, primary outcome, stopping rule, sample. We would rather settle them with someone else in the room than alone.
This is not an overview of the company. It is the design, the instrument, the data, and the list of what is still undecided.
Design as of 8 September 2026. Every item marked Not yet fixed will be replaced with a value, and with the date it was fixed, before the trial begins.
The question
The question, in one sentence.
In December 2026 we run a non-inferiority trial: licensed teachers supported by human mentors, against unlicensed facilitators supported by an AI mentor, with the four-to-eight person teaching team held constant in both arms.
Whether a person without a licence can raise attainment is not the open question — that has been asked before. The open question is how far the cost of developing a teacher can fall while the teaching stays equivalent. The world is short about 44 million teachers by 2030, and that is a supply problem before it is a pedagogy problem.
The design
The design, including the parts that are not decided.
Six rows below are not decided yet. They are marked. They will be fixed and registered before the first lesson is recorded, not afterwards. If the design is wrong, we would rather hear it now than in December.
| Item | What it is today |
|---|---|
| Design | Two-arm parallel non-inferiority trial. |
| Setting | The online Japanese-language school we open in December 2026 for Japanese-speaking children living outside Japan. Fifty-minute lessons, in Japanese, recorded with audio and video. |
| Control arm | Licensed teachers, supported by human mentors. Teachers hold a Japanese teaching licence and are coached by a person. |
| Intervention arm | Unlicensed facilitators, supported by an AI mentor. No teaching licence. Trained in about three months, then coached from timestamped evidence taken out of their own recorded lessons. |
| Held constant | The teaching team of four to eight people, in both arms. Same team size, same weekly rhythm, same curriculum. The licence and the source of the mentoring are what differ. |
| Allocation Not yet fixed | The unit of allocation — child, class or facilitator — and the randomisation procedure are not decided. To be registered before the trial begins. |
| Non-inferiority margin Not yet fixed | No margin has been set. We are not going to publish a number here and choose it later. To be registered before the trial begins. |
| Primary outcome Not yet fixed | Under discussion: student attainment, and the observational rating on the six-dimension rubric below. A third-party instrument for student outcomes is also under discussion, so that the effect of an AI mentor is not measured only by the people who built it. Which measure carries the non-inferiority claim is exactly the decision we would like to make with a research partner rather than alone. |
| Secondary outcomes Not yet fixed | Inter-rater reliability between the two external raters. The student survey. How long the facilitators we develop keep teaching. The cost of developing one teacher. The final list, and the analysis plan for each, are not fixed. |
| Blinding | External raters never see the teacher’s name, which arm a lesson belongs to, or the machine’s values. There is no route into a lesson except the blind scoring screen. This one is not a plan — it is what the system already enforces. |
| Stopping criteria Not yet fixed | The conditions under which we stop the trial have not been written down. They will be written before it starts and registered with the protocol, and they will not be edited afterwards. |
| Pre-registration Not yet fixed | We intend to pre-register the trial together with the cost-analysis plan; the AEA RCT Registry is the venue we are looking at. Nothing has been filed and there is no registry ID yet. |
| Sample Not yet fixed | Number of classes, number of children and the length of the observation window are not fixed. No power calculation has been done, because the primary outcome and the margin above are not fixed either. |
| Ethics and consent | Guardian consent is a precondition for a lesson entering the system at all; withdrawal deletes the child’s words; utterances are dropped past the retention window while scores and timestamps survive; every viewing of a transcript is logged. The rules are set out under The data. We have not been through ethics review, and we have not yet decided whose review we will go through. |
Six of fourteen rows are undecided. That ratio is the honest state of this trial on 8 September 2026, three months before it starts.
Allocation and outcome flow for the December 2026 trial. The amber blocks are the parts we would most like to decide with someone else in the room.
On a narrow screen, the figure scrolls sideways. Everything in it is also in the table above.
What is missing
What we do not have.
We do not have a research director or a measurement specialist. Those are the two people we most need. If the design above is wrong, we would rather be told now than in December.
-
No research lead, no measurement specialist
The two roles the design most depends on are unfilled. Everything below follows from that.
-
No margin, no primary outcome, no stopping rule
The decisions that determine whether December produces evidence or produces an anecdote.
-
No pre-registration on file
Intended, not filed. There is no registry ID to check us against yet.
-
No validated instrument
The rubric’s own status field says unvalidated, and we have left it saying that. Factor analysis and inter-rater reliability are still ahead of us.
-
No ethics review
We have not been through IRB and have not decided whose review we will go through.
-
No crosswalk to CLASS, PLATO or MQI
A placeholder sits in the specification, with nothing behind it yet.
A finished plan has no room in it for anyone else. This one has room, and we would rather say where.
The instrument
What the instrument refuses to do.
The instrument is Ranius Teacher Studio. A fifty-minute recording goes in. What comes out is not a score — it is a set of observations, each anchored to a timestamp, each carrying a confidence value, each stamped with the rubric, prompt, model and contract version that produced it.
Most products built on language models succeed by producing something plausible. This one has to keep saying "I cannot tell" and mean it. Below is what is implemented, not what is planned.
- nijin-core-1.0.0 rubric contract validated on every job
- 0 schema errors
- 0.800–0.882 confidence range carried on observations
- $1.164 / $0.505330 cost per lesson, estimated against actual
Values observed on the current build, September 2026.
Behaviour, as built
-
A null is not a zero
When overall_score is null we do not round it to zero, and the interface does not recompute score differences around the gap.
-
Versions are not compared across
Results produced under different rubric, prompt, model or contract versions are never compared. The interface refuses, and displays a comparison only after confirming the versions match.
-
"Cannot judge" is a state, not a value
A segment that cannot be judged — overlapping audio, a missing recording — does not receive a number. It moves to awaiting_human_review and waits for a person.
-
Nothing unconnected is invented
Where a source is not connected, the field says it is not connected. It does not show a guess.
-
Every observation is anchored to a timestamp
Each claim points back at the seconds of the recording it came from, so it can be checked against the tape.
-
Confidence is carried, not hidden
Observations report a confidence value. The range observed on the current build is 0.800 to 0.882.
-
The contract is validated on every job
Rubric nijin-core-1.0.0, schema errors 0.
-
Cost is measured against a ceiling
Estimated and actual, per lesson. On the current build, an estimate of $1.164 against an actual of $0.505330.
-
Human judgement is fed back in
A reviewer marks an observation as agreed, doubted, or impossible to judge, and that routes the segment to re-scoring and to human review.
-
Facilitators are not ranked against each other
The specification forbids it. The only comparison is against the same person’s own past.
-
Machine and human ratings are not mixed
Dimensions scored by a person are excluded from the machine average. The machine figure is shown as an aid, nothing more.
-
The score is arithmetic, not a model
The dimensions marked Machine are computed by fixed rules. What the language model writes is the narrative comment, nothing else. A frozen model has to return the same score for the same transcript later.
The six dimensions
| Code | Dimension | What it asks | Scored by |
|---|---|---|---|
| CI | Comprehensible input | Is the Japanese pitched where the child can follow it? | Machine |
| TD | Distribution of talk | Are turns skewed towards a few children? | Machine |
| CF | Corrective feedback and uptake | Does the correction actually land with the child? | Machine |
| NM | Negotiation of meaning | When something does not get through, does the exchange that fixes it happen? | Machine |
| SC | Scaffolding and safety | Is it a place where getting it wrong is fine? | Machine |
| OUT | Output quality and change | What did the child leave the session with? | Human |
Rated 0 to 4. The rubric of record is spec/rubric.ja.json, version 1.2.0, status: unvalidated.
What we would push back on ourselves
-
The instrument itself is not validated
There is no validated, session-level observation instrument for facilitators of small online groups, so we are building one. The factor analysis, the inter-rater reliability measurement and the crosswalk to the existing instruments (CLASS, PLATO, MQI) are all still ahead of us. What sits in the specification today for that crosswalk is a placeholder.
-
We observe the first fifteen minutes only
A single lesson observation has reliability of only 0.14 to 0.37; the first fifteen minutes carry about 60% of the reliability of the whole session at a third of the observer time (MET Project). That trade-off is a decision we are willing to defend and equally willing to have overturned.
The data
The half we have that most groups do not.
The scarce input in this question is not a model. It is recorded classroom teaching in Japanese, with consent, in a setting that already runs.
-
What is recorded
Fifty-minute classroom lessons, in Japanese, as video with audio — not text transcripts alone. Tone, pauses and silence survive.
-
Who is in the room
Children who are not attending conventional school, including children with developmental and other differences. More than 1,000 children have enrolled at NIJIN Academy in the three years since it opened, across four metaverse campuses and 48 sites across Japan.
-
What that setting has produced
97% of students have their attendance formally recognised by their home school, and more than 400 have returned to school or graduated. Around 100 instructors teach there, and about half have never taught professionally before.
-
Alongside the recordings
Rubric ratings, mentoring records and the children’s own responses, kept in a structure designed around what has to be recorded now in order to be verifiable later.
Consent, and what the system will not let us do
Every research conversation reaches this question. What follows is not a set of intentions; it is what the system running the school actually enforces. Operations that break these rules stop where they are.
-
No consent, no record
If a class roster contains a child whose guardian consent is not on file, the lesson transcript cannot be imported at all. The import names that child and stops.
-
Withdrawal means deletion
When consent is withdrawn, that child’s words are deleted. Deletion cannot be triggered from the interface; it has to be run deliberately from a terminal. That ordering exists so no family is told "it is deleted" before it is.
-
Words go, scores stay
Past the retention window — 180 days by default — a lesson keeps its scores and timestamps and loses the child’s utterances. What survives is enough to explain why a score was what it was, not what the child said.
-
Every viewing is logged
Opening a transcript is recorded, so a guardian can be told who looked and when. External raters reach a lesson only through the blind scoring screen.
-
Anonymisation for research use Not yet fixed
What a partner institution would receive — de-identified transcripts, ratings only, or video under a separate consent — is not settled. The consent we hold today covers operating the school and improving the instrument. Extending it to a named research collaboration is a conversation with families that we have not yet had, and we will not pretend otherwise to make a collaboration easier to start.
The figure above is cumulative enrolment since NIJIN Academy opened, not the number of children enrolled today. They are different quantities and we do not use one for the other. NIJIN Academy is run by NIJIN Inc., the parent company; Ranius’ own record begins with the December 2026 trial.
Contact
Where to start.
-
Who to write to
Tatsuro Hoshino, Founder and CEO, NIJIN Inc. — contact form. Japan or the United States, either is fine. To begin with it is enough to check whether the questions line up.
-
Where we will be
In the Bay Area from 26 October to 6 November 2026, through StartX.
-
The protocol as a document In preparation
The full protocol in English, as a PDF, carrying the same undecided items marked the same way. It will be linked here.
-
What we can share today
The rubric specification, the scoring rules, the version-pinning behaviour and the cost measurements, on request. The pre-registration, once filed, will be public.
-
What we cannot share yet Not yet fixed
Lesson recordings and child-level data. The procedure for granting a partner institution access — what agreement, what review, what de-identification — has not been written. See the consent note above.
-
Publication
We publish the December result whichever way it goes. That commitment is easier to keep if the stopping criteria and the margin are registered first, which is why they are marked undecided here rather than quietly assumed.
The fastest useful reply is the one that says which row of the design table is wrong.
Sources. Observation reliability, and the fifteen-minute figure: MET Project. Instrument behaviour, confidence range, contract validation and per-lesson cost: observed on the Ranius Teacher Studio build, September 2026; the rubric of record is spec/rubric.ja.json, version 1.2.0, status unvalidated. School figures — cumulative enrolment, campuses, attendance recognition, returns and graduations, instructor count: NIJIN Inc., the parent company. The teacher shortage of about 44 million by 2030 is a figure in wide international circulation; we have not traced it to a primary source ourselves, so we do not cite one. Every figure on this site carries its source, and anything we have not verified is labelled as unverified.