Skip to content
Ranius orchid + genius

Research protocol · trial begins December 2026

What we are about to measure, and what we have not settled.

What
A two-arm non-inferiority trial. Licensed teachers supported by human mentors, against unlicensed facilitators supported by an AI mentor, with the four-to-eight person teaching team held constant in both arms.
When
December 2026, at the online Japanese-language school we open for Japanese-speaking children living outside Japan. Registered before the first lesson is recorded, not after.
What we are asking for
A research lead and a measurement specialist. Five design values are still open — allocation, margin, primary outcome, stopping rule, sample. We would rather settle them with someone else in the room than alone.

This is not an overview of the company. It is the design, the instrument, the data, and the list of what is still undecided.

Design as of 8 September 2026. Every item marked Not yet fixed will be replaced with a value, and with the date it was fixed, before the trial begins.

The question

The question, in one sentence.

In December 2026 we run a non-inferiority trial: licensed teachers supported by human mentors, against unlicensed facilitators supported by an AI mentor, with the four-to-eight person teaching team held constant in both arms.

Whether a person without a licence can raise attainment is not the open question — that has been asked before. The open question is how far the cost of developing a teacher can fall while the teaching stays equivalent. The world is short about 44 million teachers by 2030, and that is a supply problem before it is a pedagogy problem.

The design

The design, including the parts that are not decided.

Six rows below are not decided yet. They are marked. They will be fixed and registered before the first lesson is recorded, not afterwards. If the design is wrong, we would rather hear it now than in December.

Item What it is today
Design Two-arm parallel non-inferiority trial.
Setting The online Japanese-language school we open in December 2026 for Japanese-speaking children living outside Japan. Fifty-minute lessons, in Japanese, recorded with audio and video.
Control arm Licensed teachers, supported by human mentors. Teachers hold a Japanese teaching licence and are coached by a person.
Intervention arm Unlicensed facilitators, supported by an AI mentor. No teaching licence. Trained in about three months, then coached from timestamped evidence taken out of their own recorded lessons.
Held constant The teaching team of four to eight people, in both arms. Same team size, same weekly rhythm, same curriculum. The licence and the source of the mentoring are what differ.
Allocation Not yet fixed The unit of allocation — child, class or facilitator — and the randomisation procedure are not decided. To be registered before the trial begins.
Non-inferiority margin Not yet fixed No margin has been set. We are not going to publish a number here and choose it later. To be registered before the trial begins.
Primary outcome Not yet fixed Under discussion: student attainment, and the observational rating on the six-dimension rubric below. A third-party instrument for student outcomes is also under discussion, so that the effect of an AI mentor is not measured only by the people who built it. Which measure carries the non-inferiority claim is exactly the decision we would like to make with a research partner rather than alone.
Secondary outcomes Not yet fixed Inter-rater reliability between the two external raters. The student survey. How long the facilitators we develop keep teaching. The cost of developing one teacher. The final list, and the analysis plan for each, are not fixed.
Blinding External raters never see the teacher’s name, which arm a lesson belongs to, or the machine’s values. There is no route into a lesson except the blind scoring screen. This one is not a plan — it is what the system already enforces.
Stopping criteria Not yet fixed The conditions under which we stop the trial have not been written down. They will be written before it starts and registered with the protocol, and they will not be edited afterwards.
Pre-registration Not yet fixed We intend to pre-register the trial together with the cost-analysis plan; the AEA RCT Registry is the venue we are looking at. Nothing has been filed and there is no registry ID yet.
Sample Not yet fixed Number of classes, number of children and the length of the observation window are not fixed. No power calculation has been done, because the primary outcome and the margin above are not fixed either.
Ethics and consent Guardian consent is a precondition for a lesson entering the system at all; withdrawal deletes the child’s words; utterances are dropped past the retention window while scores and timestamps survive; every viewing of a transcript is logged. The rules are set out under The data. We have not been through ethics review, and we have not yet decided whose review we will go through.

Six of fourteen rows are undecided. That ratio is the honest state of this trial on 8 September 2026, three months before it starts.

Allocation and outcomes One figure. The amber blocks are the parts that are not decided.
Allocation and outcome flow, December 2026 non-inferiority trial Children enrolling at the online Japanese-language school are allocated to one of two arms. The control arm is licensed teachers supported by human mentors. The intervention arm is unlicensed facilitators supported by an AI mentor. The four-to-eight person teaching team is held constant in both arms. Lessons are recorded, scored by blind external raters and by fixed computation on a six-dimension rubric, and compared against a non-inferiority margin. The unit of allocation, the sample size, the primary outcome, the non-inferiority margin and the stopping criteria are all not yet fixed. ENROLMENT Children enrolling at the online Japanese-language school, from December 2026 ALLOCATION — NOT YET FIXED Unit of allocation (child, class or facilitator): not yet fixed Randomisation procedure: not yet fixed Number of classes, children and weeks: not yet fixed ARM A — CONTROL Licensed teachers Supported by human mentors. Teachers hold a Japanese teaching licence and are coached by a person. ARM B — INTERVENTION Unlicensed facilitators Supported by an AI mentor. No teaching licence; trained in about three months, then coached from timestamped evidence. HELD CONSTANT IN BOTH ARMS The teaching team of four to eight people. Same team size, same weekly rhythm, same curriculum. WHAT IS COLLECTED Fifty-minute lesson recordings, in Japanese, from both arms. SCORED BY PEOPLE Two external raters, blind to the teacher's name, to the arm, and to the machine's values. Disagreement between the human and the machine resolves by a rule written before the trial starts. SCORED BY FIXED COMPUTATION Six-dimension rubric, rated 0 to 4. The language model writes the comment, never the score. A segment that cannot be judged is not given a number. it moves to awaiting_human_review PRIMARY OUTCOME — NOT YET FIXED To be named and registered before the trial begins. Under discussion: student attainment; the six-dimension rubric; a third-party instrument for student outcomes. SECONDARY OUTCOMES Inter-rater reliability. Student survey. How long the facilitators keep teaching. Cost of developing one teacher. The final list is not yet fixed. NON-INFERIORITY MARGIN AND STOPPING CRITERIA — NOT YET FIXED. TO BE REGISTERED BEFORE THE TRIAL BEGINS.

Allocation and outcome flow for the December 2026 trial. The amber blocks are the parts we would most like to decide with someone else in the room.

On a narrow screen, the figure scrolls sideways. Everything in it is also in the table above.

What is missing

What we do not have.

We do not have a research director or a measurement specialist. Those are the two people we most need. If the design above is wrong, we would rather be told now than in December.

  • No research lead, no measurement specialist

    The two roles the design most depends on are unfilled. Everything below follows from that.

  • No margin, no primary outcome, no stopping rule

    The decisions that determine whether December produces evidence or produces an anecdote.

  • No pre-registration on file

    Intended, not filed. There is no registry ID to check us against yet.

  • No validated instrument

    The rubric’s own status field says unvalidated, and we have left it saying that. Factor analysis and inter-rater reliability are still ahead of us.

  • No ethics review

    We have not been through IRB and have not decided whose review we will go through.

  • No crosswalk to CLASS, PLATO or MQI

    A placeholder sits in the specification, with nothing behind it yet.

A finished plan has no room in it for anyone else. This one has room, and we would rather say where.

The instrument

What the instrument refuses to do.

The instrument is Ranius Teacher Studio. A fifty-minute recording goes in. What comes out is not a score — it is a set of observations, each anchored to a timestamp, each carrying a confidence value, each stamped with the rubric, prompt, model and contract version that produced it.

Most products built on language models succeed by producing something plausible. This one has to keep saying "I cannot tell" and mean it. Below is what is implemented, not what is planned.

  • nijin-core-1.0.0 rubric contract validated on every job
  • 0 schema errors
  • 0.800–0.882 confidence range carried on observations
  • $1.164 / $0.505330 cost per lesson, estimated against actual

Values observed on the current build, September 2026.

Behaviour, as built

  • A null is not a zero

    When overall_score is null we do not round it to zero, and the interface does not recompute score differences around the gap.

  • Versions are not compared across

    Results produced under different rubric, prompt, model or contract versions are never compared. The interface refuses, and displays a comparison only after confirming the versions match.

  • "Cannot judge" is a state, not a value

    A segment that cannot be judged — overlapping audio, a missing recording — does not receive a number. It moves to awaiting_human_review and waits for a person.

  • Nothing unconnected is invented

    Where a source is not connected, the field says it is not connected. It does not show a guess.

  • Every observation is anchored to a timestamp

    Each claim points back at the seconds of the recording it came from, so it can be checked against the tape.

  • Confidence is carried, not hidden

    Observations report a confidence value. The range observed on the current build is 0.800 to 0.882.

  • The contract is validated on every job

    Rubric nijin-core-1.0.0, schema errors 0.

  • Cost is measured against a ceiling

    Estimated and actual, per lesson. On the current build, an estimate of $1.164 against an actual of $0.505330.

  • Human judgement is fed back in

    A reviewer marks an observation as agreed, doubted, or impossible to judge, and that routes the segment to re-scoring and to human review.

  • Facilitators are not ranked against each other

    The specification forbids it. The only comparison is against the same person’s own past.

  • Machine and human ratings are not mixed

    Dimensions scored by a person are excluded from the machine average. The machine figure is shown as an aid, nothing more.

  • The score is arithmetic, not a model

    The dimensions marked Machine are computed by fixed rules. What the language model writes is the narrative comment, nothing else. A frozen model has to return the same score for the same transcript later.

The six dimensions

Code Dimension What it asks Scored by
CI Comprehensible input Is the Japanese pitched where the child can follow it? Machine
TD Distribution of talk Are turns skewed towards a few children? Machine
CF Corrective feedback and uptake Does the correction actually land with the child? Machine
NM Negotiation of meaning When something does not get through, does the exchange that fixes it happen? Machine
SC Scaffolding and safety Is it a place where getting it wrong is fine? Machine
OUT Output quality and change What did the child leave the session with? Human

Rated 0 to 4. The rubric of record is spec/rubric.ja.json, version 1.2.0, status: unvalidated.

What we would push back on ourselves

  • The instrument itself is not validated

    There is no validated, session-level observation instrument for facilitators of small online groups, so we are building one. The factor analysis, the inter-rater reliability measurement and the crosswalk to the existing instruments (CLASS, PLATO, MQI) are all still ahead of us. What sits in the specification today for that crosswalk is a placeholder.

  • We observe the first fifteen minutes only

    A single lesson observation has reliability of only 0.14 to 0.37; the first fifteen minutes carry about 60% of the reliability of the whole session at a third of the observer time (MET Project). That trade-off is a decision we are willing to defend and equally willing to have overturned.

The data

The half we have that most groups do not.

The scarce input in this question is not a model. It is recorded classroom teaching in Japanese, with consent, in a setting that already runs.

  • What is recorded

    Fifty-minute classroom lessons, in Japanese, as video with audio — not text transcripts alone. Tone, pauses and silence survive.

  • Who is in the room

    Children who are not attending conventional school, including children with developmental and other differences. More than 1,000 children have enrolled at NIJIN Academy in the three years since it opened, across four metaverse campuses and 48 sites across Japan.

  • What that setting has produced

    97% of students have their attendance formally recognised by their home school, and more than 400 have returned to school or graduated. Around 100 instructors teach there, and about half have never taught professionally before.

  • Alongside the recordings

    Rubric ratings, mentoring records and the children’s own responses, kept in a structure designed around what has to be recorded now in order to be verifiable later.

Consent, and what the system will not let us do

Every research conversation reaches this question. What follows is not a set of intentions; it is what the system running the school actually enforces. Operations that break these rules stop where they are.

  • No consent, no record

    If a class roster contains a child whose guardian consent is not on file, the lesson transcript cannot be imported at all. The import names that child and stops.

  • Withdrawal means deletion

    When consent is withdrawn, that child’s words are deleted. Deletion cannot be triggered from the interface; it has to be run deliberately from a terminal. That ordering exists so no family is told "it is deleted" before it is.

  • Words go, scores stay

    Past the retention window — 180 days by default — a lesson keeps its scores and timestamps and loses the child’s utterances. What survives is enough to explain why a score was what it was, not what the child said.

  • Every viewing is logged

    Opening a transcript is recorded, so a guardian can be told who looked and when. External raters reach a lesson only through the blind scoring screen.

  • Anonymisation for research use Not yet fixed

    What a partner institution would receive — de-identified transcripts, ratings only, or video under a separate consent — is not settled. The consent we hold today covers operating the school and improving the instrument. Extending it to a named research collaboration is a conversation with families that we have not yet had, and we will not pretend otherwise to make a collaboration easier to start.

The figure above is cumulative enrolment since NIJIN Academy opened, not the number of children enrolled today. They are different quantities and we do not use one for the other. NIJIN Academy is run by NIJIN Inc., the parent company; Ranius’ own record begins with the December 2026 trial.

Contact

Where to start.

  • Who to write to

    Tatsuro Hoshino, Founder and CEO, NIJIN Inc. — contact form. Japan or the United States, either is fine. To begin with it is enough to check whether the questions line up.

  • Where we will be

    In the Bay Area from 26 October to 6 November 2026, through StartX.

  • The protocol as a document In preparation

    The full protocol in English, as a PDF, carrying the same undecided items marked the same way. It will be linked here.

  • What we can share today

    The rubric specification, the scoring rules, the version-pinning behaviour and the cost measurements, on request. The pre-registration, once filed, will be public.

  • What we cannot share yet Not yet fixed

    Lesson recordings and child-level data. The procedure for granting a partner institution access — what agreement, what review, what de-identification — has not been written. See the consent note above.

  • Publication

    We publish the December result whichever way it goes. That commitment is easier to keep if the stopping criteria and the margin are registered first, which is why they are marked undecided here rather than quietly assumed.

The fastest useful reply is the one that says which row of the design table is wrong.

Sources. Observation reliability, and the fifteen-minute figure: MET Project. Instrument behaviour, confidence range, contract validation and per-lesson cost: observed on the Ranius Teacher Studio build, September 2026; the rubric of record is spec/rubric.ja.json, version 1.2.0, status unvalidated. School figures — cumulative enrolment, campuses, attendance recognition, returns and graduations, instructor count: NIJIN Inc., the parent company. The teacher shortage of about 44 million by 2030 is a figure in wide international circulation; we have not traced it to a primary source ourselves, so we do not cite one. Every figure on this site carries its source, and anything we have not verified is labelled as unverified.