Vahag Byurat  ·  All Projects
Vanilla JavaScript · Python · Zero Dependencies

Learning Dojo

An adaptive flashcard trainer that runs directly from index.html: no framework, no build step, nothing to install. An Elo-based engine updates category mastery from self-grades, then picks the next card from recent accuracy and response time.

Elo mastery model Difficulty heuristics Matcher implemented twice 0 dependencies Runs off file://
The Premise
Eight script tags, and one of them is the deck

Open index.html straight off disk and the whole application is running: eight plain script elements, one stylesheet, progress in localStorage. There is no package manager, no bundler, and no module loader. The files load in dependency order and hang their one public object on window.

That constraint decided the deck format. The build tool writes decks/deck.js as window.DOJO_DECK = {…} rather than a JSON file, because a page served from file:// cannot fetch a sibling document: the origin is opaque and the request is blocked. A script tag is not. The identical payload is written alongside as deck.json for tooling that has a filesystem.

0Dependencies
60Sample cards
6 categories
39/60Auto-gradeable
22/22Matcher cases
passing

The sample deck is scaffolding. Sixty cards across Linux shell, IP and subnetting, DNS and DHCP, HTTP semantics, git, and testing concepts exist to exercise the engine. It is demonstration content rather than a complete course.


The Engine
It reads the clock, not just the score

Each category carries an Elo-like rating from 0 to 1000, starting at 330. Every card carries a rating too, fixed by its difficulty. An answer is scored 1 for got it, 0.5 for had to look it up, 0 for a miss, and the rating moves by K × (score − expected) where expected is the usual logistic, with a 300-point divisor instead of chess's 400. K decays from 34 to 13 as the category accumulates answers. These values are application heuristics, not a validated learning model.

The second input is time. Every difficulty carries a configured response-time budget used by the selection heuristic.

DifficultyCard ratingExpectedWhat it asks for
d119012 sSingle fact or command
d238518 sCommon working knowledge
d356527 sExplanation or short application
d474538 sMulti-step reasoning or edge cases
d590550 sDetailed application or debugging

A rolling window of the last eight answers yields two numbers: accuracy on that same 1 / 0.5 / 0 scale, and speed as the median ratio of elapsed time to the budget for that card's difficulty. Median, not mean, and only computed once at least three of the eight were actually timed: one interrupted card should not be allowed to redefine your pace.

acc ≥ 0.88 & speed ≤ 0.75×
Cruising. Push the served difficulty up hard (+0.55).
acc ≥ 0.82
Comfortable. Push, gently (+0.32).
acc ≥ 0.70 & speed ≤ 1.0×
Holding, and on the clock. A nudge (+0.12).
acc ≤ 0.42
Drowning. Back off (−0.60).
acc ≤ 0.60 or speed ≥ 1.7×
Struggling, or right but grinding. Ease down (−0.30).
bias ← prev × 0.62 + δ
Exponential smoothing, clamped to ±1.2, so the target drifts rather than lurches.

Target difficulty is then 1 + 4 × (mastery / 1000) plus the global flow bias plus a per-category bias, clamped to the 1-5 range. Card fit is a Gaussian around that target with σ = 0.95, multiplied by priority, by how badly the card has gone before, by how overdue it is, and by a final jitter of 0.86-1.18 so two identical sessions do not produce an identical queue.

Discarding abandoned timers

A card left on screen for more than 150 seconds is treated as “walked away”: the grade still counts toward mastery, but the stopwatch reading is thrown away rather than folded into your median. Without that, one coffee break would convince the engine you had become slow and it would distort the next set of difficulty choices.

A variety governor over the weakness bias

Weak categories are weighted up: the multiplier runs from 1× at full mastery to 3.2× at zero. But weak areas produce missed cards, missed cards come due in four minutes, and left alone that loop eats the whole session. So once at least eight answers are on the board, a category holding more than 45% of the last twelve has its weight cut to 0.45×; over 33%, to 0.72×. A category you are actively bleeding in is eased back another quarter rather than pushed harder.

Unseen material wins outright. While any card in the filtered pool has never been served, cards you have already answered are not candidates at all: not down-weighted, excluded. Re-serving a recent miss either rewards short-term recall or double-counts the same gap, and both poison the numbers the rest of the engine is reading.


A Design Decision
The confidence builder

Two consecutive misses, or three misses within five answers, trigger one easier card. This is a pacing heuristic chosen for the application, not a claim about learner psychology.

The next card comes from the learner's strongest category, with a difficulty ceiling set 1.4 below that category's normal target. The interface labels it ↺ confidence builder rather than presenting it as a normal adaptive choice.

Selection guards

Never twice in a row. A boost sets a flag that blocks the next one outright; the win has to be followed by real work or it stops meaning anything. Never a foregone conclusion. The easy pool must offer at least three candidates; if the strongest category cannot supply them the engine widens to any card at difficulty ≤ 2, and if that is still thin it abandons the boost entirely rather than serving something predictable. “Strongest” must be earned. A category needs at least three answered cards before it can be nominated, so a lucky first question cannot become the comfort zone. Fresh-first still applies. The confidence builder has to find its easy win among cards you have never seen, which is a real constraint rather than a cosmetic one.

It only runs in adaptive mode. The narrowed review modes (missed cards, flagged cards, weakest categories) keep their requested filters and do not use this adjustment.


Grading
Build-time validation and browser grading

Thirty-nine of the sixty cards can machine-check a typed answer. Those cards carry an accept pattern, a short token list in a sidecar file keyed by the card's index in its source deck. The same matching logic has to run in two places: in Python at build time, to prove a pattern is usable, and in JavaScript at runtime, to grade what you actually typed.

decks/accept/*.json accept tokens · 39 of 60 cards BUILD TIME Python tools/build_deck.py pattern_holds() token_matches() contains_token() drops any pattern the card’s own model answer cannot satisfy RUNTIME JavaScript app/check.js Check.grade() tokenMatches() containsToken() grades the typed answer as ok · partial · couldn’t confirm PYTHON TESTED tools/test_check.py 22 cases · JavaScript mirror not yet in harness
Python executes the 22-case suite. JavaScript mirrors the written behavior table at runtime, but a cross-runtime conformance harness is still missing.

The rules the table pins down are the ones that are easy to get subtly wrong. Matching is case-insensitive and anchored at the front of a word only. A bare substring test is wrong: the token skin matched inside asking. Anchoring both ends is also wrong, because patterns deliberately use stems: expir is meant to catch expires, expired and expiry. Tokens that begin with punctuation (/etc/fstab, -p 8080:80) skip the leading anchor entirely.

Numbers are compared as numbers, never as text. Integers must match exactly, so exit code 200 can never satisfy a card asking for 203 and 62 does not match “162 usable hosts”. Decimals match within 5% or 0.02, whichever is larger, so “about 6 dB” satisfies an answer of 6.02 and 9 dB does not.

A check that cannot pass its own answer

Every accept pattern is run at build time against the concatenation of its own card's model answer and explanation. If the card's own canonical answer would not satisfy the pattern, no human answer ever will (the check can only produce false negatives), so it is deleted from the built deck and listed in decks/REPORT.md. The card then falls back to self-grading. The current build drops none.

The grader never tells you that you are wrong

A failed match returns couldn't confirm, and a partial match says so explicitly rather than collapsing to a flat no. The verdict you act on is still your own self-grade. This is also why twenty-one cards carry no pattern at all: an explain-it-in-your-own-words answer has no checkable string, and a matcher that calls you wrong when you were right is worse than no matcher.


The Deck Pipeline
Rebuild the deck without losing your history

A card's identity is sha1(category | normalized question) truncated to twelve hex characters. Rewrite an explanation, retune a difficulty, add a tag, reorder the file: the id is unchanged and your progress on that card survives. Rewrite the question itself and you get a new card, which is the correct answer: it is now a different question. The one thing reordering does move is the accept sidecar, which is keyed by position rather than by id, and the build's self-validation is what catches it, because a pattern that has drifted onto the wrong card generally cannot match that card's own answer.

The payload deliberately carries no build timestamp. An unchanged deck produces a byte-identical deck.js, so rebuilding does not create a timestamp-only diff.

Bag-of-words collisions are rejected

The fingerprint strips code fences, lowercases, drops a stopword list, and sorts the remaining words into a set. Questions with the same normalized word set collide; the second is dropped and named in the report. This is a coarse duplicate heuristic, not semantic equivalence.

Near-duplicates are only flagged

Three-word shingle overlap within a category, reported above 72%. Flagged, not dropped: recall, apply and debug angles on one fact are legitimately different work, and the builder is not qualified to decide which.

Unknown categories are kept

A category name outside the canonical table is preserved verbatim and listed as suspect, rather than discarded. A typo in a category name should cost you a line in the report, not ten cards.

The report is the interface

Per-file kept and dropped counts, the difficulty spread per category, format mix, machine-checkable coverage, dropped checks, and coaching completeness. Currently 60 of 60 cards carry all four coaching fields.


Read-Aloud
Teaching say to pronounce sysadmin

The browser's speechSynthesis exposes only a subset of installed system voices, and Chrome quietly substitutes its own network voices. So the optional local server shells out to the macOS say binary instead: real Enhanced and Premium voices, exact rate control, entirely offline, with a SHA-1-keyed WAV cache pruned at 400 files. Premium voices render at 48 kHz because downsampling them to 22 kHz is audible. The voice name is checked against the parsed catalog rather than passed through, the rate is clamped to an integer range, and the text is handed over as a single argument after -- with no shell anywhere in the path, so a request cannot smuggle a flag into the subprocess.

Then there is the vocabulary problem: a speech synthesiser reads systemctl as a word. A pronunciation map fixes the ones that matter: systemctl becomes “system c t l”, nginx becomes “engine ex”, cidr becomes “cider”, yaml becomes “yamel”, sudo becomes “soo doo”. Code spans get read the way a person reads them aloud: -u becomes “dash u”, && becomes “and then”, >> becomes “append to”, and /etc/hosts becomes “slash etc slash hosts”.

The bug that only shows up on localhost

The server refuses to start if anything is already listening, and it probes ::1 and 127.0.0.1 separately. A leftover python3 -m http.server holds the IPv6 wildcard, which does not prevent a bind to the IPv4 loopback, so the new server would come up looking perfectly healthy while the browser, which resolves localhost to ::1 first, kept talking to the old one and 404'd every audio request. The startup message prints the lsof line to find it.


Limits
A trainer, not a course

The shipped deck is a demonstration. Sixty cards on fundamentals are enough to exercise every branch in the engine and nothing like enough to learn a subject from. Read them critically and expect to replace them. The authoring contract in CARD-SPEC.md is the real deliverable on that side: adding a subject means adding a row to the canonical category table and writing one JSON file.

Only 39 of 60 cards auto-grade

The rest are explain-it-out-loud answers with no single checkable string. The app shows you which is which rather than pretending to a confidence it does not have.

The contract runs on one side

The 22-case suite executes against the Python matcher; the JavaScript is a deliberate line-for-line mirror held to the same table. It is a written specification, not yet a cross-runtime harness.

It is not Anki

Intervals run from four minutes to a three-day cap and are biased toward this session, not a long-horizon calendar. It optimizes what to serve you next, not what to serve you in March.

localStorage and nothing else

No accounts, no sync, and storage is per-origin: progress saved on file:// does not follow you to localhost. Export and import exist precisely because of that.

High-quality read-aloud is macOS-only. Everywhere else the server still serves files and the app falls back to the browser's own voices, which work and sound like it.

Public, under an educational-use license. Roughly 340 lines of engine.js hold the mastery model, the flow controller and the selection weights; the surrounding code handles deck building, storage, grading, and the optional read-aloud server.