Working Paper

Not All Sessions Are Equal

Sixty-six years of Jūyō shinsa, graded by the record’s own second opinion

×4×2×1×0.519601970198019902000201020201982 shinsa reformfounding erathe trough (1969–81)weakly identifiedpromotion odds vs average session
Figure 1 — The session-of-origin effect on Tokubetsu Jūyō promotion odds, net of maker composition, exposure, and Tokuju-session standards. Shaded band: 90% interval. Dashed segment (2018–2024): cohorts too young to identify reliably.

Two blades hang side by side, each with a Jūyō Tōken certificate. One passed session 2, in 1958 — a cohort of which seven in ten were later elevated to Tokubetsu Jūyō. The other passed session 26, in 1979 — a cohort of which one in twenty was. Same paper, same words, same seal. Not the same meaning.

This paper measures how the meaning of a Jūyō pass has moved across sixty-six years of shinsa — using nothing but the designation record's own second opinion of itself. The figure above is the result: the promotion odds of each Jūyō session's cohort, net of what those sessions contained, from the founding sessions through 2024.


The Record's Second Opinion

Since 1971, the NBTHK has run a court of appeal over its own decisions, though it is not usually described that way. Each Tokubetsu Jūyō session re-examines the accumulated pool of Jūyō blades — thousands of them — and elevates a few dozen under stricter standards. Every promotion is a later, harder panel affirming an earlier panel's pick. Every non-promotion, repeated across decades of opportunities, is quiet dissent.

Aggregate those verdicts by the session that originally awarded the Jūyō, and the record grades itself: sessions whose picks keep getting promoted were selecting to a higher standard than sessions whose picks keep getting passed over.

Raw promotion percentages cannot be read directly, for three reasons. A 1965 blade has had twenty-seven chances at promotion; a 2018 blade has had three. Each Tokubetsu Jūyō session has a fixed handful of slots contested by an ever-growing pool. And sessions differ in what they happened to contain — a session rich in Kamakura tachi by saijō-saku smiths will promote well no matter how rigorous its panel was.

So we model the promotion process as what it actually is — a repeated competition for scarce slots — and control for composition using ratings that predate the NBTHK record entirely: each maker's Fujishiro grade and Tōkō Taikan valuation, plus the item's era, form, and signature state. What remains, after all of that, is the session's own contribution: how much a pass from that session predicts later elevation, given who made the blade. The machinery is specified in the technical annex.


What the Record Says About Itself

Three eras emerge, and they are not subtle.

The founding era, 1958–1962 (sessions 1–9). The first nine sessions carry promotion odds 2 to 4 times the historical average — the strongest effects in the entire record, with the strongest statistical support. Of the 556 blades they designated, 205 went on to Tokubetsu Jūyō. This is the era when the great collections came out: the inaugural sessions were skimming a century of accumulated masterpieces, and the designation standard sat correspondingly high. A session-2 Jūyō was, in hindsight, a Tokuju-in-waiting.

The long slide and the trough, 1969–1981 (sessions 18–28). Odds fall through the 1960s and bottom out at ×0.37–0.41 in the sessions of 1974–1979. This is the largest cohort in the record — 3,840 blades, nearly a third of everything ever designated — and it is precisely the era collectors already regard with suspicion: the years of mass submissions and loosened standards that ended in the paper scandals. The record's own second opinion agrees with the folklore, and now puts a number on it: a mid-seventies Jūyō carries roughly 40% of the promotion-implied signal of an average one.

The reform and the slow climb, 1982 onward. The trajectory turns upward beginning with the sessions of 1980–82 — and here it is worth stating plainly: the model was told nothing about NBTHK history. No scandal, no reform, no dates. It located the inflection at 1982 — the year the NBTHK abolished the compromised regional paper system and centralized shinsa in Tokyo — entirely from promotion outcomes. The recovery is gradual rather than instant: odds climb through the late 1980s and 1990s and return to par around the early 2000s, consistent with a reformed institution tightening standards year by year rather than overnight.

(The most recent sessions, 2018–2024, show elevated estimates — but these cohorts have had at most three promotion opportunities, and the estimates there lean heavily on the statistical prior. We treat the recent edge as unresolved; see the caveats.)


How To Read a Jūyō Paper Now

For a collector, the practical content of this result is that the session number printed on a Jūyō certificate is information, the same way a date on any credential is information.

  • A Jūyō from the founding sessions (1–9) historically carried elevated odds of being Tokubetsu Jūyō material — those cohorts have been affirmed by later panels at rates no other era matches.
  • A Jūyō from the trough (sessions 18–28, 1969–1981) certifies less, as a pass. The designation was a wider net in those years, and the record's later verdicts reflect it.
  • A Jūyō from the post-reform decades sits near the long-run average, with standards tightening gradually into the 2000s.

Two readings of this must be firmly separated. This is a statement about the average information content of the pass, not about any individual blade. Exceptional blades passed through the trough sessions — hundreds of them were later promoted — and their quality was never diminished by the company they kept. What changes with session is the strength of the inference a pass licenses, exactly as a saijō-saku rating means one thing among Kamakura smiths and another among Shintō smiths. Session numbers are the eras of the designation record.


What This Does Not Say

  • It does not grade individual blades. A blade's workmanship, condition, and attribution are what they always were. Session effects are cohort averages.
  • The recent edge is unresolved. Sessions from 2018 onward show high estimates, but young cohorts are the hardest to identify — age-since-designation and session-of-origin are nearly the same variable at the edge, and the smoothing prior extrapolates. Session 70 (2024) has zero promotions so far, and its estimate is essentially borrowed from its neighbors. Revisit in a decade.
  • Owner behavior is unobserved. Promotion requires an owner choosing to submit. We assume submission propensity is unrelated to which session a blade passed decades earlier, given the controls — plausible, unproven.
  • A small share of promotions have unobservable origins. 90 of the ~1,300 Tokubetsu Jūyō objects (7%) lack a digitized Jūyō sibling record, so their session of origin is unknown. They are excluded; we assume this missingness is origin-neutral.

One incidental finding deserves its own sentence, because it bears on a standing debate. Among comparable works, a den attribution carries no promotion penalty whatsoever (the effect is −0.03 on the log-odds scale — statistically zero). Unsigned works as a class promote somewhat less than signed ones, but the den qualifier itself costs nothing: the panels' later verdicts treat den-attributed works as full members of their maker's corpus. Anyone arguing that den works should be discounted in maker rankings now argues against the NBTHK's own revealed behavior.


Toward Standing

Our artisan standing system currently weights every Jūyō designation identically, whichever session awarded it. These results justify session-aware weights — a founding-era Jūyō is demonstrably not a trough-era Jūyō — and the machinery to apply them is a one-line multiplication in the designation factor. We have deliberately not yet adopted them: a change that moves published rankings deserves to have its evidence published first, examined, and argued with. This paper is that evidence. If it survives scrutiny, session weighting follows.


Technical Annex

The Data

12,369 blades designated Jūyō across sessions 1–70 (1958–2024), deduplicated to physical objects via sibling records; 1,129 of them observed promoted across the 28 Tokubetsu Jūyō sessions (1971–2024). Analysis restricted to blades (tosogu are almost never elevated to Tokubetsu Jūyō and would only load on controls). Designation dates are taken from each record's shinsa date; session-to-year mapping is derived empirically from the records themselves, not assumed. 90 Tokubetsu Jūyō objects lacking a Jūyō sibling record are excluded (7% — assumed origin-neutral missingness).

The Model

Promotion is a repeated competition: at each Tokubetsu Jūyō session tt, a risk set RtR_t of all not-yet-promoted Jūyō blades contests a few dozen slots. The exact likelihood for such data is the risk-set choice model — the Cox partial likelihood. Because per-set promotion probabilities are small, it is well-approximated by logistic regression on item × opportunity rows (222,644 of them) with Tokubetsu-Jūyō-session fixed effects:

ηit=μ+αs(i)+γt+f(ageit)+βXi\eta_{it} = \mu + \alpha_{s(i)} + \gamma_t + f(\text{age}_{it}) + \beta^\top X_i

  • αs\alpha_s — the session-of-origin effect, the estimand. Adjacent sessions share panels and standards, so the α\alpha sequence carries a random-walk smoothing penalty, plus a weak level penalty for identifiability. Sparse sessions shrink toward their neighbors; a session estimate deviates only when its data insists.
  • γt\gamma_t — Tokubetsu Jūyō session fixed effects. These absorb slot scarcity, pool size, and any drift in Tokuju standards, which therefore never need to be modeled.
  • f(age)f(\text{age}) — a spline in years since designation, absorbing seasoning and depletion dynamics (the fitted hazard declines with age: most promotions happen young).
  • XiX_i — composition controls: era, form, signature state, a den indicator, and the maker's Fujishiro grade and Tōkō Taikan bucket. The last two are the identification linchpin: they are fixed expert judgments from 1935 and 1977, causally upstream of the designation record, so controlling on them cannot leak our own designation-derived scores back into the session estimates. αs\alpha_s is thereby a within-maker-stratum deviation — "items by comparable makers promote less when they came through session 23."

Validation

Composition controls behave. Every control points where domain knowledge says it should: Fujishiro is monotone at the top (saijō-saku +0.70, jōjō-saku +0.42, jō-saku −0.19), Tōkō Taikan is monotone, kotō dominates (Muromachi −1.44, Edo and later −1.18 against a Kamakura reference), signed works promote more than mumei (−0.42), tachi outperform wakizashi. A misspecified model would not reproduce the field's known structure this cleanly.

Permutation nulls (two). Shuffling session labels globally (exposure structure intact) collapses the α trajectory: real SD 0.63 against a null of mean 0.13, max 0.17 across permutations — the observed trajectory carries roughly five times the variance of noise, outside every permutation. A stricter within-decade shuffle, which by construction preserves decade-scale structure and tests only fine session-to-session signal, is also exceeded by the real data in every permutation.

Temporal holdout. Trained on Tokubetsu Jūyō sessions 1–21 and evaluated on the held-out sessions 22–28 (74,927 rows, 283 promotions), the session effects improve out-of-sample log-likelihood over the same model with α zeroed. The effects are predictive, not just descriptive.

The 1982 inflection was not supplied to the model in any form and constitutes an out-of-band historical validation: the estimated trajectory turns at the year of the NBTHK's shinsa reform.

Full Results

α is the session-of-origin effect on the log-odds of promotion; ×mult = eαe^\alpha, the odds multiplier against an average session. Estimates for sessions 64–70 are weakly identified (few promotion opportunities; see caveats).

SessionYearnPromotedRaw %αSE×mult
11958261142.3%+1.260.28×3.51
21958332369.7%+1.390.25×4.03
31959442454.5%+1.290.24×3.62
41960502040.0%+1.080.23×2.96
51960662334.8%+0.980.23×2.67
61961792835.4%+0.900.23×2.47
71961842631.0%+0.810.23×2.24
81962872731.0%+0.690.22×2.00
91962872326.4%+0.570.23×1.77
1019631593220.1%+0.390.22×1.47
1119632282812.3%+0.080.21×1.08
1219642813913.9%+0.020.21×1.02
1319651852211.9%−0.160.22×0.85
141966368297.9%−0.400.21×0.67
151967284248.5%−0.420.21×0.65
161967247187.3%−0.460.22×0.63
171968323299.0%−0.510.21×0.60
181969293144.8%−0.710.21×0.49
1919703994010.0%−0.570.19×0.57
2019713533710.5%−0.520.20×0.60
2119733713710.0%−0.640.19×0.53
221974349267.4%−0.880.20×0.41
231975491244.9%−0.950.20×0.38
241976472224.7%−0.960.20×0.38
251977342205.8%−0.970.20×0.38
261979370184.9%−0.990.20×0.37
271980224146.2%−0.710.22×0.49
281981176116.2%−0.640.22×0.53
29198213186.1%−0.620.22×0.54
301983187137.0%−0.520.22×0.60
311984221177.7%−0.550.21×0.58
32198510587.6%−0.550.23×0.58
33198719952.5%−0.690.22×0.50
341988138107.2%−0.580.22×0.56
35198921094.3%−0.560.21×0.57
3619902042210.8%−0.330.21×0.72
371991194126.2%−0.480.21×0.62
381992151149.3%−0.200.22×0.82
3919931491711.4%−0.150.21×0.86
4019941391510.8%−0.090.22×0.92
411995176137.4%−0.320.22×0.72
42199612054.2%−0.480.23×0.62
43199714564.1%−0.560.23×0.57
441998148138.8%−0.370.23×0.69
451999155106.5%−0.430.22×0.65
462000202136.4%−0.340.22×0.71
472001190136.8%−0.350.22×0.71
48200218894.8%−0.200.22×0.82
4920032273113.7%+0.090.21×1.10
5020041881910.1%+0.330.22×1.39
512005202199.4%+0.210.22×1.23
52200611597.8%+0.260.24×1.29
532007154149.1%+0.280.23×1.32
542008891112.4%+0.330.25×1.39
55200910832.8%+0.070.25×1.08
5620105323.8%+0.220.27×1.24
57201132618.8%+0.460.28×1.59
582012481020.8%+0.540.27×1.71
5920131041110.6%+0.280.25×1.32
60201412654.0%+0.060.26×1.06
61201516895.4%−0.050.25×0.95
622016150128.0%+0.190.26×1.20
63201714042.9%+0.160.26×1.17
642018138107.2%+0.620.27×1.86
6520191031211.7%+0.830.27×2.29
662020121108.3%+0.980.28×2.67
67202110754.7%+0.860.29×2.37
6820226723.0%+0.900.33×2.46
6920235647.1%+0.930.36×2.53
7020245000.0%+0.870.42×2.39

Reproducibility

The full pipeline — data assembly, model fit, both permutation nulls, and the temporal holdout — runs from a single deterministic script (scripts/_session-difficulty-analysis.ts, fixed seed) in under a minute against the live corpus. Estimates will drift slightly as new Tokubetsu Jūyō sessions land and as digitization of the remaining records completes; that is the point — like everything in the standing system, this analysis is recomputed from a living record.

This is a working paper: the methodology and the underlying data curation are active and ongoing, and the numbers described here are subject to revision as the dataset grows.

Companion papers. For the system these results feed into — the full artisan rating framework — read Artisan Standing; for the complete derivation of the designation factor, the Elite Factor working paper.