Standing is NihontoWatch's rating of a swordsmith or fittings maker's historical importance — the tier ("Foremost," "Leading," "Major") shown on every artist profile. It is read directly from Japan's designation record: every National Treasure, every Jūyō and Tokubetsu Jūyō designation, and the classical honors of the sword world, combined into one conservative measure. Because the record grows — a new Jūyō session nearly every year, Tokubetsu Jūyō every other — standing is recomputed from the living record and does not go out of date.
This page introduces the idea in plain terms first. The statistical machinery lives in the technical annex below, and the full derivation in the companion Elite Factor working paper.
What Standing Measures
Every collector eventually asks the same question about a maker — swordsmith or fittings artist: how important is this maker, really? Standing answers a precise version of that question: how deeply has Japan's institutional record recognised this artisan's work? — the depth, breadth, and rarity of the designations and honors attached to their surviving oeuvre.
It is deliberately not a judgment of the work itself — not the steel of a blade, not the carving of a tsuba. It is a measure of recognition: what the expert panels of the NBTHK, the pre-war government, and the classical connoisseurship tradition have collectively said about a maker's work, distilled into a single defensible rating.
A Century of Rating Smiths
Collectors have never lacked smith ratings. Two references sit on every serious bookshelf, and each embodies a distinct — and opposite — philosophy.
Fujishiro: relative to one's era. Fujishiro Yoshio's Nihon Tōkō Jiten (日本刀工辞典, first published 1935, revised for decades afterward by his brother Matsuo, the celebrated polisher) grades roughly 1,500 smiths on a five-step scale: chū-saku (中作), chūjō-saku (中上作), jō-saku (上作), jōjō-saku (上々作), and saijō-saku (最上作) at the summit. The crucial — and most often misunderstood — feature of these grades is that they are normalized to the smith's own era: graded within the Kotō and Shintō volumes, against contemporaries. A saijō-saku smith of the Muromachi period is not the equal of a saijō-saku Kamakura master; the grade says "supreme among his peers," not "supreme in all of history." This is a feature, not a flaw: Fujishiro tells you where a smith stands in his own world. (It is also a demanding scale — of the ~1,500 smiths rated, only around 65 achieved saijō-saku, the great majority of them kotō smiths.)
Tōkō Taikan: one absolute scale. Dr. Tokunō Kazuo's Tōkō Taikan (刀工大鑑, 1977) takes the opposite approach. Each smith receives a single valuation in yen, drawn from Dr. Tokunō's observation of the actual market — the price a hypothetical perfect example would command: signed, ubu, in fine polish, from the smith's best period. Because it is one absolute axis, it lets you place a Kamakura master and a Shintō smith on the same line — precisely what Fujishiro declines to do. (Hawley's English-language point ratings, familiar to Western collectors, track the Tōkō Taikan loosely and serve the same absolute-scale purpose.)
Both books are indispensable, and both share the same structural limits. Each is a snapshot of one expert's judgment, frozen at press time: Fujishiro's grades predate the NBTHK designation program entirely, the Tōkō Taikan's yen figures describe the market of 1977, and neither can register anything the sword world has learned since. And both rate swordsmiths only — the makers of tsuba, menuki, and kozuka never received a comparable standard scale.
The Record That Keeps Score
Since 1958 there has been another rating system accumulating in plain sight — one that no single expert wrote, and that updates every year.
The NBTHK's Jūyō shinsa is not a pass/fail authenticity check; it is a competition. Each session, submitted works — blades and fittings alike — contend for a limited number of designations, judged by a rotating panel against the best of what exists. Above Jūyō sits Tokubetsu Jūyō (held since 1971, roughly every other year), and above that the government's own ladder: Jūyō Bijutsuhin, Jūyō Bunkazai, and Kokuhō — plus the Gyobutsu of the Imperial Collection. Six tiers, each dramatically rarer than the last: of the ~25,000 designated objects in our corpus, some 13,600 are Jūyō and only about 107 are Kokuhō.
Count these designations per maker and an emergent ranking appears — not one scholar's opinion, but the accumulated verdict of independent expert panels across nearly seventy years. Collectors already reason this way informally: "Norishige has over a hundred Jūyō, dozens of them Tokujū" is a statement about Norishige's stature, and everyone understands it as such. Darcy Brockbank was the first to treat this record systematically, and standing builds directly on his foundation.
Two things must be added before the record can be read fairly:
Statistical care. Raw counts reward the prolific over the excellent. Naive ratios do worse: a maker with one Tokujū out of one submission scores a "perfect" 100% — ahead of Masamune. And any measure that only counts the elite tiers scores every tosogu master at zero, since fittings are almost never elevated past Jūyō. Standing weights each tier by its rarity, and treats thin evidence with explicit skepticism — a one-work fluke cannot outrank a maker with a deep, high-tier corpus. (How, exactly, is the subject of the annex.)
The classical honors. Designation counting alone would bury legends. The smiths of the Tenka-Goken (天下五剣, the "Five Swords Under Heaven"), Go-Toba's Goban-Kaji (御番鍛冶, the retired emperor's rotation smiths), the Sansaku (三作, the three greatest masters of tradition), the makers named in the Kyōhō Meibutsu-chō (名物帳, the shogunate's register of famed swords) — some of these artisans have only a handful of surviving designated works, yet their esteem in the tradition is beyond question. Standing folds these honors in as floors: a maker of one of the Tenka-Goken is Foremost, whatever the statistics of his thin surviving corpus would say alone.
This combination — the full six-tier designation record, weighted and read conservatively, plus the classical honors — is what standing measures. And because the NBTHK adds a Jūyō session nearly every year and a Tokubetsu Jūyō session every other, the record standing is computed from keeps growing. Where the 1935 and 1977 books are fixed in time, standing is recomputed from the live record: when a maker's works pass future sessions, their standing rises with them.
(The record can even grade itself. Reading Tokubetsu Jūyō promotions as the panels' second opinion on earlier Jūyō picks reveals that not all sessions certified to the same standard — see the companion paper Not All Sessions Are Equal.)
The Importance Ladder
A raw decimal score would read as false precision, so standing is expressed as a five-tier ladder — in the spirit of Fujishiro's saku grades, but derived from the designation record rather than one scholar's judgment:
| Rank | 漢字 | English | Reached by |
|---|---|---|---|
| 1 | 随一 zuiichi | Foremost | apex of the designation factor · Tenka-Goken · Sansaku · Kiku-Gyosaku |
| 2 | 屈指 kusshi | Leading | Kokuhō · Goban-Kaji · very strong designation record |
| 3 | 有数 yūsū | Major | Tokubetsu Jūyō · Meibutsu-chō · strong designation record |
| 4 | 著名 chōmei | Notable | at least one Jūyō designation |
| 5 | 列伝 retsuden | Recorded | in the index, no designation yet |
Two kinds of criteria appear in the right-hand column, and both can lift a maker. The statistical bands reward depth: a corpus of many high-tier designations earns Foremost or Leading on the numbers alone. The floors reward singular achievement: one Kokuhō guarantees at least Leading, one Tokubetsu Jūyō at least Major, one Jūyō at least Notable — and the great classical honors guarantee the top tiers outright. A maker's tier is simply the highest that any criterion earns.
Note what "Notable" means on this scale: at least one Jūyō — a work that won a competitive designation. Like Fujishiro's chū-saku, the lower rungs of this ladder are not faint praise; the mass of makers with no designated work at all sits below them.
Absolute and Relative, Both
Fujishiro chose to rate smiths relative to their era. Dr. Tokunō chose one absolute scale. Standing declines to choose — the same artisan is rated in up to four nested frames:
- Global — against all of history, on one absolute axis. This is the Tōkō Taikan's view, and here the Kamakura golden age dominates, as it should.
- Era — against Kotō, Shintō, Shinshintō, or Gendai peers.
- Period — against the finer historical period (Kanbun Shintō, late Muromachi, …).
- Tradition — within the five classical traditions (gokaden: 備前 Bizen, 山城 Yamashiro, 大和 Yamato, 相州 Sōshū, 美濃 Mino).
The era, period, and tradition frames are Fujishiro's insight, made explicit. They are what lets a great Shinshintō master legitimately read "新々刀随一 — Foremost of Shinshintō" even though, on the global axis, the Kamakura kotō cohort towers over everyone. An artist profile shows the most flattering frame a maker has honestly earned, and the full ladder behind it.
(A safeguard keeps the frames honest: no frame is awarded with fewer than ten designated members — there is no "Foremost of four.")
What Standing Is Not
Standing is not a verdict on artistry. We keep two axes separate and never conflate them:
| Standing | Artistry | |
|---|---|---|
| Question | How much has the institutional record recognised this artisan? | What is the aesthetic character of the work, and how fine is the workmanship? |
| Source | the designation record — Kokuhō through Jūyō, honors, provenance | the setsumei — the NBTHK's own descriptive prose |
| Nature | external · institutional · social | intrinsic · perceptual |
The two are correlated but not identical, and the honest caveats are worth stating plainly:
- The corpus is designated work only. An undesignated masterpiece is invisible to standing until it passes shinsa.
- Structural facts leak in. Survival, signature, and condition help a blade pass shinsa; smiths whose works survive ubu and signed are advantaged over those whose tachi were shortened centuries ago. Standing carries those confounds honestly rather than pretending to correct for them.
- Every score is a floor, not a guess. Standing reports what the evidence can defend, not a best estimate — so it will rise, never fall, as more of a maker's record is catalogued and as new sessions designate more of their work.
In one line: standing reads the ever-growing designation record — weighted by rarity, tempered by statistical skepticism, rescued by the classical honors — and expresses it as a Fujishiro-style ladder on both absolute and era-relative axes.
Technical Annex
Everything above can be verified from what follows. This annex specifies the estimator, the weights, and the assignment rules; the complete derivation — corpus construction, the seven source publications, worked examples, and validation against the Tōkō Taikan and Fujishiro ratings — is in the Elite Factor working paper.
The Problem, Formally
Of roughly 13,500 artisans in the directory, most have one or two designated works; a few have hundreds. Any ranking must reward genuine excellence without being fooled by a one-of-one "100%" fluke.
Raw counts favour the prolific regardless of quality. Simple ratios of elite to total designations overweight small samples — a smith with one Tokujū out of one total scores a perfect 100%, ahead of Masamune, ahead of Tomonari. A binary "elite vs standard" split ignores the graded hierarchy entirely: a Kokuhō is far rarer than a Tokubetsu Jūyō, yet both count equally under a binary model. And tosogu makers with dozens of Jūyō — decades of NBTHK recognition — would score zero simply because fittings are rarely elevated to Tokubetsu Jūyō.
The answer to every one of these failure modes is the same statistical idea.
Bayesian Shrinkage
Every standing metric uses Bayesian shrinkage: small samples are pulled toward a skeptical prior, and the uncertainty of a thin record becomes its penalty. No raw ratios, no point estimates. Where a frequentist would report "1 of 1 = 100%," we report a conservative lower bound that knows a single observation proves almost nothing.
Concretely, each metric reports a lower credible bound — the bottom of the range the evidence can defend — rather than its best-guess centre. A deep, high-tier corpus survives the shrinkage; a one-item fluke collapses toward zero. This single principle is applied identically to designation depth, elite rate, and provenance.
The Designation Factor
The designation factor (df) is the active stature metric. Rather than treating all designations equally, it weights each tier by its rarity, measured as self-information — the rarer the tier, the more a single object in it tells us.
For each tier, the weight is the negative log of its frequency in the corpus:
where is the count of all designated objects and the count in that tier. Each physical object contributes only its single highest designation (a blade both Jūyō and Tokujū counts once, as Tokujū). An artisan's evidence is then aggregated and passed through the shrinkage estimator:
Under the current population the tier weights are:
| Tier | 漢字 | Weight |
|---|---|---|
| Kokuho | 国宝 | 5.08 |
| Gyobutsu | 御物 | 3.97 |
| Juyo Bunkazai | 重要文化財 | 3.09 |
| Juyo Bijutsuhin | 重要美術品 | 2.78 |
| Tokubetsu Juyo | 特別重要刀剣 | 2.57 |
| Juyo | 重要刀剣 | 0.24 |
The shrinkage means a one-item record lands near zero while a deep, high-tier corpus is rewarded. The NihontoWatch stature value is then .
Why rarity-weighting rather than an elite count. A metric that only counted Kokuhō and Tokubetsu Jūyō scored every tosogu master at zero — their ceiling is usually Jūyō, so Yasuchika came out an unknown. Weighting Jūyō by its own rarity gives fittings a real, meaningful gradient instead of a flat zero.
The Gyobutsu discount. Imperial Collection (Gyobutsu) provenance for post-Nanbokuchō artisans reflects politics and patronage more than assessed quality, so those works are excluded from the Gyobutsu weight. For pre-1392 smiths, whose works entered the Imperial Collection as marks of genuine historical esteem, Gyobutsu counts in full.
Elite Factor and Provenance Factor
Two companion metrics share the same shrinkage philosophy.
The elite factor is the original stature metric and the direct ancestor of the designation factor. It models an artisan's elite-designation rate as a Beta-Binomial posterior with a skeptical prior — a base rate near 10% — and reports the 5th-percentile lower credible bound. It extends Brockbank's classical "pass factor" by adding uncertainty quantification, so a thin record cannot post a confident high score. The designation factor has since superseded it as the primary number, but the Beta-Binomial treatment remains the canonical way we handle a binary rate at small sample size. Its full methodology is documented here.
The provenance factor is a separate signal entirely: a Bayesian lower-bound prestige score derived from the named collectors and transmission history (denrai / 伝来) attached to an artisan's works. It captures social standing that the designation tiers alone miss, and it uses the same conservative shrinkage so a single illustrious former owner does not dominate.
The Assignment Rule
The continuous df is bucketed into the five-tier ladder shown above. The assignment rule is:
The designation and honor floors exist to rescue supreme-but-thin makers the prior would otherwise crush: Mitsuyo, maker of one of the Tenka-Goken, scores near zero on corpus depth alone, and is Foremost by floor. We are aggregating the importance the designation record already conferred (重要 = "important"), not re-grading craft.
The band cutoffs tighten or loosen with the reference frame: globally, Foremost is the top 1%, Leading 4%, Major 10%; within an era, period, or tradition, the bands open to 2% / 7% / 23%. Dropping the absolute designation floors at the non-global levels is what lets the greatest Shintō and Shinshintō smiths legitimately read "Foremost of Shinshintō," even though the Kamakura kotō cohort dominates in absolute terms.
Two safeguards keep the frames honest. A minimum-cohort guard suppresses any frame with fewer than ten designated members (so there is no "Foremost of four"), and a deterministic tiebreaker (df descending, then maker id) makes every ranking reproducible to the byte.
Cohort Versus Period
A subtle but important distinction underlies the labels.
A cohort is a ranking bucket, assigned by the midpoint year of an artisan's active span. The period cohorts are deliberately coarse — for example, the "Sue-Momoyama" cohort spans 1467–1595, grouping late-Muromachi (sue-kotō) smiths together with Momoyama masters into one ranking group. A cohort label therefore names the whole bucket, not the individual.
Rendered naively as a smith's period, that mislabels everyone in the early part of the bucket. Kiyomitsu (1532–1572) and Yosozaemon Sukesada (1504–1551) are pure late-Muromachi — roughly 80% of the Sue-Momoyama cohort never reached Momoyama at all — yet a cohort-as-period rendering would stamp them "… / Momoyama."
The fix is a per-artisan display period, derived from the artisan's own dates rather than the bucket's. It carries an English and Japanese label and a year span, and it follows the historical-period taxonomy (Heian; Kamakura ≤1333; Nanbokuchō ≤1392; Muromachi ≤1572; Momoyama 1573–1602; Edo ≥1603; Modern). So the period shown on a profile reflects when the smith actually worked, while the cohort label is reserved for the ranking frame it belongs to ("ranked among sue-kotō / Momoyama").
Safeguards and Honest Limitations
Safeguards. Bayesian shrinkage is applied everywhere — no point estimate is trusted at small n. Every band cutoff, floor, and cohort boundary lives in a configuration table the engine reads, so the system has a single source of truth and can be retuned in one place. Every ranking has an explicit tiebreaker and is reproducible to the byte.
Limitations, stated plainly. The corpus is truncated to designated work — every score is computed within an already-elite set, and an undesignated masterpiece is invisible to it. Standing is dominated by structural facts (signature, condition, era, fame) that correlate with, but are not identical to, quality. Each metric is a lower bound: it is the floor the evidence can defend, not a best guess, and it will rise as more of an artisan's record is catalogued. And cohort labels are coarse ranking buckets — which is precisely why the period pill is rendered from an artisan's own dates instead.
This is a working paper: the methodology and the underlying data curation are active and ongoing, and the numbers described here are subject to revision as the dataset grows.