← Research Notes中文
Technology · 2026-04-22 · 10 min read · Deep study

Automatically Judging Day-Master Strength in BaZi: Using Blind Human Review to Surface Two Kinds of Systematic Error

Notes from an algorithm-independent blind human review: using the three factors of season, ground, and support to build a ground-truth baseline for judging day-master strength in BaZi, and using it to expose two systematic errors in automated judgment — under-rating by one grade, and the more serious case of getting the direction backwards.

When you turn BaZi chart-casting into a program, "deciding whether the day master is strong or weak" is just about the easiest step to get wrong. It cannot be settled by looking something up in a table: the same set of stems and branches will yield different conclusions from different schools, and any single automated scoring rule will systematically go astray on certain structures. This article records one hands-on exercise — ignore the algorithm's output, judge a batch of charts blind using purely traditional methods, build a "ground-truth baseline" for strength, and then use it to expose two kinds of systematic error in the algorithm.

Note

This article belongs to the study of traditional Chinese metaphysics and culture. What it discusses is the engineering reliability of a set of judgment rules — it is not fortune-telling, and it does not constitute advice on health, marriage, or any life decision. All charts in the text are synthetic examples, not the charts of real individuals.

Why "strong vs. weak day master" is hard to judge automatically

Day-master strength is not a binary label. Traditionally it is divided into at least several grades — strong, leaning strong, balanced, leaning weak, and weak — and it is the foundation on which everything downstream ("selecting the useful god, fixing the structure") is built. Get the foundation off by one grade and everything above it goes wrong. The difficulty is that strength emerges from a contest of several forces, and no single factor can call it on its own.

The most common approach to automated judgment is to score "supporting the self" (Seal, Peer/Rob) and "draining the self" (Officer/Killing, Output, Wealth) separately and then compare the totals. There is nothing wrong with the approach itself; the trouble lies in the weights. Once one factor (the month command above all) is given too much weight, the program will reliably misjudge an entire class of structures — not err randomly, once in a while. To catch this kind of systematic bias, you need a ruler that the algorithm cannot influence.

The three factors: season, ground, and support

Judging strength by hand, the classical method looks at three things: getting the season, getting the ground, and getting the support.

The key lesson from experience is this: a strong month command does not mean a strong day master, and a weak month command does not mean a weak day master. A day master born in a month that restrains and drains it, yet sitting on two strong roots with Peer/Rob revealed on the stems, can perfectly well be strong. Treating "the month command restrains and drains" as the decisive vote is precisely the pit that automated judgment most easily falls into.

Building a ground-truth baseline that does not depend on the algorithm

The method is plain: from several hundred test charts, pick the batch where human intuition and algorithm output might diverge, then cover up the algorithm's conclusion and judge each one independently again using only the three factors — season, ground, support — plus structure-and-useful-god logic, writing down the reasoning. The set of "human verdicts" you get this way becomes the ground-truth baseline for evaluating the algorithm.

Judging blind is deliberate: if you look at the algorithm's answer before judging, you will unconsciously be anchored and lose the point of the check. In this review, out of 15 candidate charts flagged as divergent, 13 were confirmed after blind judgment to be genuine, hard disagreements — while another 2 were cases where I misjudged on my own first pass and the algorithm was in fact right. Those two get their own discussion later; they turn out to be the most valuable part of this whole method.

Two kinds of systematic error

The 13 real disagreements split cleanly into two classes pointing in completely opposite directions.

Pattern A: under-rated by one grade

The algorithm has the direction right, it just pushed the grade down one level — what should have been judged strong was judged leaning strong, and what should have been leaning strong was judged balanced. The shared structure of these charts is: the day master has a strong root at the Prosperity-post, Blade, or Growth level, and the stems also reveal Seal or Peer/Rob, yet the algorithm failed to recognize the final outcome of "strong."

Take a synthetic example (not a real chart) to illustrate this structure:

A Jiǎ-Wood day master, born in the Yín month (Jiǎ's Prosperity-post, Establishing-Prosperity commanding the season), with two more Yín (Prosperity-post) stacked in the day and hour branches, and the year branch Mǎo as the Blade; the month stem Jiǎ reveals a Peer and the year stem reveals a Seal — three Prosperity-posts and one Blade, with both Peer and Seal revealed. This is a textbook strong day master.

A chart like this is strong at a glance; no human would dispute it. But because its scoring runs conservative, the algorithm only gives it leaning strong. Right direction, off by one grade — the impact on the user's impression is relatively limited, and it falls into the class you can rescue by gradually tightening the threshold.

Pattern B: direction reversed

This class is far more serious: the algorithm judged a chart that should have been leaning strong, or even strong, as leaning weak — not off by a grade, but with the direction entirely flipped. The shared structure is: the day master sits on a strong root at the Prosperity-post, Emperor-Prosperity, or Blade level, and the stems reveal two or more Peer/Rob, only the month command restrains or drains it (Officer/Killing, Output, or Wealth commands the season).

Synthetic example (not a real chart):

A Wù-Earth day master, with the month command Yǒu (Wù-Earth's Death place, Hurting-Officer commanding the season), seemingly losing the season and being drained; but the day and hour branches are both Wǔ (Wù's Emperor-Prosperity), and the year branch Xū is an Earth-vault root; three Wù characters line up across the year, day, and hour, with two Peer/Rob revealed — two Emperor-Prosperities plus a vault root, plus a chart full of Peer/Rob, and only the month command draining all the way through. Human verdict: Seal and Peer roots extremely stable, day master leaning strong.

The algorithm, however, weighted the "month command restrains and drains" vote far too heavily, overwhelming the weight of the twin Emperor-Prosperities beneath and the Peer/Rob revealed on the stems, and its conclusion tipped straight over to leaning weak. The root cause of this kind of error can almost certainly be pinned on the scoring rule under-counting the weight of "Peer/Rob revealed on the stems," or letting the month command's restrain-and-drain term override it.

Error typeShared structureSeverity
Pattern A · under-rated by one gradeProsperity-post / Blade / Growth strong root + Seal or Peer/Rob revealed on stems; direction right but grade too lowMilder, small impact on impression
Pattern B · direction reversedProsperity-post / Emperor-Prosperity / Blade-level strong root beneath + two Peer/Rob revealed, yet judged leaning weak because the month command restrains and drainsSerious, conclusion reversed

An unexpected bonus: even the "ground truth" itself can be wrong

I mentioned earlier that there were 2 charts where I misjudged during blind review and the algorithm was right instead. They deserve their own discussion, because they are a reminder that so-called "human ground truth" is not inherently reliable — it steps into pits just the same.

What these two self-corrections have in common is the same cognitive bias: attention gets captured by the one or two most conspicuous strong signals (a single Prosperity-post, two revealed stems), without weighing the whole board's restraining-and-draining and its real-versus-empty roots all together. The algorithm, by contrast, dodged this bias precisely because it "honestly counts the books." This also shows that when building a ground-truth baseline, beyond blind judgment you still need a second review against the written reasoning to squeeze your own confirmation bias out.

The priority order for fixes

Once you have these two classes of error in hand, the order of handling is clear:

  1. Fix the reversed direction (Pattern B) first. A wrong direction is far more serious than a wrong grade. The diagnostic method is to export and compare the intermediate quantities of these charts — root strength, the number of satisfied judgment conditions, the ratio of support-the-self to restrain-and-drain, and so on — to find the one anomalous variable they all share, which is very likely the down-weighted "Peer/Rob revealed on the stems" term.
  2. Then deal with the under-rating by one grade (Pattern A). Since the direction is already right, you only need to loosen the entry threshold for the "strong" grade appropriately, so that charts with "enough strong roots and enough support conditions satisfied" can cross smoothly into strong, rather than getting stuck at leaning strong.

There are two more general lessons. First, for this kind of judgment born of a multi-factor contest, do not let any single factor (not even the month command) hold a veto vote; the weights need to be regression-calibrated against a batch of human ground truth, not set by gut feel. Second, the evaluation baseline itself must be guarded against contamination — judge blind, write the reasoning, then review again; drop any one of those links and what you expose may be only your own bias, not the algorithm's error.

Amos · research.xishe.ai · Please credit when sharing