Pooled models across 16 seed companies and 719,758 published plot entries. Pick a crop, pick two varieties, put a field on the map: the model answers how often one beats the other there, and how sure it is.
Calibration is the number to read first. When a page says 70%, the first variety wins about 70% of the time on a fresh trial. Across all 102 fitted responses the median gap between stated and realised probability is 0.025, and on the two biggest crops it is 0.017.
The intervals are a little tight. Checked against held-out trials, a stated 95% interval contains the truth about 87% of the time, and a stated 50% about 45%. The shortfall is the same in every crop and at every sample size, so read the direction and rough size of a difference as sound and the interval as slightly optimistic.
Good hybrids, or a good showcase? These are the companies' own published trials, so two very different things get mixed together: a brand can score well because its genetics are strong, or because the plot was arranged to flatter it. They can be separated.
A company fills most of its own plot with its own hybrids — Bayer 85%, Pioneer 88% — so a brand's margin at home is measured largely against itself. And what a visiting hybrid scores depends mostly on whose plot it is standing in: the average guest is 4.01 bu/ac below the mean in Bayer's plots and 1.48 above it in Syngenta's. Comparing each brand only with the other guests in the same plot cancels both, because the host's lineup and the host's choice of who to invite apply equally to everyone visiting. That is the neutral column below. Beside it is what the same brand scores in its own company's trials, against the competitors that company chose to bring in.
| Brand | On neutral ground | In its own plots | Showcase gain | Verdict |
|---|---|---|---|---|
| Pioneer | +1.09 | +4.20 | +3.11 | genuinely strong |
| DeKalb | +0.97 | +5.40 | +4.43 | genuinely strong |
| Beck's | +0.20 | +1.79 | +1.59 | about average |
| Wyffels | +0.06 | +3.42 | +3.36 | average, flattered at home |
| Channel | +0.04 | +4.35 | +4.31 | average, flattered at home |
| Croplan | −0.27 | +1.56 | +1.83 | about average |
| AgriGold | −0.38 | +0.59 | +0.97 | about average |
| LG Seeds | −0.85 | −0.01 | +0.84 | slightly behind |
| Golden Harvest | −3.62 | −4.54 | −0.92 | behind, honestly shown |
| NK | −6.58 | −7.39 | −0.81 | behind, honestly shown |
Read the first column for whether the hybrids are good and the third for how much the company's own trials add. Pioneer and DeKalb are the only two clearly ahead on neutral ground, by about a bushel over the average competitor entry, and DeKalb holds that edge just as steadily in Pioneer's plots as in Syngenta's. Channel and Wyffels are the widest gap between reputation and performance: both are within a tenth of a bushel of average once the host effect is removed, yet their own trials show them three to four bushels up. Syngenta's two brands are the only ones that look worse at home than away — whatever else their corn is, their plots are not arranged to flatter it.
None of this is a claim that anyone is cheating. Showing your own lineup in depth against a handful of competitor checks is what a company plot programme is for. It does mean a single company's published trial is a poor guide to how its hybrids compare, and pooling all sixteen is what makes the comparison possible at all.
Most varieties are thinly observed, though most data is not. About 57% of corn and soybean varieties appear in two trials or fewer — but those account for only 4–7% of plot entries. A comparison between two widely grown hybrids rests on a great deal of evidence; one involving a tail variety rests on very little, and the interval will say so. Anything seen in fewer than two trials is not offered.
Moisture is close to unpredictable for canola and soybeans. Both are harvested to a target, so varieties barely differ within a trial — within-trial spread of 0.26 and 0.35 points against corn's 1.08. The moisture answers on those two crops rest on almost no signal. Sorghum's sources publish no soil at all, so its model carries no soil term.
Everything shared by a trial — the field, the season, the management — cancels from a variety-to-variety comparison, so the answer comes from within-trial contrasts rather than from comparing yields across places. What remains is conditioned on things knowable before the season: thirty-year growing-degree-day and water-balance normals read off a 0.2° Daymet lattice at the point you click, soil as a continuous clay-and-sand composition taken from the USDA triangle, tillage and irrigation. That season's weather is in the fit to explain error, never to predict — nobody knows it in advance.