2026 MLB Season
Methodology
Toe the Slab turns every start into one comparable score, then uses recent form and matchup context to rank probable starters before first pitch.
How rankings work
context-v8
Start line
IP / K / ER / traffic
Length and missed bats lift the score; runs, hits, and walks pull it back.
Context
Park / opponent
The same line gets adjusted for run environment and matchup quality.
Display
20-80 GS+
The raw total is calibrated onto a scouting-style range so slate ranks are easy to scan.
Formula transparency
Raw GS+ starts at 45, adds length and strikeouts, subtracts runs, hits, and walks, then applies context and a public display transform.
The display cap keeps GS+ on the familiar 20-80 scouting scale. When a start reaches the cap, large score surfaces show the frozen raw value beside the displayed 80 so capped starts can still be compared.
Completed starts use line, park, opponent, and verified pitch-event context when available. Hitter-friendly parks add context credit for equivalent lines, and pitcher-friendly parks trim it. When a start settles, GS+ freezes with the context available at settle; later league-context updates do not move that final score. Upcoming cards use MLB team hitting splits vs the starter's handedness for OPS, K%, BB%, and ISO matchup context.
FAQ
The 20-80 scale is baseball's oldest shared grading language, credited to Branch Rickey's front office in the 1950s. It works like a standard score: 50 is major league average, and each 10 points represents one standard deviation away from it. In a normal distribution, three standard deviations on either side of the mean cover 99.7 percent of the population. That is why the scale runs 20 to 80 instead of 0 to 100: beyond three deviations there is almost nobody left to grade.
Scouts have used this vocabulary for decades. A 60 is plus, a 70 is plus-plus, and an 80 is reserved for the rare extreme: the best fastball, the best tool, the best of the best. Because the language is shared, a grade carries meaning instantly, with no legend required. GS+ adopts the scale so a start's number reads the same way a scouting grade does.
The cap is part of the scale's definition, not a limitation we regret. An 80 already means the extreme, so GS+ never displays beyond it. When a start is good enough to hit the cap, the raw pre-calibration score is shown beneath the 80, so historic starts still separate from merely great ones. On the other end, 20 is the floor for the same reason.
Further reading: FanGraphs: Scouting Explained, the 20-80 scale
Change note
Jul 2: exhibition games were removed from the dataset, regular-season archive rows were rebuilt, and league baselines were recalculated. Some recent GS+ scores moved slightly because the regular-season pool is now centered around league average.
Jul 2: settled starts now freeze GS+ and adjustment context at post-game reconciliation. A one-time season sweep applies that rule to completed starts so final scores stay fixed between polls.
Jul 4: completed-start GS+ moved to context-v8. Park adjustment now credits equivalent lines in higher run environments and trims equivalent lines in lower run environments. Starts settled before v8 retain their frozen pre-v8 park context until the P0-3 sweep. The x12 weight is unchanged; broader calibration remains a separate season-store review.
Calibration bridge
Game Score v2 is the familiar box-score benchmark. The standard formula uses runs; Toe the Slab substitutes earned runs from the pitcher line so unearned runs do not penalize the pitcher. It starts at 40, rewards outs and strikeouts, and subtracts hits, walks, and earned runs from the pitcher line. Toe the Slab stores GSv2 on the same canonical start record as GS+ so every page can show the same comparison.
The adjustment label is shown as GS+ minus GSv2. Positive adjustment means GS+ liked the start more than the box-score baseline after context; negative adjustment means the context and GS+ components pulled it down. GSv2 uses earned runs from the pitcher line so unearned runs do not corrupt the box-score benchmark.
On the homepage ticker, ▲ means a live provisional GS+ is at or above that starter's pregame projected GS+. ▼ means it is below the projection. If no projection is available, the comparison falls back to league-average 50.
Decision context
Decisions are context, not ranking inputs. A hard-luck flag marks a loss or no-decision with GS+ 60 or better. A vulture flag marks a win with GS+ 35 or worse. The chips exist to explain the story around the line, while the ranked order stays driven by GS+ and its visible line tiebreakers.
Data sources
Upcoming strikeout projection starts with the pitcher's season K/9, multiplies it by projected innings, and rounds to one decimal. Projected innings use recent workload when available, fall back to season innings per start, and are capped from 3.5 to 7.5 innings. Likely openers use a 2.0 inning profile.
K prop lines come from PropLine or The Odds API snapshots written by cron and read during render. Once a starter's game begins, the line stops updating and remains the last pre-first-pitch capture. Edges are projection minus line, shown at one decimal, with no pick language or recommendation.
Lines PropLine or The Odds API · 21+ only. For help call 1-800-GAMBLER
Watch Score Confidence
Watch score confidence uses the same recent-start window as form. HIGH means both probable starters have at least 3 qualified starts in the payload. MEDIUM means one side is below 3; LOW means both sides are below 3.
When a side is below the threshold, that side's form-derived watch component is multiplied by 0.85 before the game score is composed. The card then shows LIMITED or LOW CONFIDENCE beside the watch score. Baseline projected GS+ values are tagged BASELINE so placeholder-derived values do not read like measured form.
GS+ grades a single start on the 20-80 scouting scale, league average near 50.It starts with the pitcher's line, then adjusts for workload, traffic, runs, strikeouts, walks, park, opponent, and slate context. Daily boards rank qualified starts of at least 2.0 IP; openers and short outings are listed separately.
Form is a rolling view of recent GS+ across a pitcher's qualified starts. Heat Check bands highlight starters who are running above, near, or below their recent baseline.
Heat Check season rankings require roughly one start per 16 team games played. Arms below that bar remain visible below the leaderboard but do not receive ranked positions.
Watch score ranks probable matchups before the game. It combines the strongest arm in the game, the quality of the pairing, and the matchup context so the best games rise first.
Duel scoring rewards games where both starters bring strong current form. Mismatch scoring highlights games with the largest gap between the two probable starters.
Start here
The daily leaderboard is the canonical archive for completed starts and links each ranked line to its start log, pitcher page, and score breakdown.
View ranked starts