Bill Baran.

The model judges.
The server does everything else.

Every morning, scheduled agents post a reading digest of Substack and Medium to my Discord. Twice a week they build a Spotify playlist of new music. They run on Hermes Agent on a home server, using three MCP servers I wrote. The first versions were long prompts that did everything. The current ones give the model only judgment calls and handle everything else in code.

$3.26 → $0.27Medium digest, per run
25 → 8 minMedium digest, wall time
$0.00Substack digest & Spotify lanes, on a local Qwen 27B
By refthe server resolves every ref, so a summary can't land on the wrong post

From feed to Discord

Sources

Substack
Medium
Spotify + open web

MCP servers code

substack-reader-mcp
medium-reader-mcp
spotify-discovery-mcp
job-ledger: state, locks, refs, dedup

Hermes Agent models

cron → skill
main model: Sonnet or local Qwen
subagents: local Qwen 27B
Jev headline ranking (~1¢/day)

Output

Discord digests
Private playlist hermesYYYYMMDD
No-model report job

What the model sees, and what the server does

These are the worked examples from each server's docs, with fictional posts and tracks. A test checks them against real tool output, so every call and every result below is exactly what the tools produce.

Every morning · Sonnet main session, local Qwen subagents

Scheduler

Hermes cron starts the medium-digest skill

The model has no file tools in this job. It can only call the server and delegate reading.

Three MCP servers, one design

substack-reader-mcp ↗

Reads my Substack subscriptions, posts, chats and DMs, and runs the daily digest.

get_feed read_post search_posts list_chats read_chat_thread interests_evidence digest_begin digest_finish

  • Handles rate limits: fetches 3 publications at a time, backs off 2/5/10 s with jitter, and retries the ones that failed one at a time. A publication that keeps failing is retried next run.
  • Detects paywalls from the paywall-jump marker, because a truncated post still comes back as HTTP 200.
  • Gets a session on custom domains by reproducing the browser sign-in handoff. Never calls the endpoints that mark things as seen.

medium-reader-mcp ↗

Reads Medium Following and For-you feeds, including full member-only posts, with optional follows, lists and claps.

get_feed read_post rate_headings list_reading_lists interests_evidence digest_begin digest_finish

  • Mapped all 975 positions of For you. The top 25 are site-wide picks and positions 25–250 are personal. It never pages past 250, because doing that rebuilds your homepage.
  • rate_headings uses MCP sampling to have the client's own model rate headlines.
  • Jev headline ranking costs about a cent a day. Against 160 hand-labelled headlines, 10 of its top 10 were posts I wanted; a local 27B model got 7.

spotify-discovery-mcp ↗

Runs a discovery playlist for a scheduled agent. The code handles feeds, verification, dedup and state. The model only judges the music.

discovery_begin verify_tracks discovery_review discovery_finish discovery_status forget

  • Checks web finds with exact field searches, then compares the label on the ℗ line with what the source claimed.
  • Dedups by URI, by ISRC, and by a normalized artist + title key, which catches Spotify relinking.
  • Enforces pick limits by dropping from the end rather than refusing the call. Lanes learn which labels keep producing.

TypeScript on the official MCP SDK, with zod schemas and vitest. All three are MIT-licensed. The reader servers use each site's own web API with my session, read-only by default. State writes take a lock, go through a temp file and a rename, and keep a .bak, because Hermes runs two copies of each server at once.

Ten days of Medium runs

RunDesignCostTime
Sep 29One big prompt, Sonnet everywhere$3.2625 min
Oct 2Local subagents, same prompt$1.5319 min
Oct 3Digest tools: refs + server-side state$0.4213 min
Oct 4+ Jev headline ranking$0.278 min

Headline ranking benchmark (160 hand-labelled headlines)

JevLocal Qwen 27B
AUCArea under the ROC curve: the chance that a post I wanted is ranked above one I didn't, picking one of each at random. 50 is a coin flip and 100 is perfect.7675
NDCG@20Normalized discounted cumulative gain over the top 20: how close the first 20 headlines come to the best possible order. Positions near the top count most, and a must-read counts more than a read. 100 is perfect.7263
Wanted posts in top 1010/107/10
Time for 160 headlines3 s178 s

What I'd pass on

  1. Measure where the money goes. In v1, the orchestrator's bookkeeping cost 5× as much as the subagents doing the reading.
  2. Give the model refs, not data. Anything the model copies, it can corrupt. Once a digest went out with 10 of 12 summaries on the wrong posts.
  3. Save state inside the final tool call. A cron run ends at the last message, so a "then update state" step never ran.
  4. Fix reliability with tools, not instructions. Each new tool took a decision away from the model, and each of those decisions had failed before.
  5. Label your own data before tuning. Ten minutes of hand labels showed that my written interest profile was wrong, not the classifier.

Who acts at each step

A digest run. The model acts at steps 2, 3, 4 and 6.
One digest run, split between the model and the server
Modeljudgment only: which posts, which stars, what each one saysServer (code)fetching, ranking, refs, layout, state1
digest_begin
Fetches everything new, drops what was already reported, ranks every headline with Jev, gives each post a ref: F1 T2 Y3
2
Shortlist
Picks the refs worth reading, starting from the top of the ranked list
3
Subagents read
On the local model, five refs each. They never see a URL
3b
read_post "F3"
Looks the ref up in the run file and fetches that exact post
4
Pick and summarize
Chooses the ⭐ posts and writes each gist, by ref
5
digest_finish
Checks every ref, saves state, then lays out the finished message
6
Reply
Sends the message unchanged. State is already saved

A digest run. The model acts at steps 2, 3, 4 and 6.

A Spotify lane. Delivery is a separate job with no model.
One Spotify lane run, split between the models and the server
Modelsthe main model picks; a research subagent searches the webServer (code)Spotify, verification, pick limits, playlist, state, delivery1
discovery_begin
Today's playlist (under a lock), new releases from known labels and artists, refs K1 C1
2
Web research
Subagent looks for labels and artists the feed doesn't cover
2b
verify_tracks
Exact Spotify search, real label from the ℗ line, dedup, refs W1
3
discovery_review
Everything pickable, from the server's own records
4
Pick
Chooses tracks by ref, best first
5
discovery_finish
Applies the pick limits, adds only new tracks, saves state, writes the report
6
"Done: a-dnb"
The lane job delivers nothing itself
7
Report job
No model. Every 5 min, posts each finished report once

A Spotify lane. Delivery is a separate job with no model.