This lesson0%

Latest tricks

Hands-onAutomationGitHub ActionsLLM API⏱ ~3 min readSep 28, 2026

Build a bot that collects AI news automatically

How this site's AI News section really works. 20 sources on a schedule, event clustering with cross-checking, five-factor LLM scoring and an automatic morning brief, all on GitHub Actions.

This site’s AI News section has no human editor. A scheduled bot does all of it. Here is the full setup so you can build your own.

The whole loop

A run starts every 6 hours and does five things:

  1. Collect 20 RSS feeds plus recent posts from a set of X accounts
  2. Dedupe by link and title
  3. Cluster reports of the same story from different outlets into one event
  4. Score each item with an LLM, which also writes a Chinese title and summary
  5. Publish the results as JSON committed to the repo, which triggers a site rebuild

One more run at 08:05 Beijing time rolls up the previous day’s picks into a daily brief.

Step 1: pick sources

The list lives in scripts/sources.json, one entry per feed with its name, RSS URL and owning entity. The 20 in use today:

  • Official: OpenAI, Anthropic, Google DeepMind, Google AI, Hugging Face, NVIDIA, AWS Machine Learning Blog
  • Media: TechCrunch AI, The Verge AI, MIT Technology Review, Jiqizhixin, QbitAI, InfoQ China, 36Kr
  • People and communities: Simon Willison, Latent Space, Hacker News
  • Papers: arXiv cs.CL, arXiv cs.AI, Hugging Face Daily Papers

WeChat official accounts have no RSS. The open-source WeWe RSS, self-hosted on your own server, turns them into feeds you can add like any other source.

One rule: go to the original source, not to aggregators. Many aggregators only allow personal reading, and republishing their data needs a license.

The entity field matters. Google DeepMind and Google AI both count as Google, so if both post the same story it counts as one outlet, not two.

Step 2: merge the same story into one event

The clusterer compares titles pairwise and merges close matches. After merging:

  • “Reported by N” is the cross-check. Each extra independent outlet adds 10 points, up to 20
  • Related X posts attach to the event as discussion signal
  • Every event keeps a stable URL, so later merges never break old links

Step 3: score with an LLM

I use a Doubao model on Volcano Ark (OpenAI-compatible, so switching providers means changing a URL and a model name). The model rates five factors from 0 to 10, weighted into 0–100:

Factor Weight Question
Impact 35% How many people, how far
Novelty 20% Is it actually new
Usefulness 20% Can readers use it
Credibility 15% Is the source reliable
Timeliness 10% Is it fresh

Weights live in a config file. Off-topic items and ads are capped below 15 and dropped; 60 and up goes into Picks, the rest into All.

High scorers also get a structured brief: one-line takeaway, what happened, key facts, background, why it matters, quotes and what to watch.

Key design choice: the pipeline never stops when the model fails. On API errors it falls back to rule-based scoring by source weight and keywords.

Step 4: keep heat separate from quality

Quality answers “is it worth reading”; heat answers “what is everyone talking about now”.

  • Picks sort by quality
  • Trending sorts by outlet count, X discussion and freshness, with heat halving every 24 hours
  • Events with no new reports for 48 hours are marked as settled

Step 5: run it on GitHub Actions

The workflow is .github/workflows/hot.yml; the schedule is two lines:

on:
  schedule:
    - cron: '0 */6 * * *'   # collect every 6 hours
    - cron: '5 0 * * *'     # 00:05 UTC = 08:05 Beijing, daily brief

API keys live in repo Secrets and are read from environment variables. After each run the new data is committed back, and the host redeploys the site.

Lessons learned

  • Push conflicts: if main moved during a run, the push fails. Rebase before pushing.
  • Foreign-language titles: readers skip English headlines, so the model translates them and the original stays in small print.
  • Minutes: a free GitHub account gets 2,000 minutes a month. Four runs a day plus the brief, at ten-odd minutes each, gets tight. Watch usage.
Lesson progress
0%