Research Notes

AI tools, certifications, books and small experiments, written down while learning. 61 posts.

Same Model, One Leaderboard Says 62.7, Another Says 99.9. Which One Do You Trust?
14 views

Same Model, One Leaderboard Says 62.7, Another Says 99.9. Which One Do You Trust?

Caleb Writes Code takes apart the GPT-6 Astra launch: of fourteen benchmarks in the announcement, only one overlaps with the popular composite index, and the same model on the same test jumps from 62.7 to 99.9 depending on whose harness runs it. He offers a new yardstick: token-efficient is not the same as cost-efficient. I checked my own week of usage and found output tokens are 8% of my bill. The slice the leaderboards measure is the slice that hurts me least. Educational notes and extension, not a purchase recommendation.

Four AI Subscriptions, One Line of Quotas: How I Decide Who Gets the Work
5 views

Four AI Subscriptions, One Line of Quotas: How I Decide Who Gets the Work

Paying for four AI subscriptions means four kinds of quota windows: five hours, a week, a month. Dispatching from memory left one account empty by Wednesday and another half unused at month end. I wired the open-source CodexBar meter into a single line of traffic lights that appears before every message, and now the line decides who does the mechanical work, who writes code, and who makes the judgment calls. Four rules that ignore quotas entirely are in here too. Personal workflow notes, educational.

Ten Dollars for One Request: The Day GPT6 Astra Drained Lunchuizhe's Balance, I Went Back and Checked My Own Bill
12 views

Ten Dollars for One Request: The Day GPT6 Astra Drained Lunchuizhe's Balance, I Went Back and Checked My Own Bill

Lunchuizhe put GPT6 Astra to the test: three chat questions cost a dollar, then one small Codex task turned a seven-dollar balance negative. His verdict was that the model isn't expensive, we're just poor. I pulled a full day of my own AI usage and found the bill tracks how much the model reads, not how much it answers. Educational notes and extension, not a buying recommendation.

A Subscription Beats Pay-Per-Use by a Mile. So Why Won't He Hand It Six Tasks at Once?
10 views

A Subscription Beats Pay-Per-Use by a Mile. So Why Won't He Hand It Six Tasks at Once?

Pushed by his commenters, 掄錘者 re-ran the same GPT6 task on a subscription instead of the API: nine dollars yesterday, 4% of a weekly quota today. He conceded the subscription wins, then admitted he no longer dares to fire six tasks at once the way he does with cheap models. I priced my own week of usage at list rates and found that a subscription doesn't save you money. It converts money into a clock that resets. Educational notes and extension, not a purchase recommendation.

Terence Tao Says Probability Is Only the Right Ruler When Things Repeat. What About the Decisions I Only Get to Make Once?
8 views

Terence Tao Says Probability Is Only the Right Ruler When Things Repeat. What About the Decisions I Only Get to Make Once?

Terence Tao spent eighty minutes on Big Think walking through his new book, Six Math Essentials. Four things stayed with me: where probability stops applying, why doubling down is bankruptcy compressed into a small number, why traffic keeps jamming after the accident is cleared, and how a correct model that fits worse at first gets killed by the data. I checked each against my own prediction ledger and my settlement-fan calibration, and every one of them landed. Educational notes and extension, not investment advice.

The Answer He Wanted Was No. The AI Gave Him Several Pages of Yes.
9 views

The Answer He Wanted Was No. The AI Gave Him Several Pages of Yes.

Terence Tao says mathematicians want two things: answers, and understanding. For centuries the two were inseparable, because getting an answer required understanding first. That link has broken — problems are being solved by people with no domain expertise, the answers are correct and verifiable, and nobody knows what happened. What AI lacks, he argues, is not generation but deletion. I checked one broken alarm and a 16 KB lessons file of my own against him.

I Bought an Expensive US Market Dataset This Year, Then Found This Public One
10 views

I Bought an Expensive US Market Dataset This Year, Then Found This Public One

An introduction to a public US market dataset: daily prices, earnings call transcripts, financial statements, company news, and delisted companies kept in full. Notes from a paying subscriber, with how to start and one assumption to know about. Educational and methodological.

AI Improving AI Doesn't Have to Mean Changing Its Brain
8 views

AI Improving AI Doesn't Have to Mean Changing Its Brain

A ten-minute video from Xiaotian asks whether AI self-improvement must mean changing model weights. He uses EvoX's swarm mode to argue that civilisations improve through institutions and cooperation, and AI can too. I put our own two 'AI improving AI' loops next to his claim: detection works, self-change never happened, and the missing piece of a swarm hit me personally this morning. Educational notes and extensions.

What Is Graph Engineering? Managing a Team of AI Agents Takes Two Diagrams — One Stays for Months, One Is Thrown Away After Use
11 views

What Is Graph Engineering? Managing a Team of AI Agents Takes Two Diagrams — One Stays for Months, One Is Thrown Away After Use

Graph Engineering blew up on X in July: one tweet, a Google PM's definition ninety minutes later, a 48-hour flood of posts. Kelly Tsai's video sorts it out — loops let an agent's behavior be written down, graphs let an agent organization be written down. In practice it's two diagrams: a stable org chart plus a disposable work plan. I checked it against how I run my own crew of AIs; even the way the division of labor grew is the same.

Codex vs DeepSeek Harness vs Hermes: The Real-World Answer Is There's No Best Agent, Only the Right Seat
13 views

Codex vs DeepSeek Harness vs Hermes: The Real-World Answer Is There's No Best Agent, Only the Right Seat

A hands-on video throws three AI agent tools at real work: building a forum theme, fixing bugs, daily ops and fending off an attack. Codex develops best, Hermes runs a month without breaking, and DeepSeek Harness — the prettiest codebase — is the one that crashes. Pairing a local Qwen3.8 27B for grunt work with cloud DeepSeek V4 Flash for finishing cost about 30 RMB for a full day; all-premium would run 100+. Two months of the difference buys a 4090.

Terence Tao's Coin Game: AI Did Four Things, Humans Did Four Things — I Checked My Own Division of Labor Against His
10 views

Terence Tao's Coin Game: AI Did Four Things, Humans Did Four Things — I Checked My Own Division of Labor Against His

Terence Tao walks through a real case of humans and AI solving a math problem together: Gemini filled a missing proof step, AlphaEvolve computed up to 16 piles, Aristotle translated proofs into Lean for machine checking — while posing the problem, making the conjecture, understanding the proof, and noticing two problems were the same all came from humans. The dividing line matches the rule I use to delegate work to AI: only hand over what a machine can verify.

The Cheap One Is Enough, the Expensive One Sits Idle: How Lunchuizhe Picks Models and Agents
7 views

The Cheap One Is Enough, the Expensive One Sits Idle: How Lunchuizhe Picks Models and Agents

In eighteen minutes, the YouTuber Lunchuizhe lines up every model and agent he has tested over the past few months: locally, only Qwen3.8 27B; online, DeepSeek V4 Flash first; Hermes as the daily assistant, Codex for code; Gemini and Grok, avoid. His ranking method matters more than his ranking: ask how much you use per month and how many steps a task runs before you decide whom to pay. I checked his conclusions against a month of my own measurements—two agree, one flips.

Local LLMs: Buy a GPU or Rent One? RunPod at NT$24 an Hour, After Actually Trying It
9 views

Local LLMs: Buy a GPU or Rent One? RunPod at NT$24 an Hour, After Actually Trying It

Yesterday I rented a cloud GPU on RunPod for the first time to run a small batch job: a 4090 at about NT$24 an hour billed by the second; four attempts, 567 GPU-seconds, a bill of about NT$3.4. This piece lays the two paths—rent and buy—side by side: when renting wins, when buying wins, and which common local-development uses—document vector indexing, batch screenshot reading, a local model as an agent's brain, fine-tuning, image and video generation—belong on which side. The code was co-developed and tested with my good friend Fred.

A Numerator Without a Denominator: Terence Tao Says AI Solves Problems Like a Tipsy Genius, and We May Be Optimizing the Wrong Thing
10 views

A Numerator Without a Denominator: Terence Tao Says AI Solves Problems Like a Tipsy Genius, and We May Be Optimizing the Wrong Thing

In an interview clip, Terence Tao lays out where AI in mathematics actually stands: four years from middle-school problems to cracking a few that humans were stuck on, but the success rate, the money burned, and how many problems were scanned before one fell are all a black box. His image is a well-read, slightly drunk person throwing out ideas nonstop; point it at a thousand problems and it solves fifty—not necessarily the fifty you wanted. I checked my own factor-validation ledger and found the same mistake: nobody had been recording the denominator.

Passing Taiwan's AICE AI Engineering Literacy Exam in One Week
12 views

Passing Taiwan's AICE AI Engineering Literacy Exam in One Week

An educational write-up: what the AICE exam covers, the concepts people mix up, how we prepared, and the trade-offs that got us past the line.

AICE Must-Know Concepts: The Whole Foundation on One Page
12 views

AICE Must-Know Concepts: The Whole Foundation on One Page

An educational cheat sheet for Taiwan's AICE exam: evaluation metrics, machine learning, the big-data ecosystem, neural networks, generative AI, and AI ethics, each with a memory hook.

Your Video Memory Shouldn't Live on Someone Else's Cloud
6 views

Your Video Memory Shouldn't Live on Someone Else's Cloud

Field notes from moving an entire 'watch a video, get a searchable multimodal memory' pipeline onto my own machine — capture, embedding, storage, search. One evening, three pitfalls, zero cloud services.

Is Qwen3.8 27B on Your Own Machine Right for You?
15 views

Is Qwen3.8 27B on Your Own Machine Right for You?

Notes after watching three hardware reviews from two YouTubers. A payback calculation, the three real costs of running a large model at home, and how to tell whether this is for you. A thinking exercise about tools, not buying advice.

When Exit Codes Start Lying: An Alarm That Rang for Seven Days and Nobody Understood It
15 views

When Exit Codes Start Lying: An Alarm That Rang for Seven Days and Nobody Understood It

Silent failure in automated systems: a tool that answered correctly then reported failure, and an alarm that fired daily under the wrong name. The full reasoning trail from one debugging session, plus three checks you can run on your own system today.

When the Genius Walks Away: Grok Catches Up, Google Loses Its Legends, Meta Returns to Open Source
17 views

When the Genius Walks Away: Grok Catches Up, Google Loses Its Legends, Meta Returns to Open Source

Notes after listening to TechWave EP150. Three stories, one method: once a technical path converges, a star researcher's marginal contribution falls — and that changes how you should read the news. Educational content, not investment advice; no stock recommendations or price targets.

Why Hardware Is Hard: The Three Most Valuable Lines from One Conversation
9 views

Why Hardware Is Hard: The Three Most Valuable Lines from One Conversation

No field updates. An over-eager supplier is a warning sign. Money is recoverable, time is not. Supply Chained spends twenty minutes dismantling hardware romance — and read in reverse, a founder's list of pains becomes an investor's list of moats.

A Thousand Sails Pass the Sunken Boat: A Freelance Engineer on the AI Content Factory
12 views

A Thousand Sails Pass the Sunken Boat: A Freelance Engineer on the AI Content Factory

A Chinese freelance engineer's YouTube channel takes a rare break from hardware to talk about its own industry: collapsing creator traffic. Short-form video is draining watch time while AI content factories let one person run ten channels; the platform's ability to detect AI content has an expiry date and is already wrongly banning humans. His conclusion: you aren't competing with AI, you're competing with people who wield it. His plan: use the three-to-five years in which being human is still provable. My extension: content can be mass-produced, liveness cannot; faces expire, records don't.

herdr: if you run a herd of AI agents, sooner or later you must answer "who's working, who's stuck?"
13 views

herdr: if you run a herd of AI agents, sooner or later you must answer "who's working, who's stuck?"

What herdr is, how to install and use it, and why we ended up borrowing its two best design ideas instead of wiring the tool into production.

Generative AI Certification — Complete Exam Cheat Sheet: 71 Key Terms + Chapter-by-Chapter Study Points
15 views

Generative AI Certification — Complete Exam Cheat Sheet: 71 Key Terms + Chapter-by-Chapter Study Points

A last-hours review tool for Taiwan's III Generative AI certification: a full 71-term key-terms reference in five groups, plus the high-frequency points and 'see X, pick Y' distinctions across all four areas — fundamentals, prompting, applied skills, ethics and law. A companion to my prep-method post. Educational sharing, not a question bank.