Agentic AI bookmarks, 2026
Links I bookmarked on agentic AI · 94 posts dated 2026 · each opened and read, tagged by theme
Read before trusting the numbers. Row 31 was deleted (X returns a 404). The jev cluster (16, 18, 22, 25) and a few others (59, 63, 64) are promotional, funneling to a paid article; their figures are the author's own and unverified, tagged
promo.| # | Date | Title | Author | Link | Summary | Tags |
|---|---|---|---|---|---|---|
| 1 | 2026-10-03 | Karpathy's loop applied to agent harnesses | @0xCodila | open ↗ | An NVIDIA/MIT paper (arxiv 2609.20519) runs Karpathy's propose → implement → review → test loop on the agent harness instead of the prompt. Four changes survived: fuse an edit and its test into one tool call; compact context only when the saving beats the cost of rewriting the cache; archive large outputs and keep a handle plus excerpt; hand log reading to a cheaper model and verify its evidence. Claims 50–54% lower API cost than native Codex and Claude Code on EdgeBench, and a third cheaper than Pi at 45–49% fewer tokens for ~94% of its score. A reply notes the missing 6% may be the task you care about. | agentsharnesscost |
| 2 | 2026-10-03 | The multi-agent software factory at Warp | @BHolmesDev | open ↗ | Ben Holmes diagrams the pipeline his team runs today. Work arrives mostly as Slack threads, sometimes Linear issues. A triage agent decides whether to build or ask questions; an implementation agent writes the code, sometimes drafting a spec with a subagent first; a verification subagent tests end to end; a code review agent cycles with the implementer before a human reviews and ships. Monitoring automation files issues from alerts. They are benchmarking whether triage and implementation need separate agents. A follow-up links Warp's "Factories as Code" page. | software-factoryagentsworkflow |
| 3 | 2026-10-03 | Semantic filters in SQL: Postgres vs. the analytical case | @sh_reya | open ↗ | Shreya Shankar quotes Avi Chawla's demo of asking semantic questions inside a database query (is this article about software engineering, how deep is it), and says Postgres, a transactional store, is the right shape for that use. Her own project, Quail (github.com/fsdatalab/quail), targets the analytical side. Replies are mixed: one calls the demo's implementation poor on fanned-out API calls, broken cancellation and credential handling; another asks how Quail holds freshness and transaction boundaries. | data |
| 4 | 2026-10-03 | Prompt cache control mod for Claude Code | @dani_avila7 | open ↗ | A Claude Code mod that surfaces the five-minute prompt cache. It adds a status bar for the cache and sends notifications so you can keep it warm, which he says saves more tokens than any other mod and lets sessions run longer. Install with `npx claude-code-templates@latest --mod observability/prompt-cache-control`. The post includes a short video of the bar in use. | costobservability |
| 5 | 2026-10-02 | Sandboxing and rogue agents | @vboykis | open ↗ | A one-line recommendation, not an argument of its own. Vicki Boykis points to a post on the Cryptography Engineering blog asking whether sandboxing is enough to contain rogue agents, and says she wishes AI discussion held to that standard. The substance sits in the linked article. | securitysandboxagents |
| 6 | 2026-10-02 | /retro: mine your agent sessions for repo friction | @mattpocockuk | open ↗ | A one-line prompt: have the agent read your last 10 coding sessions and find ways to make the repo easier to navigate: where it took too long to find information, where it leaned on stale docs. His argument is that navigability is an underrated way to cut token spend. The prior day's prompt in the same series, `/codebase-design`, hunts shallow modules and deletion candidates. | workflowcostcontext |
| 7 | 2026-10-02 | The "one true SaaS layout" | @thenanyu | open ↗ | Nan Yu posts a single layout diagram and claims it is what peak SaaS usability looks like. The argument lives entirely in the image; the text is two lines. Heavy engagement and a long reply thread arguing about it. | design |
| 8 | 2026-10-02 | NVX, Microsoft's micro-VM sandbox for agents | @gerhart_x | open ↗ | NVX is an ultra-light micro-VM sandbox for running untrusted workloads with hardware-enforced isolation. It is built on OpenVMM and runs Linux as a guest. Khudyaev adds that it appears to work with Windows Hypervisor Platform and Linux MSHV. Repo at github.com/microsoft/nvx. | sandboxinfrasecurity |
| 9 | 2026-10-01 | Make the model's output easier to understand | @karpathy | open ↗ | Karpathy's ladder for reading model output. Ask for ASD-STE100, the controlled English written for aerospace maintenance docs, or "80% of the way there" since the spec is strict. Better: ask for a diagram. Better still: ask for HTML, since models are good at frontend. Best: bespoke explainer videos, 3b1b style, with a narration API key. His framing is that as models do more legwork, our work rises into oversight, and cheap code makes large discardable artifacts worth building. | writingworkflow |
| 10 | 2026-10-01 | Google Mantis: security review as a skills pack | @dani_avila7 | open ↗ | Install with `npx skills add google/mantis`. Commands: `/mantis-threat-model` builds a threat model from the codebase, `/mantis-researcher` scans for vulnerabilities, `/mantis-review` filters false positives, `/mantis-reproduce` writes a PoC and runs it in a sandbox, `/mantis-patch` applies a fix and confirms it blocks the PoC, `/mantis-report` writes the report. He notes the reproduce and patch steps run generated code, so sandbox it. | securityskillscode-review |
| 11 | 2026-09-30 | gh-aw: agents inside GitHub Actions | @LLMpsycho | open ↗ | GitHub's agentic workflows extension runs Claude Code, Codex and Copilot agents inside GitHub Actions. 5,204 stars at the time of the post. One line plus a repo link, no commentary on how it behaves in practice. | ci-cdagents |
| 12 | 2026-09-30 | Design constraints for AGENTS.md | @tvykruta | open ↗ | A C/C++ veteran argues Claude writes the same spaghetti lazy programmers always wrote, just 200x faster: reaching through objects, over-DRY, speculative hooks, deep inheritance, god functions. His fix is a block for AGENTS.md listing ten design principles in priority order: separation of concerns first, then encapsulation, cohesion and coupling, DRY without over-DRY, KISS, single responsibility, depend on abstractions, YAGNI, composition over inheritance, open/closed. Hard rule: refactor to the principle first, then change behavior. | configagentscode-review |
| 13 | 2026-09-27 | jevgrep: context collection as a cheaper CLI | @dzhng | open ↗ | A research agent CLI powered by jev, claiming 40% lower coding agent cost verified on SWE-bench. It ships a skill so the coding agent knows to call `jg` for context collection rather than reading files itself. Repo at github.com/dzhng/jevgrep. Benchmark claim is the author's own. | jevcontextcost |
| 14 | 2026-09-25 | smolvm v1.19: credential vending by interception | @binsquares | open ↗ | A release note. The highlight is an interception mechanism for secure credential vending: grep and sed over outgoing requests, swapping placeholders for real credentials or other sensitive values before they leave. Short post, repo link only. | securityinfra |
| 15 | 2026-09-25 | 2,500 PRs in a month, the talk | @poteto | open ↗ | Lauren recorded the talk she had prepared for Cursor Compile in London and posted it free on X instead. It covers how she shipped 2,500 PRs to production in one month. Video post; no written summary in the thread. | workflow |
| 16 | 2026-09-25 | Route decisions to jev, writing to Opus | @starmexxx | open ↗ | The pitch: the model that decides should not be the model that writes. A router sends every yes/no/route/score call to jev, which returns a typed verdict with a confidence score in ~70ms and bills no output tokens; Opus 5.5 only wakes when something needs prose or confidence is too low. Components are a router, a policy file, a live cost log and an append-only ledger. Claims 1,204,882 decisions in a day with 6 escalations. Funnels to a paid article; treat the numbers as marketing. | jevcostagentspromo |
| 17 | 2026-09-24 | An AI coding dictionary, 71 terms | @undefinedKi | open ↗ | Matt Pocock's dictionary of AI coding vocabulary, as a repo and an interactive site. Each term gets a definition, links to related terms, and a "heard in the wild" line showing real usage; the whole thing also renders as a clickable knowledge graph. Seven sections: the model, sessions and context windows, tools and environment, failure modes, handoffs, memory and steering, patterns of work. Daily Claude Code users are pointed at failure modes and handoffs. | referenceskills |
| 18 | 2026-09-24 | jev as the decision layer under Opus | @Av1dlive | open ↗ | Claims jev plus Opus 5.5 cut his cost and time by roughly 80%. The division of labor: jev picks from options the harness prepares and validates: which project notes to load, when to route a task to a faster worker, which recovery path to take when a tool fails, which focused checks to run before the full suite, while Opus does the hard reasoning. Links to his builder's guide; the detail is behind it. | jevcostagentspromo |
| 19 | 2026-09-24 | The NVIDIA/MIT harness result, with the caveats | @undefinedKi | open ↗ | The clearest write-up of the same work as row 1. An agent rewrote another agent's harness, trying ~150 ideas across 500 environments; 4 survived, each required to hold quality within tolerance and cut at least one cost metric. The four: merge file edit and test run into one request; compact context only when it pays for itself; show tool outputs over 10KB in full twice then swap for a 1KB excerpt plus a handle; let a cheap model pull key lines from logs with a code check and full-log fallback. Caveats he names: Terminal-Bench 4 solved 15 tasks against 18, and it was tuned on GPT-5.6 Sol, so the tricks fire less on Opus 5. Code at NVlabs/SoL-Pi. | harnesscostevaluations |
| 20 | 2026-09-24 | What a software factory stack should look like | @zachlloydtweets | open ↗ | Warp's CEO argues for open, composable, defined-in-code infrastructure, layer by layer. Factories-as-code puts agent config, runners, repo access, integrations, webhooks and triggers in a version-controlled `factory.yaml`, which buys benchmarking, A/B testing and rollbacks. Above it: a context layer you host yourself with zero data retention; a compute layer of pausable, resumable remote dev environments; an inference layer that takes any model or harness; improvement infrastructure tracking cost-per-PR and LLM-as-judge quality; and an orchestration control plane. | software-factoryinfraconfig |
| 21 | 2026-09-21 | HelixDB: graph and vectors in one engine | @agenticgirl | open ↗ | An open-source Rust database for AI memory, knowledge graphs and RAG. Instead of embeddings in one store and relationships in another, it keeps graph and vector data in the same engine, and also handles KV, document and relational data. A single query can combine traversal, filtering and vector search; the Rust, TypeScript, Go and Python SDKs all send the same structured query AST. Runs in memory, on disk, or against S3-compatible storage. 6k stars, Apache 2.0. | knowledge-graphdatamemory |
| 22 | 2026-09-20 | jev founder talk: heavy promotion | @0xCodez | open ↗ | Promotes a 36-minute talk by jev's founder, Diogo Almeida, claiming 200x faster, 400x cheaper, zero hallucination and no human in the loop, and that Claude Code and Codex belong to a soon-ending "assistance era". The post calls the talk worth more than a Stanford ML degree and routes to an article on becoming a "Jev Engineer". No technical detail in the post; the claims are unverified. | jevpromo |
| 23 | 2026-09-20 | Tensorlake: snapshot the agent's whole machine | @agenticgirl | open ↗ | Tensorlake gives agents stateful Linux sandboxes on Firecracker microVMs, then snapshots them once the expensive setup is done: repo cloned, dependencies installed, Postgres running, failing test reproduced. The snapshot keeps filesystem, memory and running processes, and restores into one new sandbox or a hundred, so parallel attempts start from the same checkpoint. Underneath, a block overlay tracks dirty blocks at 4KiB granularity: their benchmark copies a 100GiB disk in 101 seconds, but takes an incremental snapshot of a 100MB change in 167ms. Their numbers, but the design is clear: snapshot cost follows what changed. | sandboxinfra |
| 24 | 2026-09-20 | Google's ax: Kubernetes rethought for agents | @rakyll | open ↗ | Jaana Dogan's first reveal of what her team has been building: Kubernetes reinvented for agentic workloads with statefulness and fast resumption, plus an open agentic orchestrator and runtime for Google. Repo at github.com/google/ax. Announcement only: no architecture detail in the post. | infraagents |
| 25 | 2026-09-19 | Prompt for finding your own jev-shaped decisions | @Av1dlive | open ↗ | He shares the prompt he handed Codex and Claude to scan his workflow and list every decision point where a cheap decision model could replace a full LLM call. The prompt is the whole post. Quote-tweets his own article on building an agentic memory factory with Kimi K3. | jevcostpromo |
| 26 | 2026-09-18 | Fine-tuning as a core skill | @sairahul1 | open ↗ | A short argument, not a how-to: fine-tuning is becoming one of the most valuable AI engineering skills, not because most people need a custom model but because understanding how models learn from data lets you build things others can't. The substance is in a linked full guide; the post is a teaser. | fine-tuningskills |
| 27 | 2026-09-16 | MCP over CLIs for most integrations | @trq212 | open ↗ | Thariq says he did not expect this, but MCP now beats CLIs for most integrations: tool calling has improved, tools can be deferred, and MCP is stateless. His practical note is that if you need to compose or filter data, add parameters like `query` to the MCP tools rather than reaching for a CLI. | mcp |
| 28 | 2026-09-16 | Graphify: map the codebase once | @techNmak | open ↗ | Turns a project into a queryable knowledge graph: functions, classes, files, SQL schemas, infrastructure, docs, PDFs, images and video as connected nodes. Not RAG: no embeddings, no vector DB, no LLM; code is parsed locally with tree-sitter and calls, imports and inheritance become edges. Every relationship is marked EXTRACTED, INFERRED or AMBIGUOUS so the agent knows what was guessed. It installs hooks for Claude Code, Codex, Cursor, Gemini CLI, Copilot and others that nudge the agent to query the graph before grepping; the graph commits to git and is exposed over MCP. | knowledge-graphcontextmcp |
| 29 | 2026-09-15 | Crawl, walk, run to cloud agents | @zachlloydtweets | open ↗ | A staged path from local interactive agents to automated cloud development, for teams who find the full software factory too much to adopt at once. The tweet is a pointer; the steps are in the attached article. | software-factoryworkflow |
| 30 | 2026-09-15 | A proposal for funding open source | @tannerlinsley | open ↗ | Tanner Linsley recommends Laurie Voss's article on paying for open source, saying that after years of trying to make his own projects sustainable he has not seen a proposal that comes this close to working at scale. He calls it atomic and says it could actually work. The argument is in the linked post. | reference |
| 31 | 2026-09-15 | (deleted) | @tspy | open ↗ | The post no longer exists: X returns a 404. Email subject was "Karpathy", so it was probably another entry in that week's Karpathy-method cluster, but there is nothing left to read. | deleted |
| 32 | 2026-09-15 | OpenAI's software factory | @GergelyOrosz | open ↗ | Gergely Orosz describes how OpenAI's agentic software factory works today, based on conversations with engineers there, with the full account in The Pragmatic Engineer. He flags "Perf Factory" as the piece he finds most interesting. The tweet carries the diagram and the link. | software-factoryagents |
| 33 | 2026-09-15 | tgrep: index once, search in milliseconds | @Fluyeporlaweb | open ↗ | In Spanish. Microsoft released tgrep so agents stop searching a repo like it's 2012 grep: index once, file watcher, millisecond search, written in Rust, open source, already used inside Copilot CLI. His line: if your agent takes longer to find the file than to edit it, the bottleneck is not the LLM. He pairs it with RTK: one trims output, the other trims search. | contextinfra |
| 34 | 2026-09-15 | What 80% AI-written code does to CI | @addyosmani | open ↗ | Claude now writes 80% of Anthropic's code and engineers ship 8x more per quarter. The side effect is the interesting part: tests grew 10x and CI jobs 25x in six months. Links Anthropic's write-up of what they did to keep that scaling. | ci-cdagentscost |
| 35 | 2026-09-13 | OpenViking: agent memory as a filesystem | @agenticgirl | open ↗ | An open-source context database holding knowledge, user memories and reusable experience in one place, organized as a virtual filesystem the agent browses with `ls`, `tree` and `read`. Rather than loading everything, summaries let the agent decide which files to open in full. Their benchmarks report 80–83% memory accuracy with input tokens down between 34% and 91%. | memorycontextagents |
| 36 | 2026-09-12 | DeepSeek v4.1 Flash vs. the original Transformer, in 3D | @petergostev | open ↗ | Peter Gostev had Astra read the DeepSeek v4.1 Flash paper and build an interactive 3D comparison against the original Transformer architecture, element by element, with zoom. His takeaway line is that things have changed quite a bit. The demo is hosted and linked; the post is the demo. | models |
| 37 | 2026-09-12 | alibaba/open-code-review | @ChrisShort | open ↗ | A code review tool Alibaba runs at their own scale, now open. Hybrid architecture: deterministic pipelines alongside an LLM agent, producing precise line-level comments. Ships a fine-tuned ruleset covering null pointer exceptions, thread safety, XSS and SQL injection. Works against OpenAI and Anthropic APIs. | code-reviewsecurity |
| 38 | 2026-09-12 | Kimi K3 on one CPU in 8GB of RAM | @techNmak | open ↗ | A 176KB C program runs a 2.78T-parameter model on one CPU with 8.24GB of RAM, and the checkpoint is still 1.56TB, sitting on NVMe. Most of it is MoE: 92 routed layers of 896 experts, 16 used per token, 82,432 experts totalling 1.447TB, or 93% of the file. The ~108.81GB of always-used weights are repacked into a contiguous trunk so each layer reads in one run, pinned to RAM where it fits and streamed where it doesn't. 69 of 93 layers use KDA, which carries fixed-size recurrent state instead of a growing KV cache. It is slow: ~26.5 seconds per token at 8GB, but RAM stops deciding whether the model runs at all. | modelsinfra |
| 39 | 2026-09-10 | Compiled wiki vs. filing cabinet | @v1lrok | open ↗ | A thread riffing on Karpathy's memory diagram. The argument: retrieval answers questions, compilation builds understanding. A notes folder only accumulates; a compiled wiki keeps raw source untouched, has the model read it once, connect it to what is already compiled, and file the synthesis, so later questions are lookups rather than searches. CLAUDE.md sits at the center as the file read before every session. Written as engagement bait and light on specifics: the architecture is in the linked article. | memorycontext |
| 40 | 2026-09-09 | Edge0: a 35B model on an iPhone | @SamuelZengML | open ↗ | Open-sourcing a framework for running large models fully on-device. The demo claim: a 35B language model on an iPhone at 1–2.5GB peak memory, with no cloud, remote server or desktop GPU. Announcement post; the mechanism is in the repo and the attached video. | modelsinfra |
| 41 | 2026-09-09 | What is actually inside a harness | @mardehaym | open ↗ | Notes in three tiers. The harness is everything wrapped around the model: instructions, scoped context, tools, a verifier, guardrails, and the model is the smallest, most swappable part. The hard-mode version has six layers: trigger (work starts from an event, not a button), orchestration (loop, memory, retries, limits as config), tools (numbers that must not vary run as fixed logic so the model doesn't improvise your KPIs), trusted context (called ~80% of agent success), control (golden sets, approvals, a named human decides), runtime (traces, cost dashboards, audit). His conclusion: the harness is the moat, because models and platform services change underneath it. | harnessagentscontext |
| 42 | 2026-09-09 | MiniCPM5-2B and the Densing Law | @itsPaulAi | open ↗ | A 2B-parameter open model that runs agents offline on a laptop or phone in 2GB of RAM, and handles coding and tool use well enough to browse Hugging Face and produce a CSV of top models unaided. Scores 23 on Artificial Analysis Intelligence Index v4.1.1, first among open models under 4B. The framing is density over size: Tsinghua and ModelBest's "Densing Law" claims the capability density of open base models roughly doubled every 3.5 months over the period studied. The release also opens training methods, agent data and the RL stack. | modelsfine-tuning |
| 43 | 2026-09-06 | ripwire: repo context without embeddings | @agenticgirl | open ↗ | From Red Hat Emerging Technologies. A zero-dependency C++23 binary parses 21 languages with tree-sitter and builds a deterministic structural map of a codebase: no embeddings, vector database, LLM indexer or daemon. It ranks symbols for the task at hand and attaches call relationships, complexity, git churn, change amplification and test coverage. Ask it what matters for "incremental cache invalidation" and you get the symbols, their callers, likely blast radius and the tests to run, inside a token budget. | contextinfra |
| 44 | 2026-09-05 | Portal cut Spotify's Claude Code tokens by 90% | @rseroter | open ↗ | Richard Seroter points at Spotify Engineering's post on Portal, which they report cut Claude Code token usage by 90%. His one-line read on why: they use the model for reasoning rather than for I/O. The detail is in the Spotify post. | cost |
| 45 | 2026-09-04 | One prompt over your whole work context | @rileybrown | open ↗ | Riley Brown gave Astra context over everything: email, Slack, business texts, Notion, every meeting note, about 20 projects, then asked one rambling planning question covering month, quarter and year. The prompt asks for flaws and strengths, which activities waste the most time, which to do more of, who is most dependable, where he is least dependable, the one skill to build this year, and how to reorganize the company, with charts only where they earn space. He shares it verbatim and calls the answer the most useful thing an AI has given him. | contextworkflow |
| 46 | 2026-09-03 | Semantic layer, context layer, ontology | @motherduck | open ↗ | Three things people collapse into one. The semantic layer is rigid logic that compiles to SQL: metrics and dimensions. The context layer is the unstructured human knowledge, docs and wikis, that agents need to decide anything. The ontology is the digital twin mapping real business entities. Points to Simon Späti's primer on arranging the three for agentic work. | datacontextknowledge-graph |
| 47 | 2026-09-03 | Warehouses are hitting the AI architecture problems first | @sethrosen | open ↗ | Snowflake, Databricks, ClickHouse, BigQuery and MotherDuck are adding models and agents to mature systems that were never designed around LLMs. Seth Rosen's point: that forces them to solve the same architecture problems the rest of software faces, and they may get there first. Links Josh Rosen's article working through the lessons. | data |
| 48 | 2026-08-31 | The post-AI data stack, and UI for verification | @sh_reya | open ↗ | On Ian Macomber's piece: analysis codegen is free now, so the hard part is getting agents to answers that are both correct and consistent over messy, fragmented data. The data team's job becomes encoding expert judgment into infrastructure so agents produce correct analysis without a data scientist in the room. Shreya extends it to interfaces: dashboards were UI-for-exploration, built for human pattern-spotting, but if the agent is the one reading the data, the UI's job may become helping humans verify the agent's reasoning. | dataevaluationsdesign |
| 49 | 2026-08-28 | Skill evolution through a persistent wiki | @dair_ai | open ↗ | A Google paper separating three things skill-evolution systems usually merge: raw execution traces, a persistent wiki of accumulated knowledge, and the executable skills. Experience consolidates into the wiki, and every later skill update builds on the wiki rather than on a scattered optimization history. Ablations show the wiki carries much of the gain. Two results stand out: smaller models with evolved skills beat substantially larger ones without, and skills evolved by one model transfer across families, sometimes beating self-evolved ones. arxiv 2608.27454. | skillsmemoryevaluationsresearch |
| 50 | 2026-08-28 | LLM cliché highlighter, 38 patterns | @simonw | open ↗ | Simon Willison's small browser tool that marks the tells in AI-written prose, now up to 38 patterns. One line and a link to tools.simonwillison.net. Useful as a checklist of what to strip from model output. | writing |
| 51 | 2026-08-24 | Models are overqualified hires | @JayaGup10 | open ↗ | The analogy: an overqualified employee reopens settled decisions, adds complexity, gets bored and leaves. A frontier model does the first two and never leaves, so the extra variance and cost run indefinitely. Her example: a password reset does not get better because the agent weighs twelve explanations, opens a security investigation and writes a personalized essay. Links her "Right-Sizing Your Intelligence Spend" piece. A reply raises the real difficulty: routing is hard when you cannot cheaply tell whether the cheap model's answer was good enough. | costmodels |
| 52 | 2026-08-19 | AI homework help, exam scores down | @paulnovosad | open ↗ | Diff-in-diff plots from a paper by Stromberg, Lei and Wu: students who lean on AI finish homework faster and score higher on it, then get crushed on exams, with exam scores falling almost in lockstep with homework effort shirked. Novosad's corollary for teaching: take-home work should be at most a tiny share of a course grade. He argues every syllabus should carry these graphs. | research |
| 53 | 2026-08-17 | Fix AI prose with the Google style guide | @natebjones | open ↗ | If you are tired of Claude-lish or Chat-lish, have the model read the Google Developer Documentation Style Guide and build a skill from it. He has tried both and prefers it to an ASD-STE100 Simplified Technical English skill. His aside to the labs: an agentic model that cannot write plain English is much less useful. He later posted his own skill repo, and a reader posted another. | writingconfigskills |
| 54 | 2026-08-17 | Anthropic's cost optimization cookbook (via a reply) | @Xxi5olc | open ↗ | The link lands on a one-line reply rather than the post worth reading. The parent, from @dani_avila7, points at Anthropic's cost optimization cookbook: a real agent going from $0.29 per task down 90% without losing accuracy, with model downgrade as the last lever rather than the first. The reply asks the fair question: if these approaches work, why are they not built into the harness? | cost |
| 55 | 2026-08-12 | Slimming down Claude Code | @EXM7777 | open ↗ | A long checklist for cutting token spend. Keep the global CLAUDE.md under 200 lines and rebuild it from scratch every few months, since long files get ignored. Run `/context` in a fresh session to see what fills the window before you type, and `/usage` to find which MCP server, skill or plugin is burning tokens. Prefer CLIs to MCPs, because a command returns only the lines you ask for. Toggle off unused MCP servers, mark rare skills manual-only, write "never do this" rules as permission settings rather than prose, reserve high effort for planning, `/clear` at checkpoints, and push heavy exploration to subagents so your window only sees the summary. | costconfigcontext |
| 56 | 2026-08-07 | kimi-k3-in-c, and the pushback | linkedin.com (Linas Beliūnas) | open ↗ | The same project as row 38, in LinkedIn form: a 176KB pure-C99 engine keeps the dense trunk in memory and streams experts from disk, 16 of 896 active per token, original MXFP4 weights, identical output from 8GB to 224GB. The top comment is worth more than the post. It points out the engine is CPU-only by design, and that even with 128GB+ of RAM the fastest it runs is 5.6 seconds per token, because generating one token means reading the ~108GB trunk at DDR5 speeds. A fine CS result; not a practical way to serve the model. | modelsinfra |
| 57 | 2026-08-01 | Semantica: knowledge graphs with provenance | github.com | open ↗ | A graph-native platform that turns enterprise data into queryable knowledge graphs carrying decision provenance. It reasons deterministically: forward chaining, Rete, Datalog, SPARQL, rather than leaning on embeddings and vector similarity, so the context an agent uses stays auditable. Covers ingestion from multiple sources, semantic extraction, conflict detection, deduplication and export to W3C PROV-O and RDF. Built for regulated industries that need explainability, self-hosted, meant to sit alongside an existing LLM stack. | knowledge-graphsecuritydata |
| 58 | 2026-07-29 | A prompt banning the AI prose tells | @0xPia | open ↗ | A paste-in list of prohibitions, and the most concrete version of this idea in the collection. No antithesis, corrective negation, paragraph pinning, parataxis, summary beats, negative parallelism or anaphora, contrasting pairs, rule of three, em dashes, throat-clearing openers, landing sentences, setup/payoff constructions, or parallel structures inside a paragraph. Also: vary sentence length unpredictably, drop stacked noun phrases, filler intensifiers, corporate verbs like leverage and underscore, nominalizations and hedges. Write for the spoken voice. | writingconfig |
| 59 | 2026-07-25 | CLAUDE.md config: promotional, no content | @cyrilXBT | open ↗ | Says a CLAUDE.md configuration making the rounds changed how he uses Claude, and that the difference was immediate. The post contains no configuration, no example and no link to one. The attached article is actually about running Kimi K3 as a coding agent. Safe to skip. | configpromo |
| 60 | 2026-07-18 | microsoft/Ontology-Playground | @thisdudelikesAI | open ↗ | Two links and nothing else: a free, open-source web app for learning about ontologies, plus a live preview hosted on GitHub Pages. The post is the pointer; the playground is the thing. | knowledge-graphdata |
| 61 | 2026-07-18 | An MIT professor's last lecture | @heyrohitai | open ↗ | A pointer post. Fifty years of teaching compressed into one recorded hour, made shortly before the professor died. No argument or summary in the tweet, just the framing and the video. | reference |
| 62 | 2026-07-11 | OpenRouter's 100-trillion-token study | @AnjneyMidha | open ↗ | Offered as the rigorous empirical answer to a question people keep asking: how much open vs. closed frontier model usage is there actually. The OpenRouter study covering 100 trillion tokens is up on arXiv, led by Maika Thoughts, Alex Atallah, cclark and team. | researchmodelscost |
| 63 | 2026-07-04 | Loop engineering, behind clickbait | @0xNoryxx | open ↗ | The framing is invented: an engineer supposedly fired over an 11-page PDF, but the loop underneath is sound. Schedule → Discover → Build → Verify → Repeat. Discovery has the agent find its own work from failing CI, open issues and recent commits rather than a handed list. Verification uses a second agent told to assume the code is broken, on the premise that an agent grading its own work always praises it. Results persist to disk rather than a context window that gets flushed. Each task gets an isolated git worktree so parallel agents don't collide. | harnessworkflowagentspromo |
| 64 | 2026-07-03 | Four jobs for your tokens | @0xCodila | open ↗ | The idea: rather than spending more tokens, split the same budget across four roles: execute does the work, advise checks the direction, grade passes or fails against a rubric, dream inspects, learns, writes to memory and sharpens the next round. One AI doing all four becomes four doing one each, at the same cost. The quoted 15%-to-90% accuracy jump is attributed to Anthropic's product team but not sourced; treat it as marketing. | costagentspromo |
| 65 | 2026-07-02 | CRUX 2: can agents do open-ended AI research | linkedin.com (Sayash Kapoor) | open ↗ | An update on CRUX, their open-world long-horizon evals. Most research-automation work keeps a human in the loop, picks narrow verifiable problems, or uses a scaffold tuned to one question type, so strong results may say more about the scaffold. CRUX 2 tries the broad version. To avoid contamination they partnered with researchers at UK AISI, Toronto and Princeton who pose open questions from papers not yet public; the agent must produce a NeurIPS-quality paper and a reproducible codebase that those authors review. Week-long horizons, VMs and GPUs, $3,000 in API credits per paper, and the agent manages its own budget. | evaluationsagentsresearch |
| 66 | 2026-06-27 | Using local coding agents | @rasbt | open ↗ | Sebastian Raschka's link to his own article on running coding agents locally. A follow-up tweet in his thread, so the post itself is just the pointer; the write-up is in his magazine. | agentsmodels |
| 67 | 2026-06-25 | Codex data on the shift to agentic AI | @daveholtz | open ↗ | David Holtz announces the first public output of his part-time stint as a visiting economics researcher at OpenAI: a study using Codex data to document how fast work is moving to agentic AI. The tweet opens a thread; the findings are in the replies and the paper. | researchagents |
| 68 | 2026-06-17 | Shape suffixes for tensor code | @vboykis | open ↗ | Vicki Boykis points to the only thing Noam Shazeer has blogged publicly, naming tensor variables with their shape as a suffix, and says she uses it all the time. A small, durable coding habit rather than a framework. | reference |
| 69 | 2026-06-16 | Cutting enterprise Anthropic spend | @JayaGup10 | open ↗ | Jaya Gupta crowdsources best practices before presenting to the C-suite of a top-30 global company on reducing their Anthropic spend. The replies are the value here. Links her "Token Budget Wars" piece, whose argument is that enterprise AI has moved from adoption to allocation, and every function is now asked to quantify its AI ROI. | cost |
| 70 | 2026-06-13 | code2lora | @liliana_hotsko | open ↗ | The resource post at the end of a thread rather than the explanation: paper at arxiv 2606.06492 and data plus models at huggingface.co/code2lora. To understand what code2lora does you need the parent tweets or the paper itself. | fine-tuningmodels |
| 71 | 2026-05-23 | A Tufte skill for Claude's charts | @draparente | open ↗ | Frustrated with Claude's charts, she fed Tufte's book to Claude and had it generate a Tufte skill, which she says immediately produced simpler, better visualizations. The gist is linked. Quote-tweets Anjney Midha telling anyone running an AI lab to have their team read Tufte before publishing charts. | designskills |
| 72 | 2026-05-22 | CLAUDE.md template from Karpathy's rules | @PrajwalTomar_ | open ↗ | A template circulated on Reddit, with an anecdotal claim of accuracy going from 65% to 94% on one codebase. The rules: no filler openers, match response length to task complexity, show two or three approaches before anything significant, flag uncertainty rather than filling gaps with plausible text, touch only files related to the current task, describe before rewriting, ask before deleting or overwriting, re-confirm deploys and database drops every time, keep a MEMORY.md of decisions and what was rejected, and an ERRORS.md so failed approaches are not retried. | config |
| 73 | 2026-05-09 | Status in AI, a personal essay | @nrmehta | open ↗ | Nick Mehta riffs on Jaya Gupta's argument that a company's talent identity becomes its long-term moat in services industries, then writes mostly about status. His upbringing with Harvard and Einstein posters over his bed, the "college is dead" years of 2016–2024, status going quiet during COVID, and its hard return now that everyone in high-growth tech is in San Francisco comparing notes at the same parties, with status rising and falling on each model release. | reference |
| 74 | 2026-04-20 | arXiv link, no commentary | @HowToPrompt__ | open ↗ | A bare link to arxiv.org/pdf/2603.19312 with no description at all. Filed under "Jepa" in the original list, so presumably a JEPA paper, but the post says nothing about it. You will have to open the PDF. | researchmodels |
| 75 | 2026-04-17 | Architecture diagram generator as a Claude skill | github.com | open ↗ | Describe a system in plain English and the skill produces a dark-themed interactive HTML diagram of components, connections and data flows. Exports to PNG, PDF or clipboard. Uses semantic color coding by component type and emits responsive SVG in a single self-contained file that opens in any modern browser. No design experience assumed. | skillsdesign |
| 76 | 2026-04-13 | Marcus Hutchins on Mythos doing vuln research | @ananayarora | open ↗ | Ananay shares what he calls the best take going on Mythos finding vulnerabilities, from the researcher who stopped WannaCry. The argument is in the video clip; the tweet is a single line of endorsement. Pairs with row 81, which tests the claims empirically. | security |
| 77 | 2026-04-13 | Who owns the harness | @vikvang1 | open ↗ | Bouncing between Claude Code and Codex is normal, since the models are good at different things, but it means paying several subscriptions, rebuilding context each time, and being exposed when a provider changes the deal. His point: switching models is easy when things are stateless, and agents are not. The harness owns context, memory, preferences, compaction, tools and workflows. So the question is not which model you use but who owns the harness around it. | harnessagents |
| 78 | 2026-04-12 | One CLAUDE.md file, 15K stars | @akshay_pachaar | open ↗ | Derived from Karpathy's coding rules. The premise: LLM coding mistakes are predictable (over-engineering, ignoring existing patterns, adding dependencies nobody asked for), and predictable mistakes can be prevented with instructions. So one markdown file in the repo root gives the agent behavioral guidelines for the whole project. No framework, no tooling. His closing observation is that the best tools in the Claude Code ecosystem are often not software. | config |
| 79 | 2026-04-12 | Harness, memory, context fragments | @Vtrivedy10 | open ↗ | Working notes, not a finished argument. The harness's main job is routing data into the context window, and every loaded object is a "context fragment": an explicit choice by the designer about what the model needs right now. Agent memory differs from human memory in one way that matters: it accumulates across agents that can be forked and duplicated. As agents run for years, the volume of data they produce grows hyper-exponentially, which makes search, distillation and owning that data yourself the real problems. | harnessmemorycontext |
| 80 | 2026-04-09 | Silicon Valley running on Chinese open models | @petergyang | open ↗ | The receipts he lists: Cursor confirmed Composer 2 is built on Moonshot's Kimi K2.5; Cognition's SWE-1.6 is likely post-trained on Zhipu's GLM; Shopify saved $5M a year moving to Alibaba's Qwen; Airbnb's Brian Chesky called Qwen good, fast and cheap. He adds that Zhipu's GLM-5.1 now performs close to Opus on coding benchmarks. More in his post on the Anthropic/OpenClaw situation and what he saw in China. | models |
| 81 | 2026-04-09 | Small open models found the same bugs | @ClementDelangue | open ↗ | Quotes Aisle's test of Anthropic's Mythos vulnerability-research announcement. They isolated the code behind the showcased vulnerabilities and ran it through small, cheap open-weight models. Eight of eight detected the flagship FreeBSD exploit, including one with 3.6B active parameters at $0.11 per million tokens, and a 5.1B-active model recovered the core chain of the 27-year-old OpenBSD bug. The useful counterweight to row 76. | securitymodelsevaluations |
| 82 | 2026-04-09 | Alternatives to GitHub Actions | @vboykis | open ↗ | A question, not an answer: are there genuinely viable CI/CD alternatives outside GitHub and Actions right now. Prompted by Astral's post on open source security. The replies are where any answer would be. | ci-cd |
| 83 | 2026-04-06 | A multi-agent job search system, open-sourced | @PrajwalTomar_ | open ↗ | Built on Claude Code over a weekend, used across 740+ roles, and the author landed a Head of Applied AI job with it. Paste a job URL and it scores compatibility on 10 criteria so you only apply where you fit, rewrites the CV per role and generates ATS-optimized PDFs through Playwright, scans 45+ company career pages, prepares STAR answers specific to the description, and tracks the pipeline in a terminal dashboard. 14 skill modes, MIT licensed. | agentsworkflow |
| 84 | 2026-04-05 | DESIGN.md files from 31 real sites | @ihteshamali | open ↗ | awesome-design-md by VoltAgent collects DESIGN.md files extracted from Stripe, Vercel, Notion, Supabase, Linear, NVIDIA, Apple and others. Drop one into the project root, tell the agent to build a page that looks like this, and it has colors, typography, spacing, buttons, cards, shadows and responsive rules to work from. No Figma exports, JSON schemas or special tooling: DESIGN.md is a plain-text design system, a concept from Google Stitch, that models read natively. Aimed at escaping Inter, purple gradients and card grids. | configdesign |
| 85 | 2026-03-30 | Semantics belongs next to the schema | @kirsten_lum_ | open ↗ | A short reply, no link. Her claim from decades of practice: the only place semantics reliably got used was the schema itself, because nothing else would get adopted. Once the semantic layer is decoupled from the schema but still integrated with it, humans and agents can work at the semantic level and let the data platform compile down. | data |
| 86 | 2026-03-22 | HTML slides demo and themes | @Kangwook_Lee | open ↗ | A follow-up tweet pointing to the demo and theme gallery for his HTML slides work. Original list filed it under "Html slides Claude code skill", but this particular post carries only the link: the skill and its explanation are elsewhere in his thread and on his site. | skillsdesign |
| 87 | 2026-03-13 | How LLMs actually work, 42 slides | @paraschopra | open ↗ | Paras Chopra gave a two-hour talk on the mechanics of LLMs and posted all 42 slides. The tweet is two lines; the deck is the content. Useful as teaching material. | referencemodels |
| 88 | 2026-03-09 | Technical debt plus cognitive debt | @infinitehumanai | open ↗ | Connects AI coding to Peter Naur's 1985 essay "Programming as Theory Building": a program is a shared mental theory living in the people who work on it, and the code is only a lossy written representation you cannot rebuild the theory from. Before AI, building something gave you the theory for free, as a byproduct of the work. AI breaks that coupling: you can produce code without building the theory, so you now accrue cognitive debt alongside technical debt. The cognitive kind is worse because it hides: you can't tell you cannot reason about your own program. | referenceagents |
| 89 | 2026-02-08 | Memory and scheduling for Claude Code | @tom_doerr | open ↗ | One line and a repo link (github.com/ascorbic/macro…). A tool that adds memory and scheduling to Claude Code. No description of the mechanism or how well it works; you will need to open the repo. | memoryconfig |
| 90 | 2026-02-03 | Column storage for the AI era | @andrewlamb1111 | open ↗ | Andrew Lamb's talk on the AI use cases pushing changes in Apache Parquet and driving newer formats. He calls it somewhat academic. Recording on YouTube and slides on Google Docs, both linked. Relevant if you are thinking about what replaces Parquet for AI workloads. | data |
| 91 | 2026-01-29 | Agent harness architectures and memory | @aparnadhinak | open ↗ | Arize's post, drawn from building their own agent and working with many customers. The core argument: files plus Unix tools make a fixed context feel effectively infinite, and bash composes into complex work without needing tool-definition JSON, since the model already knows the commands. The historical parallel is 1980s CPU memory hierarchies: caches made memory feel fast, virtual memory made it feel infinite, with the filesystem now playing both roles for agents. | harnessmemorycontext |
| 92 | 2026-01-26 | Spec-driven development as the declarative limit | @karpathy | open ↗ | A reply, short but worth it. Karpathy calls spec-driven development the limit of the imperative-to-declarative transition: being declarative entirely. He points to Drew Breunig's "A Software Library with No Code" as an extreme and early example that he found inspiring. | agentsworkflow |
| 93 | 2026-01-17 | The shorthand guide to Claude Code | @giyu_codes | open ↗ | An endorsement rather than the guide. The underlying piece, by @affaan, is a full setup after ten months of daily use: skills, hooks, subagents, MCPs, plugins and what actually works. Read the quoted article, not the tweet. | configreference |
| 94 | 2026-01-13 | skills.md as a contract for Python repos | @tdhopper | open ↗ | Add a skills.md to the repo telling the model how your codebase actually works: style, patterns and footguns. Tim Hopper frames it as a small contract rather than documentation. Links the Dagster team's playbook and examples at pydevtools.com. | skillsconfig |