Agentic AI bookmarks, 2026

Links I bookmarked on agentic AI · 94 posts dated 2026 · each opened and read, tagged by theme

Read before trusting the numbers. Row 31 was deleted (X returns a 404). The jev cluster (16, 18, 22, 25) and a few others (59, 63, 64) are promotional, funneling to a paid article; their figures are the author's own and unverified, tagged promo.
#DateTitleAuthorLinkSummaryTags
12026-10-03Karpathy's loop applied to agent harnesses@0xCodilaAn NVIDIA/MIT paper (arxiv 2609.20519) runs Karpathy's propose → implement → review → test loop on the agent harness instead of the prompt. Four changes survived: fuse an edit and its test into one tool call; compact context only when the saving beats the cost of rewriting the cache; archive large outputs and keep a handle plus excerpt; hand log reading to a cheaper model and verify its evidence. Claims 50–54% lower API cost than native Codex and Claude Code on EdgeBench, and a third cheaper than Pi at 45–49% fewer tokens for ~94% of its score. A reply notes the missing 6% may be the task you care about.agentsharnesscost
22026-10-03The multi-agent software factory at Warp@BHolmesDevBen Holmes diagrams the pipeline his team runs today. Work arrives mostly as Slack threads, sometimes Linear issues. A triage agent decides whether to build or ask questions; an implementation agent writes the code, sometimes drafting a spec with a subagent first; a verification subagent tests end to end; a code review agent cycles with the implementer before a human reviews and ships. Monitoring automation files issues from alerts. They are benchmarking whether triage and implementation need separate agents. A follow-up links Warp's "Factories as Code" page.software-factoryagentsworkflow
32026-10-03Semantic filters in SQL: Postgres vs. the analytical case@sh_reyaShreya Shankar quotes Avi Chawla's demo of asking semantic questions inside a database query (is this article about software engineering, how deep is it), and says Postgres, a transactional store, is the right shape for that use. Her own project, Quail (github.com/fsdatalab/quail), targets the analytical side. Replies are mixed: one calls the demo's implementation poor on fanned-out API calls, broken cancellation and credential handling; another asks how Quail holds freshness and transaction boundaries.data
42026-10-03Prompt cache control mod for Claude Code@dani_avila7A Claude Code mod that surfaces the five-minute prompt cache. It adds a status bar for the cache and sends notifications so you can keep it warm, which he says saves more tokens than any other mod and lets sessions run longer. Install with `npx claude-code-templates@latest --mod observability/prompt-cache-control`. The post includes a short video of the bar in use.costobservability
52026-10-02Sandboxing and rogue agents@vboykisA one-line recommendation, not an argument of its own. Vicki Boykis points to a post on the Cryptography Engineering blog asking whether sandboxing is enough to contain rogue agents, and says she wishes AI discussion held to that standard. The substance sits in the linked article.securitysandboxagents
62026-10-02/retro: mine your agent sessions for repo friction@mattpocockukA one-line prompt: have the agent read your last 10 coding sessions and find ways to make the repo easier to navigate: where it took too long to find information, where it leaned on stale docs. His argument is that navigability is an underrated way to cut token spend. The prior day's prompt in the same series, `/codebase-design`, hunts shallow modules and deletion candidates.workflowcostcontext
72026-10-02The "one true SaaS layout"@thenanyuNan Yu posts a single layout diagram and claims it is what peak SaaS usability looks like. The argument lives entirely in the image; the text is two lines. Heavy engagement and a long reply thread arguing about it.design
82026-10-02NVX, Microsoft's micro-VM sandbox for agents@gerhart_xNVX is an ultra-light micro-VM sandbox for running untrusted workloads with hardware-enforced isolation. It is built on OpenVMM and runs Linux as a guest. Khudyaev adds that it appears to work with Windows Hypervisor Platform and Linux MSHV. Repo at github.com/microsoft/nvx.sandboxinfrasecurity
92026-10-01Make the model's output easier to understand@karpathyKarpathy's ladder for reading model output. Ask for ASD-STE100, the controlled English written for aerospace maintenance docs, or "80% of the way there" since the spec is strict. Better: ask for a diagram. Better still: ask for HTML, since models are good at frontend. Best: bespoke explainer videos, 3b1b style, with a narration API key. His framing is that as models do more legwork, our work rises into oversight, and cheap code makes large discardable artifacts worth building.writingworkflow
102026-10-01Google Mantis: security review as a skills pack@dani_avila7Install with `npx skills add google/mantis`. Commands: `/mantis-threat-model` builds a threat model from the codebase, `/mantis-researcher` scans for vulnerabilities, `/mantis-review` filters false positives, `/mantis-reproduce` writes a PoC and runs it in a sandbox, `/mantis-patch` applies a fix and confirms it blocks the PoC, `/mantis-report` writes the report. He notes the reproduce and patch steps run generated code, so sandbox it.securityskillscode-review
112026-09-30gh-aw: agents inside GitHub Actions@LLMpsychoGitHub's agentic workflows extension runs Claude Code, Codex and Copilot agents inside GitHub Actions. 5,204 stars at the time of the post. One line plus a repo link, no commentary on how it behaves in practice.ci-cdagents
122026-09-30Design constraints for AGENTS.md@tvykrutaA C/C++ veteran argues Claude writes the same spaghetti lazy programmers always wrote, just 200x faster: reaching through objects, over-DRY, speculative hooks, deep inheritance, god functions. His fix is a block for AGENTS.md listing ten design principles in priority order: separation of concerns first, then encapsulation, cohesion and coupling, DRY without over-DRY, KISS, single responsibility, depend on abstractions, YAGNI, composition over inheritance, open/closed. Hard rule: refactor to the principle first, then change behavior.configagentscode-review
132026-09-27jevgrep: context collection as a cheaper CLI@dzhngA research agent CLI powered by jev, claiming 40% lower coding agent cost verified on SWE-bench. It ships a skill so the coding agent knows to call `jg` for context collection rather than reading files itself. Repo at github.com/dzhng/jevgrep. Benchmark claim is the author's own.jevcontextcost
142026-09-25smolvm v1.19: credential vending by interception@binsquaresA release note. The highlight is an interception mechanism for secure credential vending: grep and sed over outgoing requests, swapping placeholders for real credentials or other sensitive values before they leave. Short post, repo link only.securityinfra
152026-09-252,500 PRs in a month, the talk@potetoLauren recorded the talk she had prepared for Cursor Compile in London and posted it free on X instead. It covers how she shipped 2,500 PRs to production in one month. Video post; no written summary in the thread.workflow
162026-09-25Route decisions to jev, writing to Opus@starmexxxThe pitch: the model that decides should not be the model that writes. A router sends every yes/no/route/score call to jev, which returns a typed verdict with a confidence score in ~70ms and bills no output tokens; Opus 5.5 only wakes when something needs prose or confidence is too low. Components are a router, a policy file, a live cost log and an append-only ledger. Claims 1,204,882 decisions in a day with 6 escalations. Funnels to a paid article; treat the numbers as marketing.jevcostagentspromo
172026-09-24An AI coding dictionary, 71 terms@undefinedKiMatt Pocock's dictionary of AI coding vocabulary, as a repo and an interactive site. Each term gets a definition, links to related terms, and a "heard in the wild" line showing real usage; the whole thing also renders as a clickable knowledge graph. Seven sections: the model, sessions and context windows, tools and environment, failure modes, handoffs, memory and steering, patterns of work. Daily Claude Code users are pointed at failure modes and handoffs.referenceskills
182026-09-24jev as the decision layer under Opus@Av1dliveClaims jev plus Opus 5.5 cut his cost and time by roughly 80%. The division of labor: jev picks from options the harness prepares and validates: which project notes to load, when to route a task to a faster worker, which recovery path to take when a tool fails, which focused checks to run before the full suite, while Opus does the hard reasoning. Links to his builder's guide; the detail is behind it.jevcostagentspromo
192026-09-24The NVIDIA/MIT harness result, with the caveats@undefinedKiThe clearest write-up of the same work as row 1. An agent rewrote another agent's harness, trying ~150 ideas across 500 environments; 4 survived, each required to hold quality within tolerance and cut at least one cost metric. The four: merge file edit and test run into one request; compact context only when it pays for itself; show tool outputs over 10KB in full twice then swap for a 1KB excerpt plus a handle; let a cheap model pull key lines from logs with a code check and full-log fallback. Caveats he names: Terminal-Bench 4 solved 15 tasks against 18, and it was tuned on GPT-5.6 Sol, so the tricks fire less on Opus 5. Code at NVlabs/SoL-Pi.harnesscostevaluations
202026-09-24What a software factory stack should look like@zachlloydtweetsWarp's CEO argues for open, composable, defined-in-code infrastructure, layer by layer. Factories-as-code puts agent config, runners, repo access, integrations, webhooks and triggers in a version-controlled `factory.yaml`, which buys benchmarking, A/B testing and rollbacks. Above it: a context layer you host yourself with zero data retention; a compute layer of pausable, resumable remote dev environments; an inference layer that takes any model or harness; improvement infrastructure tracking cost-per-PR and LLM-as-judge quality; and an orchestration control plane.software-factoryinfraconfig
212026-09-21HelixDB: graph and vectors in one engine@agenticgirlAn open-source Rust database for AI memory, knowledge graphs and RAG. Instead of embeddings in one store and relationships in another, it keeps graph and vector data in the same engine, and also handles KV, document and relational data. A single query can combine traversal, filtering and vector search; the Rust, TypeScript, Go and Python SDKs all send the same structured query AST. Runs in memory, on disk, or against S3-compatible storage. 6k stars, Apache 2.0.knowledge-graphdatamemory
222026-09-20jev founder talk: heavy promotion@0xCodezPromotes a 36-minute talk by jev's founder, Diogo Almeida, claiming 200x faster, 400x cheaper, zero hallucination and no human in the loop, and that Claude Code and Codex belong to a soon-ending "assistance era". The post calls the talk worth more than a Stanford ML degree and routes to an article on becoming a "Jev Engineer". No technical detail in the post; the claims are unverified.jevpromo
232026-09-20Tensorlake: snapshot the agent's whole machine@agenticgirlTensorlake gives agents stateful Linux sandboxes on Firecracker microVMs, then snapshots them once the expensive setup is done: repo cloned, dependencies installed, Postgres running, failing test reproduced. The snapshot keeps filesystem, memory and running processes, and restores into one new sandbox or a hundred, so parallel attempts start from the same checkpoint. Underneath, a block overlay tracks dirty blocks at 4KiB granularity: their benchmark copies a 100GiB disk in 101 seconds, but takes an incremental snapshot of a 100MB change in 167ms. Their numbers, but the design is clear: snapshot cost follows what changed.sandboxinfra
242026-09-20Google's ax: Kubernetes rethought for agents@rakyllJaana Dogan's first reveal of what her team has been building: Kubernetes reinvented for agentic workloads with statefulness and fast resumption, plus an open agentic orchestrator and runtime for Google. Repo at github.com/google/ax. Announcement only: no architecture detail in the post.infraagents
252026-09-19Prompt for finding your own jev-shaped decisions@Av1dliveHe shares the prompt he handed Codex and Claude to scan his workflow and list every decision point where a cheap decision model could replace a full LLM call. The prompt is the whole post. Quote-tweets his own article on building an agentic memory factory with Kimi K3.jevcostpromo
262026-09-18Fine-tuning as a core skill@sairahul1A short argument, not a how-to: fine-tuning is becoming one of the most valuable AI engineering skills, not because most people need a custom model but because understanding how models learn from data lets you build things others can't. The substance is in a linked full guide; the post is a teaser.fine-tuningskills
272026-09-16MCP over CLIs for most integrations@trq212Thariq says he did not expect this, but MCP now beats CLIs for most integrations: tool calling has improved, tools can be deferred, and MCP is stateless. His practical note is that if you need to compose or filter data, add parameters like `query` to the MCP tools rather than reaching for a CLI.mcp
282026-09-16Graphify: map the codebase once@techNmakTurns a project into a queryable knowledge graph: functions, classes, files, SQL schemas, infrastructure, docs, PDFs, images and video as connected nodes. Not RAG: no embeddings, no vector DB, no LLM; code is parsed locally with tree-sitter and calls, imports and inheritance become edges. Every relationship is marked EXTRACTED, INFERRED or AMBIGUOUS so the agent knows what was guessed. It installs hooks for Claude Code, Codex, Cursor, Gemini CLI, Copilot and others that nudge the agent to query the graph before grepping; the graph commits to git and is exposed over MCP.knowledge-graphcontextmcp
292026-09-15Crawl, walk, run to cloud agents@zachlloydtweetsA staged path from local interactive agents to automated cloud development, for teams who find the full software factory too much to adopt at once. The tweet is a pointer; the steps are in the attached article.software-factoryworkflow
302026-09-15A proposal for funding open source@tannerlinsleyTanner Linsley recommends Laurie Voss's article on paying for open source, saying that after years of trying to make his own projects sustainable he has not seen a proposal that comes this close to working at scale. He calls it atomic and says it could actually work. The argument is in the linked post.reference
312026-09-15(deleted)@tspyThe post no longer exists: X returns a 404. Email subject was "Karpathy", so it was probably another entry in that week's Karpathy-method cluster, but there is nothing left to read.deleted
322026-09-15OpenAI's software factory@GergelyOroszGergely Orosz describes how OpenAI's agentic software factory works today, based on conversations with engineers there, with the full account in The Pragmatic Engineer. He flags "Perf Factory" as the piece he finds most interesting. The tweet carries the diagram and the link.software-factoryagents
332026-09-15tgrep: index once, search in milliseconds@FluyeporlawebIn Spanish. Microsoft released tgrep so agents stop searching a repo like it's 2012 grep: index once, file watcher, millisecond search, written in Rust, open source, already used inside Copilot CLI. His line: if your agent takes longer to find the file than to edit it, the bottleneck is not the LLM. He pairs it with RTK: one trims output, the other trims search.contextinfra
342026-09-15What 80% AI-written code does to CI@addyosmaniClaude now writes 80% of Anthropic's code and engineers ship 8x more per quarter. The side effect is the interesting part: tests grew 10x and CI jobs 25x in six months. Links Anthropic's write-up of what they did to keep that scaling.ci-cdagentscost
352026-09-13OpenViking: agent memory as a filesystem@agenticgirlAn open-source context database holding knowledge, user memories and reusable experience in one place, organized as a virtual filesystem the agent browses with `ls`, `tree` and `read`. Rather than loading everything, summaries let the agent decide which files to open in full. Their benchmarks report 80–83% memory accuracy with input tokens down between 34% and 91%.memorycontextagents
362026-09-12DeepSeek v4.1 Flash vs. the original Transformer, in 3D@petergostevPeter Gostev had Astra read the DeepSeek v4.1 Flash paper and build an interactive 3D comparison against the original Transformer architecture, element by element, with zoom. His takeaway line is that things have changed quite a bit. The demo is hosted and linked; the post is the demo.models
372026-09-12alibaba/open-code-review@ChrisShortA code review tool Alibaba runs at their own scale, now open. Hybrid architecture: deterministic pipelines alongside an LLM agent, producing precise line-level comments. Ships a fine-tuned ruleset covering null pointer exceptions, thread safety, XSS and SQL injection. Works against OpenAI and Anthropic APIs.code-reviewsecurity
382026-09-12Kimi K3 on one CPU in 8GB of RAM@techNmakA 176KB C program runs a 2.78T-parameter model on one CPU with 8.24GB of RAM, and the checkpoint is still 1.56TB, sitting on NVMe. Most of it is MoE: 92 routed layers of 896 experts, 16 used per token, 82,432 experts totalling 1.447TB, or 93% of the file. The ~108.81GB of always-used weights are repacked into a contiguous trunk so each layer reads in one run, pinned to RAM where it fits and streamed where it doesn't. 69 of 93 layers use KDA, which carries fixed-size recurrent state instead of a growing KV cache. It is slow: ~26.5 seconds per token at 8GB, but RAM stops deciding whether the model runs at all.modelsinfra
392026-09-10Compiled wiki vs. filing cabinet@v1lrokA thread riffing on Karpathy's memory diagram. The argument: retrieval answers questions, compilation builds understanding. A notes folder only accumulates; a compiled wiki keeps raw source untouched, has the model read it once, connect it to what is already compiled, and file the synthesis, so later questions are lookups rather than searches. CLAUDE.md sits at the center as the file read before every session. Written as engagement bait and light on specifics: the architecture is in the linked article.memorycontext
402026-09-09Edge0: a 35B model on an iPhone@SamuelZengMLOpen-sourcing a framework for running large models fully on-device. The demo claim: a 35B language model on an iPhone at 1–2.5GB peak memory, with no cloud, remote server or desktop GPU. Announcement post; the mechanism is in the repo and the attached video.modelsinfra
412026-09-09What is actually inside a harness@mardehaymNotes in three tiers. The harness is everything wrapped around the model: instructions, scoped context, tools, a verifier, guardrails, and the model is the smallest, most swappable part. The hard-mode version has six layers: trigger (work starts from an event, not a button), orchestration (loop, memory, retries, limits as config), tools (numbers that must not vary run as fixed logic so the model doesn't improvise your KPIs), trusted context (called ~80% of agent success), control (golden sets, approvals, a named human decides), runtime (traces, cost dashboards, audit). His conclusion: the harness is the moat, because models and platform services change underneath it.harnessagentscontext
422026-09-09MiniCPM5-2B and the Densing Law@itsPaulAiA 2B-parameter open model that runs agents offline on a laptop or phone in 2GB of RAM, and handles coding and tool use well enough to browse Hugging Face and produce a CSV of top models unaided. Scores 23 on Artificial Analysis Intelligence Index v4.1.1, first among open models under 4B. The framing is density over size: Tsinghua and ModelBest's "Densing Law" claims the capability density of open base models roughly doubled every 3.5 months over the period studied. The release also opens training methods, agent data and the RL stack.modelsfine-tuning
432026-09-06ripwire: repo context without embeddings@agenticgirlFrom Red Hat Emerging Technologies. A zero-dependency C++23 binary parses 21 languages with tree-sitter and builds a deterministic structural map of a codebase: no embeddings, vector database, LLM indexer or daemon. It ranks symbols for the task at hand and attaches call relationships, complexity, git churn, change amplification and test coverage. Ask it what matters for "incremental cache invalidation" and you get the symbols, their callers, likely blast radius and the tests to run, inside a token budget.contextinfra
442026-09-05Portal cut Spotify's Claude Code tokens by 90%@rseroterRichard Seroter points at Spotify Engineering's post on Portal, which they report cut Claude Code token usage by 90%. His one-line read on why: they use the model for reasoning rather than for I/O. The detail is in the Spotify post.cost
452026-09-04One prompt over your whole work context@rileybrownRiley Brown gave Astra context over everything: email, Slack, business texts, Notion, every meeting note, about 20 projects, then asked one rambling planning question covering month, quarter and year. The prompt asks for flaws and strengths, which activities waste the most time, which to do more of, who is most dependable, where he is least dependable, the one skill to build this year, and how to reorganize the company, with charts only where they earn space. He shares it verbatim and calls the answer the most useful thing an AI has given him.contextworkflow
462026-09-03Semantic layer, context layer, ontology@motherduckThree things people collapse into one. The semantic layer is rigid logic that compiles to SQL: metrics and dimensions. The context layer is the unstructured human knowledge, docs and wikis, that agents need to decide anything. The ontology is the digital twin mapping real business entities. Points to Simon Späti's primer on arranging the three for agentic work.datacontextknowledge-graph
472026-09-03Warehouses are hitting the AI architecture problems first@sethrosenSnowflake, Databricks, ClickHouse, BigQuery and MotherDuck are adding models and agents to mature systems that were never designed around LLMs. Seth Rosen's point: that forces them to solve the same architecture problems the rest of software faces, and they may get there first. Links Josh Rosen's article working through the lessons.data
482026-08-31The post-AI data stack, and UI for verification@sh_reyaOn Ian Macomber's piece: analysis codegen is free now, so the hard part is getting agents to answers that are both correct and consistent over messy, fragmented data. The data team's job becomes encoding expert judgment into infrastructure so agents produce correct analysis without a data scientist in the room. Shreya extends it to interfaces: dashboards were UI-for-exploration, built for human pattern-spotting, but if the agent is the one reading the data, the UI's job may become helping humans verify the agent's reasoning.dataevaluationsdesign
492026-08-28Skill evolution through a persistent wiki@dair_aiA Google paper separating three things skill-evolution systems usually merge: raw execution traces, a persistent wiki of accumulated knowledge, and the executable skills. Experience consolidates into the wiki, and every later skill update builds on the wiki rather than on a scattered optimization history. Ablations show the wiki carries much of the gain. Two results stand out: smaller models with evolved skills beat substantially larger ones without, and skills evolved by one model transfer across families, sometimes beating self-evolved ones. arxiv 2608.27454.skillsmemoryevaluationsresearch
502026-08-28LLM cliché highlighter, 38 patterns@simonwSimon Willison's small browser tool that marks the tells in AI-written prose, now up to 38 patterns. One line and a link to tools.simonwillison.net. Useful as a checklist of what to strip from model output.writing
512026-08-24Models are overqualified hires@JayaGup10The analogy: an overqualified employee reopens settled decisions, adds complexity, gets bored and leaves. A frontier model does the first two and never leaves, so the extra variance and cost run indefinitely. Her example: a password reset does not get better because the agent weighs twelve explanations, opens a security investigation and writes a personalized essay. Links her "Right-Sizing Your Intelligence Spend" piece. A reply raises the real difficulty: routing is hard when you cannot cheaply tell whether the cheap model's answer was good enough.costmodels
522026-08-19AI homework help, exam scores down@paulnovosadDiff-in-diff plots from a paper by Stromberg, Lei and Wu: students who lean on AI finish homework faster and score higher on it, then get crushed on exams, with exam scores falling almost in lockstep with homework effort shirked. Novosad's corollary for teaching: take-home work should be at most a tiny share of a course grade. He argues every syllabus should carry these graphs.research
532026-08-17Fix AI prose with the Google style guide@natebjonesIf you are tired of Claude-lish or Chat-lish, have the model read the Google Developer Documentation Style Guide and build a skill from it. He has tried both and prefers it to an ASD-STE100 Simplified Technical English skill. His aside to the labs: an agentic model that cannot write plain English is much less useful. He later posted his own skill repo, and a reader posted another.writingconfigskills
542026-08-17Anthropic's cost optimization cookbook (via a reply)@Xxi5olcThe link lands on a one-line reply rather than the post worth reading. The parent, from @dani_avila7, points at Anthropic's cost optimization cookbook: a real agent going from $0.29 per task down 90% without losing accuracy, with model downgrade as the last lever rather than the first. The reply asks the fair question: if these approaches work, why are they not built into the harness?cost
552026-08-12Slimming down Claude Code@EXM7777A long checklist for cutting token spend. Keep the global CLAUDE.md under 200 lines and rebuild it from scratch every few months, since long files get ignored. Run `/context` in a fresh session to see what fills the window before you type, and `/usage` to find which MCP server, skill or plugin is burning tokens. Prefer CLIs to MCPs, because a command returns only the lines you ask for. Toggle off unused MCP servers, mark rare skills manual-only, write "never do this" rules as permission settings rather than prose, reserve high effort for planning, `/clear` at checkpoints, and push heavy exploration to subagents so your window only sees the summary.costconfigcontext
562026-08-07kimi-k3-in-c, and the pushbacklinkedin.com (Linas Beliūnas)The same project as row 38, in LinkedIn form: a 176KB pure-C99 engine keeps the dense trunk in memory and streams experts from disk, 16 of 896 active per token, original MXFP4 weights, identical output from 8GB to 224GB. The top comment is worth more than the post. It points out the engine is CPU-only by design, and that even with 128GB+ of RAM the fastest it runs is 5.6 seconds per token, because generating one token means reading the ~108GB trunk at DDR5 speeds. A fine CS result; not a practical way to serve the model.modelsinfra
572026-08-01Semantica: knowledge graphs with provenancegithub.comA graph-native platform that turns enterprise data into queryable knowledge graphs carrying decision provenance. It reasons deterministically: forward chaining, Rete, Datalog, SPARQL, rather than leaning on embeddings and vector similarity, so the context an agent uses stays auditable. Covers ingestion from multiple sources, semantic extraction, conflict detection, deduplication and export to W3C PROV-O and RDF. Built for regulated industries that need explainability, self-hosted, meant to sit alongside an existing LLM stack.knowledge-graphsecuritydata
582026-07-29A prompt banning the AI prose tells@0xPiaA paste-in list of prohibitions, and the most concrete version of this idea in the collection. No antithesis, corrective negation, paragraph pinning, parataxis, summary beats, negative parallelism or anaphora, contrasting pairs, rule of three, em dashes, throat-clearing openers, landing sentences, setup/payoff constructions, or parallel structures inside a paragraph. Also: vary sentence length unpredictably, drop stacked noun phrases, filler intensifiers, corporate verbs like leverage and underscore, nominalizations and hedges. Write for the spoken voice.writingconfig
592026-07-25CLAUDE.md config: promotional, no content@cyrilXBTSays a CLAUDE.md configuration making the rounds changed how he uses Claude, and that the difference was immediate. The post contains no configuration, no example and no link to one. The attached article is actually about running Kimi K3 as a coding agent. Safe to skip.configpromo
602026-07-18microsoft/Ontology-Playground@thisdudelikesAITwo links and nothing else: a free, open-source web app for learning about ontologies, plus a live preview hosted on GitHub Pages. The post is the pointer; the playground is the thing.knowledge-graphdata
612026-07-18An MIT professor's last lecture@heyrohitaiA pointer post. Fifty years of teaching compressed into one recorded hour, made shortly before the professor died. No argument or summary in the tweet, just the framing and the video.reference
622026-07-11OpenRouter's 100-trillion-token study@AnjneyMidhaOffered as the rigorous empirical answer to a question people keep asking: how much open vs. closed frontier model usage is there actually. The OpenRouter study covering 100 trillion tokens is up on arXiv, led by Maika Thoughts, Alex Atallah, cclark and team.researchmodelscost
632026-07-04Loop engineering, behind clickbait@0xNoryxxThe framing is invented: an engineer supposedly fired over an 11-page PDF, but the loop underneath is sound. Schedule → Discover → Build → Verify → Repeat. Discovery has the agent find its own work from failing CI, open issues and recent commits rather than a handed list. Verification uses a second agent told to assume the code is broken, on the premise that an agent grading its own work always praises it. Results persist to disk rather than a context window that gets flushed. Each task gets an isolated git worktree so parallel agents don't collide.harnessworkflowagentspromo
642026-07-03Four jobs for your tokens@0xCodilaThe idea: rather than spending more tokens, split the same budget across four roles: execute does the work, advise checks the direction, grade passes or fails against a rubric, dream inspects, learns, writes to memory and sharpens the next round. One AI doing all four becomes four doing one each, at the same cost. The quoted 15%-to-90% accuracy jump is attributed to Anthropic's product team but not sourced; treat it as marketing.costagentspromo
652026-07-02CRUX 2: can agents do open-ended AI researchlinkedin.com (Sayash Kapoor)An update on CRUX, their open-world long-horizon evals. Most research-automation work keeps a human in the loop, picks narrow verifiable problems, or uses a scaffold tuned to one question type, so strong results may say more about the scaffold. CRUX 2 tries the broad version. To avoid contamination they partnered with researchers at UK AISI, Toronto and Princeton who pose open questions from papers not yet public; the agent must produce a NeurIPS-quality paper and a reproducible codebase that those authors review. Week-long horizons, VMs and GPUs, $3,000 in API credits per paper, and the agent manages its own budget.evaluationsagentsresearch
662026-06-27Using local coding agents@rasbtSebastian Raschka's link to his own article on running coding agents locally. A follow-up tweet in his thread, so the post itself is just the pointer; the write-up is in his magazine.agentsmodels
672026-06-25Codex data on the shift to agentic AI@daveholtzDavid Holtz announces the first public output of his part-time stint as a visiting economics researcher at OpenAI: a study using Codex data to document how fast work is moving to agentic AI. The tweet opens a thread; the findings are in the replies and the paper.researchagents
682026-06-17Shape suffixes for tensor code@vboykisVicki Boykis points to the only thing Noam Shazeer has blogged publicly, naming tensor variables with their shape as a suffix, and says she uses it all the time. A small, durable coding habit rather than a framework.reference
692026-06-16Cutting enterprise Anthropic spend@JayaGup10Jaya Gupta crowdsources best practices before presenting to the C-suite of a top-30 global company on reducing their Anthropic spend. The replies are the value here. Links her "Token Budget Wars" piece, whose argument is that enterprise AI has moved from adoption to allocation, and every function is now asked to quantify its AI ROI.cost
702026-06-13code2lora@liliana_hotskoThe resource post at the end of a thread rather than the explanation: paper at arxiv 2606.06492 and data plus models at huggingface.co/code2lora. To understand what code2lora does you need the parent tweets or the paper itself.fine-tuningmodels
712026-05-23A Tufte skill for Claude's charts@draparenteFrustrated with Claude's charts, she fed Tufte's book to Claude and had it generate a Tufte skill, which she says immediately produced simpler, better visualizations. The gist is linked. Quote-tweets Anjney Midha telling anyone running an AI lab to have their team read Tufte before publishing charts.designskills
722026-05-22CLAUDE.md template from Karpathy's rules@PrajwalTomar_A template circulated on Reddit, with an anecdotal claim of accuracy going from 65% to 94% on one codebase. The rules: no filler openers, match response length to task complexity, show two or three approaches before anything significant, flag uncertainty rather than filling gaps with plausible text, touch only files related to the current task, describe before rewriting, ask before deleting or overwriting, re-confirm deploys and database drops every time, keep a MEMORY.md of decisions and what was rejected, and an ERRORS.md so failed approaches are not retried.config
732026-05-09Status in AI, a personal essay@nrmehtaNick Mehta riffs on Jaya Gupta's argument that a company's talent identity becomes its long-term moat in services industries, then writes mostly about status. His upbringing with Harvard and Einstein posters over his bed, the "college is dead" years of 2016–2024, status going quiet during COVID, and its hard return now that everyone in high-growth tech is in San Francisco comparing notes at the same parties, with status rising and falling on each model release.reference
742026-04-20arXiv link, no commentary@HowToPrompt__A bare link to arxiv.org/pdf/2603.19312 with no description at all. Filed under "Jepa" in the original list, so presumably a JEPA paper, but the post says nothing about it. You will have to open the PDF.researchmodels
752026-04-17Architecture diagram generator as a Claude skillgithub.comDescribe a system in plain English and the skill produces a dark-themed interactive HTML diagram of components, connections and data flows. Exports to PNG, PDF or clipboard. Uses semantic color coding by component type and emits responsive SVG in a single self-contained file that opens in any modern browser. No design experience assumed.skillsdesign
762026-04-13Marcus Hutchins on Mythos doing vuln research@ananayaroraAnanay shares what he calls the best take going on Mythos finding vulnerabilities, from the researcher who stopped WannaCry. The argument is in the video clip; the tweet is a single line of endorsement. Pairs with row 81, which tests the claims empirically.security
772026-04-13Who owns the harness@vikvang1Bouncing between Claude Code and Codex is normal, since the models are good at different things, but it means paying several subscriptions, rebuilding context each time, and being exposed when a provider changes the deal. His point: switching models is easy when things are stateless, and agents are not. The harness owns context, memory, preferences, compaction, tools and workflows. So the question is not which model you use but who owns the harness around it.harnessagents
782026-04-12One CLAUDE.md file, 15K stars@akshay_pachaarDerived from Karpathy's coding rules. The premise: LLM coding mistakes are predictable (over-engineering, ignoring existing patterns, adding dependencies nobody asked for), and predictable mistakes can be prevented with instructions. So one markdown file in the repo root gives the agent behavioral guidelines for the whole project. No framework, no tooling. His closing observation is that the best tools in the Claude Code ecosystem are often not software.config
792026-04-12Harness, memory, context fragments@Vtrivedy10Working notes, not a finished argument. The harness's main job is routing data into the context window, and every loaded object is a "context fragment": an explicit choice by the designer about what the model needs right now. Agent memory differs from human memory in one way that matters: it accumulates across agents that can be forked and duplicated. As agents run for years, the volume of data they produce grows hyper-exponentially, which makes search, distillation and owning that data yourself the real problems.harnessmemorycontext
802026-04-09Silicon Valley running on Chinese open models@petergyangThe receipts he lists: Cursor confirmed Composer 2 is built on Moonshot's Kimi K2.5; Cognition's SWE-1.6 is likely post-trained on Zhipu's GLM; Shopify saved $5M a year moving to Alibaba's Qwen; Airbnb's Brian Chesky called Qwen good, fast and cheap. He adds that Zhipu's GLM-5.1 now performs close to Opus on coding benchmarks. More in his post on the Anthropic/OpenClaw situation and what he saw in China.models
812026-04-09Small open models found the same bugs@ClementDelangueQuotes Aisle's test of Anthropic's Mythos vulnerability-research announcement. They isolated the code behind the showcased vulnerabilities and ran it through small, cheap open-weight models. Eight of eight detected the flagship FreeBSD exploit, including one with 3.6B active parameters at $0.11 per million tokens, and a 5.1B-active model recovered the core chain of the 27-year-old OpenBSD bug. The useful counterweight to row 76.securitymodelsevaluations
822026-04-09Alternatives to GitHub Actions@vboykisA question, not an answer: are there genuinely viable CI/CD alternatives outside GitHub and Actions right now. Prompted by Astral's post on open source security. The replies are where any answer would be.ci-cd
832026-04-06A multi-agent job search system, open-sourced@PrajwalTomar_Built on Claude Code over a weekend, used across 740+ roles, and the author landed a Head of Applied AI job with it. Paste a job URL and it scores compatibility on 10 criteria so you only apply where you fit, rewrites the CV per role and generates ATS-optimized PDFs through Playwright, scans 45+ company career pages, prepares STAR answers specific to the description, and tracks the pipeline in a terminal dashboard. 14 skill modes, MIT licensed.agentsworkflow
842026-04-05DESIGN.md files from 31 real sites@ihteshamaliawesome-design-md by VoltAgent collects DESIGN.md files extracted from Stripe, Vercel, Notion, Supabase, Linear, NVIDIA, Apple and others. Drop one into the project root, tell the agent to build a page that looks like this, and it has colors, typography, spacing, buttons, cards, shadows and responsive rules to work from. No Figma exports, JSON schemas or special tooling: DESIGN.md is a plain-text design system, a concept from Google Stitch, that models read natively. Aimed at escaping Inter, purple gradients and card grids.configdesign
852026-03-30Semantics belongs next to the schema@kirsten_lum_A short reply, no link. Her claim from decades of practice: the only place semantics reliably got used was the schema itself, because nothing else would get adopted. Once the semantic layer is decoupled from the schema but still integrated with it, humans and agents can work at the semantic level and let the data platform compile down.data
862026-03-22HTML slides demo and themes@Kangwook_LeeA follow-up tweet pointing to the demo and theme gallery for his HTML slides work. Original list filed it under "Html slides Claude code skill", but this particular post carries only the link: the skill and its explanation are elsewhere in his thread and on his site.skillsdesign
872026-03-13How LLMs actually work, 42 slides@paraschopraParas Chopra gave a two-hour talk on the mechanics of LLMs and posted all 42 slides. The tweet is two lines; the deck is the content. Useful as teaching material.referencemodels
882026-03-09Technical debt plus cognitive debt@infinitehumanaiConnects AI coding to Peter Naur's 1985 essay "Programming as Theory Building": a program is a shared mental theory living in the people who work on it, and the code is only a lossy written representation you cannot rebuild the theory from. Before AI, building something gave you the theory for free, as a byproduct of the work. AI breaks that coupling: you can produce code without building the theory, so you now accrue cognitive debt alongside technical debt. The cognitive kind is worse because it hides: you can't tell you cannot reason about your own program.referenceagents
892026-02-08Memory and scheduling for Claude Code@tom_doerrOne line and a repo link (github.com/ascorbic/macro…). A tool that adds memory and scheduling to Claude Code. No description of the mechanism or how well it works; you will need to open the repo.memoryconfig
902026-02-03Column storage for the AI era@andrewlamb1111Andrew Lamb's talk on the AI use cases pushing changes in Apache Parquet and driving newer formats. He calls it somewhat academic. Recording on YouTube and slides on Google Docs, both linked. Relevant if you are thinking about what replaces Parquet for AI workloads.data
912026-01-29Agent harness architectures and memory@aparnadhinakArize's post, drawn from building their own agent and working with many customers. The core argument: files plus Unix tools make a fixed context feel effectively infinite, and bash composes into complex work without needing tool-definition JSON, since the model already knows the commands. The historical parallel is 1980s CPU memory hierarchies: caches made memory feel fast, virtual memory made it feel infinite, with the filesystem now playing both roles for agents.harnessmemorycontext
922026-01-26Spec-driven development as the declarative limit@karpathyA reply, short but worth it. Karpathy calls spec-driven development the limit of the imperative-to-declarative transition: being declarative entirely. He points to Drew Breunig's "A Software Library with No Code" as an extreme and early example that he found inspiring.agentsworkflow
932026-01-17The shorthand guide to Claude Code@giyu_codesAn endorsement rather than the guide. The underlying piece, by @affaan, is a full setup after ten months of daily use: skills, hooks, subagents, MCPs, plugins and what actually works. Read the quoted article, not the tweet.configreference
942026-01-13skills.md as a contract for Python repos@tdhopperAdd a skills.md to the repo telling the model how your codebase actually works: style, patterns and footguns. Tim Hopper frames it as a small contract rather than documentation. Links the Dagster team's playbook and examples at pydevtools.com.skillsconfig