ads

Monday, August 31, 2026

Show HN: Keel - A conductor, not an agent loop https://ift.tt/at5u3wA

Show HN: Keel - A conductor, not an agent loop https://daneb.github.io/keel/ September 1, 2026 at 03:22AM

Show HN: Guess the AI model from its web design https://ift.tt/DZc4Mvr

Show HN: Guess the AI model from its web design https://ift.tt/wTsmoOY September 1, 2026 at 02:40AM

Show HN: Corporate Mind Games – logic puzzles with a sarcastic corporate theme https://ift.tt/3uLJktU

Show HN: Corporate Mind Games – logic puzzles with a sarcastic corporate theme I have always enjoyed building and playing puzzles. I am a regular user of NYT Games and LinkedIn Games. I thought it would be fun to create some corporate-themed puzzles as a silly alternative to LinkedIn Games. So, I built Corporate Mind Games. The games are intended to be a little bit sassy and sarcastic, but I hope they are enjoyable and challenging. I was planning to post this on HN last week. However, many years ago, I made a game called Don't Wordle, which unexpectedly appeared on the front page last week. I ended up getting distracted by that, and I also figured it was more appropriate to post this game on HN on a Monday morning given the theme. So here we are. https://ift.tt/iCtb39h August 31, 2026 at 11:08PM

Sunday, August 30, 2026

Show HN: BentoPDF, Hyper Compress and Kura https://ift.tt/daKmRbU

Show HN: BentoPDF, Hyper Compress and Kura Hello. I developed an open source tool called BentoPDF. Its an open source PDF toolkit that runs in your browser. With the latest update, you can actually edit existing pdf text, images and objects right in your browser. Live Website: https://ift.tt/uTrHJIU Repository: https://ift.tt/v2hq16U Along with the latest update I would like to share with you Hyper Compress. Its a high fidelity, content preserving compression engine that preserves PDF conformance and surpasses all the other open source PDF compression tools. It runs everywhere: CLI, Node SDK, C API, self hosted service, and also in browser via WebAssembly. Live Website: https://ift.tt/xCzmZcw Repository and benchmarks: https://ift.tt/sCQmIKc Kura is a PDF standards, conversion and preflight engine. It supports: - All 11 PDF/A conformance levels: PDF/A-1a, PDF/A-1b, PDF/A-2a, PDF/A-2b, PDF/A-2u, PDF/A-3a, PDF/A-3b, PDF/A-3u, PDF/A-4, PDF/A-4e and PDF/A-4f - Accessibility: PDF/UA-1 and PDF/UA-2 - Print production: PDF/X-1a, PDF/X-3, PDF/X-4, PDF/X-4p, PDF/X-5g, PDF/X-5n and PDF/X-5pg - Engineering and variable data printing: PDF/E-1 and PDF/VT - E-invoices: Factur-X, ZUGFeRD, XRechnung and Order-X - 396 bundled print-preflight profiles It has been tested against several standards suites, including the veraPDF corpus, Isartor, BFO, Ghent Output Suite 5.0, the PDF/UA Reference Suite and Cal Poly's PDF/VT suite. Across 30,677 PDF conversions it had zero crashes and zero timeouts, with a 0.05 second median conversion time. Like Hyper, it ships as a CLI, C library, npm package, Docker image and WebAssembly build. Live Website: https://ift.tt/o4vhszj Repository and benchmarks: https://ift.tt/czxwSDV Both Kura and Hyper Compress are running the WASM build so none of your PDF is uploaded and everything runs in your browserr If possible I would like you guys to try them out and give me feedback on how it worked for you. Thank you. https://ift.tt/uTrHJIU August 30, 2026 at 10:57PM

Saturday, August 29, 2026

Friday, August 28, 2026

Show HN: Free lifetime pro access to a crunchbase alternative https://ift.tt/ONLvzwu

Show HN: Free lifetime pro access to a crunchbase alternative A month ago we launched on HackerNews and reached the front page! As a thank you we are now offering free lifetime pro subscriptions to the HN community! https://ift.tt/tEVuKzP August 28, 2026 at 11:07PM

Show HN: Get agents to do what I want with code documentation https://ift.tt/Ab9aYXl

Show HN: Get agents to do what I want with code documentation https://ift.tt/eCzsbI7 August 28, 2026 at 10:20PM

Thursday, August 27, 2026

Show HN: I built a site to share and find Claude Code passes https://ift.tt/FUtMhbm

Show HN: I built a site to share and find Claude Code passes This is an exchange site for people with paid Claude accounts who want to get free credits for sharing their passes. When someone redeems your Claude Code coupon, Anthropic credits you $10 in usage. I’ve already added mine. If you have a paid Claude account, run "/passes", get your coupons, and add them to the site too. https://ift.tt/pIDyXPT August 27, 2026 at 10:43PM

Wednesday, August 26, 2026

Show HN: A robot football league where frontier AI models manage the clubs https://ift.tt/w9W0UTb

Show HN: A robot football league where frontier AI models manage the clubs https://rfl.football August 26, 2026 at 08:53PM

Tuesday, August 25, 2026

Show HN: JeopardyGithub: GitHub Repo to Jeopardy Game https://ift.tt/I75emcR

Show HN: JeopardyGithub: GitHub Repo to Jeopardy Game Introducing https://ift.tt/PEsFoqZ convert any github repo into a multiplayer classic jeopardy board! https://ift.tt/PtTyFfE August 25, 2026 at 11:33PM

Show HN: A portable evidence record for what an AI agent ran https://ift.tt/okU62e4

Show HN: A portable evidence record for what an AI agent ran It is part of Linux Foundation now. This is the foundation for verifiability we have been working on for some time. Would love to have your thoughts and I'll watch out and reply to ur comments. https://ift.tt/JctxaPE August 25, 2026 at 11:14PM

Monday, August 24, 2026

Show HN: PicoMQ – Durable Streams over HTTP, on object storage https://ift.tt/6UGfnDm

Show HN: PicoMQ – Durable Streams over HTTP, on object storage PicoMQ is a Rust server for Durable Streams, built on Object Store. Cheap, URL-addressable, granular streams (create/append/read/long-poll/SSE), with Pico Protocol or Durable Streams Protocol as the facade. S3Stream is the stream storage primitive, used in AutoMQ, shipped as a Rust library. Coordination is a command log in Postgres. https://picomq.com/ August 24, 2026 at 11:08PM

Sunday, August 23, 2026

Show HN: Mnemosyne Local hierarchical memory engine for AI agents (MCP Native) https://ift.tt/nuwxWv5

Show HN: Mnemosyne Local hierarchical memory engine for AI agents (MCP Native) https://ift.tt/4eqRSyV August 23, 2026 at 11:43PM

Show HN: Dev // Cyber Toolkit – a free floating dev/cybersecurity widget https://ift.tt/SEiyAks

Show HN: Dev // Cyber Toolkit – a free floating dev/cybersecurity widget https://dev-cyber-toolkit.netlify.app August 23, 2026 at 11:07PM

Show HN: dav-next – Nginx module – small and fast Nextcloud file server https://ift.tt/fQ9tel6

Show HN: dav-next – Nginx module – small and fast Nextcloud file server This project is part of a much broader personal project that will eventually be published. Alpine packages are available on the project repository and, hopefully soon, on Alpine's own repository. https://ift.tt/TBCKJjs August 23, 2026 at 10:32PM

Saturday, August 22, 2026

Show HN: Zcomplete – Shell Typo Correction https://ift.tt/uEQaN73

Show HN: Zcomplete – Shell Typo Correction https://ift.tt/dcGH2lY August 22, 2026 at 02:18AM

Show HN: Get Your IP Address https://ift.tt/7NyRbsK

Show HN: Get Your IP Address I made simple webservice to get your IP4 or IP6 service in different formats. Not interesting, just a quick project that I needed myself that might be usefull to somebody else. https://ip.ysebie.be/ August 22, 2026 at 09:30PM

Friday, August 21, 2026

Show HN: Whodunit? Solve a daily AI-written Mystery https://ift.tt/AyBnmDU

Show HN: Whodunit? Solve a daily AI-written Mystery Whodunit started as a pen-and-paper game for family game night. It evolved into a web app with a new mystery every day. The mysteries are all written by LLMs. Some are better than others, but most work nicely. The web app is written in go with templ HTML templates, HTMX for on-page interactivity, SSE for content updates, and a turn-based game engine written as a Temporal workflow. Check out the mystery gen write up for more info: https://ift.tt/QOYwezV https://whodunit.rip August 21, 2026 at 11:04PM

Thursday, August 20, 2026

Show HN: Omacosy – Omarchy-style tiling desktop for macOS, no SIP https://ift.tt/ib3VwNq

Show HN: Omacosy – Omarchy-style tiling desktop for macOS, no SIP I have been using omarchy on my tower since nearly a year now, shortly after it was released first. I really love the experience I am having with it but I still use my macbook for daly work, so I wanted to recreate a similar experience on it. Thats why I created omacosy, a setup for tiling windows, custom menu bar, some themes from omarchy, focus follows mouse, focus rings around windwos, some mac flavors with trackpad events and a custom mission control overview for your workspaces. I used AeroSpace over yabai for the tiling window manager because I didnt wanted to compromise on SIP which is a mac security feature. It is supposed to be keyboard first like omarchy to move windows organize workspaces etc The setup runs around 157mb of ram and consists of AeroSpace, Karabiner (for the super key), and five small self build swift binaries. I am running it daily on my M1 max macbook, currently on macOS26. I havent tested it much on other macbooks or macOS versions. The install script creates a manifest file to backup what was installed before and what it installed itself, the uninstall script takes that into account to clean up the macbook to exactly the state it was in before. It needs quite some permissions for it sfunctionality which I layed our in the project readme. I wanted to be really transparent about which permissions it uses and for what reason. I would love to get some feedback or see people trying it out and hearing your opinion. Mostly about what still doesnt feel smooth in the experience or if you find any performance issues. https://ift.tt/PVdecaA August 20, 2026 at 09:12PM

Wednesday, August 19, 2026

Show HN: Progressive web-app to explore Paris while solving puzzles without app https://ift.tt/ZyG2rHu

Show HN: Progressive web-app to explore Paris while solving puzzles without app https://flanori.com/ August 19, 2026 at 09:18PM

Tuesday, August 18, 2026

Show HN: Voidleap Code – agentic IDE, own harness, swap models mid-conversation https://ift.tt/K2O94Uk

Show HN: Voidleap Code – agentic IDE, own harness, swap models mid-conversation We wanted to build a development environment that can make you a better agentic engineer, not a tool that makes money when you waste tokens. A tool where you can swap models between turns, edit the context, and see every agent action. To do it, we had to build our own harness. Free to use. We don't sell inference, so bring your own key. The sore spot is Claude subscriptions. We can't support them because of Anthropic's terms. The app runs on your machine, and your data never touches our servers. We've been building Voidleap Code using Voidleap Code since January. We didn't rush to launch, but instead focused on getting the architecture and package just right. There are still more features to build, and more polishing to do, but it's ready for you to try. Judge it for yourself and tell us what you think. https://voidleap.com/ August 19, 2026 at 01:40AM

Show HN: macOS data protection keychain for Electron apps https://ift.tt/lfFCMq7

Show HN: macOS data protection keychain for Electron apps Hey HN, I've been working on Hansel [1] (an encrypted personal data store you can query with agents), and there wasn't a good way to use the modern macOS Data Protection Keychain. Electron's safeStorage [2] uses the legacy file-based keychain, which allows other apps/agents to query it with the `security` CLI. Not great when you have a dozen agents running in the background! The Data Protection Keychain is nice because it limits access via code-signing access groups and lets you set access rules like Touch ID and/or password. 1: https://hansel.so/ 2. https://ift.tt/YPKrQfi https://ift.tt/xrnGYSs August 19, 2026 at 12:25AM

Monday, August 17, 2026

Show HN: Open-source comment section widget, Disqus alternative https://ift.tt/pEZuIRW

Show HN: Open-source comment section widget, Disqus alternative https://ift.tt/32aSQt0 August 18, 2026 at 01:06AM

Show HN: UL-SMF – Open-source linear-complexity ~300x KV-cache compression https://ift.tt/3ja6dt8

Show HN: UL-SMF – Open-source linear-complexity ~300x KV-cache compression https://ift.tt/Bs2XMH5 August 17, 2026 at 10:56PM

Sunday, August 16, 2026

Show HN: built my 15-year-old game idea https://ift.tt/TCd0uGj

Show HN: built my 15-year-old game idea Distro Fighter: WARS is in closed beta, but HN gets its own door: https://ift.tt/UZPQNxp - the arcade at https://ift.tt/t9kd2h7 is open to everyone. Promotions happen at midnight tonight (weekly) and the seasons (migration) are quarterly. With a massive surprise at the end of each year. August 17, 2026 at 12:35AM

Show HN: I shrank DeepSeek V4 Flash to 57GB and it wrote a compiler on my Mac https://ift.tt/gqJoVpN

Show HN: I shrank DeepSeek V4 Flash to 57GB and it wrote a compiler on my Mac I built a specialized package of DeepSeek V4 Flash 0731 (originally 284B total parameters, 13B active), preserving reasoning, tool calling and coding capabilities: https://ift.tt/8A9K2wH... I let it write a minimal C compiler targeting ARM64, then test the result with Fibonacci and FizzBuzz programs, and it succeeded in less than 1 hour, with the full recording at: https://youtu.be/XiwSilmV8B0 You can run it on Silicon Macs with my engine https://ift.tt/5msD9i6 , while one of the core libraries developed to obtain this result is available at https://ift.tt/WfuENg1 . The above recording was on a 128GB memory MacBook M3 Max, but you can also run it on 32GB MacBooks with a very usable context (128K tokens) and projected 5 tok/s. I did try it on a fanless 16GB memory MacBook Air M1 (1.39 tok/s), but unfortunately the available context was very small. How: - First, efficient quantisation: mlx-iqk takes advantage of IQ_K tensor encoding, more efficient than the ones available via llama.cpp or barebones MLX, originally designed by Iwan Kawrakow - I also changed the layout to a k-contiguous one, to make it faster, at least in this Metal setup. - Second, expert pruning: each of the 40 learned-router layers had 256 experts, and not all of them are equally important for the coding use cases. I removed 80B parameters - this is a known technique called REAP, shared at https://ift.tt/h8obxUX . - Third: balancing the cheapest IQ1_S_R4 tensor encoding (~1.5 bits per weight), selectively promoting projections to IQ2_KS or IQ2_K where the measured error reduction justified the bytes. One of the main ideas was not only to save relevant knowledge, but also to not make it forget how to... stop thinking, how to use reasoning. In the first experiments, it would sometimes reason for thousands of tokens without closing its thinking section, or it would go in loops. Then I solved this by heavily weighting tool-calling traces and structured reasoning in the calibration mix. https://ift.tt/WlGR70t August 17, 2026 at 12:13AM

Show HN: Sib - Unixy LLM Client using Git to store converastions, instead SQLite https://ift.tt/8yrOFpA

Show HN: Sib - Unixy LLM Client using Git to store converastions, instead SQLite Hi all :) The reason I like git is that the source tree snapshot and the commit structure itself are always immutable, and destructive operations are just renaming a ref file. Once you understand that, no matter how hard the CLI is to make sense of, you never hesitate to run a command. You can always get it back. LLM conversations aren't as fragile to change as a source tree, but for me an auto-generated SHA-1 hash is more comforting than an auto-generated session title ^~^ The project is in its early stages and contributions are welcome. If you've worked with git plumbing commands, it'll be easy to hack on - and I think it'll be pretty fun. Original: https://ift.tt/kj108vu... https://ift.tt/1WOA3gF August 16, 2026 at 10:35PM

Saturday, August 15, 2026

Show HN: ElevenMusic Sounds, loop library that grows as other people generate https://ift.tt/hpmQcjV

Show HN: ElevenMusic Sounds, loop library that grows as other people generate https://ift.tt/wl06tSG August 16, 2026 at 05:22AM

Show HN: Live Claude Usage HUD for a $38 Thermalright Trofeo Vision LCD https://ift.tt/jhypkr7

Show HN: Live Claude Usage HUD for a $38 Thermalright Trofeo Vision LCD https://ift.tt/FvkMKsd August 16, 2026 at 04:42AM

Show HN: Every MUNI, BART and Caltrain in SF on a Live Map https://ift.tt/mf08HNa

Show HN: Every MUNI, BART and Caltrain in SF on a Live Map https://ift.tt/wa9GdI7 August 16, 2026 at 02:09AM

Show HN: Self-hosted monitoring for AI recommendations (MIT License) https://ift.tt/c8IgVqC

Show HN: Self-hosted monitoring for AI recommendations (MIT License) https://lettertrace.com August 16, 2026 at 12:35AM

Friday, August 14, 2026

Show HN: Is AI Dumber Today? An index of AI model experience from user's opinion https://ift.tt/DGI7zUA

Show HN: Is AI Dumber Today? An index of AI model experience from user's opinion https://isaidumber.today/ August 14, 2026 at 08:50PM

Thursday, August 13, 2026

Show HN: OJCP – an open protocol for agent-consumable job data https://ift.tt/7Zt0wQh

Show HN: OJCP – an open protocol for agent-consumable job data Author here! Agents are applying to jobs for people right now, with progressively more volume, and there's nothing built for it. So they scrape career pages and fight ATS forms with Playwright/Browser Use, which breaks constantly (or they get bot blocked). Employers get buried in applications that don't fit, candidates hear nothing back, and the resume is now an AI-written thing that another AI scores (which breaks the existing model entirely, btw). OJCP is MCP tools for search and apply, a manifest at /.well-known/ojcp.json so agents can find providers, and schemas that extend schema.org instead of replacing it. The playground on the site is a live MCP endpoint, so you can throw calls at it right now. Why a spec at all when models keep getting better at figuring things out? Inference can't produce authorization. An agent can work out what a form wants. It can't establish that someone consented to this specific submission, and then employer has no way to verify who's calling. So TL;DR a more capable agent is also a more capable impersonator. In this model, trust runs both direction. Agents sign requests using the same method that CloudFlare and OpenAI are already using, providers sign their manifests, agents can check against a JWKS, and trust tiers cap how much candidate PII can go to a given provider. Validation happens at consent, so browsing costs nothing and you only pay the verify when the interaction occurs. I'm the CTO of Recruitics (job advertising) and spent time at LinkedIn before that, so I've been at the intersection of hiring and job search for a while and have felt the pain of both sides. Happy to answer any questions! https://ojcp.dev/ August 12, 2026 at 10:27PM

Show HN: At 16K features, flat autoencoders break. Curved space doesn't https://ift.tt/xZKYFCE

Show HN: At 16K features, flat autoencoders break. Curved space doesn't https://ift.tt/BDAjqdJ August 13, 2026 at 11:16PM

Wednesday, August 12, 2026

Show HN: Dynobox – A test runner for AI agent skills and workflows https://ift.tt/SnDFqQc

Show HN: Dynobox – A test runner for AI agent skills and workflows https://ift.tt/YS5V9RM August 12, 2026 at 11:14PM

Tuesday, August 11, 2026

Show HN: Pulp – Zero-allocation C11 telemetry engine (28M logs/SEC to disk) https://ift.tt/jvGIe2c

Show HN: Pulp – Zero-allocation C11 telemetry engine (28M logs/SEC to disk) https://ift.tt/35Sne8Y August 12, 2026 at 01:45AM

Show HN: HyperSAE – Hyperbolic Sparse Autoencoders for LLM Interpretability https://ift.tt/61D3Ccg

Show HN: HyperSAE – Hyperbolic Sparse Autoencoders for LLM Interpretability https://ift.tt/0MjiICg August 12, 2026 at 01:19AM

Show HN: Proxima serves 4x more requests with no hardware change on vLLM https://ift.tt/baq9wOc

Show HN: Proxima serves 4x more requests with no hardware change on vLLM hey everyone, i decided to make a vLLM plugin that implements the Star-KV paper. the results are quiet promising with a decode kernel thats faster than FA2 in higher batch sizes. would love any thoughts and recommendations https://ift.tt/5MAN4cC August 11, 2026 at 10:04PM

Monday, August 10, 2026

Show HN: Slashscore, an open developer graph built from public GitHub activity https://ift.tt/42ghm5R

Show HN: Slashscore, an open developer graph built from public GitHub activity There are millions of public signals across GitHub such as repositories, commits, contributions, languages, organizations, locations, we wanted to turn those signals into something you can explore rather than another static developer directory so we built Slashscore, the developer graph. You can explore developers and their public activity geographically, get a 0–100 score based on an open scoring formula and compare your profile with the broader developer ecosystem. We’re especially interested in whether the graph itself is useful, whether the scoring model makes sense and what we're missing. If you’re a developer, what would you change? https://ift.tt/Usw5CTl August 10, 2026 at 10:45PM

Show HN: PrivateRedact – Offline PII redaction with a local LLM, no cloud https://ift.tt/Rv6MF7E

Show HN: PrivateRedact – Offline PII redaction with a local LLM, no cloud https://ift.tt/AnLMSTc August 10, 2026 at 10:43PM

Sunday, August 9, 2026

Saturday, August 8, 2026

Friday, August 7, 2026

Thursday, August 6, 2026

Show HN: Pokémon Emerald Ported to Raspberry Pi Pico 2 https://ift.tt/78OEeTY

Show HN: Pokémon Emerald Ported to Raspberry Pi Pico 2 Pokémon Emerald ported to the RP2350 microcontroller. No emulator, 60 fps HDMI output. Recompiled from ARMv4T to Cortex-M33 and the Game Boy Advance's video hardware is reimplemented in software on the second core. https://ift.tt/0EJ3eiV August 7, 2026 at 04:49AM

Show HN: Validate your idea, know what to charge, and how to get first users https://ift.tt/OmGyRfa

Show HN: Validate your idea, know what to charge, and how to get first users I made an idea validation tool that can quickly adjust or kill off weak ideas. The big difference between Nell and any other validation tool is it compares your idea with real passed and failed examples, calculates TAM, overall buying sentiment, and 14 more important pointers that VCs value. It then also creates the most confident go to market plans to get initial users. My goal with Nell is to compress the research that takes more than a month into a few minutes so that early stage founders can spend time on polishing, preparing strong investor pitches, and executing GTM. How it works: Once you’re in. Just paste your idea and you will get a 14+ pointer diligence report having - validation of the problem you are trying to solve, funding potential based on past similar ventures, demand signals, willingness to pay, pricing research, market size, revenue ceiling, and other critical metrics. Please check it out and feel free to run a validation test on your idea. I will really appreciate feedback from the HN community. Live tool: https://nellailabs.ai https://nellailabs.ai August 7, 2026 at 12:34AM

Show HN: Silo – S3-compatible object storage, a maintained fork of MinIO https://ift.tt/X9euD14

Show HN: Silo – S3-compatible object storage, a maintained fork of MinIO https://silo.pgsty.com August 6, 2026 at 11:05PM

Wednesday, August 5, 2026

Show HN: LiminalML – Study ML or SWE at interview depth, grounded in your resume https://ift.tt/89gEzue

Show HN: LiminalML – Study ML or SWE at interview depth, grounded in your resume https://liminalml.com August 6, 2026 at 12:39AM

Show HN: Twocal – Calendar sync that verifies both sides actually agree https://ift.tt/GDpjIYx

Show HN: Twocal – Calendar sync that verifies both sides actually agree https://twocal.app/ August 5, 2026 at 09:37PM

Tuesday, August 4, 2026

Show HN: Adapt, Automatically Turns Files into REST APIs, Web UI, and MCP https://ift.tt/H8Dlr5L

Show HN: Adapt, Automatically Turns Files into REST APIs, Web UI, and MCP https://ift.tt/nJcM8H3 August 4, 2026 at 10:49PM

Monday, August 3, 2026

Show HN: Product analytics (and evals) for agent sessions on your MCP https://ift.tt/jfJ2Ih7

Show HN: Product analytics (and evals) for agent sessions on your MCP Hi HN! We’re Theodore and Louis, founders of Armature (YC P26). We reconstruct the entire session behind the MCP tool calls you receive, including what the user asked their agent to do and what the agent thought. You wrap your MCP in 3 lines of code (our SDK is available in Typescript, Python and Go) and start seeing in your dashboard: - All sessions reconstructed: it’s like reading the real conversation the user had inside Claude or ChatGPT! - A ranking of your MCP most popular use cases, built from sessions clustering - The most frequent issues your users’ agents encounter so you can fix them. Here is a quick demo: https://youtu.be/ZFlvquhyNMQ The story behind this is that we initially launched Armature as a standalone testing tool ( https://ift.tt/jdGrQ7t... ) that could naturally be used through an MCP itself. We quickly realized we had no idea how our users were using Armature MCP and if they were satisfied with it or frustrated. It’s something we had also experienced in our previous companies: Louis built MCPs exposed to millions of users and Theo was a Forward Deployed Engineer at Palantir before joining a Datadog spin-off as Founding Engineer. Both testing and product analytics had always been real pains when exposing a product to agents but we always thought there wasn’t much we could do about analytics because the conversation lived in our users’ AI client. Then it struck us: what if we asked the agents why they were making this or that tool call? And what’s the user's intent or potential frustration? So we started experimenting with MCP instrumentation and the use-cases actually surprised us! Many of our first customers had implemented workarounds for their CI to trigger new tests or for their coding agents to fetch the results efficiently. Even though we talked to our first users regularly, they had never shared this feedback with us. We then built automations to automatically cluster use-cases, identify issues frequently encountered and let our own coding agents fix them. When our CTO friends heard about this, they wanted to try it for themselves so we gave them access to a cloned version of our internal product and they started sharing feedback like they never did on our “real” product! That’s when we decided to start working seriously on MCP Analytics as a product. At first we were afraid of degrading MCP performance so we iterated until we reached the exact same success rate as without our instrumentation (89.17 % vs 89.15 % pass rate out of 870 runs). Then privacy was an obvious constraint so we applied the same methods we had learned from working with banking data or building sensitive data scanning in logs. Today, redaction runs client-side before reaching our servers. There are still a lot of things we haven’t fully figured out: not all fields are equally filled by all models, session fingerprinting for serverless / stateless MCPs isn’t perfect, and use-case clustering remains to be optimized. But we are finally launching our analytics product to everyone, self-serve at https://armature.tech with a set-up that takes less than 5 minutes and a generous free tier. And now we are working on fully closing the loop, bringing evals back in our product so we can: identify top workflows and issues -> recommend fixes and improvements -> test fixes at scale on the same workflows run by users, across all harnesses and models -> open PRs to ship fixes directly. The evals can be generated automatically from the session analytics so you can catch every regression and can test every improvement’s real impact across all models and harnesses before shipping it. Here’s an example to make it more concrete: 10 days ago, a marketing automation platform which has had early access to what we built for weeks identified thanks to MCP Analytics that users were frustrated not being able to change their target audience after campaign creation. So they shipped the feature and tested it successfully locally with Claude Code on Fable 5. Then a few days later when preparing their new MCP public release, they ran a suite of evals on Armature and realized that small models could hallucinate audience_ids which would lead their MCP to send the campaign to ALL their contacts by default (which could obviously lead to disasters in prod). This is the kind of story that makes what we are building feel so helpful! Now, the most useful feedback for us would be to know what’s still missing in our product so you can feel you are now in full control of the “Agent Experience”. And if you run an MCP in production we’d also love to know: what do you do today to know if agents succeed and if the users behind them are happy? https://armature.tech/ August 3, 2026 at 11:17PM

Sunday, August 2, 2026

Show HN: Schmess – chess with no turns; pieces freeze on cooldown after moving https://ift.tt/SGOtyIz

Show HN: Schmess – chess with no turns; pieces freeze on cooldown after moving Solo dev here. Schmess is chess with the turn structure removed. Both players move whenever they have a free slot, and every piece freezes on a cooldown after it moves. It's on a half board, with 100+ puzzle stages. If two people played this over a physical board they'd scrape each other's fingers raw - that's roughly the tempo. Plays in the browser, no signup. 29 stages free, one-time unlock for the rest ($3.99 / Rs 149). No subscription, no ads. It's the second of a few small-board variants I'm building (halfchess was the first, more coming). Building a chess engine for a game with no turns turned out to be the hard part - I've put the war stories in a comment below. https://schmess.com/ August 2, 2026 at 11:47PM

Show HN: TamedTable, AI ETL in Natural Language https://ift.tt/S1NzTmt

Show HN: TamedTable, AI ETL in Natural Language Hi HN, TamedTable is an LLM harness for data ETL. And yes, it was developed using AI, meaning you can take the entire specification and recreate it to your desires: https://ift.tt/NpDiuC7 https://ift.tt/cR2gqbT August 2, 2026 at 02:51PM

Saturday, August 1, 2026

Show HN: SteerPlane – Deterministic runtime guardrails for AI agents https://ift.tt/avByo0p

Show HN: SteerPlane – Deterministic runtime guardrails for AI agents https://ift.tt/AGPoslX August 2, 2026 at 01:09AM