ResearchAudio evidence lab
Seventeen instruments for the decision behind the demo.
Check whether a task has an error signal, audit the benchmark and launch claim, size GPU and KV-cache memory, budget tokens, price voice outcomes and latency, calculate real operating cost, inspect the production loop, stress-test the business case, then configure the coding agent.
Join 51,000+ readers for the next evidence-led briefing → New here? Get the guided starter kit → Publish here? Embed a working calculator →AI launch evidence scorecard
Seven checks expose the gap between a launch claim and decision-ready proof.
Score a launch → Instrument 02 / economicsCost per successful task
Calculate the retry-adjusted cost after model calls and human review.
Calculate real cost → Instrument 03 / reliabilityAgent loop diagnostic
Find missing termination, validation, recovery, and escalation guardrails.
Inspect a loop → Instrument 04 / business caseAI agent ROI stress test
Price failed automations, human review, recurring cost, and implementation payback.
Stress-test the ROI → Instrument 05 / token budgetLLM API cost calculator
Estimate monthly spend from traffic, input and output tokens, current prices, prompt caching, retries, and attached tools.
Build a token budget → Instrument 06 / cache economicsPrompt caching cost calculator
Model cache reads, cache writes, monthly savings, and the hit rate required to break even.
Find the break-even rate → Instrument 07 / developer workflowCodex CLI config.toml generator
Generate project document fallbacks and the instruction byte limit without displacing AGENTS.md.
Build the config → Instrument 08 / conversational latencyVoice AI latency calculator
Budget endpointing, transcription, parallel fast and slow model branches, TTS, and playout.
Map the silence → Instrument 09 / GPU fitLLM GPU memory calculator
Estimate weight VRAM, inference headroom, usable capacity, and the minimum GPU count for a model.
Size the deployment → Instrument 10 / inference cacheLLM KV cache calculator
Calculate cache memory per token, full-context sequence, concurrency level, and GQA/MHA architecture.
Size the cache → Instrument 11 / voice economicsAI voice agent cost
Calculate loaded cost per AI-resolved call after the full voice stack, fixed fees, failures, and human handoffs.
Price the outcome → Instrument 12 / benchmark protocolAI benchmark audit
Check twelve protocol details behind a model, agent, or coding benchmark before trusting the score.
Audit the benchmark → Instrument 13 / task framingAI task-fit diagnostic
Check whether the workflow has a target, independent judge, timely error signal, bounded actions, and a human owner for premise changes.
Test the task fit → Instrument 14 / Codex automationCodex exec command builder
Generate a shell-safe command for JSONL events, ephemeral runs, sandboxing, structured output, and deliberate non-Git execution.
Build the command → Instrument 15 / local inferenceLocal LLM GPU compatibility
Check 7B, 13B, 32B, and 70B weight floors against common 8 GB to 141 GB GPU profiles before adding context.
Check the hardware → Instrument 16 / model discoveryWhat LLM can I run?
Start with 8 to 141 GB of VRAM and find the largest standard model-size floor it clears at INT4, INT8, or FP16.
Find the model tier → Instrument 17 / source lookupHugging Face VRAM + badge
Paste a public model repository, calculate three source-backed weight floors, copy an attributed badge, then continue in the native Hugging Face Space or companion GitHub Action.
Inspect a repository →Daily index + search field notes
Start with the simple metric. Finish with the system.
Primary-source worksheets turn security, voice-pricing, and GPU-sizing questions into explicit controls and assumptions.
Trending local LLM hardware index
Track twelve text-generation models moving on Hugging Face, with source-backed parameter counts and transparent INT4, INT8, and BF16 planning floors.
Inspect today’s model queue → Reviewed cluster / repository-specificLocal LLM model hardware pages
Compare eight reviewed repositories with source-backed parameter totals, precision floors, card-fit matrices, architecture metadata, and editable GPU plans.
Browse the reviewed model pages → Field guide / local LLM hardwareLocal LLM GPU and VRAM requirements
Start from a model, a VRAM budget, a specific GPU, or a released checkpoint, then follow the evidence to an editable deployment plan.
Use the hardware decision guide → Field note / formulaVoice AI cost per minute
Add the five usage meters, fixed fees, failed resolutions, and human handoffs before comparing stacks.
Read the formula → Field note / worksheetAI receptionist cost
Open editable after-hours, appointment-booking, and high-volume planning scenarios.
Use the worksheet → Field note / GPU sizing7B vs 13B LLM GPU requirements
Compare six weight and 32K-cache profiles, then see how GQA versus MHA changes the fit.
Compare the smaller models → Field note / GPU sizing70B LLM GPU requirements
Separate INT4, INT8, and FP16 weight floors from explicit 32K KV cache and runtime headroom.
Compare the scenarios → Field note / 12 GB GPUWhat LLM can an RTX 3060 12GB run?
Separate the generic 14B INT4 weight tier from exact Qwen3 and gpt-oss profiles once context memory is included.
Inspect the RTX 3060 boundary → Field note / 8 GB GPUWhat LLM can an RTX 4060 8GB run?
Follow Qwen3 8B from a comfortable 4K fit to the exact 16K and 32K context-memory boundaries.
Inspect the RTX 4060 boundary → Field note / 16 GB GPUWhat LLM can an RTX 4060 Ti 16GB run?
Verify the 16GB variant, then inspect the narrow Qwen3 14B and gpt-oss-20b planning fits.
Inspect the RTX 4060 Ti boundary → Field note / 24 GB comparisonRTX 3090 vs RTX 4090 for local LLMs
See why both cards share the same one-card model ceiling, then separate capacity from measured runtime value.
Compare the 24 GB cards → Field note / 24 GB GPUWhat LLM can an RTX 4090 run?
Compare generic INT4, INT8, and FP16 tiers with exact Qwen3 and gpt-oss profiles near the 24 GB boundary.
Inspect the RTX 4090 boundary → Field note / 8 GB vs 16 GBRTX 5060 Ti 8GB vs 16GB for local LLMs
Compare exact Qwen3 and gpt-oss boundaries to see where the doubled VRAM changes the answer.
Compare the RTX 5060 Ti variants → Field note / 12 GB vs 16 GBRTX 4070 Super vs 4070 Ti Super for local LLMs
Map the extra 4GB to exact Qwen3 and gpt-oss outcomes before choosing between the two Ada cards.
Compare the RTX 4070 Super cards → Field note / Apple unified memoryMac mini M4 for local LLMs
Choose between 16GB, 24GB, 32GB, 48GB, and 64GB using a conservative shared-memory budget.
Choose the Mac mini memory tier → Field note / 16 GB GPUWhat LLM can an RTX 5080 run?
Inspect the razor-thin Qwen3 14B and short-context gpt-oss-20b paper fits before trusting a 16 GB deployment.
Inspect the RTX 5080 boundary → Field note / 32 GB GPUWhat LLM can an RTX 5090 run?
See where Qwen3 32B and a 128K gpt-oss-20b profile cross from a 24 GB near miss into a narrow 32 GB fit.
Inspect the RTX 5090 boundary → Field note / model familyQwen2.5 GPU requirements
Compare 7B, 32B, and 72B across INT4, INT8, BF16, exact 32K KV cache, and the 128K context jump.
Choose the checkpoint → Field note / dense + MoEQwen3 GPU requirements
Compare 8B, 14B, 32B, and 30B-A3B across INT4, INT8, BF16, exact 32K cache, and Qwen's observed memory.
Size the Qwen3 family → Field note / current modelDeepSeek V4 Flash GPU requirements
Start with the official 0731 release and its 155.43 GiB mixed-precision artifact, then compare 32K, 128K, 384K, and 1M context plans.
Size DeepSeek V4 Flash 0731 → Field note / current modelGLM-5.2 GPU requirements
Start with the official 703.74 GiB FP8 artifact, then compare the 8×80 GiB shortfall, 8×H200 baseline, 8×B200 full-context recipe, and BF16 floor.
Size GLM-5.2 FP8 → Field note / current modelKimi K3 GPU requirements
Start with the official 1,453.74 GiB MXFP4 artifact, then separate the 8×H200 shortfall, 12×H200 raw floor, and Moonshot's 64-plus-accelerator recommendation.
Size Kimi K3 → Field note / current familyGemma 4 GPU requirements
Compare Google's E2B, E4B, 12B, 26B A4B, and 31B loading estimates across BF16, SFP8, and Q4_0, then inspect the exact 12B files.
Size the Gemma 4 family → Field note / current modelDiffusionGemma GPU requirements
Start with the exact 48.10 GiB BF16 and 17.53 GiB NVIDIA NVFP4 artifacts, then keep 256K context outside the weight-only claim.
Size DiffusionGemma → Field note / open weightsgpt-oss hardware requirements
Compare the official 20B and 120B checkpoints, MXFP4 weight floor, 4K deployment budget, and 128K cache jump.
Size the OpenAI models → Field note / agent securityAI agent security checklist
Test twelve controls for goal hijacking, tool permissions, isolation, secrets, audit trails, shutdown, and recovery.
Inspect the authority boundary →Field guide / model economics
Keep frontier reasoning. Cut frontier typing.
The 10-page Fable 5 Cost Playbook turns model routing, effort, caching, context, batch execution, and spend caps into one practical checklist.
Lab protocol
Start with the decision.
A metric matters only if it changes what you do.
Name the hidden denominator.
Retries and failures belong in the result.
Leave with a next test.
Every score should create an inspection, not a vibe.
One teardown at a time
Get the evidence behind the next AI launch.
ResearchAudio traces primary sources, costs, limitations, and adoption questions for engineers and builders.
Prefer the hosted signup page?Join ResearchAudio free →