RAResearchAudio

Reviewed model hardware cluster

Start with the repository, not the model-size rumor.

Eight source-backed hardware pages turn public safetensors totals into comparable INT4, INT8, and BF16 floors. Every page names its license metadata, architecture, usable-memory rule, and missing deployment costs.

Browse reviewed pages →See today’s model queue →

Publication gate

Source-backed.
Public repository metadata and safetensors totals are linked on every page.

Decision-specific.
Each page resolves a distinct memory or card-tier boundary.

Not deployment proof.
Weight fit never stands in for cache, runtime support, speed, quality, or licensing review.

Initial reviewed batch

Repository-specific GPU plans.

Compare the INT4 floor and architecture at a glance, then open the page for the full precision ladder, card-fit matrix, source record, limitations, and prefilled calculator.

2.7B parameters · 1.51 GiB INT4 floor

LFM2.5-2.6B

This page earns a standalone URL because it resolves a real boundary: unlike generic 7B guidance, the source-backed 2.70B total makes even a BF16 weight floor plausible on 8 GB, while still showing why weight fit alone is not deployment proof.

First INT4 tier
8 GB
Architecture
Lfm2ForCausalLM
Open the reviewed plan →
20.21B parameters · 11.3 GiB INT4 floor

maple-preview

This page is distinct from generic 20B or 24B advice because it makes the 12 GB versus 16 GB boundary explicit using the repository's own parameter total and a declared usable-memory rule.

First INT4 tier
16 GB
Architecture
MapleForCausalLM
Open the reviewed plan →
127.49B parameters · 71.24 GiB INT4 floor

Ling-3.0-flash

This page deserves separate treatment because the 80 GB result is unusually narrow and because a mixture-of-experts name can tempt readers to size memory from active parameters instead of the source-reported total.

First INT4 tier
80 GB
Architecture
BailingMoeV3ForCausalLM
Open the reviewed plan →
34.66B parameters · 19.37 GiB INT4 floor

KAT-Coder-V2.5-Dev

The page resolves a high-value 24 GB boundary with source-specific arithmetic and keeps that decision separate from the broader Qwen3 family guide, which covers different released parameter tiers and architectures.

First INT4 tier
24 GB
Architecture
Qwen3_5MoeForConditionalGeneration
Open the reviewed plan →
5.72B parameters · 3.2 GiB INT4 floor

fuse-1-Lite

The standalone value is provenance-aware sizing: it prevents readers from applying the 35B parent name to a repository whose published safetensors total produces a very different memory floor.

First INT4 tier
8 GB
Architecture
Fuse3ForCausalLM
Open the reviewed plan →
8.03B parameters · 4.49 GiB INT4 floor

Llama-3.1-8B-Instruct

This page provides a source-specific precision ladder for the exact instruct repository and adds the gated-license boundary that generic 8B sizing pages often omit.

First INT4 tier
8 GB
Architecture
LlamaForCausalLM
Open the reviewed plan →
35.11B parameters · 19.62 GiB INT4 floor

BTL-4

BTL-4 earns a reviewed page because its 24 GB INT4 boundary, MoE architecture, and limited adoption signals create a different decision from a generic dense 35B guide.

First INT4 tier
24 GB
Architecture
Qwen3_5MoeForConditionalGeneration
Open the reviewed plan →
3.82B parameters · 2.14 GiB INT4 floor

Phi-3.5-mini-instruct

The page is independently useful because it distinguishes the 3.82B repository from generic 7B advice and shows the precise point where BF16 crosses from 8 GB to 12 GB.

First INT4 tier
8 GB
Architecture
Phi3ForCausalLM
Open the reviewed plan →

One method across the cluster

Comparable arithmetic, explicit uncertainty.

Weight calculation. Total safetensors parameters × bits ÷ eight ÷ 1024³. A shared 20% allowance is added before comparing with hardware.

Usable capacity. The first single-card tier must fit inside 90% of labeled VRAM. This creates a reproducible screen without claiming that every runtime reserves the same amount.

Progressive rollout. The batch remains intentionally small. Search indexing, useful visits, and attributed subscriptions must be measured before more reviewed pages are added.

Daily refresh. Public downloads, likes, modification dates, parameter totals, license metadata, and architecture are checked again by the scheduled source refresh. A failed or incomplete response cannot overwrite the last verified site.

The next useful model decision

Get evidence, not another release headline.

ResearchAudio translates model claims and infrastructure boundaries into practical decisions for engineers and builders.