LListicle

Best Small Models for Text Classification You Can Run on a Laptop (2026)

By @pickwiseresearched with AIUpdated

Small open models that sort text into your own labels on a laptop, for spam, intent, routing, moderation and redaction, with no API bill and nothing leaving the machine. The picks rest on independent tests by Red Hat's AI safety team, Oliver Borchers and the BTZSC benchmark. Checked October 5, 2026.

It started with a build on Made with Jev: a developer swapped Jev, TypeSafe's hosted decision model, for a DeBERTa large classifier on a base M1 Mac to decide, inside a Claude Code hook, whether a message was just a question, and found it just as fast and just as good for that job. Jev is the yardstick here. In Borchers' field test on 400 fresh arXiv papers it scored 86.3% with no training, for about $0.04 per million input tokens.

How to choose comes down to labels. With none, try Clef-flash on a 16 GB Mac with a Pro or Max chip, or GLiClass on anything else. With a few dozen examples per label, embed the text and train a logistic regression on top. With a few hundred, fine-tune DeBERTa-v3 or ModernBERT. No labels yet? Have Jev or a large LLM label a few hundred of your own examples, then train the small model on those.

Skip facebook/bart-large-mnli, still the default in Hugging Face's zero-shot pipeline: it dates from 2019, runs a full pass for every label, and trails newer models in both the 28-task zero-shot benchmark and BTZSC. The list is ordered by how well each model did in those tests for the work it takes to set up, with the specialists for PII, prompt injection and secrets last.

At a glance

PickTypeLabels neededRuns onTested on
EmbeddingGemma 308M, plus a logistic regressionEmbedding model + simple classifierA few dozen per class (shown with a bigger embedder)Phones, laptops and desktops, per GoogleMacBook Pro M1 Max: 33.5 embeddings a second (Q8)
DeBERTa-v3, fine-tuned (base or large)Encoder you fine-tune on your labelsA few hundred: 500 matched classic ML on 5,000MacBook Pro M1, CPU onlyMacBook Pro M1 CPU: 54-80 ms a prompt (base)
ModernBERT, fine-tuned (base or large)Encoder you fine-tune on your labelsHundreds; thousands for the best resultsMacBook Pro M1 CPU (Laya, a ModernBERT model)MacBook Pro M1 CPU: 119 ms a prompt (Laya)
Cloudflare Clef-flash 9BDecision model, same API as JevNone; new categories need no retrainingApple silicon Mac, 16 GB+ (4-bit MLX)16 GB base M4 MacBook: 3.9 s a decision
Qwen3-4B-Instruct-2507Small LLM you prompt or fine-tuneFew-shot prompt; tuned on 10,000 synthetic examplesMacBook Pro M3 Pro, 18 GB (4-bit MLX)M3 Pro: reads 615 tokens/s, writes 44.6 tokens/s
SetFitFew-shot fine-tuning of a small embedderAs few as 8 per classCPU: trains in a few minutes, per Hugging FaceUnnamed CPU: 716 texts a second (MiniLM-L6 body)
GLiClass v3Zero-shot classifierNone; you pass labels at run timeM4 Max CPU, or a Galaxy S26 phone (edge model)16 GB M1 Pro: 0.46 s an abstract (large)
Apple Foundation ModelsBuilt into macOS: Apple's on-device LLMNone; give it your labels as an enumMac with M1 or later (Apple Intelligence)Desktop M4 Max Mac Studio: 265 ms to first token
Ettin encoders (17M to 68M)Encoder you fine-tune on your labelsMany: credential-line-32m used 87,707x86 CPU with no GPU (CPU-only fine-tune)Desktop i7-13700K: 267/31 pairs/s (17M/68M rerankers)
OpenAI Privacy FilterPII detector (token classifier)None (8 fixed labels)A laptop or a web browser, per OpenAI16 GB M2 MacBook Pro: about 1 s a document
Llama Prompt Guard 2 (86M)Prompt-injection and jailbreak detectorNone; set the threshold on normal trafficCPU on Apple silicon (chip not named)Apple silicon CPU: 149 ms median a call
credential-line-32mSecret scanner, one line at a timeNone (fixed: credential or not)CPU; card says laptop, measured on a desktopIntel desktop CPU, 4 threads: 4.1 ms a line
  1. EmbeddingGemma 308M, plus a logistic regression

    Turn each text into an embedding and train a logistic regression on a few dozen labeled examples per class: in Borchers' test that recipe, with a larger embedder, matched Jev. EmbeddingGemma is the small embedder to use, scoring 87.6 on MTEB's English classification tasks, about nine points above bge-large. The catch is Google's Gemma license, which gates the download. Pricing: free, under Google's Gemma terms.

    Type
    Embedding model + simple classifier
    Labels needed
    A few dozen per class (shown with a bigger embedder)
    Runs on
    Phones, laptops and desktops, per Google
    Tested on
    MacBook Pro M1 Max: 33.5 embeddings a second (Q8)
    Memory
    Under 200 MB RAM when quantized
    Size
    308M parameters
    License
    Gemma terms (gated)
    Context
    2,048 tokens
  2. DeBERTa-v3, fine-tuned (base or large)

    The model behind the Made with Jev build, and fine-tuned it still edges out newer encoders on accuracy. In Red Hat's prompt-injection test a fine-tuned 200M DeBERTa scored 89.0% in 54 ms on an M1, ahead of Jev's 86.4%. It reads only 512 tokens, works in English only, and needs a few hundred labeled examples before it's useful. Pricing: free (MIT).

    Type
    Encoder you fine-tune on your labels
    Labels needed
    A few hundred: 500 matched classic ML on 5,000
    Runs on
    MacBook Pro M1, CPU only
    Tested on
    MacBook Pro M1 CPU: 54-80 ms a prompt (base)
    Memory
    Large: 0.8 GB RAM, 1.1 GB GPU on an M1 Pro
    Size
    Backbone 86M (base) or 304M (large)
    License
    MIT
    Context
    512 tokens
  3. ModernBERT, fine-tuned (base or large)

    Reads up to 8,192 tokens, so it handles long support tickets, emails, logs and agent traces that would overflow DeBERTa. Fine-tuned in Borchers' test, the large version sorted fresh arXiv papers at 87.0%, a hair above Jev. Its famous speed depends on an Nvidia GPU with Flash Attention, and on a laptop it runs about as fast as DeBERTa-v3. Pricing: free (Apache 2.0).

    Type
    Encoder you fine-tune on your labels
    Labels needed
    Hundreds; thousands for the best results
    Runs on
    MacBook Pro M1 CPU (Laya, a ModernBERT model)
    Tested on
    MacBook Pro M1 CPU: 119 ms a prompt (Laya)
    Memory
    599 MB download (base)
    Size
    149M (base) or 395M (large)
    License
    Apache 2.0
    Context
    8,192 tokens
  4. Cloudflare Clef-flash 9B

    Needs no labels, speaks Jev's API, and came closest to Jev of the open models that fit a 16 GB laptop. On a base M4 MacBook, Marco Mornati found it matched Jev on routing (90% vs 89%) and beat it on risk flags. Each decision took about 3.9 seconds there, so anything interactive wants a Pro or Max chip. Pricing: free (Apache 2.0).

    Type
    Decision model, same API as Jev
    Labels needed
    None; new categories need no retraining
    Runs on
    Apple silicon Mac, 16 GB+ (4-bit MLX)
    Tested on
    16 GB base M4 MacBook: 3.9 s a decision
    Memory
    6.2 GB download, 7 to 8.6 GB peak (4-bit MLX)
    Size
    9B parameters
    License
    Apache 2.0
    Context
    64K tokens
  5. Qwen3-4B-Instruct-2507

    If you'd rather fine-tune a small LLM than an encoder, start here. In distil labs' study of 15 small models, a LoRA-tuned Qwen3-4B-Instruct-2507 matched or beat its 120B-plus teacher on two of three classification tasks, and Apache 2.0 lets you ship it. Untuned, results swing with the prompt: one test got 53% or 78% from the newer Qwen3.5 4B depending on how the answer was read. Pricing: free (Apache 2.0).

    Type
    Small LLM you prompt or fine-tune
    Labels needed
    Few-shot prompt; tuned on 10,000 synthetic examples
    Runs on
    MacBook Pro M3 Pro, 18 GB (4-bit MLX)
    Tested on
    M3 Pro: reads 615 tokens/s, writes 44.6 tokens/s
    Memory
    3.3 GB peak (4-bit MLX)
    Size
    4.0B (3.6B non-embedding)
    License
    Apache 2.0
    Context
    262,144 tokens
  6. SetFit

    Hugging Face's few-shot trainer builds a classifier from 8 to 64 examples per class by fine-tuning a small sentence-transformer, and it trains on a laptop CPU in minutes. On a customer-review dataset in its paper, 8 examples per class rivaled RoBERTa-large trained on all 3,000. A plain embedding plus logistic regression scored about the same in the Model2Vec team's benchmark, so try that first. Pricing: free (Apache 2.0).

    Type
    Few-shot fine-tuning of a small embedder
    Labels needed
    As few as 8 per class
    Runs on
    CPU: trains in a few minutes, per Hugging Face
    Tested on
    Unnamed CPU: 716 texts a second (MiniLM-L6 body)
    Memory
    70 MB (MiniLM body) or 420 MB (MPNet body)
    Size
    110M (MPNet body) or 355M (RoBERTa-large)
    License
    Apache 2.0
    Context
    Depends on the body model
  7. GLiClass v3

    Zero-shot labels on a plain CPU: GLiClass scores every label in one pass, so a long label list barely slows it down, and the 32.7M edge model answers in about 9 ms on an M4 Max. Use it when labels change too often to train anything. Accuracy is the price: on Borchers' fresh papers the large version scored 64.3% to Jev's 86.3%. Pricing: free (Apache 2.0).

    Type
    Zero-shot classifier
    Labels needed
    None; you pass labels at run time
    Runs on
    M4 Max CPU, or a Galaxy S26 phone (edge model)
    Tested on
    16 GB M1 Pro: 0.46 s an abstract (large)
    Memory
    131 MB (edge) to 1.75 GB (large) download
    Size
    32.7M (edge) to 439M (large)
    License
    Apache 2.0
    Context
    1,024 tokens in the default pipeline
  8. Apple Foundation Models

    Built into macOS on Apple silicon Macs with Apple Intelligence, so there's no model to download and no per-token cost. Constrain the answer to an enum and it stays on your label list: in one developer's test it kept to the list in 15 of 15 runs, where free text invented categories. The context window is just 4,096 tokens, and the model changes with each OS update. Pricing: free, built into macOS.

    Type
    Built into macOS: Apple's on-device LLM
    Labels needed
    None; give it your labels as an enum
    Runs on
    Mac with M1 or later (Apple Intelligence)
    Tested on
    Desktop M4 Max Mac Studio: 265 ms to first token
    Memory
    Apple Intelligence: up to 8 to 14 GB of storage
    Size
    About 3B parameters
    License
    Apple licence, accepted once; no API key
    Context
    4,096 tokens per session
  9. Ettin encoders (17M to 68M)

    The smallest modern encoders, from Johns Hopkins, for classifiers that must run on a CPU or in a browser tab. A community fine-tune of the 17M model scored 94.2 F1 on PII extraction against GPT-4o-mini's 78.7, and credential-line-32m below is built on the 32M. The tiny sizes give up accuracy on general tasks: the 17M averages 79.2 on GLUE. Pricing: free (MIT).

    Type
    Encoder you fine-tune on your labels
    Labels needed
    Many: credential-line-32m used 87,707
    Runs on
    x86 CPU with no GPU (CPU-only fine-tune)
    Tested on
    Desktop i7-13700K: 267/31 pairs/s (17M/68M rerankers)
    Memory
    17M: 67 MB ONNX file, 17 MB in int8
    Size
    17M, 32M or 68M parameters
    License
    MIT
    Context
    Up to 8K tokens
  10. OpenAI Privacy Filter

    Finds names, contact details, account numbers, private dates and secrets so you can redact text before it reaches a cloud model or a training set. It scored 0.96 F1 on the PII-Masking-300k benchmark and handled about one document a second on a 16 GB M2 MacBook Pro. One independent test caught it missing AWS keys and database connection strings, so keep a regex scanner in front. Pricing: free (Apache 2.0).

    Type
    PII detector (token classifier)
    Labels needed
    None (8 fixed labels)
    Runs on
    A laptop or a web browser, per OpenAI
    Tested on
    16 GB M2 MacBook Pro: about 1 s a document
    Memory
    2.8 GB weights, under 6 GB RAM
    Size
    1.5B total, 50M active
    License
    Apache 2.0
    Context
    128,000 tokens
  11. Llama Prompt Guard 2 (86M)

    Meta's small classifier for jailbreaks and "ignore previous instructions" attacks hidden in prompts, documents and tool output, run before an agent acts. Out of the box it's close to useless: against 629 realistic agent attacks, the default threshold caught 6. Lowered to 0.003, the 86M caught 97 to 100% at about 5% false alarms, so tune it on your own traffic. Pricing: free under the Llama 4 Community License.

    Type
    Prompt-injection and jailbreak detector
    Labels needed
    None; set the threshold on normal traffic
    Runs on
    CPU on Apple silicon (chip not named)
    Tested on
    Apple silicon CPU: 149 ms median a call
    Size
    86M backbone; 0.3B with embeddings
    License
    Llama 4 Community License (gated)
    Context
    512 tokens
  12. credential-line-32m

    Reads one line of code, config or logs and scores whether it holds a literal credential, in about 4 ms on a desktop CPU. On its author's holdout of human-written lines it scored 0.86 F1, where detect-secrets managed 0.51. It's a one-author project that misses about one credential in five, so run it beside gitleaks rather than instead of it. Pricing: free (MIT).

    Type
    Secret scanner, one line at a time
    Labels needed
    None (fixed: credential or not)
    Runs on
    CPU; card says laptop, measured on a desktop
    Tested on
    Intel desktop CPU, 4 threads: 4.1 ms a line
    Memory
    128 MB ONNX file (fp32)
    Size
    32M parameters
    License
    MIT
    Context
    One line, up to 128 tokens

More guides like this

Best AI Task Managers (2026)

Fifteen task and project managers whose AI does real work, from planning your day to running agents on a board, picked from hands-on tests and roundups by Wirecutter, PCMag, Zapier, TechRadar and Tom's Guide. Prices checked October 5, 2026. Match it to how you work. If you want AI to plan your day, Sunsama walks you through a daily plan, Reclaim fits tasks around your meetings and has a free plan, and Motion plans and reshuffles your whole calendar. If you want a fast list with a little AI, start with Todoist. Teams should try Asana first, ClickUp if they want everything in one app, or Notion if their work already lives in docs. Software teams belong in Linear. On a budget: Todoist Pro and Reclaim Starter together come to about $15 a month billed yearly and put your Todoist tasks on your calendar, less than Motion's $19 alone. Google Gemini schedules tasks into Google Calendar for free. Watch the AI add-ons: ClickUp's Brain AI costs $9 a user on top of the plan, Notion's agents need the $20 Business plan, and Motion, Asana and monday.com meter AI with credits or request limits. Ordered by wins in named hands-on tests, then by how many independent reviewers recommend each app, then by fit for individuals and small teams and by price. Disclosure: Wand is made by Automatique, the team behind Listicle; it's rated by the same criteria as everything else.

15 itemsupdated 13h ago

Best Wispr Flow Alternatives (2026)

Fifteen dictation apps that type into any app, for people leaving Wispr Flow over its price, its cloud processing or its platforms, picked from tests and reviews by 9to5Mac, ZDNet, Zapier, TechCrunch, Lifehacker and MacStories. Prices checked October 5, 2026. Leaving over price? Handy and FluidVoice are free, VoiceInk is $25 once for a Mac, and Aqua Voice Pro is $8 a month billed yearly against Wispr Flow's $12. For privacy, Superwhisper, Handy, FluidVoice, VoiceInk and Spokenly can transcribe without your audio leaving the computer, while Aqua and Typeless are cloud services like Wispr. On Linux, try Handy, Spokenly or Typeless. On Android, Typeless or Superwhisper. Accuracy depends on the model as much as the app: in Voice-list's test, Superwhisper's own models ranged from 1.6% to 22.8% of words wrong, so try another model before giving up on an app. Free tiers differ too. Aqua's is 1,000 words in total, Willow's free plan uses a lighter model, and Superwhisper's never runs out. Ordered by wins in named tests and how many independent reviewers recommend each app, then by how well it fixes the usual reasons for leaving Wispr Flow (price, privacy, platforms), then by price. Disclosure: Chirp is made by Automatique, the team behind Listicle; it's rated by the same criteria as everything else.

15 itemsupdated 13h ago

Best Clipboard Managers for Windows (2026)

Fifteen clipboard managers for Windows 10 and 11, ranked on hands-on tests and reviews from MakeUseOf, XDA, How-To Geek, PCWorld, Zapier and Windows Central, for anyone who has outgrown Win+V. Prices checked October 5, 2026. Ordered by test wins and how many independent reviewers recommend each app, then by price, with Windows' own clipboard history near the end as the baseline and Pasta last because no independent reviews of it exist yet. Win+V keeps 25 items and forgets them on restart unless you pin them. Ditto fixes both for free and won head-to-head tests at MakeUseOf and XDA. For more, CopyQ adds tabs and scripting, ClipClip adds an editor and OCR, and Pasteboard is the simplest modern-looking choice. Already using Raycast as a launcher? Its clipboard history is free and searchable. Most picks are free at home but not at work: ClipClip charges $49 for commercial use, ClipboardFusion's free version is for personal use only (Pro starts at $19 per PC), and Clipdiary's business license is $29.99. If you also use a Mac, Raycast and CopyQ run on both. Disclosure: Pasta is made by Automatique, the team behind Listicle; it's rated by the same criteria as everything else.

15 itemsupdated 13h ago

Best Clipboard Managers for Mac (2026)

Fifteen clipboard managers for Mac, ranked on hands-on reviews and tests from Zapier, MacStories, Six Colors, Cult of Mac, iMore, 9to5Mac, XDA and PCWorld, from free apps to the history built into macOS Tahoe. Prices checked October 5, 2026. Ordered by test wins and how many independent reviewers recommend each app, then by price, with Spotlight's built-in history near the end as the free baseline and Pasta last because no independent reviews of it exist yet. To choose, try Tahoe's Spotlight history first (Command-Space, then Command-4). If a week of history isn't enough, Maccy is the free step up, Paste is the one to get if your clipboard should follow you to an iPhone, and Pastebot suits people who reshape text as they paste. Already use Raycast, Alfred, LaunchBar or Keyboard Maestro? Try the history it already has before buying anything. Subscription or one payment? Paste costs $29.99 a year, while PastePal ($14.99), ClipBook ($9.99) and CopyClip 2 ($7.99) are one-time buys, and of those three only PastePal syncs with an iPhone. Pastebot sits in between: $39 up front, then $19 for each further year of updates. Disclosure: Pasta is made by Automatique, the team behind Listicle; it's rated by the same criteria as everything else.

15 itemsupdated 13h ago

Best AI Story Writing Apps (2026)

Fifteen AI apps for writing your own fiction, from drafting partners like Sudowrite and Novelcrafter to manuscript critics like ProWritingAid, ranked on what Kindlepreneur, Reedsy, The Write Practice and other writing outlets found in their reviews. Prices checked October 5, 2026. Match the tool to the job. To draft with an AI partner, Sudowrite works out of the box, while Novelcrafter gives you more control if you bring your own AI key. To revise, ProWritingAid catches line-level problems, Fictionary checks scene structure and Marlowe sends a whole-book report. For brainstorming or notes on a full draft, a general chatbot works: Claude for prose, ChatGPT for features, Gemini for very long manuscripts. On cost: Raptor Write is free and you pay only for the AI you use, and Gemini's Plus plan is $4.99 a month. Sudowrite's Hobby plan and ProWritingAid Premium together come to about $20 a month on yearly billing. Most of these meter AI use in credits or tokens, so heavy drafting can cost more than the plan price. Ordered by how many independent reviewers tested or recommended each app, then by how well it suits writing your own fiction, then by price. Disclosure: Fable is made by Automatique, the team behind Listicle; it's rated by the same criteria as everything else.

15 itemsupdated 13h ago

Best Portable Jump Starters (2026): Cheap Ones That Actually Start a Dead Car

A portable jump starter is a battery pack with clamps: it starts a dead car without jumper cables or a second car. These picks come from hands-on tests by Car and Driver, TechGearLab and Project Farm, which jumped dead batteries in V6s, trucks and beaters. Prices checked October 4, 2026. Quick picks: the WOLFBOX MegaVolt24 is the best overall, the GOOLOO GP2000 the best value, and the Strauto the cheapest one that passed a real test. How to choose: match the pack to your engine rather than the peak-amp number on the box, which isn't comparable between brands. The NOCO GB40 is rated for gas engines up to 6 liters, the GOOLOO GP2000 up to 8 and the WOLFBOX up to 10. For a V8 truck or a diesel, get the GOOLOO GT6000. The small lithium packs fit in a glovebox; the Stanley is a heavier lead-acid unit meant for the garage. Keep it charged: a pack that's flat when you need it is no help. Top it up every few months. The GP2000 is rated to hold a charge for up to two years.

8 itemsupdated 17h ago