top of page
Search

Small Language Models are a gift for socially and environmentally responsible AI

  • mickbrawn
  • Jul 28
  • 3 min read

Small Language Models (SMLs) can run without data centres. They provide flexibility, keep your data on your PC, overcome internet latency and enable good governance.


In this blog I focus on the opportunity to reduce data centre dependence when using AI.


SLMs are designed to run on the modest hardware requirements of AI laptops and offer a viable alternative to the voracious energy and water demands of data centre AI.


A wide range of SLMs are available to choose from. SLMs cover diverse topics with good enough quality that an AI PC can now be a rational investment. The SLM market is mature, fast moving, and no longer niche.


2025–2026 produced dozens of strong models across the 1B–14B parameter range. More parameters = more capable AI, covering general chat, reasoning, coding, multilingual tasks, and even multimodal input - but more parameters also increase memory, processing and storage requirements.


• Modern 2B–4B models now deliver useful chat, summarisation, and tool use behaviour, with some (Gemma 4 e2b) even supporting image input at tiny sizes

• The strongest 3–4B models (Phi 4 mini, Qwen3 3B, SmolLM3, Llama 3.2 3B) reach MMLU scores in the mid-60s to 70, approaching GPT 3.5 level quality on consumer hardware

• Sub 4B models run comfortably on 4–8 GB RAM, delivering 30–70 tokens/sec on a CPU which is fast enough for real time chat


This is a dramatic shift from 2023–2024, when anything under 7B struggled with following basic instructions.


• A modern laptop with 8–16 GB RAM can run 3B–8B models locally with good performance. Qwen3 (1.7B–14B), Gemma 3 (1B–12B), and Phi 4 mini (3.8B) are explicitly designed for this tier

• Even low-end PCs with no GPU can run capable models like Qwen3 1.7B or Phi 4 mini at 15–40 tokens/sec

• With a mid-range GPU (e.g., 12–24 GB VRAM), you can run Phi 4 14B, which competes with much larger models on reasoning and coding


So, you no longer need cloud inference for everyday tasks as local models are fast, cheap, private, and increasingly powerful. There are strong models for:


• General chat & reasoning: Phi 4 mini, Llama 3.2 3B

• Coding: Qwen3 3B, Granite 3.x

• Multilingual tasks: Qwen3 family

• Enterprise & compliance: IBM Granite 3.x

• Mobile/embedded: Qwen3 1.7B, Gemma 4 e2b

• Multimodal (text + image): Gemma 4 e2b/e4b


This breadth means an AI PC can run multiple specialised models locally depending on the task. That is something cloud only setups cannot do without switching providers or paying for multiple APIs.


If your goal is to run AI models locally, privately, and without cloud dependency, the investment in an AI PC may now make sense.


If you want to use AI in an environmentally and socially responsible way, local models give you predictable behaviour, zero cost inference, and full control over your workflows.


Give me a call if you would value some guidance on how to select one or more SLMs for your specific needs, how to download them and the appropriate AI Runtime Framework, and how to start reducing your dependency on AI data centres.


Mick at Michael Brawn Consulting (MBC)

📱 0414 987 129


* MMLU stands for the Massive Multitask Language Understanding benchmark test used to measure how well language models can apply general knowledge and reasoning across a wide range of subjects.

 
 
 

Comments


bottom of page