The model that cleans up my drafts every morning is called GLM 5.3 Flash, and it is open source. Before you assume what that means, two facts about my setup. I pay real money every month to use it. And it does not run on my computer, because no machine I own is big enough. I rent it by the token from a provider, same as the closed models.
So this is not a story about AI becoming free. It is a story about control, not price. Anybody can download the files. Anybody with serious hardware can host them and sell you access. Two years ago open models were the consolation prize, noticeably dumber than the paid ones, useless for real work. That ended this past year. If you run a shop or a trade, nobody thought to tell you.
Open source stopped meaning second best
"Frontier" is what people call the best AI models in the world at a given moment, whatever sits at the front of the pack. For years the front of the pack was private property. You rented it from OpenAI, Anthropic, or Google. The decent chat tiers still run twenty bucks a month per person, and the open versions trailing behind were a generation dumber. If the work mattered, you rented from whoever sat at the front.
Pull up an independent scoreboard today. Artificial Analysis runs every major model through the same battery of tests, and as I write this, an open model called Kimi K3 sits at number 3 out of 114 models, ahead of almost everything the closed labs charge for. Open source, in plain terms, means the model files get published. Anyone can download them and, with the right machines, run them as their own. Nobody can take them back.
The ranking is not the business story, though. The business story is that open models crossed from "impressive for a free thing" to plain good enough. Reading the morning inbox and drafting replies that sound like you wrote them. Matching receipts to invoices. Most business work is not a frontier problem. It is Tuesday. A model a notch below the best one in the world handles Tuesday fine, and the open ones are there now.
Three open models worth knowing by name
Dozens of open models exist now, and most will never matter to you. These three do. Each comes with a note about where your data goes, because that part gets left off the marketing.
GLM 5.3 Flash, the one I work with all day
Z.ai makes it, and it is published under the MIT license, which is lawyer-speak for take it and use it in your business without paying for the files themselves. It reads text and pictures equally well, so a photo of a handwritten invoice is no problem. It can hold around two thousand pages in its head at once, which matters when the job is "read these ninety receipts." On the independent scoreboard it sits in the top handful of open models, and Z.ai claims it runs close to Anthropic's flagship on coding and agent work. Take maker claims with salt. From daily use, it does regular business work without embarrassing itself.
Here is where I have to be straight with you, because most articles skip it. The model is enormous, hundreds of gigabytes of numbers. "Free to download" is technically true and practically useless on any machine you or I will ever own. What happens in practice is I rent it from a provider at roughly $0.15 per million tokens of reading and $0.50 per million of writing. Far less per unit of work than the flagships charge, and still a line on my card that I notice at the end of the month. Across GLM and the other models I rent, it adds up to real money.
One more thing. Z.ai is a Chinese company. Run the model through their own service and your words travel to their servers. Not automatically a problem, but it should be a decision and not an accident. You do not have to run it there, and the fix comes later in this post.
DeepSeek V4 Flash, the volume one
DeepSeek built this one for throughput. Thousands of small jobs thrown at it all day, answered fast. It scores well above the average model on the independent tests while costing four and eight cents per million tokens. When a task is high volume and medium difficulty, this is the one I reach for.
Same note, same fix. DeepSeek is also a Chinese company, so their first-party service puts your data on their machines. A US provider that keeps nothing makes that disappear. More below.
Qwen 3.8 27B, the one that can live in your office
Alibaba's Qwen family makes the best small open models right now, and the 27B is small in the useful sense. On paper it runs on about 17 gigabytes of memory. It is free for commercial use, and it handles images and reasoning. In practice, plan on more machine than the paper suggests. From my own tinkering, you want at least 48 gigs of RAM before it runs well enough to work with. That rules out the laptop you bought two years ago, and probably the desktop too.
This is the paragraph I care most about. The models I rely on all day are far too big for any desktop, mine included. The dream for open source is a model that matters running on a decent machine instead of a high end one. We are not there yet. Today the local model earns its place for one reason above all: the data note. There is none. If the model runs on a machine in your office, your client's files never touch anybody's network. If you hold paperwork people get sued over, that can be the whole argument.
About that cup of coffee
You may have seen the posts: a whole month of AI agent work now costs less than a cup of coffee. I run agents every day, and I can tell you whoever writes that has never paid an agent bill. It is bullshit. Between GLM and the other models I use, open and closed, I send AI providers a real chunk of money every month. Any business that puts agents to work daily will do the same. There is no version of this where the work runs on pocket change.
Here are the real numbers. AI gets billed in tokens, little pieces of words, and a million tokens works out to roughly 1,600 printed pages. Anthropic's Opus models, the flagship tier everyone respects, run $5 per million tokens of reading and $25 per million of writing. GLM 5.3 Flash lists at about $0.15 and $0.50. DeepSeek sits in cents. Same stack of paper, very different invoices. Prices move weekly in this world, so read all of it as snapshots.
And yet the bill was never the interesting part. Because the files are public, the same model can be hosted by a crowd of competing companies: OpenRouter, Ollama, the makers themselves, and a growing list of others. If your provider gets expensive or careless with your data, you take the same model somewhere else and the answers come out identical. That was impossible under the old arrangement, where serious intelligence came from one or two landlords and raising the rent was your problem to swallow. The prices matter less than the plumbing underneath: nobody can hold your access hostage anymore. And notice, nobody on this menu asks how many employees you have, the way the chat subscriptions all did.
The treadmill never stops
One more wrinkle to plan around. The moment the open models get close, the big labs push the frontier out again. The model that was impressively smart six months ago becomes this year's budget option, and you are left with a shitty little decision that never resolves. Build your workflows on the old model, the one that already does your work fine? Or move everything to the new release so you do not fall behind?
After living through a few of these handoffs, my rule is pick per job, not per leaderboard. The Tuesday work stays on the model that already handles Tuesday. The genuinely hard problems get the new model, and the invoice shows why. What I try not to do is rebuild everything every time a scoreboard shuffles, because that is a month of moving work that was already done.
Where your data goes is still the real decision
Open source and private are two different switches, and flipping the first one does nothing to the second. A free download says nothing about whose computer runs the model for you. If that computer belongs to the model's maker in another country, your business words flow there every time the agent works.
There are two clean ways to handle it, and both are boring in the best way.
Option one: run the open models through a US-based company with a written no-keeping policy. The example right now is Ollama, an American company under California law. Its privacy policy says your prompts and answers are processed only long enough to respond. They are never stored, and never used to train anything. And where Ollama routes work to hosting partners for the biggest models, it passes the same terms on to them, no logging and zero data retention included. There is a free tier to kick the tires, and the paid plan buys usage instead of seats. For most small business busywork, this is the right answer.
Option two: for the stuff that must never leak, skip other people's computers entirely. Put Qwen 3.8 27B on a machine in your office and the data physically cannot go upstream. A no-retention policy is a promise, and promises depend on every company in the chain keeping them, this quarter and every quarter after. A machine in your own office needs no promise. Keep the fine print from earlier in mind, though. That means 48 gigs of RAM and the little sibling model, and you double as the IT department.
What to actually hand an agent
The stuff that works today on these models. Triage the inbox every morning and have draft replies waiting for your yes or no. Match receipts and invoices into the bookkeeping. Draft first-pass quotes from the request email. Answer the reviews with something better than "thanks for your feedback." I used to call AI a fast intern who never sleeps. I still do, and like a real intern, it needs supervision.
What you keep is the judgment, the kind of call that decides your quarter. Whether to fire a client, whether the proposal in front of you is honest. Nobody's model is ready for that, open or closed, and anyone telling you otherwise is selling.
So the setup most businesses land on is boring. Open models through a no-logs US provider handle the everyday volume. A paid flagship account covers the rare hard problem, because the hard problem costs less than the time you would waste wrestling a weaker model. Your own hardware waits for the files that genuinely cannot leave the building. About that last one: pointing an agent at a no-logs provider is a settings change, but running your own machine is a hobby right up until it becomes your job, updates and babysitting included. I wrote last month about owning your software instead of renting all of it. This post is the layer underneath that one, about who does the thinking and where it happens.
No, the brains did not get free. What happened is quieter. Open models got good enough for the regular work that fills your week, and because their files are public, the same model now runs behind a crowd of competing providers. Nobody can lock you in anymore. Serious open intelligence on a decent machine in your back office is still a hope. Nobody can give you a date for it. When it happens, I will write that post. Until then, what matters is what you hand over and where the data lives while it works. If sorting that out sounds like a chore you would rather delegate, well, that one is mine.