Spin Up a Hermes Agent: Agentic AI Without Handing Over Your Business
What it takes to run a Hermes agent on your own hardware, why a local model cannot do the job, and why renting the intelligence is the sensible half of the deal.
What it takes to run a Hermes agent on your own hardware, why a local model cannot do the job, and why renting the intelligence is the sensible half of the deal.

What it takes to run an agent on your own hardware, why a local model cannot do the job, and why renting the intelligence is the sensible half of the deal.
The most useful thing I heard about agentic AI this week was a man telling his computer to install a different computer.
A man in Philadelphia is being asked what he does when he needs to think. By trade he runs a media company. He keeps an agent running on a rented machine, and at the beginning of the exchange he confesses to something small and strange: when the guest walked into the room, the host was on the phone. No person was on the other end. He was describing to his agent how he wanted the day's newsletter framed, while walking around his neighbourhood.
Later he says the sentence that stays with you. He had been building the previous setup for four months, moving his company's whole workflow into it. Then somebody told him to try something else. So on a Sunday afternoon he opened his existing agent on Telegram and typed: "Hey, spin up a Hermes agent in your server and let me know when that's finished."
Twenty-four hours of setup later, he typed: "Hermes agent, shut down the open claw agent and we're going to begin working."
That is the whole story, compressed. Not a model that writes better paragraphs. A machine that installs a second machine, on its own, because you asked it to in a text message. The one he switched to is called Hermes Agent, and this essay is about how it runs, what it costs, and where the intelligence behind it should live.

The vast majority of people who use artificial intelligence are using it as a faster search box. In the same conversation, the guest from the lab that builds Hermes Agent makes a distinction worth holding onto. His own mother, he says, loves a popular answer engine and uses it constantly — and she likes it because it is a nicer version of a search box. This is not a criticism of her. That describes where the crowd stands.
The crossing happens at a different point. A chat interface gives you an answer and then waits. An agent gives you a result, and it did the work while you were elsewhere. It read the file, ran the command, checked the output and fixed the mistake it made on the first try. Afterwards it wrote down how it did that, so it can repeat the job next Tuesday. What differs is not the quality of the paragraphs. The difference is that one of them keeps working after you close the tab.
And this is where almost everyone stops, because the honest objection is not laziness. The obstacle is risk. If a chatbot gives a wrong answer, you lose a minute. If an agent with access to your files, your accounts and your machine gives a wrong answer, you lose something that does not come back.
So the interesting question is no longer whether agentic systems are impressive. The real question is what makes them safe enough to deploy. The answer turns out to be structural, and it is not technical. It is not a question of trusting harder.
It is also not the answer most people expect. The version of this that gets discussed online is a choice between running everything on your own hardware and sending everything to somebody else's. That framing is wrong in a way that leads people to spend money badly, and untangling it is most of what follows.
Worth being concrete about the tool at the centre of this, because "agent" has become one of those words that means everything and therefore nothing.
Hermes Agent is a program you install and run yourself. It sits on a machine you control, and you talk to it the way you would talk to a person: by typing. It belongs to the same family as the coding assistants that run in a terminal, and it differs from a chat website in one structural way. It has a system underneath it rather than just a conversation. It can read and write files, run programs, open a browser, keep a schedule, and connect to messaging services so that reaching it does not require sitting at the machine at all.
The demonstration at the top of this essay is what that looks like in practice. The host had told his existing agent to install the new one, and then he stopped paying attention, because there was nothing left to supervise. When he came back the work was done.
Three properties of it matter for everything that follows, and the guest names all three.
It works when you install it. That sounds like faint praise and is not. A great deal of open-source tooling in this field arrives as a collection of parts that assumes you enjoy assembling parts. The design decision here was to keep the core small and opinionated so that the first run produces something useful, and to let everything else be added rather than required. The reason that matters for a reader is simple. The barrier to trying it is an afternoon, rather than a project.
It remembers, and it improves. This is the part that changes how you use it. Most assistants are amnesiac: excellent in a session, blank at the start of the next one, and unpredictable about a task they solved last week. An agent with persistent memory writes down the path it found, so the same job gets easier the second time and can be depended on the tenth. The guest describes the failure this fixes precisely, which is that you cannot build anything on top of a system that works on Monday and cannot repeat itself on Tuesday. Once a job is reliable, other jobs can depend on it, and that is the point at which a tool becomes infrastructure.

It does not care which model you use. This is the one that matters most for the second half of this essay, and it is a deliberate design choice rather than a technical detail. The agent is the harness; the model is a separate part that plugs into it. You can point it at a local model, at a rented open-weight model, at a frontier closed model, or at several of them for different jobs, and switch between them without rebuilding anything.
That last property is the reason the rest of this essay can talk about where the intelligence lives as a choice rather than a commitment.

There is a fourth thing worth mentioning because it is the least discussed and the most revealing about who this was built for. The host does not use a desktop application or a terminal. He talks to his agent through Telegram, in a private group conversation containing only himself and the agent. He uses the group's topic threads to keep separate workstreams apart: one thread for the newsletter, one for business operations, one for a client. He does this by voice while walking around his neighbourhood, dictating instead of typing. Somebody on his team prefers the desktop application and never touches the chat. Both are the same agent. There is no single correct way in, and the tool is built so that there does not have to be.

The host offers the cleanest picture in the whole hour, and it is worth drawing because it makes the rest of this essay simple.
An agentic system, he says, stands on three legs. The first is the second brain, meaning the files. Every note, every dataset, every project, every thing you know about your own work, sitting somewhere the agent can read it. The second is the harness, the agent itself: the thing that has tools, keeps a schedule, remembers what it learned, and can touch the machine. In this case that harness is Hermes Agent, which is the part you install and keep. The third is the model, the intelligence doing the actual thinking.

Here is why the picture earns its keep: two of those three legs are already yours, and they are the two that take years to build. Nobody can sell you your own history.

The third leg — the model — is the one that changes every few months, and it is the one people are most afraid of being locked out of.
That asymmetry is the practical argument for this whole class of tooling. The host puts it bluntly, in a line that reads like a slogan and is actually a statement about portability. Being tied to a single provider is the thing to avoid. The model is the part you want to be able to swap when a better or a cheaper one appears. A harness that lets you change the model without changing everything else is not a feature. This is the thing that makes the investment in the other two legs safe.
Now the uncomfortable part, and the guest does not soften it.
Businesses are sending their internal material to models that may be trained on it. He describes it as a psychological phenomenon, one where people share enormously and are barely aware they are doing it, and then makes the sharper version of the argument. If you are a company of any size, feeding your company's information into a lab's training pipeline is handing over the thing that makes you different. The phrase he uses is secret sauce. He is not being cute. Your pricing logic, your customer list, your internal method, your accumulated judgement. That is not raw material. That is the business.
The host then pushes it one step further and lands on a word that will probably end up in a lot of board rooms: fiduciary. His point is that a director who ships the company's intellectual property to a vendor that is also building competing products may eventually have to explain that decision to shareholders. Whether or not that becomes a legal standard, the direction of the pressure is clear. The cost of sending everything out is not measured on the invoice. The bill arrives later, and it arrives as lost negotiating room.


This is the part of the hour that makes the rest actionable, because it splits a confusion that trips people constantly. There are not two options here. There are several, and they differ in where one specific thing goes: the prompt and everything attached to it.

The guest spends real time on an objection he clearly meets often, and it deserves to be quoted rather than paraphrased.
The open-weight models that made the price collapse were mostly trained outside the United States. People hear that and conclude their data is being shipped abroad. That conclusion is wrong in the specific case that matters. Connecting straight to a provider's own servers in another country is one thing. Downloading the same open weights and running them on machines you control is a completely different thing, because a model is a file. Once it is a file, it runs wherever you put it, under whatever rules apply there.
He says it in one sentence: it is a Chinese model running on American servers. That is a different object from Chinese servers. The file and the jurisdiction travelled separately. Same reason a book is not published in the country where it was printed.
Here the conversation gets concrete, and so does this essay, because "run it yourself" is easy advice to give and hard to price honestly. And the honest version starts with something that the enthusiasm around local models tends to skip.
A model you can hold in a home machine cannot do the job an agent needs done.
That is not a limitation of the software. It is a limitation of the machine. Serving a language model is work on memory bandwidth rather than on arithmetic: every token the model produces requires reading the weights, and the weights are larger than any consumer card's memory. A widely sold twelve-gigabyte card moves memory at around three hundred and sixty gigabytes per second. A card built for this work moves it at three and a third terabytes per second, which is roughly nine times more, and it holds eighty gigabytes rather than twelve.

The consequence is blunt, and it is worth stating before anyone buys hardware on the strength of the previous section. A six-gigabyte model will summarise, sort and answer questions about your documents, and that is genuinely useful. It will also lose the thread halfway through a long chain of tool calls, forget an instruction given twenty steps earlier, and produce a confident answer where it should have said it does not know. An agent lives or dies on exactly those things. Ask a small model to run your workflow unattended and you will get a workflow that looks like it is working for the first ten minutes.
The same arithmetic runs the other way as well, and this is the part that matters. The models good enough to drive an agent are open-weight models. They exist, they are downloadable, and they are the reason the guest can say that most users can do most of their work with open models — the ninety percent figure quoted earlier. Nothing about them is closed. What is closed to you is the machine that can hold one.

So the useful split has nothing to do with local against cloud. It is the agent in one place and the intelligence in another.
And the second half of that split deserves a plain sentence, because it is the part people get wrong in both directions. No ordinary person is going to host a frontier open-weight model at home.

Determination does not help here, and neither does a better card. Serving one needs memory measured in the hundreds of gigabytes and bandwidth measured in terabytes per second. That is a machine in a building with serious power and serious cooling, and pretending otherwise wastes people's money.
What an ordinary person can absolutely do is run the agent and the files on their own hardware, and reach a frontier open-weight model on someone else's rented machine. That is a private server in the sense that matters: yours, under your control, holding your data, talking to a model rather than handing over your business. That is a different act from sending the same material to a closed lab that trains on it and answers to shareholders whose interests are not yours.

Put the agent on hardware you own. That is where the value accumulates: your files, the memory it builds up, the schedules it runs, the record of what it has done for you. None of that needs a graphics card, and all of it is worth keeping. Then rent the intelligence, because intelligence is the one component that is genuinely commoditised, genuinely interchangeable, and genuinely improving every few months.
This is also the arrangement that keeps you out of the trap the guest warns about. Being tied to one provider is the danger. Using one provider's model this month because it is the strongest or the cheapest, while your files and your agent stay yours, is not being tied to anything. The harness that lets you make that swap is doing the actual work.
Now the arithmetic, with the correction that it deserves.
A machine you already own and leave switched on carries no usage bill for the agent itself. If you never buy a graphics card, your marginal cost is electricity.

For the models that can drive an agent, the hardware question answers itself. A sixty-five gigabyte file does not go into a twelve-gigabyte card, and stitching it across system memory puts you back in the territory of seconds per token. The honest conclusion is that the frontier-class open models belong on rented machines, whether that is a metered open-model service or a graphics card you rent by the hour.

Which brings the essay to the service that most people in this ecosystem already have installed. Ollama runs models on your own hardware and charges nothing for doing so, and it also offers hosted inference for open-weight models under the name Ollama Cloud. Its entry tier is free, and the cloud side is priced on credits spent per token rather than on a flat seat. Beyond the free tier the published plans are twenty dollars a month and one hundred dollars a month. Read them on the pricing page itself, because they have been restructured more than once and any number quoted second-hand ages badly.

Between the cheapest and the most expensive model on that chart lies a factor of fifty. This single fact is the entire economic argument of this essay, and it is why the guest says something that would have sounded absurd a year earlier: most users can do most of their work with open models. He puts it at ninety percent of users doing ninety percent of their work. He also names the honest exception, which is that some work genuinely wants the biggest model, and he is careful not to claim the cheap models win everywhere.
What follows from the chart is a habit rather than a purchase: send the routine work to the cheap model and keep the expensive one for the hard ten percent. That habit is only available to you if your setup can change models without rebuilding anything.
It pays to know who is behind the harness in that conversation, because it explains the shape of the thing.
The guest works at a research lab that is still small and describes itself that way. Before it, he was in the Bitcoin industry, working on distributed systems and decentralised compute, which is a background worth noticing. A number of the people now building AI infrastructure came through that world. They arrived with an argument already formed: a person should be able to own the machine that thinks for them, rather than merely rent access to somebody else's.
That is not a marketing position at a lab like this one. It is the reason the harness exists in the form it does. The lab post-trains open models and runs an agent that anyone can install and point at any model, including models made by competitors — the same Hermes Agent named in the exchange at the top of this essay. Its own hosted product is optional, and the people who never touch it are not treated as a failure case. The stated goal is to meet people wherever they are. That is an unusual thing for a company to mean literally. That is why the same tool is meant to work for one person on a laptop and for a company running it across a fleet.
Then there is the question everyone actually asks, which is whether buying the hardware beats renting it. This one is arithmetic, so let us do the arithmetic instead of asserting a conclusion.
Take a used twenty-four-gigabyte card plus a power supply and case at roughly seven hundred and fifty dollars. Such a card pulls about two hundred and fifty watts under load and close to twenty at idle. Run it four hours a day and leave it plugged in the rest of the time, and you burn about forty-two kilowatt hours a month. At a household rate near thirty-five cents that lands around twelve dollars a month in electricity. Now compare that against the published subscription tiers: twenty dollars a month at the lower tier, one hundred at the upper.

Two crossing points, and they say different things.
Against the hundred-dollar tier, the machine pays for itself at month nine. This is not a close call. Inside a year, and after that the hardware is simply yours.

Against the twenty-dollar tier the arithmetic turns round completely. The machine reaches that crossing point in month ninety-four, which is nearly eight years. Anyone who tells you local always wins is not reading their own numbers. Which is the honest answer: on price alone, a twenty-dollar subscription beats a graphics card comfortably, and buying hardware to save that subscription is a hobby with a spreadsheet attached.
But the comparison is not purely financial, and this is the part the spreadsheet misses. The hardware buys you something the subscription does not: the prompt never leaves the room, the model cannot be retired out from under you, and the price cannot change because the vendor changed its mind. Whether that is worth the difference is a judgement about your own work. What matters is that it is now a judgement you can make with numbers instead of a feeling.

Everything above is argument. Here is the part that is instruction, and it is deliberately small, because the failure mode in this field is trying to build a cathedral on the first day.
Start with one job you do every week. Something tedious and specific, rather than a project or a migration: a folder you read through, a report you assemble, a list you keep updating by hand. Write down what "done" looks like for it. If you cannot describe the finished thing in one sentence, the job is too big for a first attempt.
Put the agent somewhere it cannot hurt you. Hermes Agent installs with one command on Linux, macOS or Windows. Give it a container, a small virtual machine, an old laptop. Give it access to that one folder and nothing else. That is not paranoia. What it buys you is the confidence to leave the thing running while you go and do something else, which is the entire point.
Connect it to a model you rent rather than one you host. This is the step people get wrong, and it is the one the pricing chart above is about. Sign up for an open-model service, take the free tier, paste the key into your agent's configuration. You are now running an agent on your own machine against a model that lives on rented hardware. Two files changed. That is the whole setup.
Then let it loose on the job, and read what it produces. Read the actual output rather than the summary it hands you. The first run will be wrong somewhere, and where it is wrong tells you what you forgot to specify. That gap between what you asked for and what you meant is the real work, and it does not go away with a bigger model.
Keep what it learned. Every serious agent writes down how it solved something so it does not have to solve it twice. Give yours a place to keep that, and read it occasionally. The compounding in these systems does not come from the model getting smarter while you sleep. It comes from a solution being written down once and never rediscovered. This is the difference between a tool you use and a system that accumulates.
Only then think about hardware. Run the thing for a month first. If the only limit you hit is the bill, and you have a spare machine, then a graphics card is worth considering for the small routine work. If you have not hit that limit, you do not need it yet, and the card you buy today will be cheaper and worse than the card you buy in a year.
Nothing in that list requires a data centre, a subscription negotiation, or anyone's permission. An afternoon and one tedious job is the whole requirement.

Three honest limits, because the field is full of people who leave them out.
First, a local model is not a frontier model. On genuinely hard reasoning it will lose, and pretending otherwise is how people end up disappointed and back on the closed services within a month.

The correct frame is a division of labour rather than a replacement.
Second, "open weights" does not automatically mean private. Open describes the licence and says nothing about the deployment. If you run open weights on somebody else's servers, you get the licence benefit and none of the privacy benefit. The guest in that conversation is explicit about this trap. An open model says nothing about who is holding the machine it runs on. Read where the compute sits, because the licence alone tells you very little.
Third, this is not free of maintenance. A machine in your room needs updating, and occasionally needs you to fix it on a Sunday. The subscription does not. That time is real and it belongs in the calculation, which is exactly why the eight-year figure above deserves to be read as a warning rather than a triumph.

One more thing in that conversation is worth carrying out of the room. It answers the obvious question of why anyone at a lab would want the weights to stay in the open.
The guest's position is that if you genuinely believed this technology could go catastrophically wrong, the concentration of compute behind two or three companies would be the alarming fact rather than the existence of freely downloadable models. His argument is compact: a model derived from another model cannot surpass the model it was derived from, so open models cannot be the frontier. He then says the thing that a lot of people are thinking and few say on the record, which is that the loudest argument against open weights comes from the parties with the strongest commercial interest in a closed door.
You do not have to accept that framing. The practical conclusion underneath it is the one this essay has been circling, and after everything above it can be stated precisely. Models are becoming a commodity, which is why renting them is sensible. The harness is becoming a commodity, which is why you should not pay a premium for it. The part that is not a commodity is the pile of your own work that you have made legible to a machine, and that part only grows where you keep it.
That asset is worth more every year. It is the one part of the stack nobody can rent to you, and it is the one part that only gets built if you start.
So buy the machine that runs the agent rather than the machine that tries to be the model. Keep the files. Rent the intelligence from an open-model service, and change your mind about which model you rent whenever a better or a cheaper one appears.
That is the arrangement worth arguing for. Not because it is pure, and not because it is free of compromise. It is worth arguing for because this is the version of the technology an ordinary person can adopt today, on hardware they already own, without handing their working life to a company they will never meet. It costs less than most people assume. It is available right now. And it leaves you holding the part that matters, which is everything you have written down.
The door is open right now. Which is a good reason to walk through it while that is still true.
The conversation quoted throughout is the TFTC interview with Tommy Eastman of Nous Research, published 28 September 2026: https://youtu.be/1TYSWqf36gE.
Subscription tiers and per-million-token rates are read from the Ollama pricing page. Model file sizes come from the Ollama library. Both were read on 30 September 2026.
The bandwidth figures are the manufacturers' own specifications. The electricity and break-even figures are my own arithmetic, and the method is stated in the text so you can check it.