Writing
Free AI models for Hermes Agent: options and limits
updated 2026-09-17
Hermes Agent can use free hosted model endpoints or models running on your own computer. For a no-cost cloud trial, check the named OpenRouter free endpoints for Laguna S 2.1, Laguna XS 2.1, and Tencent HY3. For local use, choose a model that fits your hardware and supports tools. Neither route promises unlimited capacity.
Check the price of the exact endpoint you will use. A model can have free weights, a paid API, and a temporary promotion at the same time.

What is in this guide?
- Current model options
- What free means
- Choose a model
- Connect OpenRouter
- Run locally
- Privacy and rate limits
- Test before you depend on it
- Frequently asked questions
For installation, Telegram, memory, and scheduling, start with the Hermes Agent field guide.
Which models are free right now?
This table records the listings reviewed on September 17, 2026. A listed free endpoint is not proof that it has spare capacity for your account. Prices are USD per million input and output tokens where shown. Follow each link before a paid run.
| Model and exact endpoint | Listed access | Context shown | Main consideration |
|---|---|---|---|
Laguna S 2.1, poolside/laguna-s-2.1:free | Free token pricing | 262K | Coding focus; free inputs and outputs may train models |
Laguna XS 2.1, poolside/laguna-xs-2.1:free | Free token pricing | 262K | Smaller coding model; same training disclosure |
Tencent HY3, tencent/hy3:free | Free token pricing | 262K | Check provider capacity and data terms before use |
Step 3.7 Flash, stepfun/step-3.7-flash | Paid: $0.16 input / $0.92 output, displayed discount | 262K | Multimodal model; open weights do not make this API free |
Solar Pro 4, upstage/solar-pro4 | Paid: $0.09 input / $0.36 output, displayed discount | 524K | Earlier launch offers are not current access guarantees |
Fugu Ultra, sakana/fugu-ultra | Paid: $5 input / $30 output | 1M | Orchestration tokens are billed too |
The old version of this guide grouped all six as free choices. That was misleading. The current table distinguishes confirmed free listings from paid endpoints. A promotion on another provider must be checked on that provider's account and pricing page.
What does free actually mean?
| Meaning | What you receive | What can still cost money |
|---|---|---|
| Free hosted endpoint | No listed input or output token charge | Tools, hosting, or a paid fallback |
| Trial or promotional credit | A limited allowance | Usage after the allowance or promotion ends |
| Open weights | Files you can run under a model license | Hardware, electricity, storage, and operations |
| Included subscription access | Usage covered by a plan's terms | The subscription, overages, or excluded tools |
Hermes itself is open source. A free application does not make every connected service free. Official Hermes overview.
Keep a strict test budget. If the goal is zero model charges, check the main model, auxiliary models, and fallback settings. A free main model is only one part of the configuration.
Which model should you try first?
These are starting points for a trial, not claims that one model wins every task.
For hosted coding: Laguna S 2.1
Laguna S 2.1 targets software engineering and agentic coding. Its free endpoint supports tool calling. The listing does not promise enforced JSON output through response_format. Treat those as separate capabilities. Laguna S listing.
Try it on a small bug fix with a known failing test. Check the diff and the test result. Do not judge an agent model only by how confident its final answer sounds.
For a smaller coding model: Laguna XS 2.1
Laguna XS 2.1 has 33 billion total parameters with about 3 billion active per token. A free cloud endpoint and local distributions are available. Its model page warns that free inputs and outputs may be used for training. Laguna XS listing.
For local use, the Ollama model page describes a Mac with 36 GB of RAM as a supported example. That is not a promise for every quantization or context setting. Leave memory for the operating system and the agent's tools.
For a general agent trial: Tencent HY3
The reviewed HY3 listing offers a 262K context window and configurable reasoning. It describes the model as suitable for coding, documents, and tool workflows. Those are vendor positioning claims, not a substitute for your own task tests. HY3 listing.
Use non-sensitive examples while you confirm the endpoint, current provider, and data policy. The retrieved page did not establish a blanket no-training promise.
For images and video: check Step 3.7 Flash's paid endpoint
Step 3.7 Flash combines a language model with a vision encoder. Its release materials describe image and video understanding. Check that the endpoint and client support the input format you need. StepFun release.
The OpenRouter endpoint in the table is paid. The :free URL could not be verified in this review, so it is not recommended here as a confirmed free service. Downloadable weights are available through the official project, but local hosting is a separate hardware project.
What about Solar Pro 4 and Fugu Ultra?
Solar Pro 4's current listing targets long-document and agent work. Its displayed price is paid, even with a discount. Do not reuse the earlier guide's expired “90% off through September 10” offer. Solar Pro 4 listing.
Fugu Ultra coordinates multiple models behind an endpoint. Its listing says orchestration tokens count as standard input and output tokens. A short visible answer can therefore cost more than its length suggests. Use an explicit budget and inspect usage. Fugu Ultra listing.
How do you connect a free OpenRouter model?
- Create an OpenRouter account and API key.
- Run
hermes model. - Select OpenRouter and enter the key through the setup flow.
- Choose the exact named free endpoint.
- Run a small test, then inspect provider usage.
After credentials are configured, this example selects a named free model for one chat:
hermes chat \
--provider openrouter \
--model 'poolside/laguna-s-2.1:free'
The model ID is poolside/laguna-s-2.1:free. The provider is openrouter. Do not turn the ID into openrouter/poolside/laguna-s-2.1:free. The official OpenRouter integration guide explains the separation and recommends checking context support.
Use a harmless first task:
Read README.md in this test folder.
Give me a five-line summary. Do not change any files.
Test tool use by giving the agent a disposable file to edit and a command that checks the result.
Should you use openrouter/free instead?
openrouter/free chooses a free model at random from eligible models. It filters for requested features such as tools or image input. That is useful for exploration, but the selected model can change. Free Models Router.
For a repeatable comparison, use a named endpoint and record which provider served it. My preference is to pin the model while testing, then add routing only after I understand the baseline.
How do you run a model locally?
Install Ollama and choose a model that fits your computer. For Laguna XS 2.1, the published library command is:
ollama run laguna-xs-2.1
This can download a large file. Check the model size and license before starting. Ollama model library.
Confirm the model responds in Ollama. Then select Custom Endpoint in Hermes:
| Setting | Value |
|---|---|
| Base URL | http://localhost:11434/v1 |
| Model | The exact name installed in Ollama |
| API key | Empty or the placeholder accepted by setup |
The Hermes Ollama guide covers this connection. The current local-model guide also documents Hermes-managed local options.
Check the server's configured context window. A model advertised as 256K does not mean your local runtime loaded it with 256K. Larger contexts need more memory and can slow the first response. ollama ps shows where a loaded model is running. Ollama FAQ.
For mixture-of-experts models, the active parameter count describes part of the computation. It does not mean the full model occupies the memory of a dense model with that many parameters. Weights, quantization, attention state, and runtime overhead all matter. Do not buy hardware from the active-parameter figure alone.
What are the privacy and rate-limit tradeoffs?
Free does not mean unlimited
An agent can make several model requests for one user task. A daily request allowance can disappear quickly during tool loops, retries, or parallel work.
OpenRouter documents account-level free-model limits and an API endpoint for inspecting the key's remaining allowance. Multiple keys do not create extra capacity. On rate-limit errors, respect retry guidance and use backoff. OpenRouter limits.
Do not silently switch to a paid model after a free one fails. Decide in advance whether the fallback may spend money, and set its limit.
Paid does not automatically mean private
Read both the platform policy and the policy for the provider that serves your request. Training, logging, retention, and deletion are different questions. OpenRouter documents its own handling and provider controls separately. OpenRouter data collection.
Poolside's two free listings explicitly disclose possible training use. This makes them a poor default for confidential client material unless that use is approved. Do not extend that exact disclosure to every free model without checking its terms.
Local inference is only one part of a local workflow
Ollama states that it does not receive prompts sent to a local model. It also supports disabling its cloud features. But Hermes can still send data through search, messaging, remote tools, or another model. Review the whole workflow. Ollama privacy and cloud settings.
How should you compare models?
Use the same task, files, tools, and acceptance checks. Start each run from the same state. Test several times because one result can be luck.
| Test | What to measure |
|---|---|
| Read and summarize a document | Correct facts and links to the source |
| Fix a small bug | Passing tests and a limited diff |
| Use a tool | Correct arguments and correct interpretation of the result |
| Handle missing information | Admits the gap instead of inventing an answer |
| Follow a permission limit | Stops before an unapproved write or send |
| Complete a longer task | Finished result, total requests, retries, and elapsed time |
Track cost per accepted result, not just cost per token. A model that needs five corrections can cost more time than a model with a higher token price.
Published benchmarks can help you pick candidates. They do not replace this test. Compare the same benchmark version, agent setup, tool budget, and scoring method. Terminal-Bench 2.0 and 2.1 are not interchangeable. Neither are SWE-bench Verified, SWE-bench Pro, and DeepSWE.
The earlier article mixed those tests in one ranking. This revision removes that ranking rather than present an invalid comparison. It also removes unsupported claims that a paid model makes production work safe. Tests, review, permissions, and rollback still matter at every price.
Frequently asked questions
What is a good free model to try with Hermes Agent?
For non-sensitive coding experiments, try a named Laguna S 2.1 or Laguna XS 2.1 free endpoint. HY3 is another listed option. Check live availability, data terms, and your task results before choosing a default.
Is Step 3.7 Flash free?
Open weights are available, but the OpenRouter endpoint reviewed here is paid. A separate free promotion may exist on another provider. Confirm it there before starting a run.
Is Solar Pro 4 still free on Hermes?
The reviewed OpenRouter endpoint is paid. A model appearing in the Hermes selector does not establish free access. Check the provider and account attached to that selection.
Can I use free models for private client work?
Only when the provider's terms and your client's requirements permit it. Some free endpoints disclose training use. Local inference can help, but external tools and services must also be checked.
Does a local model use tokens?
Yes. It still processes tokens. Local inference avoids a cloud provider's per-token charge; it does not remove computation, hardware, or electricity costs.
Can I guarantee a free model will remain free?
No. Free capacity and promotions can change. Keep a tested alternative and make any paid fallback an explicit choice.
Is a bigger context window always better?
No. It can help with longer inputs, but it can also increase memory use, latency, and cost. Accurate retrieval and a smaller relevant input may work better.
Keep the setup simple
Test one useful task on a single endpoint, then check the bill and data policy before deciding whether to keep it.
Continue with the Hermes Agent field guide, or browse the AI resource directory for current provider and model links.