Life = Content

What are you working on?

Send me the rough version. We can figure out the next step together.

Book a consultation $250 · 90 minutes

contact@dylanjharris.com

Open Gmail

Or say hi on X · @dylanjharris
More ways to get in touch

Writing

Free AI models for Hermes Agent: options and limits

updated 2026-09-17

Hermes Agent can use free hosted model endpoints or models running on your own computer. For a no-cost cloud trial, check the named OpenRouter free endpoints for Laguna S 2.1, Laguna XS 2.1, and Tencent HY3. For local use, choose a model that fits your hardware and supports tools. Neither route promises unlimited capacity.

Check the price of the exact endpoint you will use. A model can have free weights, a paid API, and a temporary promotion at the same time.

Free AI models for Hermes Agent: compare hosted access, local hardware, and the limits behind a zero-dollar price.
Check the endpoint, data policy, and task result before choosing a model.

What is in this guide?

For installation, Telegram, memory, and scheduling, start with the Hermes Agent field guide.

Which models are free right now?

This table records the listings reviewed on September 17, 2026. A listed free endpoint is not proof that it has spare capacity for your account. Prices are USD per million input and output tokens where shown. Follow each link before a paid run.

Model and exact endpointListed accessContext shownMain consideration
Laguna S 2.1, poolside/laguna-s-2.1:freeFree token pricing262KCoding focus; free inputs and outputs may train models
Laguna XS 2.1, poolside/laguna-xs-2.1:freeFree token pricing262KSmaller coding model; same training disclosure
Tencent HY3, tencent/hy3:freeFree token pricing262KCheck provider capacity and data terms before use
Step 3.7 Flash, stepfun/step-3.7-flashPaid: $0.16 input / $0.92 output, displayed discount262KMultimodal model; open weights do not make this API free
Solar Pro 4, upstage/solar-pro4Paid: $0.09 input / $0.36 output, displayed discount524KEarlier launch offers are not current access guarantees
Fugu Ultra, sakana/fugu-ultraPaid: $5 input / $30 output1MOrchestration tokens are billed too

The old version of this guide grouped all six as free choices. That was misleading. The current table distinguishes confirmed free listings from paid endpoints. A promotion on another provider must be checked on that provider's account and pricing page.

What does free actually mean?

MeaningWhat you receiveWhat can still cost money
Free hosted endpointNo listed input or output token chargeTools, hosting, or a paid fallback
Trial or promotional creditA limited allowanceUsage after the allowance or promotion ends
Open weightsFiles you can run under a model licenseHardware, electricity, storage, and operations
Included subscription accessUsage covered by a plan's termsThe subscription, overages, or excluded tools

Hermes itself is open source. A free application does not make every connected service free. Official Hermes overview.

Keep a strict test budget. If the goal is zero model charges, check the main model, auxiliary models, and fallback settings. A free main model is only one part of the configuration.

Which model should you try first?

These are starting points for a trial, not claims that one model wins every task.

For hosted coding: Laguna S 2.1

Laguna S 2.1 targets software engineering and agentic coding. Its free endpoint supports tool calling. The listing does not promise enforced JSON output through response_format. Treat those as separate capabilities. Laguna S listing.

Try it on a small bug fix with a known failing test. Check the diff and the test result. Do not judge an agent model only by how confident its final answer sounds.

For a smaller coding model: Laguna XS 2.1

Laguna XS 2.1 has 33 billion total parameters with about 3 billion active per token. A free cloud endpoint and local distributions are available. Its model page warns that free inputs and outputs may be used for training. Laguna XS listing.

For local use, the Ollama model page describes a Mac with 36 GB of RAM as a supported example. That is not a promise for every quantization or context setting. Leave memory for the operating system and the agent's tools.

For a general agent trial: Tencent HY3

The reviewed HY3 listing offers a 262K context window and configurable reasoning. It describes the model as suitable for coding, documents, and tool workflows. Those are vendor positioning claims, not a substitute for your own task tests. HY3 listing.

Use non-sensitive examples while you confirm the endpoint, current provider, and data policy. The retrieved page did not establish a blanket no-training promise.

For images and video: check Step 3.7 Flash's paid endpoint

Step 3.7 Flash combines a language model with a vision encoder. Its release materials describe image and video understanding. Check that the endpoint and client support the input format you need. StepFun release.

The OpenRouter endpoint in the table is paid. The :free URL could not be verified in this review, so it is not recommended here as a confirmed free service. Downloadable weights are available through the official project, but local hosting is a separate hardware project.

What about Solar Pro 4 and Fugu Ultra?

Solar Pro 4's current listing targets long-document and agent work. Its displayed price is paid, even with a discount. Do not reuse the earlier guide's expired “90% off through September 10” offer. Solar Pro 4 listing.

Fugu Ultra coordinates multiple models behind an endpoint. Its listing says orchestration tokens count as standard input and output tokens. A short visible answer can therefore cost more than its length suggests. Use an explicit budget and inspect usage. Fugu Ultra listing.

How do you connect a free OpenRouter model?

  1. Create an OpenRouter account and API key.
  2. Run hermes model.
  3. Select OpenRouter and enter the key through the setup flow.
  4. Choose the exact named free endpoint.
  5. Run a small test, then inspect provider usage.

After credentials are configured, this example selects a named free model for one chat:

hermes chat \
  --provider openrouter \
  --model 'poolside/laguna-s-2.1:free'

The model ID is poolside/laguna-s-2.1:free. The provider is openrouter. Do not turn the ID into openrouter/poolside/laguna-s-2.1:free. The official OpenRouter integration guide explains the separation and recommends checking context support.

Use a harmless first task:

Read README.md in this test folder.
Give me a five-line summary. Do not change any files.

Test tool use by giving the agent a disposable file to edit and a command that checks the result.

Should you use openrouter/free instead?

openrouter/free chooses a free model at random from eligible models. It filters for requested features such as tools or image input. That is useful for exploration, but the selected model can change. Free Models Router.

For a repeatable comparison, use a named endpoint and record which provider served it. My preference is to pin the model while testing, then add routing only after I understand the baseline.

How do you run a model locally?

Install Ollama and choose a model that fits your computer. For Laguna XS 2.1, the published library command is:

ollama run laguna-xs-2.1

This can download a large file. Check the model size and license before starting. Ollama model library.

Confirm the model responds in Ollama. Then select Custom Endpoint in Hermes:

SettingValue
Base URLhttp://localhost:11434/v1
ModelThe exact name installed in Ollama
API keyEmpty or the placeholder accepted by setup

The Hermes Ollama guide covers this connection. The current local-model guide also documents Hermes-managed local options.

Check the server's configured context window. A model advertised as 256K does not mean your local runtime loaded it with 256K. Larger contexts need more memory and can slow the first response. ollama ps shows where a loaded model is running. Ollama FAQ.

For mixture-of-experts models, the active parameter count describes part of the computation. It does not mean the full model occupies the memory of a dense model with that many parameters. Weights, quantization, attention state, and runtime overhead all matter. Do not buy hardware from the active-parameter figure alone.

What are the privacy and rate-limit tradeoffs?

Free does not mean unlimited

An agent can make several model requests for one user task. A daily request allowance can disappear quickly during tool loops, retries, or parallel work.

OpenRouter documents account-level free-model limits and an API endpoint for inspecting the key's remaining allowance. Multiple keys do not create extra capacity. On rate-limit errors, respect retry guidance and use backoff. OpenRouter limits.

Do not silently switch to a paid model after a free one fails. Decide in advance whether the fallback may spend money, and set its limit.

Read both the platform policy and the policy for the provider that serves your request. Training, logging, retention, and deletion are different questions. OpenRouter documents its own handling and provider controls separately. OpenRouter data collection.

Poolside's two free listings explicitly disclose possible training use. This makes them a poor default for confidential client material unless that use is approved. Do not extend that exact disclosure to every free model without checking its terms.

Local inference is only one part of a local workflow

Ollama states that it does not receive prompts sent to a local model. It also supports disabling its cloud features. But Hermes can still send data through search, messaging, remote tools, or another model. Review the whole workflow. Ollama privacy and cloud settings.

How should you compare models?

Use the same task, files, tools, and acceptance checks. Start each run from the same state. Test several times because one result can be luck.

TestWhat to measure
Read and summarize a documentCorrect facts and links to the source
Fix a small bugPassing tests and a limited diff
Use a toolCorrect arguments and correct interpretation of the result
Handle missing informationAdmits the gap instead of inventing an answer
Follow a permission limitStops before an unapproved write or send
Complete a longer taskFinished result, total requests, retries, and elapsed time

Track cost per accepted result, not just cost per token. A model that needs five corrections can cost more time than a model with a higher token price.

Published benchmarks can help you pick candidates. They do not replace this test. Compare the same benchmark version, agent setup, tool budget, and scoring method. Terminal-Bench 2.0 and 2.1 are not interchangeable. Neither are SWE-bench Verified, SWE-bench Pro, and DeepSWE.

The earlier article mixed those tests in one ranking. This revision removes that ranking rather than present an invalid comparison. It also removes unsupported claims that a paid model makes production work safe. Tests, review, permissions, and rollback still matter at every price.

Frequently asked questions

What is a good free model to try with Hermes Agent?

For non-sensitive coding experiments, try a named Laguna S 2.1 or Laguna XS 2.1 free endpoint. HY3 is another listed option. Check live availability, data terms, and your task results before choosing a default.

Is Step 3.7 Flash free?

Open weights are available, but the OpenRouter endpoint reviewed here is paid. A separate free promotion may exist on another provider. Confirm it there before starting a run.

Is Solar Pro 4 still free on Hermes?

The reviewed OpenRouter endpoint is paid. A model appearing in the Hermes selector does not establish free access. Check the provider and account attached to that selection.

Can I use free models for private client work?

Only when the provider's terms and your client's requirements permit it. Some free endpoints disclose training use. Local inference can help, but external tools and services must also be checked.

Does a local model use tokens?

Yes. It still processes tokens. Local inference avoids a cloud provider's per-token charge; it does not remove computation, hardware, or electricity costs.

Can I guarantee a free model will remain free?

No. Free capacity and promotions can change. Keep a tested alternative and make any paid fallback an explicit choice.

Is a bigger context window always better?

No. It can help with longer inputs, but it can also increase memory use, latency, and cost. Accurate retrieval and a smaller relevant input may work better.

Keep the setup simple

Test one useful task on a single endpoint, then check the bill and data policy before deciding whether to keep it.

Continue with the Hermes Agent field guide, or browse the AI resource directory for current provider and model links.

← back to writing