Claude Code as an Agentic Client for Local Ollama Models

A Secure Hybrid AI Engineering Sandbox

AI

Jack Jalali

9/27/20267 min read

There is an interesting middle ground between using a completely cloud-hosted AI coding assistant and building an entirely bespoke local AI agent.

In my test/sandbox environment, I wanted to explore that middle ground.

The objective was straightforward:

Use Claude Code on Windows Server 2022 as the agentic development interface, while performing the actual LLM inference on an open-weight model running under Ollama on a separate Linux sandbox server. At the same time having the option of changing to a frontier model for more completed tasks or when preliminary work has been done by the opensource model.

The resulting architecture looks like this:

The distinction is important.

In option 1, Claude Code is being used as the agentic client, but a Claude model is not performing the inference. In option 1, Claude Code handles the interactive coding experience, tool invocation and agent workflow. The LLM requests are redirected to Ollama, where an open-weight model performs the reasoning and generation locally.

I tested this architecture using:

  • Windows Server 2022

  • Claude code

  • Linux based Ollama server

  • gpt-oss:20b model

  • NVIDIA GPU acceleration

  • a private network between the Windows and Linux systems

The model successfully performed inference, invoked shell commands, inspected the Windows operating system, created files, read them back, verified their contents and deleted them.

This article explains why this architecture is useful, how it was built, how it was verified that inference really was happening on the Linux server, and why I think this is a particularly compelling architecture for secure AI-assisted infrastructure engineering.

Why Build It This Way?

AI-assisted development normally presents an architectural choice.

At one extreme: Developer --- > Cloud AI service

At the other end: Developer --- > Local model

Both approaches have advantages.

Cloud frontier models provide excellent reasoning and coding capabilities. Local models provide significantly greater control over where prompts, source code and infrastructure information are processed.

The architecture described here introduces another option, the ability to use the Claude code integrated with an IDE to use frontier models and fully secured sandbox opensource models

The Security Motivation

Infrastructure repositories are particularly sensitive. A Terraform repository may reveal much more than application source code. It can contain or expose information about:

  • network topology

  • Azure subscriptions

  • resource naming conventions

  • IP addressing and CIDR ranges

  • DNS architecture

  • storage accounts

  • Key Vault architecture

  • identity configuration

  • RBAC assignments

  • firewall policies

  • virtual networks

  • peering relationships

  • management infrastructure

  • security tooling

  • deployment pipelines

Terraform state is even more sensitive and may contain secrets or sensitive values depending on the resources involved. For this reason, I wanted the ability to perform high-volume AI analysis without automatically sending the entire infrastructure repository to an externally hosted model.

With local inference, the flow can instead be:

The source code and prompts can remain within infrastructure under organisational control.

That does not make local AI automatically secure. The Ollama endpoint, operating-system permissions, network controls, model provenance, logs, credentials and agent permissions all still require proper security engineering.

But it gives the organisation considerably more control over the data path.

The Lab Architecture

The test/sandbox environment consisted of two logical components.

Windows Server 2022. Windows provided the engineering workstation environment. It ran:

  • Claude Code

  • PowerShell

  • Git

  • Terraform

  • Development tooling, e.g. IDE

Claude Code was therefore responsible for the agentic workflow.

Linux AI Sandbox

A separate Linux server provided the inference infrastructure:

Ollama

│

├── gpt-oss:20b

├── gpt-oss:120b

└── other local models

│

â–¼

NVIDIA GPUs

The Ollama service was exposed only on the required network interface, fully firewalled with no access to unauthorised networks and Internet. The final communication path was therefore:

Preparing Ollama for Remote Access

The first requirement was to make the Ollama service reachable from the Windows server.

A common Ollama installation listens only on: 127.0.0.1:11434

That is appropriate for local-only operation but prevents another server from accessing the API. I used a systemd service override to bind Ollama to the Linux server's required network interface.

For example:

with:

The second variable was important in my environment because the models were stored on a dedicated data volume.

The service was then reloaded:

sudo systemctl daemon-reload

sudo systemctl restart ollama

sudo systemctl edit ollama

[Service]

Environment="OLLAMA_HOST=172.31.x.x:11434"

Environment="OLLAMA_MODELS=/dept-s/Data/ollama_models/"

The listener can be verified with:

sudo ss -lntp | grep 11434

The desired result was a listener on the server where the Ollama was running and not the default 0.0.0.0:

172.31.x.x:11434

Binding Ollama to a specific interface is preferable where practical because it reduces unnecessary network exposure.

Verifying the Ollama Model Inventory

The network-facing API was tested directly:

curl -s http://172.31.x.x:11434/api/tags | jq -r '.models[].name'

This returned the available models, including:

gpt-oss:120b

gpt-oss:20b-32k

gpt-oss:20b

For the initial Claude Code test I selected, gpt-oss:20b. The reason was simple: prove the architecture with the smaller model before moving to the substantially larger model.

Testing Network Connectivity from Windows

Before involving Claude Code, I verified basic connectivity from Windows Server 2022.

Test-NetConnection departments02 -Port 11434

The critical result was, TcpTestSucceeded : True

I then queried the Ollama API:

$response = Invoke-RestMethod -Uri "http://departments02:11434/api/tags"

$response.models | Select-Object name, model, size

This proved that Windows could communicate directly with the network-facing Ollama service.

Redirecting Claude Code to Ollama

The next step was to configure the environment used by Claude Code.

For initial testing I deliberately used session-level PowerShell environment variables rather than immediately changing the Windows system environment.

$env:ANTHROPIC_AUTH_TOKEN = "ollama"

$env:ANTHROPIC_API_KEY = ""

$env:ANTHROPIC_BASE_URL = "http://departments02:11434"

Claude code was then launched from vscode, I have now the option of running Claude with all its features with either opensource, local models running on Ollama or utilising frontier models from Anthropic.

How Do We Prove Claude Code Is Really Using Ollama?

DNS and network path checks:

GPU utilisation can also be observed, on the Linux server monitor GPU utilisation while on Windows machine Claude is given a task:

Combined with the configured base URL, service logs, `ollama ps` and GPU activity, this provides compelling evidence that the Linux-hosted open-weight model is performing the inference.

No Production Cloud Credentials Required

A local AI analysis environment does not inherently need production Azure permissions. For infrastructure-as-code analysis, that creates an important security boundary. The sandbox can inspect:

-tf

.tfvars

modules

documentation

pipeline definitions

scripts

without possessing credentials capable of:

terraform apply

terraform destroy

Azure resource modification

production RBAC changes

The Hybrid Model

Local models are not always the best model for every task.

That is why I consider this a hybrid architecture, rather than an ideological "everything must be local" architecture.

My intended workflow is:

Local models can handle tasks such as:

* repository discovery

* documentation

* module analysis

* repetitive refactoring

* Terraform syntax inspection

* naming reviews

* variable analysis

* code summarisation

* troubleshooting

* first-pass security analysis

* generating module scaffolding

A frontier hosted model can then be reserved for selected high-value tasks such as:

* architecture review

* complex debugging

* security-critical reasoning

* identity and RBAC design

* networking architecture

* Terraform state/lifecycle review

* final pull-request review

Applying This to Terraform

Infrastructure-as-code is where I think this architecture becomes particularly interesting.

Consider a large Terraform repository.

Instead of immediately uploading the repository to a hosted model, I can create a disposable sandbox clone.

Claude Code and the local model can first analyse:

Root modules

Child modules

Providers

Variables

Outputs

tfvars

Dependencies

Naming conventions

Tagging

RBAC

Networking

Security controls

Repeated code

Technical debt

It can then perform controlled development, for example:

The sandbox itself does not need authority to deploy into production. That is an important design principle:

AI that can write infrastructure code does not automatically need permission to deploy infrastructure code.

Security Controls I Would Add Before Production Use

The proof-of-concept demonstrates that the architecture works.

A production implementation should go further.

I would add:

* host firewall rules restricting TCP/11434

* dedicated Ollama service account

* controlled filesystem permissions

* TLS or a protected reverse proxy where appropriate

* authentication in front of the inference API where required

* network segmentation

* logging and monitoring

* model provenance controls

* approved model catalogue

* Git-based change control

* secret scanning

* repository access controls

* restricted agent permissions

* disposable development branches

* CI/CD security gates

* human approval before deployment

I would also explicitly prevent the AI sandbox from accessing production Terraform state unless there is a carefully reviewed requirement.

Terraform state deserves particular protection because it may contain sensitive values even where the corresponding Terraform configuration treats those values as sensitive.

Conclusion

This experiment started with a simple question:

Can Claude Code be used as an agentic development environment while inference is performed by an open-weight model running on another server?

In my lab, the answer was yes.

More importantly, I verified it at several levels.

The Windows server could reach Ollama.

The Ollama API exposed the local models.

Remote inference worked independently of Claude Code.

Claude Code was configured to use the Ollama endpoint.

The requested model appeared in `ollama ps`.

GPU execution was visible.

The result is a compelling architecture for an AI engineering sandbox:

Claude Code for agentic interaction, Ollama for locally controlled inference, open-weight models for routine engineering workloads, and hosted frontier models retained selectively for tasks where their additional capability justifies external inference.

For infrastructure engineering in particular, this offers an attractive balance between capability, cost, experimentation and control.

The next stage is to apply the architecture to a real Terraform repository: first as a read-only codebase analysis agent, then as a controlled development agent operating on feature branches, and finally as part of a hybrid local/frontier-model review workflow.

Connect

Email

Follow
Stay up to date

Write your text here...