Kai Ase Siren

Contracting

If you are thinking about bringing me in, start here. The work I take on is platform engineering first, the infrastructure work a team already has a line in the budget for. The agent work comes last on this page, because it is the specialism I bring to that work rather than a separate offer.

The work

Platform engineering, more than ten years of it.

Each entry below is work I have already done in production. The resume is the full record, with where and when.

Infrastructure as code
Terraform, CloudFormation, Pulumi, and Ansible across AWS, GCP, and Azure. I have brought deployment tooling that had fallen behind back to current standards, and moved internal systems onto production-grade managed services.
Kubernetes and containers
Services on Kubernetes, container platforms on AWS ECS, and the Docker work underneath both. Most recently that included making a Dockerized Python service start about twice as fast.
Observability
OpenTelemetry, Datadog, Prometheus, Grafana, and New Relic. I have migrated a company's monitoring onto a customized Datadog and Fluent Bit system.
Networking across clouds
Azure OpenAI deployed through a multi-cloud BGP VPN into AWS.
Compliance work
SOC 2 control remediation and the tooling that supports it, and NIST 800-53 remediation on a government site for the US Department of Health and Human Services.
Developer tooling
Internal tools and automation that speed up daily engineering, and a maintainer stint on urfave/cli, a Go CLI framework with broad downstream use.

How an engagement is shaped

Contract shape
Every shape is open. Direct contracting through my own business, consultancy, embedded delivery, staffing-agency W-2, contract-to-hire, and agency-mediated C2C all work. Use whichever one your procurement finds easiest.
Rate
$150 an hour. One rate covers every shape above, agency arrangements included, so the paperwork you prefer does not change the price.
Availability
I plan at about 70 percent utilisation rather than trying to fill every hour, so there is room for the work to overrun without the next thing suffering for it.
Where I am
East Bay, California. Remote works, and I can be in a Bay Area room when being in the room is the thing that moves it.

The specialism

Constrain what an agent can do, and prove what it did.

Agents are the newest consumer of the platform underneath them, and this is the layer I have built for them, four times. That is one offer with four proofs rather than four services. Each project below is the same idea at a different layer. All four ship, all four are open source, and each has its own page here with its own documentation, so you can decide whether an hour on a call is worth it before you spend the hour.

umbra
You gave an agent a shell. Now name every command it can run. umbra validates argv before execve, checks a scope token per verb, routes egress through a per-invocation proxy, appends every call to an audit log, and publishes an exit-code taxonomy that separates a policy refusal from a tool failure. MIT, on brew and scoop, Linux, macOS, and Windows.
mcp-beaver
You need an MCP server for one API. Now write a Go package, a handler and a schema per tool, and a Dockerfile. mcp-beaver renders one guardfile into a running MCP server, an HTTP tool API, and a helm release. One grant becomes one tool, and an operation nobody declared has no tool and no endpoint at all rather than a guarded one. MIT, preview.
housecast
You changed what a role is allowed to do. The evaluation that checked it did not notice. housecast declares roles, personalities, and boundaries in one YAML roster, validates it, emits an immutable bundle, and runs the behavior evaluations against exactly that bundle. The graded artifact and the shipped artifact are the same thing. MIT, preview, Python.
agent-compose
You gave your agents different jobs. They all load the same prompt. agent-compose materializes a named specialist with its own charter, its own set of things it hands off instead of attempts, and its own tool inventory, as a directory of plain files you can diff before anything runs. The same pack runs unchanged on Claude Code, Codex, Goose, and OpenCode. MIT, on brew and scoop.

None of the four ships or calls a model of its own. They bound and record what an agent does, whatever is running underneath, so they carry no model-provenance question into a procurement review.

If you want umbra or agent-compose configured for your team at a fixed size instead of by the hour, that is a different product, on the setups page.