SPENCER HEYWOOD

AI TRANSFORMATION · INDUSTRIAL OPERATIONS

From AI opportunity to measurable business value.

  • Strategy
  • Use-case discovery
  • Delivery
  • Measurable value

I translate operational needs into practical AI and automation—from use-case evaluation and architecture through implementation, rollout, and production support.

CURRENT ASSIGNMENT

BEUMER Group

Software Systems Engineer II

High-throughput intralogistics systems · Enterprise AI · Systems integration

Live Lab Stats

The self-hosted stack from the AI Systems Lab runs my day-to-day agentic work—chat, agents, RAG pipelines, and the model answering on this page. These figures stream live from the LLM gateway on my home server: what the build cost, what it has saved in hosted API spend, the power it burns, and how hard one person can push a single workstation before it saturates and requests spill over to the API.

LIVE LOCAL INFERENCE RTX PRO 6000 · SELF-HOSTED
Connecting to the local token counter
Connecting to LLM gateway… – active streams Checking the local model’s activity

Lifetime activity from models running on my own hardware. Ask the model below and watch the counter move.

SYSTEM ECONOMICS · LIVE FROM THE GATEWAY Waiting for the gateway…
API cost avoided What all locally served tokens would have cost through commercial APIs (OpenRouter list prices, Sep 2026), with the usual discount for cached prompts. An estimate — no actual bills. All tokens vs hosted list pricing
Est. payback What’s left of the $12,400 build divided by today’s net daily savings (API cost avoided minus electricity), counting from when it started serving. $12,400 build vs net savings
GPU utilization How much of the last month the local GPU has been busy — shared work counts once, and requests that overflowed to APIs don’t count. Busy time · last 30 days
Served locally tokens overflowed to the API (–%)
Lifetime tokens · all routes

Input/output split pending

requests · tokens overflowed to the API

Power & energy · cost per hour

Energy totals pending

Lifetime cost so far: –

~500W average draw at 3 concurrent streams

Net of power: –

Mean throughput · tokens/sec
Peak
24h avg
Busy avg
Top models

Cost estimates value each token at that model’s own hosted list price, not invoices · power is plug-measured at 3 concurrent streams · throughput figures count generated output tokens only, prompts excluded · aggregates refresh every minute

Transformation Focus

Turning emerging AI capabilities into governed, practical systems for engineering teams, operations, and the wider business.

  • Use-case discovery
  • Business impact & feasibility
  • AI solution architecture
  • Hands-on implementation
  • Adoption & enablement
  • Production integration

Professional Experience

BEUMER Group

Software Systems Engineer II

2+ yrs

Engineering and validating high-throughput intralogistics systems across software, databases, networks, host integrations, and industrial controls.

3,000+ items / minute Production troubleshooting AI architecture & evaluation

Delaire USA

Software & Systems Engineer · Part-Time

2+ yrs, part-time

Modernizing enterprise infrastructure, security, monitoring, and software while evaluating LLM and RAG use cases.

Infrastructure modernization LLM & RAG evaluation

SERVO

Founder / Founding Software Developer

3.5 yrs

Built and operated a services marketplace from product concept through production, owning architecture and implementation.

10+ partner companies 91%+ lower cost / session

End-to-End Delivery

01

Identify & Prioritize

Structure opportunities around business impact, technical feasibility, risk, and operational fit.

02

Build the Solution

Translate requirements into automation, data pipelines, AI applications, integrations, and governed access.

03

Deploy & Enable

Move beyond prototypes with production integration, team education, documentation, and practical adoption.

04

Measure & Improve

Evaluate quality, reliability, utilization, and value to inform iteration and investment decisions.

AI Systems Lab

What started as a hobby build became a hands-on lab for learning local inference, hardware constraints, and dependable self-hosted infrastructure.

  1. Open workstation showing the initial GeForce RTX 5060 Ti build
    V1 · THE START

    Learning locally with an RTX 5060 Ti

    I started with a practical, air-cooled desktop: enough GPU to move beyond hosted demos and begin running, comparing, and troubleshooting models on my own hardware.

    VRAM
    16 GB
    GPU power
    Up to 180W
    Capability
    Local inference
    • Air cooled
    • First model benchmarks
  2. Compact AI workstation during a GPU upgrade
    V2 · MORE HEADROOM

    Upgrading to an RTX PRO 4500

    As experiments grew, memory and sustained workloads mattered more. V2 expanded the lab from basic inference into larger models, longer sessions, and more serious day-to-day use.

    VRAM
    24 GB
    GPU power
    ≈ 200W
    Capability
    Larger models
    • More VRAM
    • Sustained workloads
  3. Workstation chassis prepared for the RTX PRO 6000 upgrade
    V3 · CURRENT WORKSTATION

    RTX PRO 6000 and water cooling

    V3 became a high-capacity workstation built for ambitious local AI work. The RTX PRO 6000 upgrade also drove a move to radiator-based water cooling, with power, thermals, airflow, and reliability designed as one system.

    VRAM
    96 GB
    GPU power
    Up to 600W
    Local throughput
    5M+ tokens/week
    • High-capacity GPU
    • Water cooling
    • Thermal headroom
SUPPORTING INFRASTRUCTURE

Repurposed compute and reliable, always-on services.

  1. Tesla P40 GPUs installed in a Dell PowerEdge R520 rack server
    04 · REPURPOSING HARDWARE

    Tesla P40 meets PowerEdge R520

    Not every useful AI machine has to be new. Installing a Tesla P40 in an older Dell PowerEdge R520 became an experiment in extracting inference value from proven enterprise hardware.

  2. Dell OptiPlex 5070 home server mounted on a utility shelf
    05 · THE ALWAYS-ON FOUNDATION

    Dell OptiPlex 5070 home server

    The quietest machine may do the most work. This OptiPlex hosts my websites, LLM and MCP gateway, NAS, and VPN—turning hardware experiments into services I can depend on every day.

    Web hostingLLM/MCP gatewayNASVPN 150+ days uptime

Selected Projects

SERVO · BookServo.com

Founder / Founding Software Developer · 2021 – 2024

Built and operated a services marketplace integrating 10+ partner companies, scheduling, payments, and operational workflows.

10+partner companies
91%+lower technical cost per session
End-to-endproduct & architecture ownership
  • Ruby on Rails
  • PostgreSQL
  • AWS
  • Stripe
  • Google Calendar API
Visit BookServo.com

Earlier work

Technical Capabilities

AI & Agents

  • LLMs
  • RAG
  • MCP
  • AgentGateway
  • LM Studio
  • OpenAI-compatible APIs
  • Agent tooling
  • Model evaluation

Software

  • Python
  • C#
  • C++
  • Ruby
  • JavaScript
  • TypeScript

Frameworks

  • Ruby on Rails
  • Angular
  • React
  • Node.js

Data

  • Microsoft SQL Server
  • PostgreSQL
  • Elasticsearch / Kibana

Infrastructure

  • Linux
  • Windows Server
  • Nginx
  • AWS
  • Networking

Engineering & Automation

  • Distributed systems
  • System integration
  • Troubleshooting
  • Validation & testing
  • Power Automate
  • Office Scripts

Education

App Academy

Full-Stack Software Engineering

Contact Me