# Entuit Enterprise Solutions Inc. > Entuit Enterprise Solutions Inc. builds AI features, websites, and web apps for small and mid-size businesses in Toronto and across Canada. We put AI to work on the repetitive jobs eating a team's week — answering repeat customer questions, reading paperwork, drafting first versions, and searching your own records — then build the sites and tools around it. Fixed prices agreed before work starts, senior engineering without a full-time hire, and everything handed over documented and owned by you. We build on a modern TypeScript stack — Bun, React with TanStack, Cloudflare, PlanetScale — with AWS available via CDK when a project genuinely needs it. Entuit is a consultancy based in Toronto, Ontario, Canada, serving the Greater Toronto Area on-site and clients across Canada and worldwide remotely. We build websites, web apps, and AI-powered features, then deploy them to the platform that actually fits — managed services such as Vercel and Cloudflare, or Amazon EKS and self-hosted inference when you need more control. We handle Canadian data residency (ca-central-1, self-hosted models) and PIPEDA/Law 25 constraints when they apply. Engagements are fixed-price with documentation and a team handoff. Full markdown versions of most pages are available by requesting `Accept: text/markdown` or by appending `?format=md`. ## Quick facts - Founder: Kevin Cearns, Founder & Principal Engineer (https://www.entuit.com/about#founder) - In business since 2003: over 20 years building software for businesses. - Location: Toronto, Ontario, Canada. Service-area business with no walk-in office; in person across the GTA, remote across Canada and worldwide. - Pricing: fixed prices in Canadian dollars (CAD), agreed in writing before work starts. Most projects run $2,500–$7,500 CAD. No hourly billing. - Typical timeline: about four weeks from first call to live. - Ideal client: businesses of roughly 2–50 people without an in-house technical team. - Ownership: code, hosting, and accounts are registered to the client, with written documentation. No lock-in. - Client workspace: each client gets a private Notion account page with projects, progress updates, requests, previews, quotes and invoices, decisions, documentation, and the components they own with renewal dates. There is no separate client portal or login. On offboarding, clients receive a full copy of their records and Entuit deletes its own, except records tax law requires it to keep. - Contact: hello@entuit.com, or book a free 30-minute call at https://www.entuit.com/contact ## Services - [Technical Review ($2,500 CAD, 1 week)](https://www.entuit.com/services#audit): A straight answer about what you already have and what to do next. We look at how your site or app is built, hosted, and updated, then tell you what's worth fixing, what's fine as-is, and what it would cost. You get a written plan you can act on — with or without us. - [Launch & Update Setup ($5,000 CAD, 1-2 weeks)](https://www.entuit.com/services#pipeline): We take the risk out of updating your site or app. Every change is tested automatically before it goes live, so a small update can't take the whole thing down — and if something does break, you find out immediately instead of hearing it from a customer. - [Build Sprint ($7,500 CAD, 2 weeks)](https://www.entuit.com/services#ai-sprint): Two weeks, one thing built properly and launched. A new website, an internal tool, a customer-facing feature, or AI added to something you already run. We agree the scope first, so you know exactly what you're getting for the price. - [Monthly Support ($3,000/mo and up CAD, Ongoing)](https://www.entuit.com/services#retainer): A senior engineer on call, without hiring one — and we run the hosting and infrastructure too. Fifteen hours a month for new features, updates, and fixes, plus uptime, backups, and cloud costs managed for you, so the whole thing stays fast, live, and predictable — at roughly a fifth of what a full-time hire costs. ## Key pages - [Home](https://www.entuit.com/): Overview of what Entuit builds and how we work. - [About](https://www.entuit.com/about): Who we are, why we exist, and how we operate. - [Services & Pricing](https://www.entuit.com/services): Fixed-price packages and process. - [AI Consulting in Toronto](https://www.entuit.com/ai-consulting-toronto): Local engagements, Canadian data residency, and GTA service area. - [Contact](https://www.entuit.com/contact): Book a free 30-minute call. - [Blog](https://www.entuit.com/blog): Technical writing on AI and platform engineering. - [Guides & Tutorial Series](https://www.entuit.com/series): Long-form, hands-on builds — full systems provisioned as reproducible AWS CDK infrastructure. ## Tutorial series - [Running a Fleet of Firecracker microVMs for eve.dev Agents](https://www.entuit.com/series/eve-firecracker-fleet) (7 parts): Turn a single Firecracker host into a fleet — a control plane that places eve.dev agents onto bare-metal AWS hosts on demand, and a routing layer that gets requests back to them. Built end to end as reproducible AWS CDK infrastructure. - [Building a Hybrid LLM Platform on EKS](https://www.entuit.com/series/eks-hybrid-llm) (8 parts): Build the EKS-based hybrid LLM platform referenced across the blog — from VPC to GPU node pools to a cloud-vs-local inference router — as reproducible AWS CDK infrastructure you can deploy and tear down yourself. ## Blog posts - [How Much Does AI Cost for a Small Business in Canada?](https://www.entuit.com/blog/ai-cost-small-business-canada): Real 2026 numbers in Canadian dollars: what off-the-shelf AI tools, a custom AI project, and the monthly running costs actually add up to for a business of 2–50 people, and when it isn't worth it. - [What It Takes to Own Your Agent Platform: The Firecracker Fleet Series in One Read](https://www.entuit.com/blog/eve-firecracker-fleet-recap): A capstone to the seven-part series on running eve.dev agents on a fleet of Firecracker microVMs in AWS. The whole platform in one picture, a guided tour of the parts, and the design threads — bare-metal constraints, desired-vs-actual state, event-driven placement, and packing-as-economics — that hold it together. - [The Cost-Efficient AI Stack: Ship AI Features Without the Runaway Bill](https://www.entuit.com/blog/cost-efficient-ai-stack): Most teams overpay for AI by routing every request to a frontier model. This is the architecture we build instead — hybrid cloud+local routing, self-hosted inference, agent orchestration, and cost-per-request observability — and the single principle that ties it together: send each unit of work to the cheapest model that can do it well. - [Run Your Own Buzz: A Local Community on Hardware You Control](https://www.entuit.com/blog/self-hosting-buzz-local-community): Buzz is a self-hostable workspace where people and AI agents share the same rooms, built as a Nostr relay. A hands-on walkthrough: standing up a local community, the host-binding gotcha that will stop you cold, and adding an agent backed by OpenRouter, a local Hermes model, or Claude Code. - [Running a Fleet of Firecracker microVMs for eve.dev Agents, Part 3: Packaging an Agent as a microVM Image](https://www.entuit.com/blog/eve-firecracker-fleet-agent-images): Part 3 of the hands-on series. We turn an ordinary eve agent directory into a bootable ext4 rootfs with a build-rootfs.sh, store it in a versioned S3 artifact bucket the hosts pull from, inject secrets per-microVM with MMDS, and use Firecracker snapshots to turn a cold multi-second boot into a warm sub-second resume. - [Running a Fleet of Firecracker microVMs for eve.dev Agents, Part 4: The Control Plane](https://www.entuit.com/blog/eve-firecracker-fleet-control-plane): Part 4 of the hands-on series. We build the scheduler that turns a deploy request into a running agent — API Gateway and Lambda over a DynamoDB registry of hosts, agents, and placements, a race-safe bin-packing placement algorithm, and EventBridge decoupling the API from the work, all in AWS CDK. - [Running a Fleet of Firecracker microVMs for eve.dev Agents, Part 6: The Deploy Workflow](https://www.entuit.com/blog/eve-firecracker-fleet-deploy-workflow): Part 6 of the hands-on series. We turn the moving parts into one command — a deploy CLI that chains build, upload, place, and wait-for-URL; IAM auth on the control-plane API; and a GitHub Actions pipeline that ships an agent on merge and gives every pull request its own preview agent. - [Running a Fleet of Firecracker microVMs for eve.dev Agents, Part 2: The Host Fleet](https://www.entuit.com/blog/eve-firecracker-fleet-host-fleet): Part 2 of the hands-on series. We put the first machines into the network from Part 1 — an Auto Scaling Group of bare-metal EC2 hosts, a launch template whose user data installs Firecracker and a host-agent daemon, the IAM role each host runs under, and the capacity model that decides how many agents a host can hold. - [Running a Fleet of Firecracker microVMs for eve.dev Agents, Part 7: Fleet Operations](https://www.entuit.com/blog/eve-firecracker-fleet-operations): The final part of the series. We make the fleet operable — event-driven host autoscaling with lifecycle-hook draining, a reconciliation loop that reschedules agents off a dead host, per-agent logs and metrics rolled up across the fleet, and the FinOps view that turns packing density into cost per agent. - [Running a Fleet of Firecracker microVMs for eve.dev Agents, Part 5: Networking & Routing](https://www.entuit.com/blog/eve-firecracker-fleet-routing): Part 5 of the hands-on series. We make a placed agent reachable — per-microVM tap devices and NAT on each host, an internet-facing ALB with a wildcard certificate, a per-agent subdomain scheme, and a host-local front proxy that resolves an agent to its microVM wherever the control plane placed it. - [A 101 Guide: Running an Eve Agent in a Firecracker microVM on AWS](https://www.entuit.com/blog/eve-agent-firecracker-aws-cdk): A beginner's walkthrough of provisioning bare-metal AWS infrastructure with CDK TypeScript, then booting a Firecracker microVM to self-host a Vercel eve agent. - [Running a Fleet of Firecracker microVMs for eve.dev Agents, Part 1: Architecture & the Network Foundation](https://www.entuit.com/blog/eve-firecracker-fleet-architecture): Part 1 of a hands-on series turning a single Firecracker host into a fleet that hosts eve.dev agents on demand. We map the whole platform — a control plane that places agents onto bare-metal hosts, and a routing layer that gets requests back to them — then provision the VPC, subnets, and security groups in AWS CDK. - [A Quick Example: Building an Agent with Vercel's Eve Framework](https://www.entuit.com/blog/vercel-eve-agent-quickstart): A short, hands-on walkthrough of scaffolding, running, and deploying an AI agent with Vercel's open-source eve framework. - [The Local AI Inflection Point: What the Next Three Years Actually Look Like](https://www.entuit.com/blog/future-of-local-ai): Local AI is crossing a threshold where on-device and self-hosted models stop being cost-cutting compromises and start being the default choice. Here's what's driving that shift and what it means for how you build software. - [Building a Hybrid LLM Platform on EKS, Part 5: Serving Local Models with vLLM and KEDA](https://www.entuit.com/blog/eks-hybrid-llm-platform-inference): Part 5 of our hands-on EKS series. We deploy vLLM model servers on the GPU pool from Part 4, load Qwen2.5-7B model weights from Amazon S3 via an init container, and wire KEDA autoscaling that scales replicas with live queue depth and drives GPU nodes to zero overnight. - [Building a Hybrid LLM Platform on EKS, Part 7: Observability and Cost Telemetry](https://www.entuit.com/blog/eks-hybrid-llm-platform-observability): Part 7 of our hands-on EKS series. We instrument the TypeScript router with OpenTelemetry, upgrade Prometheus to kube-prometheus-stack for GPU and vLLM metrics, add Grafana Tempo for distributed traces, and wire Langfuse so every request shows its backend, token count, and dollar cost. - [Building a Hybrid LLM Platform on EKS, Part 6: The Hybrid Router](https://www.entuit.com/blog/eks-hybrid-llm-platform-router): Part 6 of our hands-on EKS series. We build a TypeScript/Hono router that sits in front of both vLLM and the Anthropic API, routes each request to the right backend based on model name and complexity heuristics, and falls back to cloud when the local model is cold-starting. - [Building a Hybrid LLM Platform on EKS, Part 8: Testing, Load, and Examples](https://www.entuit.com/blog/eks-hybrid-llm-platform-testing): The final part of our EKS series. We write integration tests with Vitest, load-test the ALB with k6, build three real-world TypeScript workloads that prove the hybrid routing works, and use the Grafana and Langfuse dashboards from Part 7 to verify the platform under traffic. - [Building a Hybrid LLM Platform on EKS, Part 4: Platform Add-ons, the Load Balancer Controller, and Karpenter](https://www.entuit.com/blog/eks-hybrid-llm-platform-addons): Part 4 of our hands-on EKS series. We install the two add-ons every production EKS cluster needs: the AWS Load Balancer Controller so Kubernetes Ingress objects provision real ALBs, and Karpenter for cost-aware autoscaling — including the GPU NodePool that scales to zero between inference workloads. - [Building a Hybrid LLM Platform on EKS, Part 3: Node Groups, GPU AMIs, and the NVIDIA Device Plugin](https://www.entuit.com/blog/eks-hybrid-llm-platform-node-groups): Part 3 of our hands-on EKS series. We add worker nodes to the empty cluster from Part 2: a CPU system pool for add-ons and the hybrid router, a GPU pool for vLLM model servers, the NVIDIA device plugin DaemonSet, and the taints and labels that make scheduling predictable. - [Building a Hybrid LLM Platform on EKS, Part 2: The Control Plane, IAM, and IRSA](https://www.entuit.com/blog/eks-hybrid-llm-platform-control-plane): Part 2 of our hands-on EKS series. We provision the EKS cluster into the VPC from Part 1, wire up OIDC federation and IRSA so pods authenticate without static credentials, and end with a working kubectl connection to a real cluster. - [Securing Self-Hosted LLMs and AI Agents on Kubernetes](https://www.entuit.com/blog/securing-self-hosted-llm-agents-kubernetes): Harden self-hosted vLLM and AI agents on Kubernetes: an auth/rate-limit gateway, gVisor tool sandboxing, prompt-injection guardrails, scoped secrets, and signed model weights — mapped to the OWASP LLM Top 10. - [Building a Hybrid LLM Platform on EKS, Part 1: Architecture and the Network Foundation](https://www.entuit.com/blog/eks-hybrid-llm-platform-architecture-network): Part 1 of a hands-on series building the EKS-based hybrid LLM platform referenced throughout this blog. We map out the full architecture, then provision the VPC, subnets, NAT, and VPC endpoints with AWS CDK — the network foundation every later part builds on. - [Build a Personal AI Dev Environment: Hybrid Models, Local Inference, and a Workflow That Costs Almost Nothing](https://www.entuit.com/blog/personal-ai-dev-environment): The production patterns we deploy for teams — hybrid cloud/local routing, self-hosted models, agent orchestration — scaled down to a single developer's workstation. A practical guide to building a personal AI dev environment with Ollama, Claude Code, and a local router that keeps your token bill near zero. - [The Agent Control Plane: Frontier Models Plan, Your Kubernetes Fleet Executes](https://www.entuit.com/blog/multi-agent-orchestration-frontier-local-models): How to orchestrate a fleet of AI agents using a shared task queue — frontier models like Claude handle planning and decomposition, while a local Kubernetes worker pool runs the high-volume execution tasks. Covers the task ledger, dynamic task creation, lane-based routing, and KEDA autoscaling. - [Observability for LLM Applications on Kubernetes: Tokens, Traces, and Cost per Request](https://www.entuit.com/blog/llm-observability-kubernetes): How to instrument self-hosted and hybrid LLM workloads with OpenTelemetry, Prometheus, and Langfuse — tracking time-to-first-token, tokens per second, GPU utilization, and unit economics down to the individual request. - [The Hybrid AI Playbook: Cloud Models for Thinking, Local Models for Doing](https://www.entuit.com/blog/hybrid-ai-local-cloud-models): How to cut your AI costs by 60-80% using a hybrid approach — Claude or GPT for planning and complex reasoning, local models like Llama and Qwen for execution tasks like code generation, summarization, and data extraction. - [Self-Hosting LLMs on Kubernetes: A Practical Guide](https://www.entuit.com/blog/self-hosting-llms-kubernetes): How to deploy, serve, and autoscale open-source large language models on Kubernetes with vLLM — from GPU node pools and deployment manifests to KEDA-based autoscaling and production guardrails. - [Container Security on Kubernetes: A Practical Guide with Trivy, Falco, and Kyverno](https://www.entuit.com/blog/container-security-kubernetes-trivy-falco-kyverno): Most Kubernetes clusters are running containers with known vulnerabilities, no runtime monitoring, and no policy enforcement. Here is how to fix that with three open-source tools. - [How to Cut Your AWS Bill in Half Without Changing Your Architecture](https://www.entuit.com/blog/cut-aws-bill-in-half): Most growing teams are overpaying on AWS by 30-50%. Here is the exact checklist we use in every infrastructure audit to find and eliminate wasted spend — no migrations, no rearchitecting. - [AI Data Residency in Canada: A Practical Guide for Businesses](https://www.entuit.com/blog/canadian-ai-data-residency): What Canadian businesses need to know about PIPEDA, data residency, and customer data when adopting AI — and when self-hosting is actually the right answer. - [Using AI to Monitor Kubernetes Clusters and Make Dynamic Scaling Decisions](https://www.entuit.com/blog/ai-driven-kubernetes-monitoring-scaling): How to move beyond static thresholds and use AI-driven observability to detect anomalies, predict traffic patterns, and automate scaling decisions across your Kubernetes infrastructure. - [A Practical Guide to AI for Small and Mid-Size Businesses](https://www.entuit.com/blog/ai-guide-for-small-business): No hype, no jargon — a straightforward guide for business owners evaluating where AI actually makes sense and how to adopt it without wasting money. - [Building a CI/CD Pipeline with Dagger That Deploys to Kubernetes](https://www.entuit.com/blog/cicd-pipeline-dagger-kubernetes): A practical guide to building a containerized CI/CD pipeline using Dagger's TypeScript SDK — from local Kind clusters to production EKS with GitHub Actions, AWS CDK, and multi-environment promotion. - [Building a Production Feature Flag Service with Claude Code](https://www.entuit.com/blog/building-feature-flag-service-claude-code): How we built FlagSignals, a full-stack feature flag platform with A/B testing and billing, using AI-assisted development. - [GPU Cost Optimization on Kubernetes: A Practical Guide](https://www.entuit.com/blog/gpu-cost-optimization-kubernetes): Learn how to reduce GPU infrastructure costs by up to 60% with proper Kubernetes scheduling, time-slicing, and right-sizing strategies. - [Platform Engineering for AI/ML Teams: Building the Foundation](https://www.entuit.com/blog/platform-engineering-ai-ml-teams): How platform engineering principles transform AI/ML infrastructure from artisanal setups to scalable, self-service platforms. - [FinOps for AI Infrastructure: Beyond Cloud Cost Tags](https://www.entuit.com/blog/finops-ai-infrastructure): Traditional FinOps practices fall short for AI workloads. Here's how to build a cost management strategy that accounts for GPU economics. ## Case studies - [Building a Production CI/CD Pipeline: From Local Dev to Multi-Environment EKS](https://www.entuit.com/case-studies/cicd-pipeline-dagger-kubernetes): How we replaced fragile YAML-based CI/CD with a Dagger pipeline written in TypeScript that runs the same on a laptop and in production — deploying to three environments on EKS with automated testing at every stage. ## Frequently asked questions - **Do I need to understand AI to work with you?** No. You bring the business problem (where time or money leaks) and we work out what, if anything, AI should do about it. If the answer is "nothing yet," we'll tell you that instead. - **How can you build in weeks when agencies quote months?** We use AI heavily in our own development, with automated tests checking every change, so the typing that used to take weeks takes days. Most of the remaining time is your feedback and the launch work (hosting, domain, monitoring) that makes it real. There's no account-management layer slowing things down. - **What happens after the project is finished?** That's up to you. Everything is handed over documented, in accounts you own, with a written guide. Many clients stay on Monthly Support; others take it and run. There's no lock-in. - **Is a fixed price really fixed?** Yes. We agree the scope in writing before we start, and the price on that page is the price. If you want to change the scope mid-build, we talk about it first. You never get a surprise invoice. - **What if the AI gets something wrong?** Every build has a review step, and the reviewer is a person on your team. AI drafts, you approve. For anything that runs on its own, like a support assistant, we set clear rules for when it hands off to a person. - **Can you work with our existing website or software?** Usually, yes. WordPress, Shopify, Vercel, Cloudflare, a custom app someone built years ago: we start where you are. If a rebuild makes more sense, we'll show you the costs side by side before touching anything. ## Optional - [Full content (llms-full.txt)](https://www.entuit.com/llms-full.txt): Every blog post in full markdown, concatenated. - [RSS feed](https://www.entuit.com/feed.xml): Subscribe to new posts. - [Sitemap](https://www.entuit.com/sitemap.xml): Machine-readable list of all URLs.