Konkui Logo

AI infrastructure & transformation

Become an AI-native company, on infrastructure you own

We design, build and run the layer between your business systems and modern AI models: MCP tool servers, agent-to-agent workflows, retrieval over your own knowledge, and the evaluation that keeps it honest. In the cloud, on your own hardware, or fully air-gapped.

  • MCP & A2A engineering
  • On-premise & air-gapped
  • Thai and English delivery teams

What we build

The plumbing that makes AI useful inside a company

The layer between your business systems and modern models. Built as real infrastructure — versioned, tested, documented and handed over.

  • MCP tool servers for your systems

    Your ERP, POS, CRM, booking system and internal APIs, exposed to any model as governed tools over the Model Context Protocol.

    • One MCP server per system, with typed tools instead of scraped screens
    • Per-tenant scoping, authentication and a full audit trail on every call
    • Reusable by Claude, OpenAI, local models, your own agents and ours
  • Agent-to-agent (A2A) systems

    Specialist agents that discover each other and hand work across teams and vendors, instead of one prompt trying to do everything.

    • A capability registry so an agent can find the right counterpart
    • Delegation contracts with timeouts, retries and idempotent replays
    • Human approval gates on the actions that move money or data
  • Agent runtime and orchestration

    The unglamorous layer that makes agents survive production: durable queues, isolated execution and turns that can run for minutes without dropping work.

    • Durable job queues with checkpointing and crash recovery
    • Sandboxed execution per turn, with an egress policy you approve
    • Budget, rate and capability limits enforced before the model is called
  • Retrieval over your own knowledge

    Documents, policies, product data and past conversations turned into retrieval your staff can actually trust — in Thai and English.

    • Ingestion pipelines for PDFs, spreadsheets and scanned Thai documents
    • Permission-aware retrieval: an answer never crosses a role boundary
    • Freshness and ownership rules so stale content is retired, not repeated
  • Evaluation and observability

    The difference between a demo and a system you can bet operations on is measurement, and measurement is part of every build we ship.

    • Golden test sets built from your real conversations and documents
    • Replay harnesses that score a model or prompt change before it ships
    • Per-turn traces with cost, latency and failure dashboards your team owns
  • Becoming AI-native, not AI-curious

    Infrastructure only pays off when the organisation changes around it, so delivery includes the people side of the work.

    • A prioritised roadmap tied to measurable operational outcomes
    • Enablement for your engineers: architecture, runbooks and handover
    • Workshops for the teams whose daily work the system changes

Sovereign AI

Local inference, on-premise and air-gapped

For teams whose data cannot leave the building — regulated industries, government, defence-adjacent work, and anyone who would simply rather own the stack.

  • Local inference stack

    Model selection, GPU sizing and a serving stack tuned to your latency and throughput targets — behind an OpenAI-compatible endpoint, so your applications do not need to know where the model runs.

  • Air-gapped deployment

    For environments with no outbound network at all: an offline supply chain for models and images, signed artefact transfer, internal package mirrors, and a network and hardware posture your auditors can follow.

  • Full local deployment

    The whole stack — models, vector store, queue, orchestrator, observability — on your own hardware or private cloud, deployed as infrastructure-as-code that stays in your repository.

  • Data residency and governance

    PDPA-aligned handling with tenant isolation, retention and redaction policy, key management, and audit trails that answer who asked what, and what the system did about it.

We specify the hardware and you buy it directly. No markup, no proprietary runtime, no lock-in to us.

How we engage

Assess, design, pilot, run

Every engagement starts with an assessment, because a roadmap built on guesses is the most expensive thing in AI.

Four stages, each one useful on its own

  1. 01

    Assess

    We map your processes, data and systems, then rank the opportunities by value and by how hard they are to do properly.

  2. 02

    Design

    A target architecture with an explicit deployment posture — cloud, on-premise or air-gapped — and the security model that comes with it.

  3. 03

    Build & pilot

    One workflow into production with real users, real integrations and the evaluation set that proves it is working.

  4. 04

    Deploy & run

    Roll out across teams, hand over to your engineers, and keep the stack current as models and hardware move.

Engagement

AI-native assessment

A short, evidence-based read on where AI belongs in your operations — and where it does not.

From

฿180,000per engagement

Timeline: 2–3 weeks

What you get

  • Process, data and systems audit
  • Opportunity map with expected impact per use case
  • Target architecture and deployment posture
  • Prioritised roadmap with build estimates
Talk to us
Most engagements start here

Build & pilot

The first production system: MCP servers for your key platforms and one agent workflow your team relies on.

From

฿650,000per engagement

Timeline: 6–10 weeks

What you get

  • MCP tool servers for your priority systems
  • One agent or A2A workflow live in production
  • Evaluation sets, tracing and cost dashboards
  • Architecture documentation and engineer handover
Talk to us
Engagement

Sovereign deployment & run

Your own inference stack, installed on your hardware — including fully air-gapped sites — and kept running.

From

฿15,000per month

Timeline: ongoing

What you get

  • Local inference stack on hardware you own
  • Air-gapped installation, runbooks and upgrade procedure
  • Model upgrades, capacity planning and tuning
  • Support in Thai and English with an agreed response time
Talk to us

Indicative figures for planning. Final scope and price are quoted after the assessment, and hardware is quoted separately at cost.

Questions we get

Before you book a call

No. We deploy the same architecture in three postures: managed cloud, on-premise in your own data centre, or fully air-gapped with no outbound network. The posture is a decision we make with you during the assessment, before anything is built.

Start here

Start with an assessment, not a proof of concept

Two to three weeks from now you can have a ranked roadmap, a target architecture and a straight answer on what belongs on your own hardware.