Projects

Local & Sovereign AI

ACTIVEContinuous benchmarking · updated 2026-08-08

Local AI Research

งานวิจัยเอไอแบบท้องถิ่น

Serving, routing and quantisation research for running Thai-capable models entirely on locally owned GPU infrastructure.

งานวิจัยการให้บริการและปรับแต่งโมเดลบนเซิร์ฟเวอร์ GPU ในประเทศ

View ResearchDemoArchitecturePaperDatasetCodeAPI

Problem

Sensitive Thai institutional data cannot leave the country, but local inference quality and throughput are poorly characterised.

Research Question

What quantisation and routing configuration gives the best Thai-language quality per watt on commodity GPUs?

Objectives

  • Benchmark quantisation levels for Thai-language tasks
  • Build a router that selects models per task class
  • Characterise throughput, VRAM and power

Methodology

  1. 01Standardised harness
  2. 02Repeated runs
  3. 03Power and thermal logging

Architecture

UserClient

Request entry

AI GatewayFastAPI

Auth, rate limiting

AI RouterPolicy engine

Task-class model selection

Local ModelsOllama

Quantised runtimes

GPU ServerCUDA

Scheduling and batching

Knowledge BaseVector store

Grounding corpus

Results & Benchmark

Runtimes compared
quantised variants

Limitations

  • · Benchmarks reflect a single hardware configuration