Projects
Local & Sovereign AI
ACTIVEContinuous benchmarking · updated 2026-08-08Local AI Research
งานวิจัยเอไอแบบท้องถิ่น
Serving, routing and quantisation research for running Thai-capable models entirely on locally owned GPU infrastructure.
งานวิจัยการให้บริการและปรับแต่งโมเดลบนเซิร์ฟเวอร์ GPU ในประเทศ
View ResearchDemoArchitecturePaperDatasetCodeAPI
Problem
Sensitive Thai institutional data cannot leave the country, but local inference quality and throughput are poorly characterised.
Research Question
What quantisation and routing configuration gives the best Thai-language quality per watt on commodity GPUs?
Objectives
- Benchmark quantisation levels for Thai-language tasks
- Build a router that selects models per task class
- Characterise throughput, VRAM and power
Methodology
- 01Standardised harness
- 02Repeated runs
- 03Power and thermal logging
Architecture
UserClient
Request entry
AI GatewayFastAPI
Auth, rate limiting
AI RouterPolicy engine
Task-class model selection
Local ModelsOllama
Quantised runtimes
GPU ServerCUDA
Scheduling and batching
Knowledge BaseVector store
Grounding corpus
Results & Benchmark
- Runtimes compared
- quantised variants
Limitations
- · Benchmarks reflect a single hardware configuration