Language models on Swedish GPUs

Some data is too sensitive to leave Sweden. For those workloads we run the strongest open models in our own data center. The API is OpenAI compatible, so your existing code works as soon as you point it at our endpoint. Your data never leaves Sweden and is never used for training.

0+ bn

tokens every month

Our own team and the solutions we operate for customers consume hundreds of billions of tokens every month, so we know what it takes to run language models in production.

Same API, full control
The platform exposes an OpenAI-compatible API. Existing tools, libraries and applications work out of the box. Behind the API we run open models we have benchmarked ourselves, on our own GPUs and with the same reliability as the rest of Liminity Cloud.

What you get

API inference

OpenAI-compatible API to the strongest open models.

POST /v1/chat/completions
"model": "gpt-oss-120b"
200 OK

Your data is yours

Nothing is stored longer than necessary. Nothing is used for training.

GDPR and the AI Act

All processing happens in Sweden. We help you document compliance.

GDPR
AI Act

Dedicated capacity

Your own NVIDIA RTX PRO 6000 with guaranteed throughput and your choice of model.

MIG ×4 · 96 GB

From pilot to production

Our AI consultants take you from pilot to production, with agent workflows and tailored search that make the models answer accurately from your data.

index: 2.4M chunks
recall@10: 0.94

A data lake for agents

Your agents get permission-controlled access to your data through our platform LimiLake: structured data, files and SQL in one place.

SELECT sum(total)
FROM lake.invoices
1 204 rows
AI without the legal headache
We are not against the global AI services, we use them ourselves every day. But sensitive personal data requires legal assessment, and sometimes the answer is no. Those workloads run with us in Sweden under clear data processing agreements, and the rest stays in the clouds you already use. We will gladly help you put the whole setup in place.
Your own hardware, fully configured
If you want your own physical server, with or without GPUs, we deliver it fully configured. Run it in our data center or on your own premises.

Pricing per model

You pay per token. The table below is the whole price list. All prices excl. VAT.

ModelInputOutput
GLM 5.217 kr52 kr
Kimi K2.69 kr42 kr
Qwen 3.5 397B8 kr43 kr
GPT-OSS 120B3 kr9 kr
Qwen 3.6 35B2 kr8 kr
Mistral Small 3.2 24B4 kr4 kr
Gemma 4 27B3 kr6 kr
E5 Large (embeddings)0.4 kr

Prices in SEK per million tokens.

Example prices. The model lineup is updated continuously.

Dedicated capacity

GPU capacity on NVIDIA RTX PRO 6000 with 96 GB VRAM. Rent a full GPU or a MIG instance. All prices excl. VAT.

MIG instance

from 3,995 SEK/month

  • Isolated slice of an RTX PRO 6000
  • Guaranteed VRAM and compute
  • Private endpoint
  • Fits smaller models and test workloads
Most popular
Dedicated GPU

from 12,995 SEK/month

  • NVIDIA RTX PRO 6000, 96 GB VRAM
  • Any model, including fine-tuned
  • Splittable into up to four MIG instances
  • Monitoring included
Enterprise

Custom quote

  • Multiple GPUs and models
  • GPU nodes in your Kubernetes cluster
  • Custom SLA
  • Data processing agreement

Example prices. Contact us for a review of your use case.

Frequently asked questions

Make your first API call the same day

Book a demo and we will show you the API and discuss how language models can create value in your business.

Do you want to be part of the journey to become Sweden's largest consulting company?

APPLY TODAY


© Liminity. All rights reserved.