Intelligent Inference
PK End-point

The i2 PK Endpoint: AI APIs with Pakistani data residency.

Access open-source models through one API while your prompts, requests, logs and billing records are processed in Pakistan, on hardware built for it.

Talk to our team

The PK NPU-native routes are opening shortly and are tagged Coming soon in the catalogue. The gateway, the keys, the ledger and the global pass-through routes are live on this endpoint today.

Same OpenAI-compatible APINPU-native routing in PakistanData stays in Pakistan
PK data residencyHuawei Ascend 910BBilled in PKRPer-key scope and capsEncrypted in transitKeys stored as hashes
NPU-native inference in Pakistan

Huawei NPUs are going to get better and be everywhere.

There is a learning-curve challenge to using Huawei's Ascend NPUs.

But Pakistan, along with other countries that face Nvidia import challenges, will mostly have NPUs in the country.

We are here to bridge that gap. We give you highly optimised open-source model endpoints, with a sovereign gateway that works on both Huawei NPUs and Nvidia GPUs.

Dedicated support, and accountability for uptime, on Huawei NPUs.

Huawei Ascend 910B
64GB NPUs, racked in country rather than rented by the hour in a foreign region.
One gateway, two kinds of silicon
The same route id serves from Ascend NPUs or Nvidia GPUs. Your code does not know which, and does not need to.
Accountable uptime
The people who wrote the Ascend serving stack are the people on call for it. Dedicated support, not a ticket queue.

Already built on CUDA

CUDA
CANN

CUDA to CANN translation.

If your model is CUDA-native and already prepared for deployment on GPUs, that is not a blocker. We translate the model code to run on CANN, Huawei’s Ascend compute stack, and serve your inference on NPUs with the performance numbers to show for it.

Residency, control, accountability

Built for teams that need to say where their AI ran, and in what currency it was paid for.

The PK Endpoint is for organisations that need guarantees about where AI data is processed, how a caller is constrained, and how the money moves. Each guarantee below is a property of the deployment as it stands, not a setting someone has to remember to switch on.

PK data residency

On the PK NPU-native routes, prompts, requests and outputs are processed on Ascend NPUs racked in Pakistan. Request telemetry, the audit log and the ledger live on infrastructure in Pakistan for every route.

A sovereign gateway

Keys, budgets and model scope are enforced on the gateway, in country. A key can be scoped to PK routes alone, so a caller cannot reach a pass-through route by accident.

Billing residency

Prices are quoted per million tokens in rupees, debited from a prepaid balance and settled in the same currency. No foreign card, no dollar invoice, no exchange-rate surprise.

Security and control

Traffic is encrypted in transit. i2 stores only a SHA-256 hash of each API key, revocation takes effect on the next request, and every administrative change is written to an append-only audit log you can export.

Built-in guardrails

Support Pakistani AI. Keep your data in Pakistan.

The endpoint is built to grow NPU-native inference in Pakistan without cutting you off from global models when you need them. The goal is simple: more of the country’s AI served in the country, with residency you can enforce rather than hope for.

  1. 01

    NPU-native first

    The catalogue highlights the routes that run on Ascend in Pakistan and tags them plainly, so a team choosing where its tokens are served can see it at a glance.

  2. 02

    Pass-through, labelled

    Global routes are relayed to partner capacity outside Pakistan and are marked Global everywhere they appear: the catalogue, the playground picker, the dashboard's model card.

  3. 03

    Your choice, per key

    Scope a key to PK routes only and the gateway answers 403 for anything else. Residency becomes a property of the key, enforced on every request, not a checkbox someone has to remember.

FAQ

Questions we are usually asked.

What is the i2 PK Endpoint?
It is i2's OpenAI-compatible API, served through a sovereign gateway that runs in Pakistan, with a pool of routes that execute on Huawei Ascend NPUs racked in country. It is for teams that need to know where their AI requests are processed and in what currency they are billed.
What is the endpoint URL?
https://api.intelligentinference.ai/v1 for every route. Chat completions are at https://api.intelligentinference.ai/v1/chat/completions.
Are the PK NPU-native routes available today?
They are listed in the catalogue with a Coming soon tag and are opening shortly. The global pass-through routes are live on the same endpoint now, with the same keys and the same billing, so an integration built today moves to the PK pool by changing the model string.
How is this different from calling a foreign API?
Three ways. The gateway and its records (keys, budgets, audit log, ledger, telemetry) run in Pakistan. The PK NPU pool executes the model in Pakistan. And the bill is in rupees against a prepaid balance, not a dollar invoice on a foreign card.
What happens when I call a Global route?
The gateway relays the request to partner capacity outside the PK NPU pool and bills it in rupees like every other route. Those routes are marked Global wherever they appear. If a caller must never reach one, scope its key to PK routes and the gateway refuses the rest with a 403.
Which models are on the PK Endpoint?
Today: gpt-oss 120B, Kimi K3, GLM-5.2 on the global pass-through pool. Opening soon on the PK NPU pool: GLM-5.3, GLM-5.3-Flash, GLM-5.2 Fast, Kimi K2.7 Code, DeepSeek-V4-Flash-0731, DeepSeek V4 Pro 0813, DeepSeek V4 Pro, Nvidia Nemotron 3 Ultra, Inkling, Inkling-Small, gpt-oss 120B, Qwen3.8 27B, Qwen3 Embedding 8B. The full list, with per-million-token prices in rupees, is on the Models section and in the dashboard.
Do I need to change my integration?
Only the base URL and the model string. The request and response shapes are OpenAI-compatible, streaming included, so an existing client library works as it is.
Can I limit what a key can spend or reach?
Yes. Each key can carry an expiry, a lifetime credit limit in rupees and a list of allowed models. The gateway enforces all three: 401 after expiry, 402 at the cap, 403 outside scope. The rest of the organisation's balance is untouched by any one key hitting its limit.
What does the gateway record about my requests?
Request telemetry for your own analytics and billing: timings, token counts, the route and the key, kept in Pakistan and shown back to you on the dashboard. Retention terms for request content are published in the documentation.
What currency am I billed in?
Pakistani rupees, end to end. Rates are quoted per million tokens in PKR, the balance is prepaid in PKR, and the gateway reports the exact rupee cost of each request in its debug metrics.
Can I reserve NPUs for my organisation alone?
Yes, as dedicated inference: Ascend NPUs held for one organisation, billed per hour of reserved capacity and quoted per organisation. See Dedicated inference.
Who is the PK Endpoint for?
Teams with residency, procurement or currency constraints: anyone handling customer data, internal knowledge or regulated workflows who needs to say where the model ran, and anyone whose finance team would rather not settle an AI bill in dollars.

Let’s start

Start building with AI while your data stays in Pakistan.

One OpenAI-compatible API, a sovereign gateway with per-key scope and caps, NPU-native routing in Pakistan, and a bill in rupees. Get a key now, or be first to know when the PK NPU pool opens.

Same API. NPU-native routing. Data stays in Pakistan. Questions: info@intelligentinference.ai