Skip to content

Connect LlamaIndex to Gonka Broker

LlamaIndex is a framework for building RAG pipelines and agents over your own data. Its OpenAI-compatible LLM class accepts a custom base URL, so you can point it at Gonka Broker and run retrieval and agents on open-source models: the only change from a standard OpenAI setup is the base URL, the key, and the model id.

The integration is the same in Python and JavaScript/TypeScript; only the package name and the constructor differ. Pick your language below.

  • A Gonka API key (starts with gnk-prx-). See Create a Gonka API Key.
  • Python 3.9+ or Node.js 18+, depending on your stack.

Install the OpenAI-compatible LLM integration:

Terminal window
pip install llama-index-llms-openai-like

Point OpenAILike at Gonka Broker:

from llama_index.llms.openai_like import OpenAILike
llm = OpenAILike(
model="moonshotai/Kimi-K2.6",
api_base="https://proxy.gonkabroker.com/v1",
api_key="gnk-prx-your-api-key",
is_chat_model=True,
context_window=128000,
)
response = llm.complete("Say hello from Gonka Broker.")
print(response)

Use OpenAILike (not the plain OpenAI class); the standard class only accepts OpenAI’s own model names. Keep is_chat_model=True so calls go to the chat endpoint, and set context_window to your model’s window.

The model must match a Gonka-supported id exactly, for example moonshotai/Kimi-K2.6 or MiniMaxAI/MiniMax-M2.7 (see Supported Models). Once the LLM is configured, pass it to a Settings.llm / index / query engine and the rest of LlamaIndex works unchanged.

For the retrieval side of a RAG pipeline, use BAAI/bge-m3 through the same base URL and key (see Embeddings & RAG).

Install the OpenAI-like embedding integration:

Terminal window
pip install llama-index-embeddings-openai-like

Point OpenAILikeEmbedding at Gonka Broker and pass it to your index:

from llama_index.core import VectorStoreIndex, Document
from llama_index.embeddings.openai_like import OpenAILikeEmbedding
embed_model = OpenAILikeEmbedding(
model_name="BAAI/bge-m3",
api_base="https://proxy.gonkabroker.com/v1",
api_key="gnk-prx-your-api-key",
)
index = VectorStoreIndex.from_documents(
[Document(text="Gonka Broker bills in USD per token.")],
embed_model=embed_model,
)
print(index.as_retriever().retrieve("How am I charged?")[0].text)

As with the LLM, use OpenAILikeEmbedding (not the plain OpenAIEmbedding class, which only accepts OpenAI’s own embedding-model names), and note the parameter is model_name, not model.

Run the snippet above. A printed reply confirms LlamaIndex is reaching Gonka Broker through your key.

  • 401 / invalid API key: wrong or paused key. Create a fresh one from Create a Gonka API Key.
  • Model not found / unsupported: the model must match a model Gonka Broker serves exactly (see Supported Models).
  • Connection errors: confirm the base URL is exactly https://proxy.gonkabroker.com/v1.
  • Output is truncated: raise context_window (Python) to match your model; the default is small.
  • “‘BAAI/bge-m3’ is not a valid OpenAIEmbeddingModelType”: the plain OpenAIEmbedding class validates model names against OpenAI’s list. Use OpenAILikeEmbedding from llama-index-embeddings-openai-like (see the embeddings section above).
  • Looking for the model’s thinking: reasoning models return their thinking in a separate reasoning response field, never inside the response text — LlamaIndex output is clean answer text. To skip thinking entirely (shorter, cheaper responses), see reasoning control.