Secrets management
Every key ai-studio-secrets needs, the vault mapping, and per-cloud SecretStore configuration.
Genesis Downloads · Platform values
A production genesis-platform values file runs to several
hundred lines, but roughly fifty values are actually
environment-specific. Define those once as YAML anchors at the
top; everything below references them. This page separates the values
you must change from the ones you should leave alone.
ai-studio-secrets — see
Secrets management. What lives in
the values file is where things are: hostnames, registry, vault
URL, deployment names. If you find yourself typing a secret into this
file, it belongs in your vault instead.
Three parts, in this order. Understanding the split is what keeps the file maintainable.
| Block | Contains | How often you touch it |
|---|---|---|
x-common |
Anchor definitions only — every environment-specific fact, declared once | Every new environment. This is the block you edit |
global |
Cross-service settings: routing, database, redis, secrets, identity. Mostly aliases into x-common |
Rarely — only to change a platform-wide behaviour |
genesis-* |
One block per service: waves, databases, env vars, probes, tolerations | Rarely — only to enable a feature or diverge one service |
x-common key is not a chart value — Helm ignores
it. It exists purely to hold anchors, so a fact is stated once and
aliased everywhere it applies. That gives you a useful escape hatch:
to diverge a single service, replace its alias with a
literal and leave every other reference on the anchor.
The pattern looks like this:
x-common:
routingHost: &routingHost "genesis.<env>.<your-domain>"
global:
routing:
mode: gateway-api
host: *routingHost # alias — follows the anchor
genesis-idp:
extraEnv:
- name: KC_HOSTNAME
value: *keycloakExternalUrl
If you copy a values file and change nothing else, change these. Every one is a fact about your infrastructure, and a stale value here fails the install or silently points the platform at someone else's environment.
| Anchor | Shape | What it drives |
|---|---|---|
routingHost | genesis.<env>.<your-domain> | The single ingress host. Keycloak's issuer is derived from it, so getting this wrong breaks auth everywhere |
baseUrl | https://<routingHost> | The https:// form. YAML cannot build it from routingHost, so it is stated separately — change both together |
keycloakExternalUrl | https://<routingHost>/auth | Browser-facing Keycloak base |
ccFePublicUrl | https://<routingHost>/command-center | Command Center frontend |
tenantMgmtPublicUrl | https://<routingHost>/tenant-mgmt | Tenant-management UI |
temporalCodecUrl | https://<routingHost>/studio/temporal/codec | Payload codec for the Temporal UI |
feRedirectUris | list: https://<routingHost>/* | Keycloak client redirect allow-list. A stale entry here is an auth loop with no useful error |
These sit in global rather than x-common, and
every chart default here is a local-development placeholder.
| Value | Chart default | Set it to |
|---|---|---|
global.environment | "development" | One of development, sprint, dev, integration, production. Easy to leave on development in a UAT or production environment, because nothing enforces it |
global.routing.mode | "single-host" | gateway-api if you route with Gateway API, otherwise single-host |
global.routing.host | "ai-studio.local.dev" | *routingHost — alias it, don't restate it |
global.ingress.className | nginx | Your controller: alb for Application Gateway for Containers, gce for GKE Gateway, nginx for ingress-nginx |
global.ingress.tlsSecretName | unset | The Secret holding your TLS certificate |
enabled: false with
className: nginx and an empty host. If you route through
Gateway API you leave ingress off deliberately; if you expect an Ingress
and never set enabled: true, the platform comes up with no
external route and nothing reports it as an error.
| Anchor | Shape | Notes |
|---|---|---|
imageRegistry | <your-registry> | The registry holding every service image |
initImage | <your-registry>/curl:8.5.0 | Wait-for and Keycloak-setup init container. Needs curl and a shell |
dbInitImage | <your-registry>/postgres:16 | Schema-creation init container. Pin a real tag — :latest works but makes the init step irreproducible across syncs |
initImage and dbInitImage embed the registry
host in a full image reference, and YAML cannot interpolate one anchor
into another. Changing imageRegistry alone leaves
both init images pointing at the previous registry, which
fails as an ImagePullBackOff on an init container —
so the pod never starts and the real service logs nothing. Change all
three in one edit. Some charts embed it a fourth time in an image
repository field; grep for the old host before you sync.
| Anchor | Shape | Notes |
|---|---|---|
dbHost | <your-server>.postgres.database.azure.com | Your managed Postgres endpoint |
dbPort | 5432 (integer) | Used by global.database.port |
dbPortStr | "5432" (string) | Used by env vars. Not interchangeable with dbPort |
dbUsername | <admin-login> | Server admin login |
dbAuthMethod | password or managed identity | Flip here to move the whole platform between auth modes |
dbMigrateUser | <migrate-role> | The role Alembic runs migrations as |
dbAiStudio | ai_studio | Shared database for backend, runtime, workers, connectors |
dbAuthz | genesis_authz | Database for authz and the OpenFGA store |
The last two are database names. They only need changing if your naming convention differs — but they must match the databases preflight created.
Two are anchors; the rest sit in global.redis. They are all
Tier 1 because the Redis tier you provisioned as a prerequisite
determines every one of them, and the chart's defaults match
only the simplest case.
| Value | Chart default | Set it from |
|---|---|---|
redisHost (anchor) → global.redis.host | "" | Your instance's endpoint |
redisAuthMethod (anchor) | password | Access-key auth. A separate fact from dbAuthMethod — they need not agree |
global.redis.port | 6379 | Your endpoint. See the table below — the default is correct on AWS and in-cluster, wrong on Azure and GCP |
global.redis.tls | false | true for any managed instance. Prerequisites require TLS in transit |
global.redis.sslVerify | "false" | Whether to verify the server certificate. A string, not a boolean |
global.redis.mode | unset | standalone or cluster. Not in the chart defaults at all, so a clustered instance needs it added explicitly |
global.redis.db | — | Logical database index. Clustered Redis supports only 0 |
tls: false and
sslVerify: "false", while prerequisites ask for
auth plus TLS in transit. Those two are wrong
for every managed product — leave them and the platform
talks plaintext to a TLS-only endpoint.
6379 is correct for
AWS ElastiCache and for in-cluster Redis, and wrong for both Azure
products and GCP. So do not reason from “TLS means not
6379” — take the port from the table below.
Values by product, each taken from a working deployment. The endpoint tells you which row you are on — the two Azure Redis products in particular are different services with different ports, and picking the wrong row is the most common Redis misconfiguration.
| Product | Host suffix | Port | tls | mode |
|---|---|---|---|---|
| In-cluster Redis — dev only (the chart default) | service DNS | 6379 |
false |
unset |
| AWS ElastiCache, transit encryption on | .cache.amazonaws.com |
6379 |
true |
standalone |
| Azure Cache for Redis (classic) | .redis.cache.windows.net |
6380 |
true |
standalone |
Azure Managed Redis (Microsoft.Cache/redisEnterprise) |
.redis.azure.net |
10000 |
true |
cluster |
| GCP Memorystore, TLS enabled | private IP | 6378 |
true |
standalone |
.redis.cache.windows.net:6380 is classic Azure Cache for
Redis. .redis.azure.net:10000 is Azure Managed Redis, a
different resource type. If you provisioned the enterprise tier and
copied the classic port — or the reverse — the connection
fails at startup. Match the port to the host suffix your endpoint
actually has. Clustered instances also require db: 0;
no other index is supported.
6379, the same port as plaintext Redis. On AWS the
port stays at the default and only tls changes, so a
“the port looks like the default, TLS must be off” reading
is wrong there.
The reference customer values file starts at port: 6380,
tls: true, auth: true — one managed shape,
not the only one. Change the port to match your own endpoint.
Leave enabled: true: the platform degrades badly without
Redis, which handles sessions, rate limiting and SSE streaming.
These chart values take Redis as host plus port.
If you are also setting the CLI's own config, that takes a
rediss:// URL instead and has its own formatting trap —
see the Redis notes in the install guide.
global.redis.caCert while sslVerify is on
is caught at render time and fails the sync —
deliberately. The gateway's rate-limit config is rendered by a subchart
that derives verification from sslVerify alone, so it would
verify against its image's system trust store, fail, and
silently stop counting daily API quotas. The guard
exists because that failure is invisible. Resolve it as the error message
directs — either let the umbrella render that config so
global.redis.gatewaySslVerify takes effect, or turn
sslVerify off.
| Anchor | Shape | Notes |
|---|---|---|
keyVaultUrl | https://<your-vault>.vault.azure.net/ | The vault the ExternalSecret pulls from |
secretName | ai-studio-secrets | The Secret the ExternalSecret writes and every service reads. Rarely changed — but if you do, it must change everywhere |
serviceAccountName | <your-sa> | A pre-existing ServiceAccount, created outside the chart (create: false), carrying the workload-identity annotations |
azureClientId / azureTenantId | GUIDs | Workload-identity federation. Annotation-only, and inert while the ServiceAccount is create: false |
| Anchor | Shape | Needed when |
|---|---|---|
azOpenAiEndpoint | https://<your-openai>.openai.azure.com/ | Azure OpenAI is your LLM backend |
azSearchEndpoint | https://<your-search>.search.windows.net | Azure AI Search is your vector store |
azDocIntelEndpoint | https://<your-cognitive>.cognitiveservices.azure.com/ | Document Intelligence handles OCR |
storageAccount | <your-storage-account> | Object storage. Not the same fact as imageRegistry, even when the names look alike |
storageContainer | <container> | The container inside that account |
otelHttpEndpoint points at your trace collector. It is
easy to leave pointing at a shared or previous environment's collector,
which is not an error — traces simply land somewhere you are not
looking.
These are not environment identity, so a copied file often "works". But they encode a contract between values, and a mismatch fails at runtime rather than at install.
| Anchor | Shape | The constraint |
|---|---|---|
llmProvider | azure_openai | Sets the default and primary provider, and the retrieval provider. One value, several consumers |
azOpenAiApiVersion | e.g. 2025-01-01-preview | Must be an API version your Azure OpenAI resource actually serves |
azEmbeddingDeployment | text-embedding-3-small | A deployment name in your resource, not a model name |
vectorSize | "1536" for text-embedding-3-small | Must equal the embedding model's dimension. A mismatch is accepted at install and then fails on every write to the vector store |
genesisEmbeddingsModel | e.g. BAAI/bge-large-en-v1.5 | Served by the embeddings service. The service's MODEL_ID and every client must name the same model |
modelGpt41, modelGpt41Mini, … | deployment names | Each names one deployment that must exist in your resource. Anchored per model, not per role, so every reference moves together |
builderModel | <provider>:<model> | A compound value. YAML cannot build it from the provider and model anchors, so it must be updated by hand when either changes |
global.capacityGate is the one Tier 2 value that
fails the render rather than failing later. It declares
how many pods will consume the database, and the chart asserts your
declaration against the footprint it actually renders.
| Value | Chart default | Change it when |
|---|---|---|
expectedApiPods | 4 | Your API replica counts differ from the subchart defaults (genesis-be 2 + genesis-be-runtime 2) |
expectedWorkerPods | 2 | You enable or disable worker variants — the per-profile io / mixed / cpu splits are off by default and each one you enable joins the sum |
externalDbConsumers | 55 | You share the database with services beyond the platform. The default covers Temporal (~40 connections) and Keycloak's pool (~15) |
The floor it enforces at pod startup:
floor = (expectedApiPods × api_ceiling)
+ (expectedWorkerPods × worker_ceiling)
+ externalDbConsumers
+ 23 reserved
helm template,
helm install and ArgoCD sync, and refuses to render
when your declared numbers diverge from the rendered
footprint. That is deliberate: it catches drift inside your
cluster rather than in our CI, which matters when you overlay your own
values onto the shipped chart. If you enable autoscaling you
must raise these to the maxReplicas peak summed
across enabled subcharts — not the steady-state count. And if you
add your own subchart that consumes either pool, add it to
apiShapeCharts or workerShapeCharts, or the
gate will not count it.
Values that look environment-specific but are not. Editing them is how a working file starts drifting.
| Value | Why it stays |
|---|---|
In-cluster service URLs — http://genesis-authz:8000, http://genesis-kc:8000/kc/api/v1, http://genesis-embeddings:7997 | Kubernetes service DNS, fixed by the charts. They are not environment identity, and rewriting them to public URLs routes internal traffic out through your gateway and back |
| Gateway upstreams | Same reasoning — they name in-cluster services |
deploymentOrder.wave | Encodes the real startup dependency order. Services init-wait on Keycloak and on each other; reordering causes timeouts that look like unrelated failures |
| Probe paths and thresholds | Tuned to each service's real startup time. A long startupProbe.failureThreshold is deliberate, not a leftover |
Feature flags — KAFKA_ENABLED, OTEL_EVENTHUB_ENABLED, APIM_ENABLED | Off by default. Turning one on requires its own configuration and its vault keys; flipping the flag alone gets you a service that fails to reach a backend that was never set up |
enabled: false on unused services | Deliberately off. Enabling one pulls in its dependencies and its required secret keys |
| Rate limits, replica counts, resource requests | Sized for a working deployment. Change them for capacity reasons, not as part of environment setup |
Each of these has produced a failure that looked like something else.
a.b.c: value does not set the
nested key a → b → c.
Helm reads it as a single key literally named a.b.c, finds
nothing expecting it, and ignores it silently. There is
no warning and no error — your setting simply never applies.
Always write the nesting out.
https://<host>/path from a host anchor, and you
cannot build <registry>/curl:8.5.0 from a registry
anchor. Every compound value restates its parts in full, which is why
the public URLs, the init images and builderModel all have
to be changed alongside the anchor they appear to derive from.
<<:, and each variant
states only its differences. Add a shared setting to the base block,
not to each variant — and remember an explicit key in a variant
overrides the merged one.
x-common:
devWorkerBase: &devWorkerBase
enabled: true
replicaCount: 2
devWorkerEnv: &devWorkerEnv
REDIS_ENABLED: "true"
DEFAULT_PROVIDER: *llmProvider
genesis-dev-worker-io:
<<: *devWorkerBase # everything shared
env:
<<: *devWorkerEnv # shared env
WORKER_PROFILES: "io" # ...then only what differs
env entry overrides the same name arriving
via envFrom on the platform Secret. That is occasionally
intended — but it also means a leftover literal in a service block
can quietly shadow the value your vault is supplying.
Some values are neither environment identity nor a safe default —
they are open questions that a copied file carries forward invisibly.
Mark them in your own file (a consistent
# ACTION: comment works well) so they can be searched for
and closed.
| Pattern | Why it needs a decision |
|---|---|
| A URL pointing at another environment | Perfectly valid YAML, and it resolves. Shared or upstream services are sometimes deliberately cross-environment — but a copied file is the usual reason, and it is worth confirming which |
| Placeholder tokens | A dummy bearer token renders and deploys. The dependent feature then fails on first use, well after the install is declared green. Either supply the real value through your vault or disable the feature |
| Object-storage settings left unset | Account and container are set, but the storage-type selector is not, so the service falls back to a local path. Nothing errors; files just do not land where you expect |
| Hard-coded identifiers inside a URL | A client or copilot id embedded in a base URL is environment-specific in a place nobody looks. Pull it into an anchor so it is visible |
| A secret mapping pinned to a subset | Trimming secrets.data to the keys your vault holds is a legitimate way to avoid an all-or-nothing sync failure — but every reference you drop must stay optional: true on the consuming service. See Secrets management |
| A localhost origin in a CORS allow-list | Useful in development, and it should not reach production |
| Check | What it catches |
|---|---|
| Grep the file for the previous environment's hostname, registry and resource names — expect zero hits | The single highest-yield check. Compound values that restate an anchor are exactly what a partial rename leaves behind |
| All three registry-bearing values name the same registry | ImagePullBackOff on an init container, where the service itself logs nothing |
| Every Azure OpenAI deployment name exists in the resource | A healthy platform whose model calls all 404 |
vectorSize matches the embedding model's dimension | Vector writes failing after a clean install |
| The redirect allow-list contains the new host | An auth redirect loop with no useful error |
| Vault URL, ServiceAccount name and workload-identity ids are yours | An ExternalSecret that never syncs, so every pod sits in CreateContainerConfigError |
Redis port matches your endpoint, and tls is on | tls: false is wrong for every managed product. The port varies: 6379 on AWS, 6380 or 10000 on the two Azure products, 6378 on GCP |
global.environment is not still development | Nothing enforces it, so it silently ships |
global.ingress.className matches your controller, or ingress is off on purpose | A platform with no external route and no error |
capacityGate numbers match your replica counts and autoscaling peaks | A render that refuses to sync — the one failure here that is loud |
No dotted keys — search for . in key position | Settings that are silently ignored |
| No credentials anywhere in the file | A secret committed to git history |
| Every open decision is marked and triaged | Cross-environment URLs and placeholder tokens shipping to production |
helm template with your values catches schema errors and
renders the manifests, but none of the checks above — each one is
valid YAML that renders cleanly and fails later.
Every key ai-studio-secrets needs, the vault mapping, and per-cloud SecretStore configuration.
How this values file is delivered — ArgoCD multi-source, with values held in your own git repo.
The Postgres, Redis, DNS, TLS and registry facts this file names must exist first.