Skip to content

Bootstrap prod Azure inference + deploy the 3 models at a high rate limit #34

Description

@Kenny-Heitritter

Question

In qbraid-infrastructure terraform/environments/prod/azure/ (today foundation-only: RG qbraid-prod + Log Analytics, no cognitive accounts), bootstrap a prod AI Foundry cognitive account (qbraid-prod-ai-services) mirroring staging's main.tf / monitoring.tf / outputs.tf, and deploy Sol/Terra/Luna with a much higher sku_capacity than staging.

Decide the prod rate-limit target per model, bounded by the quota ceiling from the discovery ticket. Apply through the (gated) prod CI pipeline.

Not user-facing: the gateway will not route prod traffic to these deployments in this effort — this ticket provisions prod, it does not expose it.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    wayfinder:taskWayfinder ticket: manual unblocking task

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions