OCI OKE · Kubernetes · GitOps · FinOps · Multi-Tenant SaaS

CloudOps Insight — GitOps & FinOps Platform

A production multi-tenant CloudOps and FinOps platform that securely connects customer AWS accounts, synchronizes cloud inventory, health, metrics, and cost data, and turns the result into operational dashboards, budgets, incidents, recommendations, and resource-level insight.

Production runtime: OCI Kubernetes Engine in Frankfurt. AWS is integrated as the customer cloud being monitored through short-lived cross-account access.

Project overview

CloudOps Insight is designed as a real platform rather than a single dashboard. It combines application engineering, cloud integrations, asynchronous workers, multi-tenant authorization, GitOps delivery, Infrastructure as Code, recovery controls, and FinOps workflows in one end-to-end system.

3 separate repositories
4 workspace RBAC roles
2 container CPU architectures
1 production OKE platform

Cloud operations

Connect AWS accounts, synchronize resource inventory, inspect health and operational state, search resources, and surface incidents and recommendations.

FinOps

Synchronize AWS Cost Explorer data and present service-level costs, budget information, cloud usage analysis, and optimization-oriented recommendations.

Platform engineering

Build and release immutable multi-architecture images, promote desired state through GitOps, reconcile with Argo CD, and manage the OCI foundation with Terraform.

System architecture

The platform separates user traffic, application services, asynchronous cloud synchronization, customer AWS access, GitOps delivery, and infrastructure ownership.

User Browser → Cloudflare Tunnel → React + TypeScript → FastAPI
FastAPI → PostgreSQL + Redis → Celery Workers + Celery Beat
Customer AWS Account ← STS AssumeRole + External ID ← boto3 Integration
AWS Inventory + CloudWatch + Cost Explorer → CloudOps Data Model

Application repository

Owns the React frontend, FastAPI API, authentication and tenant logic, AWS integrations, Celery jobs, domain models, migrations, tests, and application CI.

GitOps repository

Owns Kubernetes desired state, Helm values, Argo CD applications, Sealed Secrets, environment configuration, observability integration, and production releases.

Infrastructure repository

Owns Terraform for OCI OKE, VCN networking, network security, persistent storage, Object Storage, Vault-aware infrastructure, and backup/recovery foundations.

GitOps delivery model

Application delivery is intentionally separated from cluster reconciliation. Application CI produces immutable artifacts; GitOps controls what production should run.

Git Push → GitHub Actions → Tests + Trivy → Multi-Arch Build
linux/amd64 + linux/arm64 → GHCR SHA Image → GitOps Desired State
GitOps Repository → Argo CD → Helm Release → OCI OKE

Immutable releases

Application images are released using commit-SHA identifiers rather than mutable production tags, improving traceability and rollback confidence.

Multi-architecture delivery

CI publishes linux/amd64 and linux/arm64 images so the same release process can support traditional x86 environments and the ARM-based OCI production worker.

Declarative reconciliation

Argo CD continuously reconciles the cluster with reviewed GitOps configuration, keeping runtime changes traceable to version-controlled desired state.

Multi-tenant security and AWS onboarding

The security model protects both tenant boundaries inside CloudOps Insight and cross-account access to customer AWS environments.

Workspace authorization

  • Owner / admin / member / viewer workspace RBAC
  • Workspace invitations and membership lifecycle
  • Tenant-isolated AWS accounts, resources, costs, incidents, and recommendations
  • Email verification and password reset lifecycle
  • Account deletion and cloud-integration lifecycle controls

AWS cross-account access

  • STS AssumeRole instead of stored customer access keys
  • Generated External ID per integration
  • Read-only customer IAM role model
  • Short-lived AWS credentials
  • Validation before synchronization

Platform security

  • Sealed Secrets for GitOps-managed Kubernetes secrets
  • Cloudflare Tunnel for HTTPS application exposure
  • Secure cookies and controlled CORS configuration
  • Rate limiting and authenticated application flows
  • Trivy scanning in CI

Tenant isolation audit

The platform was explicitly tested to prove that one customer workspace cannot access another customer's AWS accounts, inventory, costs, incidents, or recommendations.

Customer credential strategy

CloudOps Insight does not require customers to store AWS access keys in the application. The onboarding model uses role assumption and External IDs instead.

FinOps, asynchronous operations, and recovery

The application combines background synchronization with operational and financial workflows while maintaining explicit backup and restore controls.

AWS synchronization

Celery workers execute cloud synchronization tasks outside the request path, while Celery Beat schedules periodic work for resource, metric, and cost updates.

Cloud inventory and metrics

The AWS integration collects resource inventory across services such as EC2, RDS, ECS, ELB, and S3 and enriches operational views with CloudWatch-derived information.

FinOps workflows

Cost Explorer synchronization supports cost analytics, service-level spend, budgets, usage analysis, and optimization-oriented recommendations.

PostgreSQL logical backup

Kubernetes CronJob-based PostgreSQL logical backups protect application data and integrate with OCI storage workflows for recoverability beyond the running database.

Infrastructure-level protection

OCI Block Volume backup policy adds storage-level protection, complementing the logical database backup path.

Restore drill

Recovery was not treated as complete when a backup merely existed. A restore drill was executed to validate that PostgreSQL data could actually be restored.

User lifecycle validation

A fresh-user flow was tested end-to-end: signup, email verification, login, workspace creation, AWS connection, role validation, synchronization, resource and cost views, disconnect, and account deletion.

Technology and version snapshot

Representative versions from the production application and infrastructure repositories at the time of this case study.

Backend

  • Python 3.13
  • FastAPI 0.141.1
  • SQLAlchemy 2.0.51
  • Celery 5.6.3
  • boto3 1.43.72
  • Alembic 1.19.1

Frontend

  • Node.js 22
  • React 19.2.8
  • React Router 7.18.2
  • Vite 8.2.0
  • TypeScript 6.0.3
  • Tailwind CSS 4.3.3
  • ECharts 6.1.0

Runtime and infrastructure

  • PostgreSQL 17-alpine
  • Redis 7-alpine
  • Terraform CI 1.15.8
  • OCI Provider 8.29.0
  • Kubernetes IaC target v1.35.2
  • Cloudflared 2026.9.0
  • Sealed Secrets chart 2.20.0

Engineering decisions and trade-offs

The current production platform is deliberately cost-conscious. The case study documents both the engineering strengths and the places where a larger commercial deployment would use more redundancy.

Three repositories instead of one deployment repository

Application code, Kubernetes desired state, and cloud infrastructure are separated. This creates clearer ownership boundaries between software delivery, GitOps reconciliation, and Terraform-managed infrastructure.

OCI runtime + AWS customer integrations

The SaaS control plane runs on OCI OKE while customer AWS accounts are monitored through cross-account IAM. This keeps the platform runtime independent from the customer cloud integration model.

ARM production worker

Production runs on an OCI Ampere A1 worker, which made multi-architecture container builds a real delivery requirement instead of only a CI exercise.

Cost-conscious stateful services

PostgreSQL and Redis currently run inside the cluster rather than as managed, highly available services. This reduces cost but is documented as a production limitation rather than hidden behind a "production-ready" label.

Current limitations and roadmap

The platform is production live, but the current deployment intentionally favors a low-cost student-operated footprint over full high availability.

Current limitations

  • Single ARM worker node in the current production cluster
  • In-cluster PostgreSQL rather than a managed HA database
  • Redis is not deployed as a durable managed cache/service
  • The runtime is not currently multi-node / multi-zone highly available
  • Disaster recovery is validated, but not an automated cross-region failover system

Next engineering steps

  • Move production toward multi-node high availability
  • Externalize PostgreSQL and Redis to managed production-grade services
  • Add stronger SLO/SLA-oriented observability and alerting
  • Expand audit trails and tenant-level operational reporting
  • Continue commercial hardening, billing, and platform reliability work