Cloud operations
Connect AWS accounts, synchronize resource inventory, inspect health and operational state, search resources, and surface incidents and recommendations.
A production multi-tenant CloudOps and FinOps platform that securely connects customer AWS accounts, synchronizes cloud inventory, health, metrics, and cost data, and turns the result into operational dashboards, budgets, incidents, recommendations, and resource-level insight.
Production runtime: OCI Kubernetes Engine in Frankfurt. AWS is integrated as the customer cloud being monitored through short-lived cross-account access.
CloudOps Insight is designed as a real platform rather than a single dashboard. It combines application engineering, cloud integrations, asynchronous workers, multi-tenant authorization, GitOps delivery, Infrastructure as Code, recovery controls, and FinOps workflows in one end-to-end system.
Connect AWS accounts, synchronize resource inventory, inspect health and operational state, search resources, and surface incidents and recommendations.
Synchronize AWS Cost Explorer data and present service-level costs, budget information, cloud usage analysis, and optimization-oriented recommendations.
Build and release immutable multi-architecture images, promote desired state through GitOps, reconcile with Argo CD, and manage the OCI foundation with Terraform.
The platform separates user traffic, application services, asynchronous cloud synchronization, customer AWS access, GitOps delivery, and infrastructure ownership.
Owns the React frontend, FastAPI API, authentication and tenant logic, AWS integrations, Celery jobs, domain models, migrations, tests, and application CI.
Owns Kubernetes desired state, Helm values, Argo CD applications, Sealed Secrets, environment configuration, observability integration, and production releases.
Owns Terraform for OCI OKE, VCN networking, network security, persistent storage, Object Storage, Vault-aware infrastructure, and backup/recovery foundations.
Application delivery is intentionally separated from cluster reconciliation. Application CI produces immutable artifacts; GitOps controls what production should run.
Application images are released using commit-SHA identifiers rather than mutable production tags, improving traceability and rollback confidence.
CI publishes linux/amd64 and linux/arm64 images so the same release process can support traditional x86 environments and the ARM-based OCI production worker.
Argo CD continuously reconciles the cluster with reviewed GitOps configuration, keeping runtime changes traceable to version-controlled desired state.
The security model protects both tenant boundaries inside CloudOps Insight and cross-account access to customer AWS environments.
The platform was explicitly tested to prove that one customer workspace cannot access another customer's AWS accounts, inventory, costs, incidents, or recommendations.
CloudOps Insight does not require customers to store AWS access keys in the application. The onboarding model uses role assumption and External IDs instead.
The application combines background synchronization with operational and financial workflows while maintaining explicit backup and restore controls.
Celery workers execute cloud synchronization tasks outside the request path, while Celery Beat schedules periodic work for resource, metric, and cost updates.
The AWS integration collects resource inventory across services such as EC2, RDS, ECS, ELB, and S3 and enriches operational views with CloudWatch-derived information.
Cost Explorer synchronization supports cost analytics, service-level spend, budgets, usage analysis, and optimization-oriented recommendations.
Kubernetes CronJob-based PostgreSQL logical backups protect application data and integrate with OCI storage workflows for recoverability beyond the running database.
OCI Block Volume backup policy adds storage-level protection, complementing the logical database backup path.
Recovery was not treated as complete when a backup merely existed. A restore drill was executed to validate that PostgreSQL data could actually be restored.
A fresh-user flow was tested end-to-end: signup, email verification, login, workspace creation, AWS connection, role validation, synchronization, resource and cost views, disconnect, and account deletion.
Representative versions from the production application and infrastructure repositories at the time of this case study.
The current production platform is deliberately cost-conscious. The case study documents both the engineering strengths and the places where a larger commercial deployment would use more redundancy.
Application code, Kubernetes desired state, and cloud infrastructure are separated. This creates clearer ownership boundaries between software delivery, GitOps reconciliation, and Terraform-managed infrastructure.
The SaaS control plane runs on OCI OKE while customer AWS accounts are monitored through cross-account IAM. This keeps the platform runtime independent from the customer cloud integration model.
Production runs on an OCI Ampere A1 worker, which made multi-architecture container builds a real delivery requirement instead of only a CI exercise.
PostgreSQL and Redis currently run inside the cluster rather than as managed, highly available services. This reduces cost but is documented as a production limitation rather than hidden behind a "production-ready" label.
The platform is production live, but the current deployment intentionally favors a low-cost student-operated footprint over full high availability.