Enhancing GitOps with Keptn for Autonomous Operations
Problem Statement
To enable autonomous operations, the existing GitOps framework must expand to accommodate closed-loop automation without compromising core GitOps principles.
Currently, Cloud Native (CN) workloads and their operating environments lack the integrated metrics and control loops needed for self-management. Implementing autonomous operations capabilities necessitates extending and enhancing existing GitOps work-flows. Incorporating structured data flows following the pattern metrics input --> inferencing logic --> configuration action allows for adding closed control-loops to existing K8s-FluxCD-GIT environments, thusly enabling automomous operations of CN workloads.
Description
Develop a Kubernetes Operator that compiles the logic of suitable Keptn tasks from sets of largely predefined configuration items (metrics, baseline values, decision making logic, etc.) so as to create the 'closed control loops' facilitating the autonomous operations of CN workloads.
Project Details
Leader: Deutsche Telekom (M. Sewera)
List of people/organization interested to join:
- n.n.
Consultancy relating to Keptn provided by:
- Cuemby, L.L.C (A. Ramirez / A.-W. Jagau) *
Use Cases
Automate Post-Deployment Functional and Performance Verification for 5GC Software (suggested by Telekom Deutschland)
After deploying a new 5G Core (5GC) software release via the GitOps pipeline, we want to ensure the network functions are fully operational and handling traffic correctly. We want to accomplish this by using Keptn to orchestrate an automated validation process that bridges the gap between basic Kubernetes pod health and 5G protocol-level verification through the following three-step verification loop:
- Triggering the Functional Tests (Pre-Conditions): Once the GitOps controller rolls out the updated AMF, SMF, or UPF pods, Keptn intercepts the deployment process before it is finalized. Keptn automatically executes localized task definitions to trigger specific testing tools (i.e. traffic simulators or probes), forcing the network to execute end-to-end control and user plane signaling, including AMF registration, PDN session establishment, and user plane traffic forwarding via the UPF.
- Evaluating 5GC KPIs (Quality Gates): While the test traffic is running, Keptn serves as an automated quality gate by directly querying the monitoring system (such as Prometheus). It utilizes declarative metrics evaluation definitions to scrape actual 5G network performance indicators over the test window, focusing specifically on session establishment rates, the percentage of successful session setups, and packet delivery success.
- Automated Judgment & GitOps Rollback: Keptn analyzes the gathered performance metrics against strict, pre-defined SLA thresholds. If all 5G network KPIs meet the required quality criteria, Keptn approves the deployment and allows the GitOps state to finish successfully; if any protocol procedures or success percentages fall below the baseline, Keptn fails the quality gate, flagging the environment to trigger an immediate, automated GitOps fallback to the previous stable software version.
Make the Deployment of 5G-Core Services Predictable and Reliable
NOTE: The feasibility of this use case has been checked with requests sent to the AI agents Claude Code, ChatGPT, and Gemini.
UC Summary
Before installing or upgrading a 5G-Core service (AMF, `SMF, NSSF, UPF, etc.), we want to be sure that the service is ready for deployment and that the cloud environment is ready and prepared to receive it. Once the service is up, we want to make sure that it has all required resources allocated to it and that it has started processing payload traffic. And, to make sure the operating service stays healthy, we want to periodically check KPIs and, if neccessary, adjust the day-2 configuration.
To ensure this end-to-end deployment scenario is entirely covered, we distinguish three phases in the proposed use-case:
- "Cold Start" -- Deployment & Initialization
- "Warm Up" -- Integration & Initial Traffic
- "Full Swiing" -- Steady State Operations
In fact, each of these three phases could be regarded as a keptn-focussed use-case in their own right.
UC Details
Using the NSSF 5G-Core service as a specimen for the implementation, the paragraphs in this section detail further each one of the phases named above:
Deployment & Intializaiton: To ensure the NSSF service is ready for deployment, we want for the system to automatically check and verify a number of key deployment pre-conditions (e.g., the updated core service image is loadable, its day-0/day-1 configuration is complete, it will be deployed to the appropriate cluster, it will operate in ‘Canary’ mode).
Similarly, automatic checks are run to ensure that the targeted cluster features sufficient compute and memory resources. And that the networking configuration meets the requirements of the considred NSSF 5G-Core service.
Of course, the above is but an exemplary selection of conceivable pre-deployment checks, that could be performed prior to clearing an NSSF service for deployment.
Integration & Initial Traffic: Once the deployment of the updated NSSF core-service is complete, we want the system to automatically check that all deployment pre-conditions have indeed been met and that the updated services is operational and healthy, i.e., it delivers the expected metrics.
Suitable checks may include but are not limited to evaluating HTTP response codes or the rate at which requests are coming in. Verifying the actual consumption of compute or memory resouces are yet further examples of checks for confirming that the NSSF service has come up and is fully operational.
Steady State Operations: Leaving the overall deployment phase and entering a continuous service monitoring mode, we want that the system automatically starts gathering and evaluating a set of predefined metrics. Metrics that are designed to establish that the NSSF service continues to operate in a healthly and performant fashion.
Depending on the capabilities of the 5G-Core environment available for implementing the use-case, the NSSF replication balance (Do all NSSF replicas handle an equal amount of work?) may be considered. Even in a bare-minimum 5G-Core environment, comparing the resource utilization of the NSSF service against the volume of incoming request should give some indication of the NSSF operating health.
UC Tooling and Lab Environment
The following tools shall be required specifically for implementing the use-case:
- an 'Open 5G-Core' solution like
free5GCinstalled and configured, - a traffic generator like
UERANSIMinstalled and configured, as well as - an HTTP request generator like
Postmaninstalled and configured.
For further details concerning the environment required for implementing the use-case see Project Lab below.
UC-Assumptions
This use-case requires some prior knowledge about 5G mobile networking and the NSSF 𝛍-service underlying a '5G-Core' solution.
Furthermore, having worked with keptn and/or Prometheus would be a definite plus. In the least, project team members working on this use-case should have familiarized themselves with these two tools prior to the Hackathon.
UC-Implementation Hints
A review of the characteristics associated with the NSSF 5G-Core service suggests that the input data and the logic required for implementing the keptn decision making for the three deployment phases can easily be derived from fairly basic knowledge about a service based architecture.
For this reason, we recommend that the focus, while implementing this use-case, be placed on the creation of the required keptn and PromQL artefacts.
Project Elements
Project Lab
The project will require a vanilla Kubernetes-based cloud environment featuring FluxCD and access to a Git data base. Furthermore, tools such as Keptn V.2, Prometheus or OpentTel, and KubeBuilder shall be installed and operational.
The lab environment shall be set up and configured before the start of the Hackathon. And it shall provide access for all Hackathon participants with a minimum of administrative overhead and commonly used access credentials (e.g., SSH private key access).
Keptn Task Design
A good portiton of the time alloted to the Hackathon will be spent on conceptualizing the closed control loop patterns required for enabling the autonomous deployment of selected cloud-native workloads. This work will yield specifications for Keptn tasks and observability related objects (data scraping agents, PromSQL queries, etc.)
Use of AI Platforms to Accelerate Coding and Testing
To ensure the delivery of demonstrable artefacts as proof-of-concept for the agreed Keptn task designs and related objects, the project team shall rely on AI tools such as Claude Code or Gemini for the creation of program code and associated test cases.
Hackathon Objectives
- Show-case the use of Keptn 'pre-deployment' and 'post-deployment' tasks for making CN workload deployments predictable
- Demonstrate the use of Keptn 'evaluation' tasks for identifying and subsequently mitigating CN workload operating deviations
- Highlight how the Keptn enabled enhancement integrate seamlessly into existing GitOps work-flows
- Make the Deployment of 5G-Core Services Predictable and Reliable
- Explore how Kubernetes Operators can be leveraged for creating the three types of
keptntasks.