# Welcome

Get started with the Cisco Crosswork NSO documentation.

[**Cisco Crosswork NSO**](https://www.cisco.com/c/en/us/products/collateral/cloud-systems-management/network-services-orchestrator/network-orchestrator-so.html) is a premier orchestration tool for hybrid networks. As a Linux-based application, it allows detailed control of network devices and can manage the configuration of both physical and virtual networks. It automates lifecycle services to help you create and deliver quality services more quickly.

<details>

<summary>NSO History and Background</summary>

A quick look at NSO and how it fits into the Cisco Crosswork suite:

* <i class="fa-clock">:clock:</i> Origin: NSO began its journey as Network Control System (NCS) to simplify network management. Since its accretion by Cisco, NSO has grown into a powerful tool, supporting thousands of devices and automating complex services like 5G and VPNs for companies worldwide.
* <i class="fa-puzzle">:puzzle:</i> Crosswork: In 2018, Cisco launched Crosswork, a broader platform to make networks smarter and self-managing. Crosswork uses NSO as its core engine to handle network setup tasks, but it adds tools for monitoring, analyzing, and fixing issues automatically.
* <i class="fa-link-simple">:link-simple:</i> How NSO and Crosswork Mesh: NSO is the “doer” in Crosswork, following instructions to configure devices. Crosswork builds on NSO by adding real-time insights and auto-fixes, making networks faster and more reliable. NSO can work alone for specific automation tasks, but with Crosswork, it’s part of a bigger, smarter system.
* <i class="fa-rocket">:rocket:</i> How It Matters: NSO and Crosswork make network management easier, saving time and reducing errors. Whether you’re new to automation or managing a huge network, NSO’s flexibility and Crosswork’s intelligence have you covered!

</details>

### Explore What's New in NSO

{% embed url="<https://nso-docs.cisco.com/guides>" %}

### Start Exploring NSO

<table data-view="cards"><thead><tr><th></th><th></th><th data-hidden data-card-cover data-type="image">Cover image</th></tr></thead><tbody><tr><td><strong>NSO Essentials</strong></td><td><ul><li><a href="/pages/q4B2SRfDpBrmfetjmLtF">NSO at a Glance</a></li><li><a href="/pages/m3Al3JrYS2yPsslMv789">Common Use Cases</a></li><li><a href="/pages/nAVsSECCbvm5DmKTfAYB">FAQs</a></li></ul></td><td><a href="/files/L9SJZHKAbtpcXPfEEJ7c">/files/L9SJZHKAbtpcXPfEEJ7c</a></td></tr><tr><td><strong>Quick Start</strong></td><td><ul><li><a href="/pages/YG3OEn6tnD1KC8fINk2d">NSO Quick Start Guide</a></li><li><a href="https://cisco-tailf.gitbook.io/nso-docs/guides/development/introduction-to-automation">Intro to Automation</a></li><li><a href="https://software.cisco.com/download/home/286331591/type/286283941/release">Evaluate NSO</a></li></ul></td><td><a href="/files/uzpzfmWaT9L0bSHesvbJ">/files/uzpzfmWaT9L0bSHesvbJ</a></td></tr><tr><td><strong>Changelogs &#x26; Explorers</strong></td><td><ul><li><a href="https://developer.cisco.com/docs/nso/changelog-explorer/">NSO Changelog Explorer</a></li><li><a href="https://developer.cisco.com/docs/nso/ned-changelog-explorer/">NED Changelog Explorer</a></li><li><a href="https://developer.cisco.com/docs/nso/ned-capabilities-explorer/">NED Capabilities Explorer</a></li></ul></td><td><a href="/files/kpOaa9htEesGKsFdGQrn">/files/kpOaa9htEesGKsFdGQrn</a></td></tr><tr><td><strong>User Guides by Role</strong></td><td><ul><li><a href="https://nso-docs.cisco.com/guides/administration/get-started">Administration</a></li><li><a href="https://nso-docs.cisco.com/guides/operation-and-usage/get-started">Operation &#x26; Usage</a></li><li><a href="https://nso-docs.cisco.com/guides/development/get-started">Development</a></li></ul></td><td><a href="/files/vfcFPIhJOEVNMNTLGu6A">/files/vfcFPIhJOEVNMNTLGu6A</a></td></tr><tr><td><strong>Developer Resources</strong></td><td><ul><li><a href="https://nso-docs.cisco.com/guides/developer-reference/pyapi">Developer Reference</a></li><li><a href="https://nso-docs.cisco.com/guides/resources/man">Manual Pages</a></li><li><a href="https://github.com/NSO-developer/nso-examples">NSO Example Collection</a></li></ul></td><td><a href="/files/TyYMv6AXuRmvZwR6hV86">/files/TyYMv6AXuRmvZwR6hV86</a></td></tr><tr><td><strong>Learning Resources</strong></td><td><ul><li><a href="/pages/dlb7UkS3GFXGFclKIYXf">Learning Paths</a></li><li><a href="https://developer.cisco.com/learning/tracks/get_started_with_nso/">Learning Labs</a></li><li><a href="https://devnetsandbox.cisco.com/DevNet">Sandbox</a></li></ul></td><td><a href="/files/T6h5fo2rJJ4Ozn6hrTOG">/files/T6h5fo2rJJ4Ozn6hrTOG</a></td></tr></tbody></table>

### **Cisco DevNet Resources**

<table data-view="cards"><thead><tr><th></th></tr></thead><tbody><tr><td><ul><li><a href="https://developer.cisco.com/">Cisco DevNet</a></li><li><a href="https://github.com/CiscoDevNet">DevNet on GitHub</a></li><li><a href="https://developer.cisco.com/site/sandbox/">Sandboxes</a></li></ul></td></tr><tr><td><ul><li><a href="https://developer.cisco.com/iot/">IoT Dev Center</a></li><li><a href="https://developer.cisco.com/site/networking/">Networking Dev Center</a></li><li><a href="https://developer.cisco.com/site/data-center/">Data Center Dev Center</a></li></ul></td></tr><tr><td><ul><li><a href="https://developer.cisco.com/site/collaboration/">Collaboration Dev Center</a></li><li><a href="https://developer.cisco.com/site/security/">Security Dev Center</a></li><li><a href="https://developer.cisco.com/cx/">CX Dev Center</a></li></ul></td></tr></tbody></table>


# NSO at a Glance

A brief product overview of NSO, its architecture, and core concepts.

***

**Welcome to the Cisco Crosswork NSO Documentation**

On this page, you'll find a brief introduction to NSO to help you learn the basics of the product, its features, architecture, components, and how it helps you tackle network management challenges.

***

## What is NSO?

Cisco Crosswork Network Services Orchestrator (NSO) enabled by Tail-f is an industry-leading orchestration platform for hybrid networks. As a Linux application, it allows fine-grained control of physical and virtual network devices and can powerfully orchestrate the configuration life cycle of networks they live in. It provides comprehensive lifecycle service automation to enable you to design and deliver high-quality services faster and easier.

{% hint style="info" %}
The terms 'ncs' and 'tail-f' are used extensively in file names, command-line command names, YANG models, application programming interfaces (API), etc. Throughout this documentation, we use NSO to refer to the product.
{% endhint %}

## Key Features

At its heart, NSO makes network orchestration possible by leveraging the following features:

* **Multi-vendor device configuration management:** Uses the native protocols of the network devices to manage a wide variety of network devices.
* **Configuration Database (CDB):** Manages synchronized configurations for all devices and services in the network domain.
* **A rich set of northbound interfaces:** Includes human interfaces like web UI and a CLI, programmable interfaces including RESTCONF, NETCONF, JSON-RPC, and language bindings including Java, Python, and Erlang.
* **A central point of access to manage NSO:** Manages entire networks for network engineers using the NSO CLI or web UI. Although this documentation illustrates the use cases using CLI examples, it is important to understand that any northbound interface can be used to achieve the same functionality.
* **Network Element Drivers (NEDs):** Used as software packages to facilitate telnet, SSH, or API interactions with the devices that it manages.

## Background: The Orchestration Challenge

The industry is rapidly moving towards a service-oriented approach to network management, where multi-vendor devices, physical and virtual, support complex services. To manage these, operators are starting a transition from manually managing devices towards a situation where an operator is actively managing the various aspects of services.

Configuring the services and the affected devices is among the largest cost drivers in provider networks. Still, the common orchestration and configuration management practice involves pervasive manual work or ad hoc scripting. Why do we still apply these sorts of techniques to the configuration management problem? Two primary reasons are the variations of services and the constant change of devices. These two underlying characteristics are, to some degree, blocking automated solutions, since it takes too long to update the solution to cope with daily changes.

Time-to-market requirements are critical for a new service to be deployed quickly and the delay in configuring the corresponding tools has a significant impact on revenue. There is an unserved need in provider networks for tools that address these complex and sometimes contradictory challenges while constructing service configurations.

### How NSO Addresses the Orchestration Challenge <a href="#how-nso-addresses-the-orchestration-challenge" id="how-nso-addresses-the-orchestration-challenge"></a>

Creating and configuring network services is a complex task that often requires multiple configuration changes to all devices participating in the service. Additionally, changes generally need to be made concurrently across all devices with the changes being either completely successful or rolled back to the starting configuration. And, configurations need to be kept in sync across the system and the network devices. NSO approaches these challenges by acting as an interface between people or software that want to configure the network and the devices in the network.

NSO enables service providers to dynamically adopt the orchestration solution according to changes in the offered service portfolio. This is enabled by using a model-driven architecture where service definitions can be changed on the fly. Rather than a hard-coded orchestrator, NSO learns from the service models. Service models are written in YANG (RFC 6020).

NSO delivers an automated orchestration solution toward a hybrid multi-vendor network. The network can be a mix of traditional equipment, virtual devices, and SDN Controllers. This flexibility is managed by a Network Element Driver, NED, layer that abstracts the device interfaces and the Device Manager which enables generic device configuration functions.

At the core of NSO is the configuration datastore, CDB, that is in sync with the actual device and service configuration. It also manages relationships between services and devices and can handle revisions of device interfaces.

All devices and services in the network can be accessed and configured using the NSO CLI, making it a powerful tool for network engineers. The CLI also provides an easy way to define roles and associated authorization policies limiting the engineer's view of the devices under NSO control. Policies and integrity constraints can also be defined to ensure the configuration adheres to operator standards.

The typical workflow when using the NSO CLI is as follows:

1. The user logs in to the CLI and thereby starts a new session. The session provides a (logical) copy of the running configuration of CDB as a scratch pad for edits.
2. Changes made to the scratch pad are optionally validated at any time against policies and schemas using the "validate" command. Changes can always be viewed and verified before committing them.
3. The changes are committed, meaning that the changes are copied to the NSO database and deltas are pushed out to the network devices that are affected by the change. Changes that violate integrity constraints or network policies will not be committed but produce validation errors. The changes to the devices are done in a distributed and atomic transaction across all devices in parallel.
4. Changes either succeed and remain committed to device configuration or fail and are rolled back as a whole returning the entire network to the prior state.

## NSO Architecture

***

**Related Learning**: Building Blocks of NSO

{% embed url="<https://youtu.be/dxPO3BxX3IU>" %}
**Building Blocks of NSO**
{% endembed %}

***

NSO has two main layers, the Device Manager and the Service Manager. They serve different purposes but are tightly integrated with a transactional engine and database.

<div data-with-frame="true"><figure><img src="/files/Mo5B7O8bpirlhfjP5EVP" alt=""><figcaption><p>NSO Architecture</p></figcaption></figure></div>

NSO uses a dedicated built-in storage Configuration Database (CDB) for all configuration data. NSO keeps the CDB in sync with the real network device configurations. Audit, to ensure configuration consistency, and reconciliation, to synchronize configuration with the devices, functions are supported. It also maintains the runtime relationships between service instances and the corresponding device configurations.

NSO uses Network Element Drivers, NEDs, to communicate with devices. NEDs are not closed hard-coded adapters. Rather, the device interface is modeled in a data model using the YANG data modeling language. NSO can render the required commands or operations directly from this model. This includes support for legacy configuration interfaces like device CLIs. This means that the NEDs can easily be updated to support new commands just by extending the data models with the appropriate model constructs which avoid any programming tasks as part of the change cycle.

NSO also comes with tooling for simulating the configuration aspects of a network. The netsim tool is used to simulate management interfaces like Cisco CLI and NETCONF for NSO examples and service development.

The main components of NSO are described below.

### **Service Manager**

The Service Manager makes it possible for an operator to manage high-level aspects of the network that are not supported by the devices directly or are supported in a cumbersome way. With the appropriate service definition running in the Service Manager, an operator could for example configure the VLANs that should exist in the network in a single place, and the Service Manager compute the specific configuration changes required for each device in the network and push them out. This covers the whole life cycle for a service: creation, modification, and deletion. NSO has an intelligent and easy way to use a mapping layer so that network engineers can define how a service should be deployed in the network.

The Service Manager addresses the following challenges:

* Transaction-safe activation of services across different multi-vendor devices.
* What-if scenarios, (dry-run), showing the effects on the network for a service creation/change.
* Maintaining relationships between services and corresponding device configurations and vice versa.
* Modeling of services
* Short development and turn-around time for new services.
* Mapping the service model to device models.

### Device Manager

The purpose of the Device Manager is to manage device configurations in a transactional manner. It supports features like fine-grained configuration commands, bidirectional device configuration synchronization, device groups and templates, and compliance reporting.

The Device Manager supports the following overall features:

* Deploy configuration changes to multiple devices in a fail-safe way using distributed transactions.
* Validate the integrity of configurations before deploying to the network.
* Apply configuration changes to named device groups.
* Apply templates (with variables) to named device groups.
* Easily roll back changes, if needed.
* Configuration audits: Check if device configurations are in sync with the NSO database. If they are not, what is the diff?
* Synchronize the NSO database and the configurations on devices, in case they are not in sync. This can be done in either direction (import the diff to the NSO database or deploy the diff on devices).

### Network Element Drivers (NEDs)

NED, or Network Element Driver, represents a key NSO component that makes it possible for the NSO's core system to communicate southbound with the network devices in most deployments. NSO has a built-in client that can be used to communicate southbound with NETCONF-enabled devices. The vast majority of existing network devices are, however, not NETCONF-enabled. There are two main categories of NEDs: the CLI NEDs, and the Generic NEDs. Both categories have the same basic components in common. A Cisco-provided NED is an NSO package containing a bundle of YANG models together with a driver element implemented in Java.

### UI and APIs

NSO provides user interfaces as well as northbound APIs for integration into other systems. The main user interface is the NSO network-wide CLI which gives a unified CLI towards the complete network including the network services. This documentation illustrates most of the functions using the CLI. NSO also provides a Web UI.

The northbound APIs are available in different language bindings (Java, Python), and as protocols, like NETCONF and REST.

To support dynamic updates of functionality as added or modified service models, support for a new device type, etc, NSO manages extensions as well-defined packages. Every NED is its own package with its own release life cycle. Every service model with the corresponding mapping is also a package of its own. These can be upgraded without upgrading NSO itself.

When running NSO against real devices, (not just the NSO network simulator for educational purposes), make sure you have the correct NED package version from the delivery repository.

To learn how to use NSO and also to simplify development using NSO, NSO comes with a network simulator, `ncs-netsim`. Many of the examples will use netsim as the network.

### HA and Clustering

NSO supports a 1:N high-availability mode. One NSO system can be primary and can have any number of secondaries. Any configuration write has to go through the primary node. The configuration changes are replicated to the read-only secondaries. The replication can be done in asynchronous or synchronous mode. In the synchronous mode, the transaction returns when the secondaries are in sync.

For large networks, the network devices can be clustered across NSO systems. Say you have 100,000 devices split into two continents. You may choose to have 50,000 devices in one NSO and 50,000 in another. There are several options on how to configure clusters to see the whole network. The most common is a top NSO where services are provisioned, and the top NSO sees the whole network.

## Core Concepts

These are some of the core concepts that make NSO special. The following picture gives a high-level view of the components involved.

<div data-with-frame="true"><figure><img src="/files/pw6HYdel9yh1vVVCbmPr" alt="" width="500"><figcaption><p>NSO Components</p></figcaption></figure></div>

### NETCONF/YANG <a href="#netconfyang" id="netconfyang"></a>

As network programmability was starting to grow in importance it [was realized](https://tools.ietf.org/html/rfc3535) that configuration of network elements needed a modern model-driven interface to enable that programmability. This led to the development of [YANG](https://tools.ietf.org/html/rfc6020), a modeling language for configuration.

NSO uses YANG as the overall modeling language to manage devices and services.\
YANG models describe all NSO configurations, including device configuration and service configuration. This means that everything that is done in NSO is model-driven, this allows all interfaces to be automatically rendered.

YANG was originally paired with [NETCONF](https://tools.ietf.org/html/rfc6241), but as REST became a more popular interface\
[RESTCONF](https://tools.ietf.org/html/rfc8040) was standardized as well. Both of the protocols let you define your API using YANG, giving you a model-driven interface; this means that the basic working of the protocol is defined by the standards and the application-specific details can be derived from the loaded YANG models. The RESTCONF protocol provides a compatible subset of the functionality of the NETCONF protocol but does not provide the full transactional interface of the latter.

### Configuration Database (CDB) <a href="#the-configuration-database-cdb" id="the-configuration-database-cdb"></a>

At the core of NSO is the Configuration Database (CDB). This is a tree-structured database that is defined by a YANG schema. This means that all of the information stored inside of NSO is validated against the schema.

Every transaction towards CDB exhibits [ACID](https://en.wikipedia.org/wiki/ACID) properties, which among other things means either the transaction as a whole ends up on all participating devices (as well as in the NSO CDB), or otherwise, the whole transaction is aborted and all changes are automatically rolled back.

The CDB always contains NSO's view of the complete network configuration. To handle out-of-band changes operations are available to check if a device is in sync, write NSO's view to the device, or read the device configuration into NSO.

### Service Algorithm - FastMap <a href="#the-service-algorithm---fastmap" id="the-service-algorithm---fastmap"></a>

As a Service Developer, you need to express the mapping from a YANG service model to the corresponding device YANG models. This is a declarative mapping in the sense that no sequencing is defined. Observe that irrespective of the underlying device type and corresponding native device interface, the mapping is towards a YANG device model, not the native CLI for example. This means that as you write the service mapping, you do not have to worry about the syntax of different devices' CLI commands or in which order these commands are sent to the devices. This is all taken care of by the NSO device manager.

NSO reduces this problem to a single data-mapping definition for the "create" scenario. At run-time\
NSO will render the minimum change for any possible change like all the ones mentioned below. This is managed by the FASTMAP algorithm.

FASTMAP covers the complete service life-cycle: creating, changing, and deleting the service. The solution requires a minimum amount of code for mapping from a service model to a device model.

FASTMAP is based on generating changes from an initial 'create'. When the service instance is created the reverse of the resulting device configuration is stored together with the service instance. If an NSO user later changes the service instance, NSO first applies (in a transaction) the reverse diff of the service, effectively undoing the previous results of the service creation code. Then it runs the logic to create the service again and finally executes a diff to the current configuration. This diff is then sent to the devices.

<div data-with-frame="true"><figure><img src="/files/K5AATM7gkvsgNEwmvvi8" alt="" width="563"><figcaption><p>The Service Algorithm - FastMap</p></figcaption></figure></div>

### Accessing the Network (NEDs) <a href="#accessing-the-network-neds" id="accessing-the-network-neds"></a>

The NSO device manager is the center of NSO. The device manager maintains a flat list of all managed devices. NSO serves as a "source of truth" and keeps a copy of the configuration for each managed device in the CDB. Whenever a change is done to the device configuration copies in the CDB, the device manager will partition this "network configuration change" into the corresponding changes for the actually managed devices. The device manager passes on the required changes to the NEDs, Network Element Drivers. A NED needs to be installed for every type of device OS, like Cisco IOS NED, Cisco XR NED, Juniper JUNOS NED, etc. The NEDs communicate through the native device protocol southbound. The NEDs fall into the following categories:

* NETCONF capable device. The Device Manager will produce NETCONF edit-configuration RPC operations for each participating device.
* SNMP device. The Device Manager translates the changes made to the configuration into the corresponding SNMP SET PDUs
* A device with Cisco CLI. The device has a CLI with the same structure as Cisco IOS or XR routers. The Device Manager and a CLI NED are used to produce the correct sequence of CLI commands which reflects the changes made to the configuration.
* For other devices that do not fit into any of the above-mentioned categories, a corresponding Generic NED is invoked. Generic NEDs are used for proprietary protocols like REST and for CLI flavors that are not resembling IOS or XR. The Device Manager will inform the Generic NED about the made changes and the NED will translate these to the appropriate operations toward the device.


# Common Use Cases

Common use cases for NSO.

Cisco Network Services Orchestrator (NSO) is a model-driven, vendor-agnostic automation and orchestration platform used to design, deploy, and operate network and service infrastructure across physical, virtual, and cloud environments.

NSO abstracts device- and vendor-specific complexity through YANG-based service and device models, enabling consistent service lifecycle management across multi-vendor and multi-domain networks. It integrates with existing OSS/BSS systems, domain controllers, and DevOps toolchains, acting as a central automation and orchestration engine rather than a solution tied to a specific technology or product.

The following sections describe automation as well as other use cases and deployment patterns for NSO.

***

## Automation Use Cases Across Network and Service Domains

Cisco NSO is commonly deployed as a cross-domain orchestration layer, enabling consistent automation patterns across different network technologies and operational domains.

Figure below illustrates example automation use cases across multiple domains orchestrated by NSO. The following sections highlight representative scenarios.

<div data-with-frame="true"><figure><img src="/files/PUoynTXNzPO9HIEK9Iha" alt=""><figcaption></figcaption></figure></div>

***

### Transport Network Automation

NSO is widely used to automate IP and transport networks.

Representative use cases include:

* Zero-touch provisioning of transport devices
* L2VPN and L3VPN service provisioning
* QoS policy application for VPN services
* Bandwidth-on-demand and service modification
* BGP and LSP optimization workflows
* Coordinated OS and software upgrades

***

### Data Center and Virtualized Infrastructure Automation

NSO orchestrates both physical and virtual infrastructure components in data center environments.

Common scenarios include:

* Leaf–spine switch provisioning
* Virtualized edge and gateway services
* E-LAN and E-Line service orchestration
* Infrastructure configuration for private cloud environments
* Integration with virtualization and cloud platforms

***

### Enterprise and Campus Services Automation

NSO supports scalable service deployment and lifecycle management for enterprise and managed service environments.

Typical use cases include:

* Enterprise switch and CPE provisioning
* Dedicated Internet Access (single-homing and dual-homing)
* Enterprise voice and IPTV services
* Perimeter security and access policy enforcement

***

### Optical Network Automation

NSO integrates with optical network elements and controllers to orchestrate services alongside IP and transport layers.

Representative use cases include:

* Optical service provisioning
* Topology discovery and inventory synchronization
* Software version compliance and lifecycle management
* Coordinated IP-over-optical workflows

***

### Mobility and 5G Service Automation

NSO is used as an orchestration layer in mobility and 5G environments.

Representative scenarios include:

* Automation of cloud-native deployment platforms
* Gateway and policy service configuration
* Virtual network function and cloud service instantiation
* Cell site service provisioning and migration workflows
* Integration with external IPAM and infrastructure systems

***

### Integration with DevOps and Automation Ecosystems

NSO integrates with external systems through APIs and event mechanisms, enabling its use within modern DevOps and automation workflows.

Common scenarios include:

* Git-driven service definitions
* Automated testing and validation
* CI/CD-driven infrastructure and service deployment

This allows network automation to follow the same principles used for application infrastructure.

***

## Other Common Use Cases

### Automated Device Onboarding and Turnup (Day-0 / Day-1)

Simplify the deployment of new network devices such as routers, switches, and virtual network elements.

NSO automates device onboarding by generating configurations from reusable service models and templates. Device-specific parameters are applied to a common model, ensuring consistency across vendors and platforms.

Key capabilities include:

* Day-0 and Day-1 configuration generation
* Configuration preview and validation before deployment
* Transactional configuration delivery with rollback on failure
* Bulk onboarding of multiple devices in a single operation
* Support for greenfield and brownfield environments

This approach reduces manual configuration effort and minimizes deployment risk.

***

### Service Lifecycle Orchestration

Orchestrate the full lifecycle of network and service offerings, including creation, modification, repair, and decommissioning.

NSO uses declarative service models to describe the desired end state of a service. Its orchestration engine automatically computes and applies the minimal set of changes required to converge the infrastructure to that state.

Services can span:

* Physical network devices
* Virtualized network functions
* Cloud and container-based environments
* External controllers and management systems

This enables rapid, repeatable service delivery while maintaining consistency and control across domains.

***

### SD-WAN and Overlay Service Orchestration

Automate the lifecycle of Software-Defined Wide Area Network (SD-WAN) and overlay network services across multi-vendor environments.

NSO orchestrates configurations and workflows across devices and external SD-WAN controllers, abstracting vendor-specific differences through NEDs and integration packages. A single service definition can be applied across heterogeneous SD-WAN solutions.

Typical use cases include:

* Hub and branch service provisioning
* Policy updates and service modification
* Lifecycle operations at scale
* Integration with existing SD-WAN management systems

NSO acts as a lifecycle orchestration layer rather than replacing domain-specific SD-WAN controllers.

***

### Policy-Based Configuration and Governance (ACLs, QoS, and More)

Simplify the management of network-wide policies such as Access Control Lists (ACLs), Quality of Service (QoS), and standardized configuration rules.

NSO allows operators to define high-level policy intent using service models. These policies are translated into the appropriate vendor-specific configuration syntax and applied consistently across the network.

Examples include:

* Security and access policies
* QoS shaping and policing rules
* Standardized configuration blocks by device role

This ensures policy consistency while reducing operational complexity.

***

### Unified Interfaces for Network Automation (CLI, API, UI)

Provide consistent operational interfaces across multi-vendor environments.

Network engineers and automation systems interact with NSO using:

* A model-driven CLI
* Northbound APIs (JSON-RPC, RESTCONF, NETCONF)
* Web-based user interfaces

NSO’s APIs enable seamless integration into external automation systems and CI/CD pipelines, allowing network services to be programmatically deployed, validated, and modified as part of broader automation workflows.

Operators work with service abstractions instead of device-specific CLIs. NSO ensures the correct syntax is generated for each device using the appropriate NED.

***

### Configuration Compliance and Drift Management

Ensure that device configurations conform to defined network-wide standards.

NSO continuously compares the running configuration of devices with intended state definitions (often referred to as *golden configurations*). Deviations caused by manual or out-of-band changes can be detected, reported, and optionally remediated.

Key capabilities include:

* Configuration audits across large device fleets
* Detection of configuration drift
* Centralized compliance reporting
* Automated reconciliation to intended state

This helps maintain operational consistency and reduce configuration-related incidents.

***

### Software and OS Lifecycle Management

Centralize and automate software and operating system upgrades across network infrastructure.

NSO can be used by customers as a framework for building software maintenance workflows. Common implementations include:

* Coordinated upgrades during maintenance windows
* Parallel execution across large device sets
* Transactional execution with rollback on failure
* Integration with external inventory and lifecycle systems

This reduces operational overhead and improves reliability during maintenance activities.

## Summary

Cisco NSO is not tied to any specific vendor, device type, or network architecture. Its strength lies in its ability to abstract complexity, enforce consistency, and automate service lifecycles across diverse environments.

The use cases described on this page represent common deployment patterns. NSO’s extensible architecture allows organizations to tailor automation to their specific operational and business requirements.


# FAQs

Frequently Asked Questions on NSO.

## Problems Solved <a href="#problems-solved-1" id="problems-solved-1"></a>

<details>

<summary>What are the key functions of NSO?</summary>

Unlike other configuration managers on the market, NSO focuses not only on reading but also on writing deep fine-grained configurations to the network. NSO provides configuration management for both devices and network services. Any detailed parameter can be changed and NSO will generate the minimum configuration changes to the devices (not the entire config files).

NSO applies distributed transactions to all network changes. That is, if there is an error on any device when committing a change to the network, none of the other devices in that [transaction will be changed](#what-is-unique-about-the-way-that-nso-manages-transactions). This includes non-transactional devices like CLI and SNMP. When NSO writes the configuration changes it is capable of generating any reverse operations to keep the network and NSO in a consistent state (these reverse operations are stored in rollback files).

</details>

<details>

<summary>When does it make financial sense to purchase NSO?</summary>

Cisco research and customer experience shows 50-70% cost savings can be achieved within Service & Network Ops using NSO and model driven orchestration. Connect with your Cisco representative to get more information about how NSO can help your operations.

</details>

## General NSO Configuration Management Principles <a href="#general-nso-configuration-management-principles-1" id="general-nso-configuration-management-principles-1"></a>

<details>

<summary>Which types of network devices can NSO manage?</summary>

Any device that can be remotely configured can be managed by NSO. This includes routers, switches, load balancers, firewalls, web servers, and a whole host of other devices (both virtual and physical). Since it is not limited to one type of device or particular vendor, NSO allows for the management of the entire network from a single pane of glass.

</details>

<details>

<summary>How does NSO handle interfaces to devices?</summary>

The device interfaces are managed by NSO Network Element Drivers (NEDs) and the accompanying toolkit. Cisco provides NEDs to Juniper, Cisco, Alcatel-Lucent, Ericsson, A10, F5, Brocade, HP, Huawei, and others. Additional NEDs can be ordered from Cisco or developed by end-customers and integrators.

As part of the support contract, Cisco will support any device upgrades with the same NED functionality.

</details>

<details>

<summary>Can customers change the Network Element Driver?</summary>

In theory, customers can create or enhance Network Element Drivers, NEDs, since the NED source (YANG, etc.) and developer guide come with the system.

But in practice, there should be no need for a customer to do so.

If a customer chooses to create or enhance a NED, this will mean that:

* The NED cannot be supported by Cisco TAC.
* The upgrade path to official NED versions will break.

The only time customer development of NEDs should be considered is:

* In case of emergency.
* For one-off devices/systems.

However, it is always advised to use a qualified partner to create or enhance a NED.

</details>

<details>

<summary>Can NSO orchestrate VNFs according to ETSI NFV MANO Specification?</summary>

NSO with the NFVO package is an ETSI-compliant implementation of NFV orchestration. The Cisco NFVO is designed to enable easy integration with Specialized-VNFMs by offering a flexible interface or “dock” into which a third-party VNFM can be integrated without having to make changes to the NFV-O itself. Cisco Elastic Services Controller, a multi-vendor VNFM, is integrated via this interface to provide a full NFV orchestration platform.

</details>

<details>

<summary>How does NSO deal with different software versions of a device?</summary>

When NSO connects to a device it discovers the device version and the appropriate data model version. The NSO Web interface, CLI, database, and APIs are version-aware, so the correct model will be used. When committing changes to the network, a user can choose different strategies (require all devices to support all changes or skip changes to devices not supporting them). The network engineer using NSO does not have to keep device versions in mind, NSO resolves all of this.

</details>

<details>

<summary>Can NSO validate configurations?</summary>

Yes, NSO supports deep fine-grain configuration validations. Validations can be applied to both devices and services. Validation constraints can be added at run-time to specify uniqueness, dependencies between configuration items and devices, and more.

Some examples of validations include:

* All URLs must be unique.
* All devices should have a management interface m0 with status up.
* The MTU for ATM interfaces must be between 64 and 17966.

</details>

<details>

<summary>What happens if NSO is temporarily down or unreachable?</summary>

NSO can reconcile any configuration change that happened out-of-band and allows a network engineer to determine which configuration is correct.

</details>

<details>

<summary>Can NSO do service configuration?</summary>

Absolutely! NSO is typically used for provisioning network services like VPNs, ACLs, BGP Peers, etc. It is of great value to have one product and one embedded database that provisions services and device configurations as one atomic transaction.

NSO supports fine-grained service updates: the user can change any aspect of a service and NSO will calculate the resulting minimum device configuration changes. NSO automatically cleans up the network when services are deleted.

NSO can also perform network audits to detect if any device configuration has changed with respect to the desired service configuration. The diff can be displayed and analyzed and the service can be re-deployed if needed.

</details>

<details>

<summary>Can NSO do compliance reporting?</summary>

NSO has built-in support for compliance reporting that will check that all devices and services are configured as expected. It also shows details for any discrepancies, such as a misconfigured VPN on an interface. The report also includes details about all changes that have been performed in the network.

</details>

<details>

<summary>Why don't we have NED roadmaps?</summary>

Any NED supports the features that have been used by previous customers and POCs for that NED. No customer should expect any NED to be "complete" and cover all possible use cases for the managed device. NEDs are developed incrementally as and when needed. There is also no NED roadmap. This is possible because of the extremely short turnaround times for NEDs and NED enhancements. This is beneficial to everyone. From a customer perspective, they will get the NEDs and NED features that they require for when they need them. The question "Does NED x support feature y or OS version x?" becomes moot - the answer is yes - if it is required then it will, provided that the customer has purchased the NED, that they have a current support agreement, that the request is for "normal" use of the managed device (i.e. a configuration use case that is not extremely customer specific), and that the NED team can access devices for test purposes, which may require access to devices in the customer's network.

The normal SWSS contract covers NEDs purchased from the price list and entitles access to all future versions of the NEDs. They can request enhancements to the NEDs, via the account team, as per the NED request process specified in the NED index slides on DevNet and Field Portal, and we will deliver these at no additional cost, provided that the enhancements represent reasonably mainstream use of the managed devices, and can be added without adverse impact on other customers. NED enhancements will normally be delivered within 2-6 weeks after we have received the required data regarding the configuration use cases and commands, depending on the volume and complexity of the enhancements, and also sometimes on being able to access the customer devices in the cases where device/OS versions and features are not available in our labs.

</details>

<details>

<summary>How long does it take to introduce new Service models?</summary>

Introducing new service types (VPN, Firewall Policies, etc.) into NSO Service models usually takes a matter of days. As with devices, services are defined by data models. Once a new service model is created, it can automatically be loaded and used for new service instances.

</details>

## Vendor-Specific Device Management <a href="#vendor-specific-device-management-1" id="vendor-specific-device-management-1"></a>

<details>

<summary>How does NSO configure Juniper devices?</summary>

The Juniper NED, with a complete JunOS Schema, is available for purchase. Upgrades are managed by reading the JunOS schema from the upgraded device and the NSO toolkit will re-render all the NSO functions from the newly read JunOS schema, including the user interfaces and the database schema.

</details>

<details>

<summary>How does NSO configure Cisco devices?</summary>

Multiple Cisco CLI NEDs are available for purchase, for IOS (IOS XE), NX-OS, IOS -XR and others. The NEDs include YANG data models representing subsets of the devices’ configuration. From these data models, NSO renders Cisco CLI sequences and parses CLI outputs without any code. This means that adding support for new commands is a matter of adding the command to the model. Furthermore, the model covers dependencies so NSO can generate commands and roll-backs.

</details>

<details>

<summary>How does NSO configure SNMP Devices?</summary>

NSO can load SNMP MIBs and generate configuration commands directly from the MIBs. Before loading the MIBs an integrator can annotate the MIB with ordering dependencies so that a user does not need to know that a certain variable needs to have a certain value before changing another variable, which is common in SNMP. NSO also handles table relationships automatically. NSO allows for transaction-based management of SNMP devices. The NSO diff-engine can generate the reverse operations to reach the previous state, thus supporting roll-backs also of SNMP devices.

</details>

## Transactions and roll-backs <a href="#transactions-and-roll-backs-1" id="transactions-and-roll-backs-1"></a>

<details>

<summary>What is unique about the way that NSO manages transactions?</summary>

NSO does not just fire off commands to the network but rather confirms that all changes in the transaction are deployed correctly at the device level. If at any point of the series a device cannot be changed, the entire transaction is automatically roll-backed. This ensures that there is always a consistent network state. Complex rollback scenarios like the correct ordering of CLI commands are handled. The NSO database, the services and the devices configurations are all part of the same transaction.

Users can load network-wide rollback files to undo network changes.

</details>

<details>

<summary>How can NSO do network transactions over non-transactional devices like those supporting SNMP and CLI?</summary>

NSO performs any configuration change as a diff operation against the current configuration. If a device does not support rollbacks, NSO can use the diff engine to calculate the diff between current state and the previous state before the configuration change is applied. NSO knows the capabilities for different network devices and will issue a config change or just a rollback command depending on the network device capabilities. The data models for the devices specify dependencies explicitly as directed references – this ensures that NSO knows in which order to create and delete items.

</details>

<details>

<summary>How many roll-backs do you keep?</summary>

That is up to the user of NSO – it is configurable.

</details>

## Usability <a href="#usability-1" id="usability-1"></a>

<details>

<summary>How can a network engineer work with NSO?</summary>

Most network engineers are hard-core CLI users. NSO is the only product on the market that is designed with the network engineer in mind. NSO provides a network-wide CLI to a multi-vendor network. As a real-time abstraction of the network, it doesn’t matter which access methods are used, the network is [always in a consistent state](#what-are-the-key-functions-of-nso). Some of our customers use NSO for network automation by hooking it into their workflow systems, while at the same time having the CLI access for network engineers to quickly review and approve changes to device configurations.

</details>

<details>

<summary>How does NSO support different user categories?</summary>

NSO supports different interfaces for different users with different needs and skills:

* CLI for power users.
* Web interface for ad-hoc users.
* Programmable interfaces: REST, scripting, Java, JavaScript.

</details>

<details>

<summary>Can NSO do dry-run and what-if scenarios?</summary>

Before a transaction is committed, a user can inspect the actual effect on the network in two ways:

* `Compare running brief` – this will show the difference between the desired configuration and the current configuration.
* `Dry-run` – this will also involve any service provisioning logic and show the resulting device configurations.

</details>

<details>

<summary>What are the training requirements for NSO?</summary>

Limited training is required to get started. NSO is designed to be user-friendly to network engineers by providing common tools that they are already familiar using, such as a CLI and a Web interface.

</details>

<details>

<summary>How complex is it to install NSO?</summary>

NSO installs in seconds. There are no third-party requirements.

</details>

## The Cisco NSO Difference <a href="#the-cisco-nso-difference-1" id="the-cisco-nso-difference-1"></a>

<details>

<summary>Does NSO provide a centralized configuration database?</summary>

The NSO-embedded configuration database and transaction engine are at the heart of the product. Whenever an NSO user changes a device or service configuration this change is applied as [one single transaction](#transactions-and-roll-backs-1) including the NSO database as well as the devices. This means that the database will always represent the true state of the network. In case of out-of-band configurations, boot-strap scenarios, etc., NSO supports synchronization and comparison tools.

The database is an integral part of NSO, so schemas and database administration are managed automatically without requiring a database administrator.

The database can be imported, exported, and queried using XML, XPath, and the CLI and Web interfaces. Programmatic interfaces to the database are also provided.

</details>

<details>

<summary>How does NSO compare to other Configuration Management tools?</summary>

If you look at the marketing material, there seems to be a lot of configuration management tools on the market. However, when you dig into the actual functionality available for writing configurations to the devices, it is disappointing. The practice for these kinds of tools is to push and pull configuration files, store them, diff them, and restore them. This works fine for static environments. The above tools do not actually understand the configurations, they treat them as opaque text files.

NSO on the other hand understands the configurations and represents them as fine-grained data structures in the embedded configuration database. Thereby NSO can provide configuration-aware tools like the CLI, configuration validation, an auto-rendered web interface, etc. And most importantly NSO can do real-time fine-grained changes rather than diffing and pulling configuration files.

NSO can also do [service provisionin](#can-nso-do-service-configuration)g.

Some vendors have fantastic tools for managing their specific devices. This works great if you have a single-vendor network. However, single-vendor networks are rare.

</details>

<details>

<summary>How does NSO compare with Service Activation tools?</summary>

NSO has unique features to ensure shorter time-to-market and lower cost of ownership than any other service activation tools on the market.

There are several stove-pipe activation tools (single-vendor, single service-type) that work very well for that specific vendor or service type. Service activation with these tools breaks down when a service requires different vendors or multiple service types. NSO is capable of handling multiple service types for multi-vendor networks.

NSO is unique in the way it communicates with the devices. Many activation tools require device “adaptors” that are custom-coded and expensive to maintain. The [NSO Network Element Driver (NED) technology](#how-does-nso-handle-interfaces-to-devices) makes device integration simple.

Another difference is that NSO applies transactions to the activation chain. Many other activation tools are workflow-centric, which makes it very hard to recover from faults, and in many cases, they escape to manual work orders.

NSO also provides detailed knowledge of the device configurations: it can make sure that the service instance is consistent with the actual device configuration. NSO provides a network and service CLI as a power tool for network engineers. The visibility of the activation chain, using the CLI, service dry-run, etc, helps the network engineers trust the tool (this is harder to do with tools that just provide a magic OK button).

</details>

## Integration, Support, and Hardware requirements <a href="#integration-support-and-hardware-requirements-1" id="integration-support-and-hardware-requirements-1"></a>

<details>

<summary>Does NSO support OpenStack?</summary>

The Havanna release of OpenStack introduced an NSO Mechanism Driver that maps the OpenStack Neutron network model changes to NSO REST calls. NSO can then map these changes to a multi-vendor network.

</details>

<details>

<summary>What type of database does NSO require?</summary>

NSO ships with an [embedded special-purpose database](#does-nso-provide-a-centralized-configuration-database), so no external database server is needed.

</details>

<details>

<summary>Which northbound interfaces does NSO support?</summary>

The following interfaces are auto-rendered for all services and devices:

* CLI: for network engineers that prefer a Juniper or Cisco-style command-line interface.
* Web interface: for network engineers that prefer a graphical interface.
* REST: for programmatic access (exactly the same feature set as the CLI and Web interface).
* Java/C: for building custom applications and service provisioning logic.
* JavaScript: for embedding NSO functions in portals.
* NETCONF: for importing and exporting XML configurations.
* SNMP: for reading status and receiving NSO alarms.
* Python: for scripting network-wide configuration changes.

</details>

<details>

<summary>Does NSO support XML import and export?</summary>

NETCONF is an XML-based protocol. So by using the northbound NETCONF interface configuration data as well as operational data can be imported and exported. The NSO database has XML import and export tools.

</details>

<details>

<summary>Does NSO support syslog?</summary>

NSO supports generating BSD and RFC 5424 Syslog messages.

</details>

<details>

<summary>Which logs does NSO generate?</summary>

Developer logs, audit trail logs, web interface logs, device communication logs, and XPath logs.

</details>

<details>

<summary>Does NSO support AAA?</summary>

NSO contains a rich AAA engine that spans all interfaces. NSO supports role-based access control lists providing different access privileges to different users. These privileges may apply to a group of devices, or to individual devices, or to any subset of any device configuration. External authentication can be used via the Pluggable Application Module (PAM).

</details>

<details>

<summary>How does NSO guarantee high Availability?</summary>

NSO can be configured to run in HA mode: one master and several slaves. All slaves are updated at every transactional border. The slaves can be used for reading data. The slaves are hot-standby and a fail-over to a slave can be made at any time. In the clustered solution described above, each cluster node is typically a HA configuration.

</details>

<details>

<summary>What is the typical hardware requirement?</summary>

NSO requires a standard Unix box

* Medium (< 1000 devices): 4 Core CPU and 32 GB RAM.
* Large (> 1000 devices): 8 Core CPU and 64 BG RAM.
* Linux/x86 – Linux/x86 64.
* MacOS 10.6/x86 64.

</details>

## NFV and SDN <a href="#nfv-and-sdn-1" id="nfv-and-sdn-1"></a>

<details>

<summary>How does NSO relate to NFV and SDN?</summary>

In the NFV architecture NSO, with the NSO NFVO component, fulfills the NFV orchestrator (NFVO) role, and also partially replaces the EMS/OSS layer. The Tail-f NSO architecture enables it to act as the central orchestrator and has the ability to control both SDN Controllers and NFV Managers. Tail-f NSO interfaces to SDN Controllers and NFV Managers via Network Equipment Drivers (NEDs). Similar to packages, NEDs are decoupled from the Tail-f NSO product and have their own release cycle enabling agile feature growth.

</details>


# Learning Paths

Explore different paths to learning NSO.

There are many different ways of learning NSO. This site has many resources available that can help you along the way.

<table data-view="cards"><thead><tr><th></th><th></th><th data-hidden data-card-target data-type="content-ref"></th></tr></thead><tbody><tr><td><strong>Product Documentation</strong></td><td>The product documentation should be your go-to place for learning and questions about the product.</td><td><a href="https://cisco-tailf.gitbook.io/nso-docs/guides">https://cisco-tailf.gitbook.io/nso-docs/guides</a></td></tr><tr><td><strong>Interactive Learning Labs</strong></td><td>The interactive learning labs and sandboxes are good tools for self-study, allowing you to learn at your own pace.</td><td><a href="https://developer.cisco.com/learning/tracks/get_started_with_nso/">https://developer.cisco.com/learning/tracks/get_started_with_nso/</a></td></tr><tr><td><strong>Formal Training</strong></td><td>Formal training is an excellent way to get an instructor-led introduction to NSO.</td><td><a href="https://learninglocator.cloudapps.cisco.com/#/search-results/NSO">https://learninglocator.cloudapps.cisco.com/#/search-results/NSO</a></td></tr><tr><td><strong>NSO Developer Hub</strong></td><td>The NSO Developer Hub is a great place to ask questions and collaborate with other NSO users.</td><td><a href="https://community.cisco.com/t5/nso-developer-hub/ct-p/5672j-dev-nso">https://community.cisco.com/t5/nso-developer-hub/ct-p/5672j-dev-nso</a></td></tr><tr><td><strong>NSO Playground</strong></td><td>The NSO Playground is a new interactive platform to play with NSO examples from the convenience of your browser.</td><td><a href="https://blogs.cisco.com/developer/nsoplayground01">https://blogs.cisco.com/developer/nsoplayground01</a></td></tr><tr><td><strong>Reservable Sandbox</strong></td><td>Use the Sandbox environment to explore NSO APIs and develop automation packages.</td><td><a href="https://devnetsandbox.cisco.com/DevNet">https://devnetsandbox.cisco.com/DevNet</a></td></tr></tbody></table>

With NSO, we often talk about an automation journey. Similarly, there is a journey for a Service Developer as well. Several skills come together when using NSO:

* **NSO skills**: You must understand how NSO works and how to implement your services in NSO.
* **Networking**: You have to have an understanding of the service that is being configured.
* **Software development skills**: Depending on your language of choice, a bit of Java or Python knowledge is desirable. Java or Python multiprocessing skills for faster service deployment are helpful.
* **DevOps skills**: Understanding the processes around automating the network, setting up delivery pipelines, and continuous integration systems.

The focus of the material we are collecting here is on the NSO skills, but as you dive deeper into NSO, additional skills will also be needed in the other areas.

## Your Learning Journey with NSO

{% stepper %}
{% step %}
**The First Day**

We recommend that you start by going through the [NSO at a Glance](https://cisco-tailf.gitbook.io/nso-docs/nso-basics/nso-at-a-glance), [Installation and Deployment](https://cisco-tailf.gitbook.io/nso-docs/guides/administration/installation-and-deployment), [Introduction to Automation](https://cisco-tailf.gitbook.io/nso-docs/guides/development/introduction-to-automation), and [Learning Labs](https://developer.cisco.com/learning/tracks/get_started_with_nso/) while following along on your own NSO instance. Refer to the [NSO documentation](https://cisco-tailf.gitbook.io/nso-docs/guides) for any other things needed.
{% endstep %}

{% step %}
**Next Steps**

After the first day, look at the additional interactive learning labs, sandboxes, and the extensive collection of examples available with the NSO distribution.

Also, consider [formal training from Cisco](https://learninglocator.cloudapps.cisco.com/#/search-results/NSO).
{% endstep %}

{% step %}
**Continue to Learn**

Learning never stops. Once you feel confident with the basics, join our community and read our blogs to follow the development of the community. One good resource for examples and inspiration is the [NSO GitHub](https://github.com/NSO-developer/) and [NSO GitLab](https://gitlab.com/NSO-developer/) pages.
{% endstep %}
{% endstepper %}


# Quick Start Guide

Quick start instructions to get started with NSO.

***

{% hint style="info" %}
This quick start guide uses a Local Install of NSO for those just getting started to try or evaluate NSO. The [NSO Installation and Deployment Guide](https://cisco-tailf.gitbook.io/nso-docs/guides/administration/installation-and-deployment) details handling Local and System installations.
{% endhint %}

<details>

<summary>Local vs. System Install</summary>

Before you install NSO onto your system, you need to decide whether to do a System or a Local installation. Here's a simple breakdown of the two:

* Use **System Install** when installing NSO for a centralized, "always-on" production-grade purpose. System installs configure NSO as a system daemon that starts and ends with the underlying operating system. Linux PAM is used instead of NSO local authentication using the admin and oper default users, and the file structure is distributed.
* Use **Local Install** for development, lab, and evaluation purposes. It unpacks all the application components, including docs and examples. The developer can use local installs to run multiple unrelated instances of NSO for different labs and demos on a single workstation.

</details>

***

## **Download Your NSO Free Trial Installer and Cisco NEDs** <a href="#download-your-nso-free-trial-installer-and-cisco-neds" id="download-your-nso-free-trial-installer-and-cisco-neds"></a>

This evaluation copy has been provided under the terms of the Cisco NSO Evaluation License. There are two versions of the NSO installer for macOS and Linux systems, respectively:

* [x] [NSO for Linux and MacOS (Darwin), including NED examples](https://software.cisco.com/download/home/286331591/type/286283941/release)

## Requirements <a href="#requirements" id="requirements"></a>

For development purposes, choose between the following:

* Linux for x86\_64 or arm64.
* macOS Darwin for x86\_64 or arm64.

A detailed list of NSO installation requirements can be found in the [NSO Installation and Deployment Guide](https://nso-docs.cisco.com/guides/administration/installation-and-deployment).

## Installation <a href="#installation" id="installation"></a>

{% hint style="info" %}
The recommended guide for installing NSO is the [NSO Installation and Deployment Guide](https://nso-docs.cisco.com/guides/administration/installation-and-deployment). The description below is a quick-start version.
{% endhint %}

### Operating System <a href="#operating-system" id="operating-system"></a>

Cisco NSO can run on macOS or Linux systems. If you are a Windows user (or do not wish to install it natively on your laptop), you can install NSO on a Linux virtual machine or in a container.

If you use Docker, there are pre-built system install images available. See the [Containerized NSO](https://nso-docs.cisco.com/guides/administration/installation-and-deployment/containerized-nso) guide.

There are [NSO Playgrounds](https://developer.cisco.com/codeexchange/github/repo/CiscoDevNet/NSO-Playground-Local-Install/) available to dive right in and try out examples in a user-friendly browser-based integrated development environment (IDE).

## Performing a Local Install <a href="#performing-a-local-installation" id="performing-a-local-installation"></a>

Once you've installed the prereqs and downloaded the NSO installation file for your operating system, you are ready to do your installation:

{% stepper %}
{% step %}
Open a terminal and navigate to the directory where you downloaded the installer.

> **Ensure that you have the correct installer binary for your OS**, `darwin` is for macOS, and `linux` for Linux distributions.

{% code overflow="wrap" %}

```bash
cd ~/Downloads
ls -l nso*.bin
-rw-r--r--@ 1 user  staff   199M Dec 15 11:45 nso-6.0.darwin.x86_64.installer.bin
-rw-r--r--@ 1 user  staff   199M Dec 15 11:45 nso-6.0.darwin.x86_64.signed.bin
```

{% endcode %}

> If your file is a signed.bin file, this means that Cisco has digitally signed the file you downloaded, and when you execute it, you'll verify the signature and unpack the `installer.bin`. If you have the `installer.bin` , skip down past the signed steps.
> {% endstep %}

{% step %}
Use the `sh` command to "run" the `signed.bin` to verify the certificate and extract the installer binary and other files.

{% code overflow="wrap" %}

```
sh nso-6.0.darwin.x86_64.signed.bin 

# Output
Unpacking...
Verifying signature...
Downloading CA certificate from http://www.cisco.com/security/pki/certs/crcam2.cer ...
Successfully downloaded and verified crcam2.cer.
Downloading SubCA certificate from http://www.cisco.com/security/pki/certs/innerspace.cer ...
Successfully downloaded and verified innerspace.cer.
Successfully verified root, subca and end-entity certificate chain.
Successfully fetched a public key from tailf.cer.
Successfully verified the signature of nso-6.0.darwin.x86_64.installer.bin using tailf.cer
```

{% endcode %}
{% endstep %}

{% step %}
If it all comes back green, you're in good shape and ready to install.

{% code overflow="wrap" %}

```bash
ls -l

# Output
-rw-r--r--  1 user  staff   1.8K Nov 29 06:05 README.signature
-rw-r--r--  1 user  staff    12K Nov 29 06:05 cisco_x509_verify_release.py
-rwxr-xr-x  1 user  staff   199M Nov 29 05:55 nso-6.0.darwin.x86_64.installer.bin
-rw-r--r--  1 user  staff   256B Nov 29 06:05 nso-6.0.darwin.x86_64.installer.bin.signature
-rwxr-xr-x@ 1 user  staff   199M Dec 15 11:45 nso-6.0.darwin.x86_64.signed.bin
-rw-r--r--  1 user  staff   1.4K Nov 29 06:05 tailf.cer
```

{% endcode %}

Here is what was unpacked:

* The NSO installer `nso-VERSION.OS.ARCH.installer.bin`.
* Signature generated for the NSO image `nso-VERSION.OS.ARCH.installer.bin.signature`.
* An enclosed Cisco-signed `tailf.cer` x.509 end-entity certificate containing the public key that is used to verify the signature.
* `README.signature` file which briefs you on more details on the unpacked content and the steps on "How to run the signature verification program". If you would like to manually verify the signature, please refer to the steps in this file.
* `cisco_x509_verify_release.py` python program that can be used to verify the 3-tier x.509 certificate chain and signature.
  {% endstep %}

{% step %}
First, check out the `--help` on the installer binary using the `sh nso-6.0.darwin.x86_64.installer.bin --help` command. Notice the two options for `--local-install` or `--system-install`.

```
sh nso-6.0.darwin.x86_64.installer.bin --help

# Output
This is the NCS installation script.

Usage: ./nso-6.0.darwin.x86_64.installer.bin [--local-install] LocalInstallDir

Installs NCS in the LocalInstallDir directory only.
This is convenient for test and development purposes.

Usage: ./nso-6.0.darwin.x86_64.installer.bin --system-install [--install-dir InstallDir]
    [--config-dir ConfigDir] [--run-dir RunDir] [--log-dir LogDir]
    [--run-as-user User] [--keep-ncs-setup] [--non-interactive]

Does a system install of NCS, suitable for deployment.
Static files are installed in InstallDir/ncs-<vsn>.
The first time --system-install is used, the ConfigDir,
RunDir, and LogDir directories are also created and
populated for config files, run-time state files, and log files,
respectively, and an init script for start of NCS at system boot
and user profile scripts are installed. Defaults are:

InstallDir - /opt/ncs
ConfigDir  - /etc/ncs
RunDir     - /var/opt/ncs
LogDir     - /var/log/ncs

By default, the system install will run NCS as the root user.
If the --run-as-user option is given, the system install will
instead run NCS as the given user. The user will be created if
it does not already exist.

If the --non-interactive option is given, the installer will
proceed with potentially disruptive changes (e.g. modifying or
removing existing files) without asking for confirmation.
```

{% endstep %}

{% step %}
For the installation directory or `LocalInstallDir`, the recommendation is to install it into your `$HOME` directory in a folder called `~/nso-VERSION`. So if our version is `6.0`, our directory will be `~/nso-6.0`.
{% endstep %}

{% step %}
Run the installer with the argument `--local-install ~/nso-6.0` to install it into your home directory.

{% code overflow="wrap" %}

```bash
sh nso-6.0.darwin.x86_64.installer.bin --local-install ~/nso-6.0

# Output
INFO  Using temporary directory /var/folders/90/n5sbctr922336_0jrzhb54400000gn/T//ncs_installer.93831 to stage NCS installation bundle
INFO  Unpacked ncs-6.0 in /Users/user/nso-6.0
INFO  Found and unpacked corresponding DOCUMENTATION_PACKAGE
INFO  Found and unpacked corresponding EXAMPLE_PACKAGE
INFO  Found and unpacked corresponding JAVA_PACKAGE
INFO  Generating default SSH hostkey (this may take some time)
INFO  SSH hostkey generated
INFO  Environment set-up generated in /Users/user/nso-6.0/ncsrc
INFO  NSO installation script finished
INFO  Found and unpacked corresponding NETSIM_PACKAGE
INFO  NCS installation complete
```

{% endcode %}
{% endstep %}

{% step %}
That's it. NSO is installed.
{% endstep %}
{% endstepper %}

## Exploring the Installation <a href="#exploring-the-installation" id="exploring-the-installation"></a>

Before we start up NSO, let's just look at what we have.

Go ahead and `cd` into the new installation directory.

```bash
cd ~/nso-6.0
```

### Documentation <a href="#documentation" id="documentation"></a>

> **Note**: Links to the online version of the [NSO Guides](https://nso-docs.cisco.com/guides) and [NSO Extension API Reference](https://developer.cisco.com/docs/nso/api/) documentation.

Along with the binaries, NSO installs a full set of documentation available in the `doc/` folder in `~/nso-6.0`.

```bash
ls -l doc/
drwxr-xr-x   5 user  staff   160B Nov 29 05:19 api/
drwxr-xr-x  14 user  staff   448B Nov 29 05:19 html/
-rw-r--r--   1 user  staff   202B Nov 29 05:19 index.html
drwxr-xr-x  17 user  staff   544B Nov 29 05:19 pdf/
```

Feel free to open up the `index.html` file in your favorite browser and poke around (you can use the `open` command in the terminal to open the file). You'll find installation, admin, user, development, and more guides available for you to jump right into.

### Examples <a href="#examples" id="examples"></a>

An NSO Local Install also comes with **A LOT** of examples of a variety of different types of ways you can use NSO. Many of these touch on advanced topics, but there are plenty of basic ones as well. Here are the high-level directories of examples in `~/nso-6.0/examples.ncs`

```bash
ls -l examples.ncs/

# Output
-rw-r--r--   1 user  staff   1.0K Nov 29 05:17 README
drwxr-xr-x   4 user  staff   128B Nov 29 04:50 datacenter/
drwxr-xr-x   3 user  staff    96B Nov 29 04:50 generic-ned/
drwxr-xr-x   4 user  staff   128B Nov 29 04:50 getting-started/
drwxr-xr-x   7 user  staff   224B Nov 29 04:50 service-provider/
drwxr-xr-x   3 user  staff    96B Nov 29 04:50 snmp-ned/
drwxr-xr-x  11 user  staff   352B Nov 29 05:17 snmp-notification-receiver/
drwxr-xr-x   5 user  staff   160B Nov 29 04:50 web-server-farm/
```

### NEDs or Network Element Drivers <a href="#neds-or-network-element-drivers" id="neds-or-network-element-drivers"></a>

In order to "talk to" the network, NSO uses NEDs as device drivers for different device types. Cisco has NEDs for hundreds of different devices available for customers, and several are included in the installer in the `~/nso-6.0/packages/neds` directory.

```bash
ls -l

# Output
drwxr-xr-x  13 user  staff   416B Nov 29 05:17 a10-acos-cli-3.0/
drwxr-xr-x  12 user  staff   384B Nov 29 05:17 alu-sr-cli-3.4/
drwxr-xr-x  13 user  staff   416B Nov 29 05:17 cisco-asa-cli-6.6/
drwxr-xr-x  12 user  staff   384B Nov 29 05:17 cisco-ios-cli-3.0/
drwxr-xr-x  12 user  staff   384B Nov 29 05:17 cisco-ios-cli-3.8/
drwxr-xr-x  13 user  staff   416B Nov 29 05:17 cisco-iosxr-cli-3.0/
drwxr-xr-x  13 user  staff   416B Nov 29 05:17 cisco-iosxr-cli-3.5/
drwxr-xr-x  13 user  staff   416B Nov 29 05:17 cisco-nx-cli-3.0/
drwxr-xr-x  13 user  staff   416B Nov 29 05:17 dell-ftos-cli-3.0/
drwxr-xr-x  10 user  staff   320B Nov 29 05:17 juniper-junos-nc-3.0/
```

Here you can see there are NEDs for Cisco ASA, IOS, IOS XR, and NX-OS. Also included are NEDs for other vendors including Juniper JunOS, A10, ALU, and Dell.

> **Note**: The NEDs included in the installer are **intended for evaluation, demonstration, and use with the** `examples.ncs` that are also included. These are **not** the latest versions available, and often don't have all the features available in production NEDs.

#### **Installing New NED Versions**

Cisco also makes additional versions of some NEDs available on DevNet for evaluation and non-production use. You can find them with the NSO downloads (scroll up!).

> **Note**: The specific file names and versions you download maybe different from this guide. Update the paths appropriately.

1. Like the NSO installer, the NEDs are `signed.bin` files that need to be run to validate the download and extract the new code.
2. First, find the downloaded files - change to the working directory where your downloads are:

> **Note**: The filenames indicate which version of NSO the NEDs are pre-compiled for (in this case NSO 6.0), and the version of the NED.

```bash
cd ~/Downloads/
ls -l ncs*.bin

# Output
-rw-r--r--@ 1 user  staff   9708091 Dec 18 12:05 ncs-6.0-cisco-asa-6.16-freetrial.signed.bin
-rw-r--r--@ 1 user  staff  51233042 Dec 18 12:06 ncs-6.0-cisco-ios-6.88-freetrial.signed.bin
-rw-r--r--@ 1 user  staff  39292052 Dec 18 12:06 ncs-6.0-cisco-iosxr-7.43-freetrial.signed.bin
-rw-r--r--@ 1 user  staff   8400190 Dec 18 12:05 ncs-6.0-cisco-nx-5.23.6-freetrial.signed.bin
```

3. Use the `sh` command to "run" the `signed.bin` to verify the certificate and extract the NED `tar.gz` and other files. Repeat for all files.

```bash
sh ncs-6.0-cisco-nx-5.23.6.signed.bin
```

Output:

{% code overflow="wrap" %}

```bash
Unpacking...
Verifying signature...
Downloading CA certificate from http://www.cisco.com/security/pki/certs/crcam2.cer ...
Successfully downloaded and verified crcam2.cer.
Downloading SubCA certificate from http://www.cisco.com/security/pki/certs/innerspace.cer ...
Successfully downloaded and verified innerspace.cer.
Successfully verified root, subca and end-entity certificate chain.
Successfully fetched a public key from tailf.cer.
Successfully verified the signature of ncs-6.0-cisco-nx-5.23.6.tar.gz using tailf.cer
```

{% endcode %}

4. You now have three tarballs (`.tar.gz`) files. These are compressed versions of the NEDs.

```bash
ls -l ncs*.tar.gz
```

Output:

{% code overflow="wrap" %}

```bash
-rw-r--r--  1 user  staff   9704896 Dec 12 21:11 ncs-6.0-cisco-asa-6.16.tar.gz
-rw-r--r--  1 user  staff  51260488 Dec 13 22:58 ncs-6.0-cisco-ios-6.88.tar.gz
-rw-r--r--  1 user  staff  39305257 Oct  7 17:47 ncs-6.0-cisco-iosxr-7.43.tar.gz
-rw-r--r--  1 user  staff   8409288 Dec 18 09:09 ncs-6.0-cisco-nx-5.23.6.tar.gz
```

{% endcode %}

5. Navigate to the `packages/neds` directory for your local install.

```bash
cd ~/nso-6.0/packages/neds
```

6. While in `~/nso-6.0/packages/neds` directory, extract the tarballs into this directory using the `tar` command with the path to where the compressed NED is located:

> Update the path and file name for the NED versions you downloaded

```bash
tar -zxvf ~/Downloads/ncs-6.0-cisco-nx-5.23.6.tar.g
tar -zxvf ~/Downloads/ncs-6.0-cisco-ios-6.88.tar.gz
tar -zxvf ~/Downloads/ncs-6.0-cisco-iosxr-7.43.tar.gz
tar -zxvf ~/Downloads/ncs-6.0-17:47-6.16.tar.gz
```

Here is a sample list of the newer NEDs extracted along with the ones bundled with the installation:

```bash
drwxr-xr-x  13 user  staff   416 Nov 29 05:17 a10-acos-cli-3.0
drwxr-xr-x  12 user  staff   384 Nov 29 05:17 alu-sr-cli-3.4
drwxr-xr-x  13 user  staff   416 Nov 29 05:17 cisco-asa-cli-6.6
drwxr-xr-x  13 user  staff   416 Dec 12 21:11 cisco-asa-cli-6.7
drwxr-xr-x  12 user  staff   384 Nov 29 05:17 cisco-ios-cli-3.0
drwxr-xr-x  12 user  staff   384 Nov 29 05:17 cisco-ios-cli-3.8
drwxr-xr-x  13 user  staff   416 Dec 13 22:58 cisco-ios-cli-6.42
drwxr-xr-x  13 user  staff   416 Nov 29 05:17 cisco-iosxr-cli-3.0
drwxr-xr-x  13 user  staff   416 Nov 29 05:17 cisco-iosxr-cli-3.5
drwxr-xr-x  13 user  staff   448 Oct  7 14:46 cisco-iosxr-cli-7.4
drwxr-xr-x  13 user  staff   416 Nov 29 05:17 cisco-nx-cli-3.0
drwxr-xr-x  14 user  staff   448 Dec 18 09:09 cisco-nx-cli-5.13
drwxr-xr-x  13 user  staff   416 Nov 29 05:17 dell-ftos-cli-3.0
drwxr-xr-x  10 user  staff   320 Nov 29 05:17 juniper-junos-nc-3.0
```

7. And now you have the newer NED versions available, as well as the demo/evaluation versions included with NSO itself!

### `ncsrc` File <a href="#ncsrc" id="ncsrc"></a>

The last thing to note is the files `ncsrc` and `ncsrc.tsch`. These are shell scripts for bash and `tsch` that set up your `PATH` and other environment variables for NSO. Depending on your shell, you will need `source` this file before starting your NSO work.

```bash
# NOTE your path may be different, double check!
 $ source $HOME/nso-6.0/ncsrc
```

Most users add `source ~/nso-6.0/ncsrc` to their `~/.bash_profile`, but you can just do it manually when you want it. Once it has been "sourced" you have access to all the NSO executable commands - which start with `ncs`.

```bash
ncs {TAB} {TAB}

# Output
ncs                      ncs-maapi                ncs-project              ncs-start-python-vm      ncs_cmd                  ncs_load                 
ncs-backup               ncs-make-package         ncs-setup                ncs-uninstall            ncs_conf_tool            ncsc                     
ncs-collect-tech-report  ncs-netsim               ncs-start-java-vm        ncs_cli                  ncs_crypto_keys 
```

## Creating an Instance of NSO <a href="#creating-an-instance-of-nso" id="creating-an-instance-of-nso"></a>

A NSO Local Install unpacks and prepares your system to run NSO, but **doesn't actually start it up**. A Local Install allows the engineer to create an NSO "instance" tied to a project, which you have not done yet.

> If you are familiar with Python, you can think of this like creating a Python virtual environment after installing Python. Within this NSO instance, you will have different inventory, configuration, and code.

### **Using `ncs-setup` to Create an NSO Instance**

One of the included scripts with an NSO installation is `ncs-setup`, which makes it very easy to create instances of NSO from a Local Install. You can look at the `--help` for full details, but the two options we need to know are: \* `--dest` defines the directory where you want to set up NSO (if the directory does not exist, it will be created) \* `--package` defines the NEDs you want to this NSO instance to have installed. You can specify this option multiple times.

> **Note**: NCS is the original name of the NSO product, so many the commands and application features will be prefaced with `ncs`. Think of `ncs` as another name for NSO.

1. Go ahead and run this command to setup an NSO instance in the current directory with the IOS, NX-OS, IOS-XR, and ASA NEDs, you only need one NED per platform that you want NSO to manage (even though you may have multiple versions in your installer `neds` directory).\\

   > You will want to use the name of the NED folder in `${NCS_DIR}/packages/neds` for the **latest** NED version you've got installed for the target platform. You can use `tab` complete after you start typing a path (or just copy and paste, though double check the NED version numbers below match what is currently on the sandbox to avoid a syntax error):

   ```bash
   ncs-setup --package ~/nso-6.0/packages/neds/cisco-ios-cli-6.44 \
     --package ~/nso-6.0/packages/neds/cisco-nx-cli-5.15 \
     --package ~/nso-6.0/packages/neds/cisco-iosxr-cli-7.20 \
     --package ~/nso-6.0/packages/neds/cisco-asa-cli-6.8 \
     --dest nso-instance
   ```
2. If you check out the `nso-instance` directory now, you'll find several new files and folders have been created. This guide won't go through them all in detail now, but a couple are handy to know about.

   * `ncs.conf` is the NSO application configuration file which is used to customize aspects of the NSO instance (for example, change ports, enable/disable features, enable Web UI, etc.). The defaults are often perfect for projects like this.
   * `packages/` is the directory that has symlinks to the NEDs that we referenced in the `--package` arguments at setup.
   * `logs/` is the directory that contains all the logs from NSO. This directory is useful when troubleshooting.

   ```bash
   $ ls nso-instance/
   logs  ncs-cdb  ncs.conf  packages  README.ncs  scripts  state
   $ ls -l nso-instance/packages/
   total 0
   lrwxrwxrwx 1 user docker 51 Mar 19 12:44 cisco-asa-cli-6.8 -> /home/user/nso-6.0/packages/neds/cisco-asa-cli-6.8
   lrwxrwxrwx 1 user docker 52 Mar 19 12:44 cisco-ios-cli-6.44 -> /home/user/nso-6.0/packages/neds/cisco-ios-cli-6.44
   lrwxrwxrwx 1 user docker 54 Mar 19 12:44 cisco-iosxr-cli-7.20 -> /home/user/nso-6.0/packages/neds/cisco-iosxr-cli-7.20
   lrwxrwxrwx 1 user docker 51 Mar 19 12:44 cisco-nx-cli-5.15 -> /home/user/nso-6.0/packages/neds/cisco-nx-cli-5.15
   $
   ```
3. Now you need to "start" your NSO instance. Navigate to the `nso-instance` directory and type the command `ncs`. It will take a few seconds to run, and you won't get any explicit output unless there is a problem.

> **Note**: You need to be in the `nso-instance` directory each time you want to start or stop NSO. If you have multiple instances, you need to navigate to each one to use the `ncs` command to start or stop each one.

```bash
ncs
```

4. You can verify that NSO is running by using the `ncs --status | grep status` command, which has a large amount of information, so we use `grep` to just search for the status:

{% code overflow="wrap" %}

```bash
$ ncs --status | grep status
status: started
        db=running id=31 priority=1 path=/ncs:devices/device/live-status-protocol/device-type
```

{% endcode %}

5. Now you should add either some netsim devices (`ncs-netsim -h`) or lab devices to NSO and get automating!

## Get Started <a href="#getting-started" id="getting-started"></a>

We recommend that you start with the [online version](https://nso-docs.cisco.com/guides) or in the doc directory `$HOME/nso-VERSION/doc/pdf/`.

There are a lot of examples in the `$HOME/nso-VERSION/examples.ncs` directory. The examples have a short description in the `$HOME/nso-VERSION/examples.ncs/README`file. Each example has a README file that explains how to run it.

To access the Web UI (usually at <http://localhost:8080>), first enable it in the `ncs.conf` file. See the [NSO Web UI Guide](https://nso-docs.cisco.com/guides/operation-and-usage/webui) for more information.

## Sandboxes <a href="#sandboxes" id="sandboxes"></a>

If you do not want to download and install Cisco NSO, please check out the [NSO Reservable Sandbox on DevNet](https://devnetsandbox.cisco.com/RM/Diagram/Index/43964e62-a13c-4929-bde7-a2f68ad6b27c?diagramType=Topology). There is also an associated tutorial called [Learn NSO the Easy Way](https://developer.cisco.com/learning/tracks/get_started_with_nso).

## Still Need Help? <a href="#still-need-help" id="still-need-help"></a>

There is a lot of material on the [NSO Developer Hub](https://community.cisco.com/t5/nso-developer-hub/ct-p/5672j-dev-nso), where you also can ask questions about anything NSO-related.


# What's New

Latest features and enhancements added in this release.

{% hint style="info" %}
Only significant new updates are listed here. To see the complete list of changes, refer to the [NSO Changelog Explorer](https://developer.cisco.com/docs/nso/changelog-explorer/?from=6.6\&to=6.7).
{% endhint %}

## Release Highlights

This release includes major enhancements in the following areas:

<details>

<summary>Southbound Datastore Subscriptions</summary>

Protocols, such as YANG-Push (RFC 8641), define a mechanism for applications to receive continuous updates of the target datastore, avoiding the need to frequently poll the remote system for latest data.

NSO 6.7 introduces native support for consuming NETCONF-based YANG-Push streams on managed devices, as well as the more general framework for managing and consuming subscribed updates. The latter allows NEDs to implement and expose the same kind of data updates by using other protocols, such as gNMI.

Documentation Updates:

* Added the [Telemetry](/guides/operation-and-usage/operations/nso-device-manager#telemetry) section documenting the new feature and potential use cases.
* Added the [Telemetry Kicker Concepts](/guides/development/advanced-development/kicker#telemetry-kicker-concepts) section on consuming telemetry data.
* Added an [example service](https://github.com/NSO-developer/nso-examples/tree/6.7/service-management/implement-a-service/iface-v6) making use of the YANG-Push to dynamically update service status.

</details>

<details>

<summary>Improved HA Transport</summary>

Raft- and rule-based HA now use a unified TLS transport for improved security and additional features:

* Rule-based HA deployment uses TLS certificates for authentication and encryption of communication between nodes, same as HA Raft.
* HA Raft leader monitors quorum and relinquishes the leader role if quorum is lost, aborting the hanging ongoing transactions. The leader also generates an alarm and releases resources, such as a shared VIP address or primary-listen ports.
* HA Raft now requires only a single listening port to be open for communication, port 4570 by default, same as rule-based HA. The port can be changed in the configuration if required.

Documentation Updates:

* Described the new transport requirements in [HA Raft](https://nso-docs.cisco.com/guides/pages/qG3CMifhI63daJ1BZfmB#ug.ha.raft) and [Rule-based HA](https://nso-docs.cisco.com/guides/pages/qG3CMifhI63daJ1BZfmB#ug.ha.builtin).
* Added a section on provisioning TLS certificates with the help of example scripts to [High Availability](/guides/administration/management/high-availability).

</details>

<details>

<summary>In-service Package Upgrades</summary>

The improved `packages reload` (and `packages ha sync and-reload`) action now supports the `optimistic` mode for in-service upgrade. In this mode, NSO keeps accepting and processing requests on the northbound interfaces while the upgrade is in progress.

In addition, the new mode supports taking an NSO backup before the upgrade commences, enabled using the newly introduced `backup` switch.

Documentation Updates:

* Updated [Package Management](/guides/administration/management/package-mgmt) and [Upgrade NSO](/guides/administration/installation-and-deployment/upgrade-nso) with the new upgrade options.

</details>

<details>

<summary>Compliance XML Templates</summary>

A new type of compliance template is introduced: compliance template specified as an XML file, that lives as part of a package or individually, under a dedicated load-directory.

The new template has enhanced flexibility by the use of processing instructions, similar to a service template. It supports more sophisticated use cases by allowing for easier integration with multiple NED-IDs and incorporating conditional if-else statements.

Additionally, compliance reports can now re-check the violating items when used with the `re-run` action.

Documentation Updates:

* Added a new section [XML Compliance Templates](/guides/operation-and-usage/operations/compliance-reporting#xml-compliance-templates) to [Compliance Reporting](/guides/operation-and-usage/operations/compliance-reporting).

</details>

<details>

<summary>Template Creation from Configuration Snippets</summary>

The `create-template` actions under `/devices`, `/services`, and `/compliance` can now consume configuration snippets directly, in addition to extracting templates from configuration already present in NSO. Snippets can be supplied either from a file on the NSO server filesystem or as inline payload data, using NETCONF-style XML wrapped in a `<config>` element, Cisco XR style CLI (`cli-c`), Juniper curly-brace CLI (`cli-j`), or Juniper set commands (`cli-j-cmd`).

Delete operations in the input, such as Cisco-style `no` commands or XML `operation="remove"` attributes, are preserved in the generated output. NSO translates them into `delete` tags in device and service templates, and into `absent` tags in compliance templates. This makes it easier to turn existing golden configurations, hardening snippets, and similar configuration samples into reusable templates.

Documentation Updates:

* Updated [Templates](/guides/development/core-concepts/templates), [NSO Device Manager](/guides/operation-and-usage/operations/nso-device-manager), and [Compliance Reporting](/guides/operation-and-usage/operations/compliance-reporting) to describe snippet-based template generation.

</details>

<details>

<summary>Changes to CDB Persistence Mode</summary>

From NSO 6.7, the default CDB persistence mode has been set to `on-demand-v1`, instead of the `in-memory-v1` mode, which has also been deprecated. If you're upgrading to NSO 6.7, the `on-demand-v1` mode will become the new default. Read more about the change in the documentation.

Documentation Updates:

* Updated the [CDB Persistence](/guides/administration/advanced-topics/cdb-persistence) section to reflect the new changes in the CDB persistence mode.

</details>

<details>

<summary>Updates to Multi-Factor Authentication Handling</summary>

MFA handling is now tied directly to the authentication method being attempted. When a method issues a challenge, NSO invokes the challenge handler associated with that method only. Package-based MFA is the preferred approach. The configuration option `/ncs-config/aaa/challenge-order` is deprecated and ignored at runtime; authentication flow is controlled solely by `/ncs-config/aaa/auth-order`.

Documentation Updates:

* Updated the [Multi-Factor Authentication](https://nso-docs.cisco.com/guides/pages/oUchYxSeOKkwhPOOefVU#ug.aaa.external_challenge) documentation in [AAA Infrastructure](/guides/administration/management/aaa-infrastructure) to cover new changes.

</details>

<details>

<summary>Secure Local IPC</summary>

NSO now uses a more secure, Unix-domain-sockets-based IPC by default. It is used for internal communication between NSO server components.

Built-in components use this IPC mechanism automatically but Java and Python code in custom packages might need an update, depending on the SDK functions used for establishing connection to the NSO. See e.g. [Java API Overview](/guides/development/core-concepts/api-overview/java-api-overview) for example code using local IPC for connections.

Documentation Updates:

* Updated [IPC Connection](/guides/administration/advanced-topics/ipc-connection) and [Authenticating IPC Access](/guides/administration/management/aaa-infrastructure#authenticating-ipc-access) with the new default.
* Updated code snippets throughout the documentation to use the new IPC mechanism where applicable.

</details>

<details>

<summary>Alarm Notification Filtering by Type</summary>

NSO 6.7 introduces `/alarms/control/filter-types` for suppressing outbound alarm notifications for selected alarm types. Matching alarms remain available in the NSO alarm list, but NSO no longer emits matching SNMP, NETCONF, or RESTCONF alarm notifications for them.

Documentation Updates:

* Updated [Alarm Manager](/guides/operation-and-usage/operations/alarm-manager).
* Updated [System Management](/guides/administration/management/system-management).

</details>

<details>

<summary>Service Bulk Actions</summary>

To facilitate operation at scale, with many service instances of differing type, bulk `re-deploy`, and `un-deploy` actions were added under `/services`. These actions invoke the corresponding service-management action on a number of service instances, such as all services of a given type or matching an XPath expression.

Documentation Updates:

* Added section [Bulk Service Actions](/guides/operation-and-usage/operations/managing-network-services#bulk-service-actions) in [Manage Network Services](/guides/operation-and-usage/operations/managing-network-services).
* Added section [Bulk Service Actions](/guides/operation-and-usage/operations/lifecycle-operations#bulk-service-actions) in [Lifecycle Operations](/guides/operation-and-usage/operations/lifecycle-operations).

</details>

<details>

<summary>Out-of-Band Change Handling Controls</summary>

NSO 6.7 adds more precise control over how out-of-band changes are handled during `confirm-network-state` operations. Broad service re-deployment is no longer implicit when out-of-band data is discovered. Instead, re-deploying all affected services is now opt-in through the new `re-deploy-all` option. In addition, out-of-band policy rules now support an `abort` action, allowing a transaction to fail immediately with an out-of-sync error when specific out-of-band changes are detected. Policy rules can also define a `default-action`, which NSO uses when no operation-specific action has been specified with `at-create`, `at-delete`, or `at-value-set`.

These changes reduce unintended blast radius during out-of-band processing while making it easier to enforce strict handling for configuration changes that must not be accepted or reconciled automatically, and simpler to define common rule behavior without repeating the same action for every operation type.

Documentation Updates:

* Updated [Out-of-band Interoperation](/guides/operation-and-usage/operations/out-of-band-interoperation) with the new `re-deploy-all` behavior and policy rule actions, including `abort` and `default-action`.
* Updated [Lifecycle Operations](/guides/operation-and-usage/operations/lifecycle-operations) to include the `re-deploy-all` option under `confirm-network-state`.

</details>

<details>

<summary>Dry-run Drift Detection</summary>

The new feature helps prevent unintended changes from being committed. If there are additional changes introduced between `commit dry-run` and the final `commit`, the system warns and prompts the user on how to proceed. Dry-run drift detection is available in NSO CLI and JSON-RPC.

Documentation Updates:

* Added section [Dry-run Drift Detection](/guides/operation-and-usage/operations/lifecycle-operations#dry-run-drift-detection) in [Lifecycle Operations](/guides/operation-and-usage/operations/lifecycle-operations).

</details>

<details>

<summary>Memory Monitoring</summary>

NSO 6.7 tracks additional memory metrics, which can be used to detect memory trends or take corrective action, such as a debug dump or raising an alarm.

Documentation Updates:

* Updated [Containerized NSO](/guides/administration/installation-and-deployment/containerized-nso) and [System Install](/guides/administration/installation-and-deployment/system-install) with the recommended Memory Monitoring setup.

</details>

<details>

<summary>OpenID Connect Support for Single Sign-On</summary>

The `cisco-nso-oidc-auth` package is now available as part of the NSO distribution, implementing OpenID Connect (OIDC) as an authentication protocol for Single Sign-On (SSO).

Documentation Updates:

* Documented the new authentication package in `$NCS_DIR/packages/auth/cisco-nso-oidc-auth/README.md`
* Added the [examples.ncs/aaa/oidc-auth](https://github.com/NSO-developer/nso-examples/tree/6.7/aaa/oidc-auth) example.

</details>

<details>

<summary>Improve <code>live-status</code> Reads with Read Intent</summary>

Reading device's `live-status` data can trigger individual requests to the device when requested data is not cached. The new read-intent set of functions gives a MAAPI user an option to announce the need for required data before-hand, allowing NSO to optimize device roundtrips.

Documentation Updates:

* Added Fetch bulk live-status via MAAPI to [Java API Overview](/guides/development/core-concepts/api-overview/java-api-overview) and [Python API Overview](/guides/development/core-concepts/api-overview/python-api-overview).

</details>

<details>

<summary>Web UI Redesign and Enhancements</summary>

This release introduces a new **Transactions** view in the NSO Web UI, along with a redesigned **Configuration Editor** for a more streamlined configuration experience. It also includes general updates across the Web UI and documentation.\
\
Documentation Updates:

* Added a new [**Transactions**](/guides/operation-and-usage/webui/transactions) page to the Web UI documentation.
* Updated the [**Config Editor**](/guides/operation-and-usage/webui/config-editor) page to align with new changes.
* Updated the Web UI documentation for general improvements.

</details>

<details>

<summary>Adaptive MCP Server</summary>

NSO now includes the Cisco NSO Adaptive MCP Server, delivered as an NSO package. The MCP server provides a standard way for MCP-compatible AI assistants and clients to interact with NSO by exposing selected NSO data and operations through MCP resources, tools, and prompts.

Documentation Updates

* Added new [NSO MCP Server](/guides/development/core-concepts/northbound-apis/nso-mcp-server) guide under Northbound APIs.

</details>


# Get Started

Administrate and manage NSO.

## Installation and Deployment

<table data-view="cards" data-full-width="false"><thead><tr><th></th><th></th><th data-hidden data-card-target data-type="content-ref"></th></tr></thead><tbody><tr><td><strong>Local Install</strong></td><td>Install NSO for test and evaluation use.</td><td><a href="/pages/yQMF45w9clXn7gkEu0oS">/pages/yQMF45w9clXn7gkEu0oS</a></td></tr><tr><td><strong>System Install</strong></td><td>Install NSO for production system-wide use.</td><td><a href="/pages/QjIujV4wIdFJehgpfTbb">/pages/QjIujV4wIdFJehgpfTbb</a></td></tr><tr><td><strong>Post-install Actions</strong></td><td>Perform post-install actions after installing NSO.</td><td><a href="/pages/IaOFAePgvpfDIn40BwzA">/pages/IaOFAePgvpfDIn40BwzA</a></td></tr><tr><td><strong>Containerized NSO</strong></td><td>Deploy NSO using Cisco-provided container images.</td><td><a href="/pages/UR7ofR6OieaS0c1WwJ9z">/pages/UR7ofR6OieaS0c1WwJ9z</a></td></tr><tr><td><strong>Dev to Prod Deployment</strong></td><td>Deploy NSO from development to production.</td><td><a href="/pages/raYfZuUIBjEwaNT0djTI">/pages/raYfZuUIBjEwaNT0djTI</a></td></tr><tr><td><strong>Upgrade NSO</strong></td><td>Upgrade NSO installation to a higher version.</td><td><a href="/pages/mKh0oXFZkQsBaoCMMyEo">/pages/mKh0oXFZkQsBaoCMMyEo</a></td></tr></tbody></table>

## Management

<table data-view="cards" data-full-width="false"><thead><tr><th></th><th></th><th data-hidden data-card-target data-type="content-ref"></th></tr></thead><tbody><tr><td><strong>System Management</strong></td><td>Configure &#x26; manage your NSO deployment.</td><td><a href="/pages/hgWUBFw1TA0R6WyLxOgc">/pages/hgWUBFw1TA0R6WyLxOgc</a></td></tr><tr><td><strong>Package Management</strong></td><td>Learn about NSO packages and how to use them.</td><td><a href="/pages/TpxOL8qGnNMkxkp0ZCeB">/pages/TpxOL8qGnNMkxkp0ZCeB</a></td></tr><tr><td><strong>High Availability</strong></td><td>Set up multiple nodes in a highly-available (HA) setup.</td><td><a href="/pages/qG3CMifhI63daJ1BZfmB">/pages/qG3CMifhI63daJ1BZfmB</a></td></tr><tr><td><strong>AAA Infrastructure</strong></td><td>Set up user authentication and authorization.</td><td><a href="/pages/oUchYxSeOKkwhPOOefVU">/pages/oUchYxSeOKkwhPOOefVU</a></td></tr><tr><td><strong>NED Administration</strong></td><td>Administer and manage Cisco-provided NEDs.</td><td><a href="/pages/vGfK2qplvC1wu0670OgS">/pages/vGfK2qplvC1wu0670OgS</a></td></tr></tbody></table>

## Advanced Topics

<table data-view="cards" data-full-width="false"><thead><tr><th></th><th></th><th data-hidden data-card-target data-type="content-ref"></th></tr></thead><tbody><tr><td><strong>Locks</strong></td><td>Understand how transaction locks work.</td><td><a href="/pages/J1IL9AFawzDgiOvHn7GD">/pages/J1IL9AFawzDgiOvHn7GD</a></td></tr><tr><td><strong>CDB Persistence</strong></td><td>Select the optimal CDB persistence mode.</td><td><a href="/pages/kTheUIXYPG8PjTY9mfeD">/pages/kTheUIXYPG8PjTY9mfeD</a></td></tr><tr><td><strong>IPC Connection</strong></td><td>Learn how client libraries connect to NSO.</td><td><a href="/pages/lrUWrnLcjFJNFak7HCTU">/pages/lrUWrnLcjFJNFak7HCTU</a></td></tr><tr><td><strong>Cryptographic Keys</strong></td><td>Encrypt and decrypt strings in NSO using crypto keys.</td><td><a href="/pages/UCcWKXJFYdRRzsHoQ8Yx">/pages/UCcWKXJFYdRRzsHoQ8Yx</a></td></tr><tr><td><strong>Service Manager Restart</strong></td><td>Configure the timeout period of Service Manager.</td><td><a href="/pages/zRSjgoss3Fuq76DHw5DT">/pages/zRSjgoss3Fuq76DHw5DT</a></td></tr><tr><td><strong>IPv6 on Northbound</strong></td><td>Use IPv6 on Northbound NSO interfaces.</td><td><a href="/pages/qI2CQ7E2XoJMLyv8ZFNl">/pages/qI2CQ7E2XoJMLyv8ZFNl</a></td></tr><tr><td><strong>LSA</strong></td><td>Learn about Layered Service Architecture.</td><td><a href="/pages/3bZU9w7gi5NLay6PIbrp">/pages/3bZU9w7gi5NLay6PIbrp</a></td></tr></tbody></table>


# Installation and Deployment

Learn about different ways to install and deploy NSO.

## Ways to Deploy NSO <a href="#d5e46" id="d5e46"></a>

* [By installation](#by-installation)
* [By using Cisco-provided container images](#by-using-cisco-provided-container-images)

### By Installation

Choose this way if you want to install NSO on a system. Before proceeding with the installation, decide on the install type.

#### Install Types

The installation of NSO comes in two variants.

{% hint style="info" %}
Both variants can be installed in **standard mode** or in [**FIPS**](https://www.nist.gov/itl/publications-0/federal-information-processing-standards-fips)**-compliant** mode. See the detailed installation instructions for more information.
{% endhint %}

<table data-card-size="large" data-view="cards"><thead><tr><th></th><th></th><th data-hidden data-card-target data-type="content-ref"></th></tr></thead><tbody><tr><td><strong>Local Install</strong></td><td>Local Install is used for development, lab, and evaluation purposes. It unpacks all the application components, including docs and examples. It can be used by the engineer to run multiple, unrelated, instances of NSO for different labs and demos on a single workstation.</td><td><a href="/pages/yQMF45w9clXn7gkEu0oS">/pages/yQMF45w9clXn7gkEu0oS</a></td></tr><tr><td><strong>System Install</strong></td><td>System Install is used when installing NSO for a centralized, always-on, production-grade, system-wide deployment. It is configured as a system daemon that would start and end with the underlying operating system. The default users of admin and operator are not included and the file structure is more distributed.</td><td><a href="/pages/QjIujV4wIdFJehgpfTbb">/pages/QjIujV4wIdFJehgpfTbb</a></td></tr></tbody></table>

{% hint style="info" %}
All the NSO examples and README steps provided with the installation are based on and intended for Local Install only. Use Local Install for evaluation and development purposes only.

System Install should be used only for production deployment. For all other purposes, use the Local Install procedure.
{% endhint %}

### By Using Cisco-Provided Container Images

Choose this way if you want to run NSO in a container, such as Docker. Visit the link below for more information.

{% content-ref url="/pages/UR7ofR6OieaS0c1WwJ9z" %}
[Containerized NSO](/guides/administration/installation-and-deployment/containerized-nso)
{% endcontent-ref %}

***

> **Supporting Information**
>
> If you are evaluating NSO, you should have a designated support contact. If you have an NSO support agreement, please use the support channels specified in the agreement. In either case, do not hesitate to reach out to us if you have questions or feedback.


# Local Install

Install NSO for non-production use, such as for development and training purposes.

## Installation Steps

Complete the following activities in the given order to perform a Local Install of NSO.

<table data-view="cards" data-full-width="false"><thead><tr><th></th><th></th></tr></thead><tbody><tr><td><strong>Prepare</strong></td><td><a href="#step-1-fulfill-system-requirements">1. Fulfill System Requirements</a><br><a href="#li.download.the.installer">2. Download Installer/NEDs</a><br><a href="#li.unpack.the.installer">3. Unpack the Installer</a></td></tr><tr><td><strong>Install</strong></td><td><a href="#li.run.the.installer">4. Run the Installer</a></td></tr><tr><td><strong>Finalize</strong></td><td><a href="#li.set.env.variables">5. Set Environment Variables</a><br><a href="#li.create.runtime.directory">6. Runtime Directory Creation</a><br><a href="#li.generate.license.token">7. Generate License Token</a></td></tr></tbody></table>

{% hint style="info" %}
**Mode of Install**

NSO Local Install can be installed in **standard mode** or in [**FIPS**](https://www.nist.gov/itl/publications-0/federal-information-processing-standards-fips)**-compliant mode**. Standard mode install supports a broader set of cryptographic algorithms, while the FIPS mode install restricts NSO to use only FIPS 140-3-validated cryptographic modules and algorithms for enhanced/regulated security and compliance. Use FIPS mode only in environments that require compliance with specific security standards, especially in U.S. federal agencies or regulated industries. For all other use cases, install NSO in standard mode.

<sup>\* FIPS: Federal Information Processing Standards</sup>
{% endhint %}

### Step 1 - Fulfill System Requirements

Start by setting up your system to install and run NSO.

To install NSO:

1. Fulfill at least the primary requirements.
2. If you intend to build and run NSO examples, you also need to install additional applications listed under Additional Requirements.

{% hint style="warning" %}
Where requirements list a specific or higher version, there always exists a (small) possibility that a higher version introduces breaking changes. If in doubt whether the higher version is fully backwards compatible, always use the specific version.
{% endhint %}

<details>

<summary>Primary Requirements</summary>

Primary requirements to do a Local Install include:

* A system running Linux or macOS on either the `x86_64` or `ARM64` architecture for development. 4 CPU cores minimum. For [FIPS](https://www.nist.gov/itl/publications-0/federal-information-processing-standards-fips) mode, OS FIPS compliance may be required depending on your specific requirements.
* GNU libc 2.24 or higher.
* Java JRE 21 or higher. Used by Cisco Smart Licensing.
* Python 3.10 or higher (3.12 recommended).
* Required and included with many Linux/macOS distributions:
  * `tar` command. Unpack the installer.
  * `gzip` command. Unpack the installer.
  * `ssh-keygen` command. Generate SSH host key.
  * `openssl` command. Generate self-signed certificates for HTTPS.
  * `find` command. Used to find out if all required libraries are available.
  * `which` command. Used by the NSO package manager.
  * `libpam.so.0`. Pluggable Authentication Module library.
  * `libexpat.so.1`. EXtensible Markup Language parsing library.
  * `libz.so.1` version 1.2.7.1 or higher. Data compression library.

</details>

<details>

<summary>Additional Requirements</summary>

Additional requirements to, for example, build and run NSO examples/services include:

* Java JDK 21 or higher.
* Ant 1.9.8 or higher.
* Python Setuptools is required to build the Python API.
* Often installed using the Python package installer pip:
  * Python Paramiko 2.2 or higher. To use netconf-console.
  * Python requests. Used by the RESTCONF demo scripts.
* `xsltproc` command. Used by the `support/ned-make-package-meta-data` command to generate the `package-meta-data.xml` file.
* One of the following web browsers is required for NSO GUI capabilities. The version must be supported by the vendor at the time of release.
  * Safari
  * Mozilla Firefox
  * Microsoft Edge
  * Google Chrome
* OpenSSH client applications. For example, the `ssh` and `scp` commands.

</details>

<details>

<summary>FIPS Mode Entropy Requirements</summary>

The following applies if you are running a container-based setup of your FIPS install:

In containerized environments (e.g., Docker) that run on older Linux kernels (e.g., Ubuntu 18.04), `/dev/random` may block if the system’s entropy pool is low. This can lead to delays or hangs in FIPS mode, as cryptographic operations require high-quality randomness.

To avoid this:

* Prefer newer kernels (e.g., Ubuntu 22.04 or later), where entropy handling is improved to mitigate the issue.
* Or, install an entropy daemon like Haveged on the Docker host to help maintain sufficient entropy.

Check available entropy on the host system with:

```bash
cat /proc/sys/kernel/random/entropy_avail
```

A value of 256 or higher is generally considered safe. Reference: [Oracle blog post](https://blogs.oracle.com/linux/post/entropyavail-256-is-good-enough-for-everyone).

</details>

### Step 2 - Download the Installer and NEDs <a href="#li.download.the.installer" id="li.download.the.installer"></a>

To download the Cisco NSO installer and example NEDs:

1. Go to the Cisco's official [Software Download](https://software.cisco.com/download/home) site.
2. Search for the product "Network Services Orchestrator" and select the desired version.
3. There are two versions of the NSO installer, i.e. for macOS and Linux systems. Download the desired installer.

<details>

<summary>Identifying the Installer</summary>

You need to know your system specifications (Operating System and CPU architecture) in order to choose the appropriate NSO installer.

NSO is delivered as an OS/CPU-specific signed self-extractable archive. The signed archive file has the pattern `nso-VERSION.OS.ARCH.signed.bin` that after signature verification extracts the `nso-VERSION.OS.ARCH.installer.bin` archive file, where:

* `VERSION` is the NSO version to install.
* `OS` is the Operating System (`linux` for all Linux distributions and `darwin` for macOS).
* `ARCH` is the CPU architecture, for example`x86_64`.

</details>

### Step 3 - Unpack the Installer <a href="#li.unpack.the.installer" id="li.unpack.the.installer"></a>

If your downloaded file is a `signed.bin` file, it means that it has been digitally signed by Cisco, and upon execution, you will verify the signature and unpack the `installer.bin`.

If you only have `installer.bin`, skip to the next step.

To unpack the installer:

1. In the terminal, list the binaries in the directory where you downloaded the installer, for example:

   ```bash
   cd ~/Downloads
   ls -l nso*.bin
   -rw-r--r--@ 1 user  staff   199M Dec 15 11:45 nso-6.0.darwin.x86_64.installer.bin
   -rw-r--r--@ 1 user  staff   199M Dec 15 11:45 nso-6.0.darwin.x86_64.signed.bin
   ```
2. Use the `sh` command to run the `signed.bin` to verify the certificate and extract the installer binary and other files. An example output is shown below.

   ```bash
   sh nso-6.0.darwin.x86_64.signed.bin
   # Output
   Unpacking...
   Verifying signature...
   Downloading CA certificate from http://www.cisco.com/security/pki/certs/crcam2.cer ...
   Successfully downloaded and verified crcam2.cer.
   Downloading SubCA certificate from http://www.cisco.com/security/pki/certs/innerspace.cer ...
   Successfully downloaded and verified innerspace.cer.
   Successfully verified root, subca and end-entity certificate chain.
   Successfully fetched a public key from tailf.cer.
   Successfully verified the signature of nso-6.0.darwin.x86_64.installer.bin using tailf.cer
   ```
3. List the files to check if extraction was successful.

   ```bash
   ls -l
   # Output
   -rw-r--r--  1 user  staff   1.8K Nov 29 06:05 README.signature
   -rw-r--r--  1 user  staff    12K Nov 29 06:05 cisco_x509_verify_release.py
   -rwxr-xr-x  1 user  staff   199M Nov 29 05:55 nso-6.0.darwin.x86_64.installer.bin
   -rw-r--r--  1 user  staff   256B Nov 29 06:05 nso-6.0.darwin.x86_64.installer.bin.signature
   -rwxr-xr-x@ 1 user  staff   199M Dec 15 11:45 nso-6.0.darwin.x86_64.signed.bin
   -rw-r--r--  1 user  staff   1.4K Nov 29 06:05 tailf.cer
   ```

<details>

<summary>Description of Unpacked Files</summary>

The following contents are unpacked:

* `nso-VERSION.OS.ARCH.installer.bin`: The NSO installer.
* `nso-VERSION.OS.ARCH.installer.bin.signature`: Signature generated for the NSO image.
* `tailf.cer`: An enclosed Cisco signed x.509 end-entity certificate containing the public key that is used to verify the signature.
* `README.signature`: File with further details on the unpacked content and steps on how to run the signature verification program. To manually verify the signature, refer to the steps in this file.
* `cisco_x509_verify_release.py`: Python program that can be used to verify the 3-tier x.509 certificate chain and signature.
* Multiple `.tar.gz` files: Bundled packages, extending the base NSO functionality.
* Multiple `.tar.gz.signature` files: Digital signatures for the bundled packages.

Since NSO version 6.3, a few additional NSO packages are included. They contain the following platform tools:

* HCC
* Observability Exporter
* Phased Provisioning
* Resource Manager

For platform tools documentation, refer to the individual package's `README` file or to the [online documentation](https://nso-docs.cisco.com/resources).

**NED packages**

The NED packages that are available with the NSO installation are NetSim-based example NEDs. These NEDs are used for NSO examples only.

Fetch the latest production-grade NEDs from [Cisco Software Download](https://software.cisco.com/download/home) using the URLs provided on your NED license certificates.

**Manual pages**

The installation program unpacks the NSO manual pages from the documentation archive in `$NCS_DIR/man`. `ncsrc` makes an addition to `$MANPATH`, allowing you to use the `man` command to view them. The manual pages are available in PDF format and from the online documentation located on [NCS man-pages, Volume 1](/guides/resources/man) in Manual Pages.

Following is a list of a few of the installed manual pages:

* `ncs(1)`: Command to start and control the NSO daemon.
* `ncsc(1)`: NSO Yang compiler.
* `ncs_cli(1)`: Frontend to the NSO CLI engine.
* `ncs-netsim(1)`: Command to create and manipulate a simulated network.
* `ncs-setup(1)`: Command to create an initial NSO setup.
* `ncs.conf`: NSO daemon configuration file format.

For example, to view the manual page describing the NSO configuration file, you should type:

```bash
$ man ncs.conf
```

Apart from the manual pages, extensive information about command-line options can be obtained by running `ncs` and `ncsc` with the `--help` (abbreviated `-h`) flag.

```bash
$ ncs --help
```

```bash
$ ncsc --help
```

**Installer help**

Run the `sh nso-VERSION.darwin.x86_64.installer.bin --help` command to view additional help on running binaries. More details can be found in the [ncs-installer(1)](/guides/resources/man/ncs-installer.1) Manual Page included with NSO.

Notice the two options for `--local-install` or `--system-install`. An example output is shown below.

```bash
sh nso-6.0.darwin.x86_64.installer.bin --help

# Output
This is the NCS installation script.
Usage: ./nso-6.0.darwin.x86_64.installer.bin [--local-install] LocalInstallDir
Installs NCS in the LocalInstallDir directory only.
This is convenient for test and development purposes.
Usage: ./nso-6.0.darwin.x86_64.installer.bin --system-install
[--install-dir InstallDir]
[--config-dir ConfigDir] [--run-dir RunDir] [--log-dir LogDir]
[--run-as-user User] [--keep-ncs-setup] [--non-interactive]

Does a system install of NCS, suitable for deployment.
Static files are installed in InstallDir/ncs-<vsn>.
The first time --system-install is used, the ConfigDir,
RunDir, and LogDir directories are also created and
populated for config files, run-time state files, and log files,
respectively, and an init script for start of NCS at system boot
and user profile scripts are installed. Defaults are:

InstallDir - /opt/ncs
ConfigDir  - /etc/ncs
RunDir     - /var/opt/ncs
LogDir     - /var/log/ncs

By default, the system install will run NCS as the root user.
If the --run-as-user option is given, the system install will
instead run NCS as the given user. The user will be created if
it does not already exist.
If the --non-interactive option is given, the installer will
proceed with potentially disruptive changes (e.g. modifying or
removing existing files) without asking for confirmation.
```

</details>

### Step 4 - Run the Installer <a href="#li.run.the.installer" id="li.run.the.installer"></a>

Local Install of NSO software is performed in a single user-specified directory, for example in your `$HOME` directory.

{% hint style="success" %}
It is always recommended to install NSO in a directory named as the version of the release, for example, if the version being installed is `6.1`, the directory should be `~/nso-6.1`.
{% endhint %}

To run the installer:

1. Navigate to your Install Directory.
2. Run the command given below to install NSO in your Install Directory. The `--local-install` parameter is optional. At this point, you can choose to install NSO in standard mode or in FIPS mode.

{% tabs %}
{% tab title="Standard Local Install" %}
The standard mode is the regular NSO install and is suitable for most installations. FIPS is disabled in this mode.

For standard NSO install, run the installer as below:

```bash
$ sh nso-VERSION.OS.ARCH.installer.bin $HOME/ncs-VERSION --local-install
```

An example output is shown below:

{% code title="Example: Standard Local Install" %}

```bash
sh nso-6.0.darwin.x86_64.installer.bin --local-install ~/nso-6.0

# Output
INFO  Using temporary directory /var/folders/90/n5sbctr922336_
0jrzhb54400000gn/T//ncs_installer.93831 to stage NCS installation bundle
INFO  Unpacked ncs-6.0 in /Users/user/nso-6.0
INFO  Found and unpacked corresponding DOCUMENTATION_PACKAGE
INFO  Found and unpacked corresponding EXAMPLE_PACKAGE
INFO  Found and unpacked corresponding JAVA_PACKAGE
INFO  Generating default SSH hostkey (this may take some time)
INFO  SSH hostkey generated
INFO  Environment set-up generated in /Users/user/nso-6.0/ncsrc
INFO  NSO installation script finished
INFO  Found and unpacked corresponding NETSIM_PACKAGE
INFO  NCS installation complete
```

{% endcode %}
{% endtab %}

{% tab title="FIPS Local Install" %}
FIPS mode restricts cryptographic operations to those provided by the CiscoSSL FIPS 140-3 validated module. **This mode should only be enabled for deployments subject to strict regulatory compliance requirements**, as it limits the available cryptographic functions to those certified under FIPS 140-3 standards.

**Installation Procedure**

To perform a FIPS-compliant NSO installation, execute the installer with the `--fips-install` flag:

```bash
$ sh nso-VERSION.OS.ARCH.installer.bin $HOME/ncs-VERSION --local-install --fips-install
```

**NSO Configuration for FIPS**

During FIPS installation, the following configurations are automatically applied:

1. **FIPS Mode Enablement**\
   The `ncs.conf` file is configured with the FIPS mode flag:

   ```xml
   <fips-mode>
       <enabled>true</enabled>
   </fips-mode>
   ```
2. **Environment Variables**\
   The `ncsrc` file is updated with FIPS-compliant environment variables:
   * `NCS_OPENSSL_CONF_INCLUDE`
   * `NCS_OPENSSL_CONF`
   * `NCS_OPENSSL_MODULES`
3. **Cryptographic Library**\
   The default `crypto.so` library is replaced with the FIPS-compliant version during installation.

**Cryptographic Algorithm Restrictions**

The CiscoSSL FIPS 140-3 validated module supports a limited subset of cryptographic algorithms compared to standard CiscoSSL. **You must configure NSO to use only FIPS-approved algorithms and cryptographic suites.**

Key configuration requirements include:

* Configuring approved algorithms in `/ncs-config/ssh/algorithm/kex` within `ncs.conf`
* Configuring device-specific algorithms in `/devices/device/ssh-algorithms/kex` within CDB

{% hint style="info" %}
The Ed25519 algorithm is **not FIPS 140-3 compliant** and must not be used in FIPS mode.
{% endhint %}

**FIPS-Approved Key Exchange Algorithms**

The following key exchange algorithms are FIPS-approved:

* `ecdh-sha2-nistp256`
* `ecdh-sha2-nistp384`
* `ecdh-sha2-nistp521`
* `diffie-hellman-group14-sha1`
* `diffie-hellman-group-exchange-sha256`

{% hint style="info" %}
Ensure that SSH keys of the correct type are configured in both `ncs.conf` and `init.xml` files.
{% endhint %}

**NED Package Compatibility**

NSO signals Network Element Drivers (NEDs) to operate in FIPS mode using Bouncy Castle FIPS libraries for Java-based components. **NED packages may require upgrading to support FIPS mode**, as older versions—particularly SSH-based NEDs—often lack:

* FIPS mode signaling capability
* Bouncy Castle FIPS library support
* Required cryptographic compliance features

Consult the NED documentation and verify compatibility before deploying in FIPS mode.
{% endtab %}
{% endtabs %}

### Step 5 - Set Environment Variables <a href="#li.set.env.variables" id="li.set.env.variables"></a>

The installation program creates a shell script file named `ncsrc` in each NSO installation, which sets the environment variables.

To set the environment variables:

1. Source the `ncsrc` file to get the environment variables settings in your shell. You may want to add this sourcing command to your login sequence, such as `.bashrc`.

   For `csh/tcsh` users, there is a `ncsrc.tcsh` file with `csh/tcsh` syntax. The example below assumes that you are using `bash`, other versions of `/bin/sh` may require that you use `.` instead of `source`.

   ```bash
   $ source $HOME/ncs-VERSION/ncsrc
   ```
2. Most users add source `~/nso-x.x/ncsrc` (where `x.x` is the NSO version) to their `~/.bash_profile`, but you can simply do it manually when you want it. Once it has been sourced, you have access to all the NSO executable commands, which start with `ncs`.

   ```bash
     ncs {TAB} {TAB}

     # Output
     ncs         ncs-maapi        ncs-project  ncs-start-python-vm  ncs_cmd
     ncs-backup  ncs-make-package ncs-setup    ncs-uninstall        ncs_conf_tool
     ncs-collect ncs-netsim       ncs-start-java-vm                 ncs_cli

     ncs_load
     ncsc
     ncs_crypto_keys-tech-report
   ```

### Step 6 - Create Runtime Directory <a href="#li.create.runtime.directory" id="li.create.runtime.directory"></a>

NSO needs a deployment/runtime directory where the database files, logs, etc. are stored. An empty default directory can be created using the `ncs-setup` command.

To create a Runtime Directory:

1. Create a Runtime Directory for NSO by running the following command. In this case, we assume that the directory is `$HOME/ncs-run`.

   ```bash
     $ ncs-setup --dest $HOME/ncs-run
   ```
2. Start the NSO daemon `ncs`.

   ```bash
   $ cd $HOME/ncs-run
   $ ncs
   ```

<details>

<summary>Runtime vs. Installation Directory</summary>

A common misunderstanding is that there exists a dependency between the Runtime Directory and the Installation Directory. This is not true. For example, say that you have two NSO local installations `path/to/nso-6.4` and `path/to/nso-6.4.1`. The following sequence runs `nso-6.4` but uses an example and configuration from `nso-6.4.1`.

```bash
  $ cd path/to/nso-6.4
  $ . ncsrc
  $ cd path/to/nso-6.4.1/examples.ncs/service-management/datacenter-qinq
  $ ncs
```

Since the Runtime Directory is self-contained, this is also the way to move between examples. And since the Runtime Directory is self-contained including the database files, you can compress a complete directory and distribute it. Unpacking that directory and starting NSO from there gives an exact copy of all instance data.

```bash
  $ cd path/to/nso-6.4.1/examples.ncs/service-management/datacenter-qinq
  $ ncs
  $ ncs --stop
  $ cd path/to/nso-6.4.1/examples.ncs/device-management/simulated-devices
  $ ncs
  $ ncs --stop
```

</details>

{% hint style="warning" %}
The `ncs-setup` command creates an `ncs.conf` file that uses predefined encryption keys for easier migration of data across installations. It is not suitable for cases where data confidentiality is required, such as a production deployment. See [Cryptographic Keys](/guides/administration/advanced-topics/cryptographic-keys) for ways to generate suitable keys.
{% endhint %}

### Step 7 - Generate License Registration Token <a href="#li.generate.license.token" id="li.generate.license.token"></a>

To conclude the NSO installation, a license registration token must be created using a (CSSM) account. This is because NSO uses Cisco Smart Licensing, as described in the [Cisco Smart Licensing](/guides/administration/management/system-management/cisco-smart-licensing) to make it easy to deploy and manage NSO license entitlements. Login credentials to the [Cisco Smart Software Manager](https://www.cisco.com/c/en/us/buy/smart-accounts/software-manager.html) (CSSM) account are provided by your Cisco contact and detailed instructions on how to [create a registration token](/guides/administration/management/system-management/cisco-smart-licensing#d5e2927) can be found in the Cisco Smart Licensing. General licensing information covering licensing models, how licensing works, usage compliance, etc., is covered in the [Cisco Software Licensing Guide](https://www.cisco.com/c/en/us/buy/licensing/licensing-guide.html).

To generate a license registration token:

1. When you have a token, start a Cisco CLI towards NSO and enter the token, for example:

   ```bash
   $ ncs_cli -Cu admin
   admin@ncs# license smart register idtoken YzIzMDM3MTgtZTRkNC00YjkxLTk2ODQt
   OGEzMTM3OTg5MG
   Registration process in progress.
   Use the 'show license status' command to check the progress and result.
   ```

   \
   Upon successful registration, NSO automatically requests a license entitlement for its own instance and for the number of devices it orchestrates and their NED types. If development mode has been enabled, only development entitlement for the NSO instance itself is requested.
2. Inspect the requested entitlements using the command `show license all` (or by inspecting the NSO daemon log). An example output is shown below.

   ```bash
   admin@ncs# show license all
   ...
   <INFO> 21-Apr-2016::11:29:18.022 miosaterm confd[8226]:
   Smart Licensing Global Notification:
   type = "notifyRegisterSuccess",
   agentID = "sa1",
   enforceMode = "notApplicable",
   allowRestricted = false,
   failReasonCode = "success",
   failMessage = "Successful."
   <INFO> 21-Apr-2016::11:29:23.029 miosaterm confd[8226]:
   Smart Licensing Entitlement Notification: type = "notifyEnforcementMode",
   agentID = "sa1",
   notificationTime = "Apr 21 11:29:20 2016",
   version = "1.0",
   displayName = "regid.2015-10.com.cisco.NSO-network-element",
   requestedDate = "Apr 21 11:26:19 2016",
   tag = "regid.2015-10.com.cisco.NSO-network-element",
   enforceMode = "inCompliance",
   daysLeft = 90,
   expiryDate = "Jul 20 11:26:19 2016",
   requestedCount = 8
   ...
   ```

<details>

<summary>Evaluation Period</summary>

If no registration token is provided, NSO enters a 90-day evaluation period, and the remaining evaluation time is recorded hourly in the NSO daemon log:

```
...
<INFO> 13-Apr-2016::13:22:29.178 miosaterm confd[16260]:
    Starting the NCS Smart Licensing Java VM
<INFO> 13-Apr-2016::13:22:34.737 miosaterm confd[16260]:
Smart Licensing evaluation time remaining: 90d 0h 0m 0s
...
<INFO> 13-Apr-2016::13:22:34.737 miosaterm confd[16260]:
  Smart Licensing evaluation time remaining: 89d 23h 0m 0s
...
```

</details>

<details>

<summary>Communication Send Error</summary>

During upgrades, if you experience the 'Communication Send Error' with license registration, restart the Smart Agent.

</details>

<details>

<summary>If You are Unable to Access Cisco Smart Software Manager</summary>

In a situation where the NSO instance has no direct access to the Cisco Smart Software Manager, one option is the [Cisco Smart Software Manager Satellite](https://software.cisco.com/software/csws/ws/platform/home) which can be installed to manage software licenses on the premises. Install the satellite and use the command `call-home destination address http <url:port>` to point to the satellite.

Another option when direct access is not desired is to configure an HTTP or HTTPS proxy, e.g., `smart-license smart-agent proxy url https://127.0.0.1:8080`. If you plan to do this, take the note below regarding ignored CLI configurations into account:

If `ncs.conf` contains a configuration for any of the java-executable, java-options, override-url/url, or proxy/url under the configure path `/ncs-config/smart-license/smart-agent/`, then any corresponding configuration done via the CLI is ignored.

</details>

<details>

<summary>License Registration in High Availability (HA) Mode</summary>

When configuring NSO in HA mode, the license registration token must be provided to the CLI running on the primary node. Read more about HA and node types in NSO [High Availability](/guides/administration/management/high-availability).

</details>

<details>

<summary>Licensing Log</summary>

Licensing activities are also logged in the NSO daemon log as described in [Monitoring NSO](/guides/administration/management/system-management#d5e7876). For example, a successful token registration results in the following log entry:

```
<INFO> 21-Apr-2016::11:29:18.022 miosaterm confd[8226]:
  Smart Licensing Global Notification:
  type = "notifyRegisterSuccess"
```

</details>

<details>

<summary>Check Registration Status</summary>

To check the registration status, use the command `show license status` An example output is shown below.

```bash
admin@ncs# show license status
Smart Licensing is ENABLED

Registration:
Status: REGISTERED
Smart Account: Network Services Orchestrator
Virtual Account: Default
Export-Controlled Functionality: Allowed
Initial Registration: SUCCEEDED on Apr 21 09:29:11 2016 UTC
Last Renewal Attempt: SUCCEEDED on Apr 21 09:29:16 2016 UTC
Next Renewal Attempt: Oct 18 09:29:16 2016 UTC
Registration Expires: Apr 21 09:26:13 2017 UTC
Export-Controlled Functionality: Allowed

License Authorization:
License Authorization:
Status: IN COMPLIANCE on Apr 21 09:29:18 2016 UTC
Last Communication Attempt: SUCCEEDED on Apr 21 09:26:30 2016 UTC
Next Communication Attempt: Apr 21 21:29:32 2016 UTC
Communication Deadline: Apr 21 09:26:13 2017 UTC
```

</details>

## Local Install FAQs

Frequently Asked Questions (FAQs) about Local Install.

<details>

<summary>Is there a dependency between the NSO Installation Directory and Runtime Directory?</summary>

No, there is no such dependency.

</details>

<details>

<summary>Do you need to source the <code>ncsrc</code> file before starting NSO?</summary>

Yes.

</details>

<details>

<summary>Can you start NSO from a directory that is not an NSO runtime directory?</summary>

No. To start NSO, you need to point to the run directory.

</details>

<details>

<summary>Can you stop NSO from a directory that is not an NSO runtime directory?</summary>

Yes.

</details>

<details>

<summary>Can we move NSO Installation from one folder to another?</summary>

Yes. You can move the directory where you installed NSO to a new location in your directory tree. Simply move NSO's root directory to the new desired location and update this file: `$NCS_DIR/ncsrc` (and `ncsrc.tcsh` if you want). This is a small and handy script that sets up some environment variables for you. Update the paths to the new location. The `$NCS_DIR/bin/ncs` and `$NCS_DIR/bin/ncsc` scripts will determine the location of NSO's root directory automatically.

</details>

***

**Next Steps**

{% content-ref url="/pages/0UTfvtVEqEJvhj44Jugi" %}
[Explore the Installation](/guides/administration/installation-and-deployment/post-install-actions/explore-the-installation)
{% endcontent-ref %}


# System Install

Install NSO for production use in a system-wide deployment.

## Installation Steps

Complete the following activities in the given order to perform a System Install of NSO.

<table data-view="cards" data-full-width="false"><thead><tr><th></th><th></th><th></th></tr></thead><tbody><tr><td><strong>Prepare</strong></td><td><a href="#step-1-fulfill-system-requirements">1. Fulfill System Requirements</a><br><a href="#si.download.the.installer">2. Download Installer/NEDs</a><br><a href="#si.unpack.the.installer">3. Unpack the Installer</a></td><td></td></tr><tr><td><strong>Install</strong></td><td><a href="#si.run.the.installer">4. Run the Installer</a></td><td></td></tr><tr><td><strong>Finalize</strong></td><td><a href="#si.setup.user.access">5. Set up User Access</a><br><a href="#si.set.env.variables">6. Review Server Configuration</a><br><a href="#si.runtime.directory.creation">7. Start Server</a><br><a href="#si.generate.license.token">8. Generate License Token</a></td><td></td></tr></tbody></table>

{% hint style="info" %}
**Mode of Install**

NSO System Install can be installed in **standard mode** or in [**FIPS**](https://www.nist.gov/itl/publications-0/federal-information-processing-standards-fips)**-compliant mode**. Standard mode install supports a broader set of cryptographic algorithms, while the FIPS mode install restricts NSO to use only FIPS 140-3-validated cryptographic modules and algorithms for enhanced/regulated security and compliance. Use FIPS mode only in environments that require compliance with specific security standards, especially in U.S. federal agencies or regulated industries. For all other use cases, install NSO in standard mode.

<sup>\* FIPS: Federal Information Processing Standards</sup>
{% endhint %}

### Step 1 - Fulfill System Requirements

Start by setting up your system to install and run NSO.

To install NSO:

1. Fulfill at least the primary requirements.
2. If you intend to build and run NSO deployment examples, you also need to install additional applications listed under Additional Requirements.

{% hint style="warning" %}
Where requirements list a specific or higher version, there always exists a (small) possibility that a higher version introduces breaking changes. If in doubt whether the higher version is fully backwards compatible, always use the specific version.
{% endhint %}

<details>

<summary>Primary Requirements</summary>

Primary requirements to do a System Install include:

* A system running Linux or macOS on either the `x86_64` or `ARM64` architecture for development. 4 CPU cores minimum. Linux for production. For [FIPS](https://www.nist.gov/itl/publications-0/federal-information-processing-standards-fips) mode, OS FIPS compliance may be required depending on your specific requirements.
* GNU libc 2.24 or higher.
* Java JRE 21 or higher. Used by Cisco Smart Licensing.
* Python 3.10 or higher (3.12 recommended).
* Required and included with many Linux/macOS distributions:
  * `tar` command. Unpack the installer.
  * `gzip` command. Unpack the installer.
  * `ssh-keygen` command. Generate SSH host key.
  * `openssl` command. Generate self-signed certificates for HTTPS.
  * `find` command. Used to find out if all required libraries are available.
  * `which` command. Used by the NSO package manager.
  * `libpam.so.0`. Pluggable Authentication Module library.
  * `libexpat.so.1`. EXtensible Markup Language parsing library.
  * `libz.so.1` version 1.2.7.1 or higher. Data compression library.

</details>

<details>

<summary>Additional Requirements</summary>

Additional requirements to, for example, build and run NSO production deployment examples include:

* Java JDK 21 or higher.
* Ant 1.9.8 or higher.
* Python Setuptools is required to build the Python API.
* Often installed using the Python package installer pip:
  * Python Paramiko 2.2 or higher. To use netconf-console.
  * Python requests. Used by the RESTCONF demo scripts.
* `xsltproc` command. Used by the `support/ned-make-package-meta-data` command to generate the `package-meta-data.xml` file.
* One of the following web browsers is required for NSO GUI capabilities. The version must be supported by the vendor at the time of release.
  * Safari
  * Mozilla Firefox
  * Microsoft Edge
  * Google Chrome
* OpenSSH client applications. For example, `ssh` and `scp` commands.
* cron. Run time-based tasks, such as `logrotate`.
* `logrotate`. rotate, compress, and mail NSO and system logs.
* `rsyslog`. pass NSO logs to a local syslog managed by `rsyslogd` and pass logs to a remote node.
* `systemd` or `init.d` scripts to start and stop NSO.

</details>

<details>

<summary>FIPS Mode Entropy Requirements</summary>

The following applies if you are running a container-based setup of your FIPS install:

In containerized environments (e.g., Docker) that run on older Linux kernels (e.g., Ubuntu 18.04), `/dev/random` may block if the system’s entropy pool is low. This can lead to delays or hangs in FIPS mode, as cryptographic operations require high-quality randomness.

To avoid this:

* Prefer newer kernels (e.g., Ubuntu 22.04 or later), where entropy handling is improved to mitigate the issue.
* Or, install an entropy daemon like Haveged on the Docker host to help maintain sufficient entropy.

Check available entropy on the host system with:

```bash
cat /proc/sys/kernel/random/entropy_avail
```

A value of 256 or higher is generally considered safe. Reference: [Oracle blog post](https://blogs.oracle.com/linux/post/entropyavail-256-is-good-enough-for-everyone).

</details>

### Step 2 - Download the Installer and NEDs <a href="#si.download.the.installer" id="si.download.the.installer"></a>

To download the Cisco NSO installer and example NEDs:

1. Go to the Cisco's official [Software Download](https://software.cisco.com/download/home) site.
2. Search for the product "Network Services Orchestrator" and select the desired version.
3. There are two versions of the NSO installer, i.e. for macOS and Linux systems. For System Install, choose the Linux OS version.

<details>

<summary>Identifying the Installer</summary>

You need to know your system specifications (Operating System and CPU architecture) to choose the appropriate NSO installer.

NSO is delivered as an OS/CPU-specific signed self-extractable archive. The signed archive file has the pattern `nso-VERSION.OS.ARCH.signed.bin` that after signature verification extracts the `nso-VERSION.OS.ARCH.installer.bin` archive file, where:

* `VERSION` is the NSO version to install.
* `OS` is the Operating System (`linux` for all Linux distributions and `darwin` for macOS).
* `ARCH` is the CPU architecture, for example`x86_64`.

</details>

### Step 3 - Unpack the Installer <a href="#si.unpack.the.installer" id="si.unpack.the.installer"></a>

If your downloaded file is a `signed.bin` file, it means that it has been digitally signed by Cisco, and upon execution, you will verify the signature and unpack the `installer.bin`.

If you only have `installer.bin`, skip to the next step.

To unpack the installer:

1. In the terminal, list the binaries in the directory where you downloaded the installer, for example:

   ```bash
   cd ~/Downloads
   ls -l nso*.bin
   -rw-r--r--@ 1 user  staff   199M Dec 15 11:45 nso-6.0.linux.x86_64.installer.bin
   -rw-r--r--@ 1 user  staff   199M Dec 15 11:45 nso-6.0.linux.x86_64.signed.bin
   ```
2. Use the `sh` command to run the `signed.bin` to verify the certificate and extract the installer binary and other files. An example output is shown below.

   ```bash
   sh nso-6.0.linux.x86_64.signed.bin
   # Output
   Unpacking...
   Verifying signature...
   Downloading CA certificate from http://www.cisco.com/security/pki/certs/crcam2.cer ...
   Successfully downloaded and verified crcam2.cer.
   Downloading SubCA certificate from http://www.cisco.com/security/pki/certs/innerspace.cer ...
   Successfully downloaded and verified innerspace.cer.
   Successfully verified root, subca and end-entity certificate chain.
   Successfully fetched a public key from tailf.cer.
   Successfully verified the signature of nso-6.0.linux.x86_64.installer.bin using tailf.cer
   ```
3. List the files to check if extraction was successful.

   ```bash
   ls -l
   # Output
   -rw-r--r--  1 user  staff   1.8K Nov 29 06:05 README.signature
   -rw-r--r--  1 user  staff    12K Nov 29 06:05 cisco_x509_verify_release.py
   -rwxr-xr-x  1 user  staff   199M Nov 29 05:55 nso-6.0.linux.x86_64.installer.bin
   -rw-r--r--  1 user  staff   256B Nov 29 06:05 nso-6.0.linux.x86_64.installer.bin.signature
   -rwxr-xr-x@ 1 user  staff   199M Dec 15 11:45 nso-6.0.linux.x86_64.signed.bin
   -rw-r--r--  1 user  staff   1.4K Nov 29 06:05 tailf.cer
   ```

{% hint style="info" %}
There may also be additional files present.
{% endhint %}

<details>

<summary>Description of Unpacked Files</summary>

The following contents are unpacked:

* `nso-VERSION.OS.ARCH.installer.bin`: The NSO installer.
* `nso-VERSION.OS.ARCH.installer.bin.signature`: Signature generated for the NSO image.
* `tailf.cer`: An enclosed Cisco-signed x.509 end-entity certificate containing the public key that is used to verify the signature.
* `README.signature`: File with further details on the unpacked content and steps on how to run the signature verification program. To manually verify the signature, refer to the steps in this file.
* `cisco_x509_verify_release.py`: Python program that can be used to verify the 3-tier x.509 certificate chain and signature.
* Multiple `.tar.gz` files: Bundled packages, extending the base NSO functionality.
* Multiple `.tar.gz.signature` files: Digital signatures for the bundled packages.

Since NSO version 6.3, a few additional NSO packages are included. They contain the following platform tools:

* HCC
* Observability Exporter
* Phased Provisioning
* Resource Manager

For platform tools documentation, refer to the individual package's `README` file or to the [online documentation](https://nso-docs.cisco.com/resources).

**NED Packages**

The NED packages that are available with the NSO installation are netsim-based example NEDs. These NEDs are used for NSO examples only.

Fetch the latest production-grade NEDs from [Cisco Software Download](https://software.cisco.com/download/home) using the URLs provided on your NED license certificates.

**Manual Pages**

The installation program will unpack the NSO manual pages from the documentation archive, allowing you to use the `man` command to view them. The Manual Pages are also available in PDF format and from the online documentation located on [NCS man-pages, Volume 1](/guides/resources/man/ncs-installer.1) in Manual Pages.

Following is a list of a few of the installed manual pages:

* `ncs(1)`: Command to start and control the NSO daemon.
* `ncsc(1)`: NSO Yang compiler.
* `ncs_cli(1)`: Frontend to the NSO CLI engine.
* `ncs-netsim(1)`: Command to create and manipulate a simulated network.
* `ncs-setup(1)`: Command to create an initial NSO setup.
* `ncs.conf`: NSO daemon configuration file format.

For example, to view the manual page describing the NSO configuration file, you should type:

```bash
$ man ncs.conf
```

Apart from the manual pages, extensive information about command line options can be obtained by running `ncs` and `ncsc` with the `--help` (abbreviated `-h`) flag.

```bash
$ ncs --help
```

```bash
$ ncsc --help
```

**Installer Help**

Run the `sh nso-VERSION.linux.x86_64.installer.bin --help` command to view additional help on running binaries. More details can be found in the [ncs-installer(1)](/guides/resources/man/ncs-installer.1) Manual Page included with NSO.

Notice the two options for `--local-install` or `--system-install`.

```bash
sh nso-6.0.linux.x86_64.installer.bin --help
```

</details>

### Step 4 - Run the Installer <a href="#si.run.the.installer" id="si.run.the.installer"></a>

To run the installer:

1. Navigate to your Install Directory.
2. Run the installer with the `--system-install` option to perform System Install. This option creates an install of NSO that is suitable for production deployment. At this point, you can choose to install NSO in standard mode or in FIPS mode.

{% tabs %}
{% tab title="Standard System Install" %}
The standard mode is the regular NSO install and is suitable for most installations. FIPS is disabled in this mode.

For standard NSO install, run the installer as below.

```bash
$ sudo sh nso-VERSION.OS.ARCH.installer.bin --system-install
```

{% code title="Example: Standard System Install" %}

```bash
$ sudo sh nso-6.0.linux.x86_64.installer.bin --system-install
```

{% endcode %}
{% endtab %}

{% tab title="FIPS System Install" %}
FIPS mode restricts cryptographic operations to those provided by the CiscoSSL FIPS 140-3 validated module. **This mode should only be enabled for deployments subject to strict regulatory compliance requirements**, as it limits the available cryptographic functions to those certified under FIPS 140-3 standards.

**Installation Procedure**

To perform a FIPS-compliant NSO installation, execute the installer with the `--fips-install` flag:

```bash
$ sudo sh nso-VERSION.OS.ARCH.installer.bin --system-install --fips-install
```

**NSO Configuration for FIPS**

During FIPS installation, the following configurations are automatically applied:

1. **FIPS Mode Enablement**\
   The `ncs.conf` file is configured with the FIPS mode flag:

   ```xml
   <fips-mode>
       <enabled>true</enabled>
   </fips-mode>
   ```
2. **Environment Variables**\
   The `ncsrc` file is updated with FIPS-compliant environment variables:
   * `NCS_OPENSSL_CONF_INCLUDE`
   * `NCS_OPENSSL_CONF`
   * `NCS_OPENSSL_MODULES`
3. **Cryptographic Library**\
   The default `crypto.so` library is replaced with the FIPS-compliant version during installation.

**Cryptographic Algorithm Restrictions**

The CiscoSSL FIPS 140-3 validated module supports a limited subset of cryptographic algorithms compared to standard CiscoSSL. **You must configure NSO to use only FIPS-approved algorithms and cryptographic suites.**

Key configuration requirements include:

* Configuring approved algorithms in `/ncs-config/ssh/algorithm/kex` within `ncs.conf`
* Configuring device-specific algorithms in `/devices/device/ssh-algorithms/kex` within CDB

{% hint style="info" %}
The Ed25519 algorithm is **not FIPS 140-3 compliant** and must not be used in FIPS mode.
{% endhint %}

**FIPS-Approved Key Exchange Algorithms**

The following key exchange algorithms are FIPS-approved:

* `ecdh-sha2-nistp256`
* `ecdh-sha2-nistp384`
* `ecdh-sha2-nistp521`
* `diffie-hellman-group14-sha1`
* `diffie-hellman-group-exchange-sha256`

{% hint style="info" %}
Ensure that SSH keys of the correct type are configured in both `ncs.conf` and `init.xml` files.
{% endhint %}

**NED Package Compatibility**

NSO signals Network Element Drivers (NEDs) to operate in FIPS mode using Bouncy Castle FIPS libraries for Java-based components. **NED packages may require upgrading to support FIPS mode**, as older versions—particularly SSH-based NEDs—often lack:

* FIPS mode signaling capability
* Bouncy Castle FIPS library support
* Required cryptographic compliance features

Consult the NED documentation and verify compatibility before deploying in FIPS mode.
{% endtab %}
{% endtabs %}

<details>

<summary>Default Directories and Scripts</summary>

The System Install by default creates the following directories:

* The Installation Directory is created in `/opt/ncs`, where the distribution is available.
* The Configuration Directory is created in `/etc/ncs`, where the `ncs.conf` file, SSH keys, and WebUI certificates are created.
* The Running Directory is created in `/var/opt/ncs`, where runtime state files, CDB database, and packages are created.
* The Log Directory is created in `/var/log/ncs`, where the log files are populated.
* System-wide environment variables are created in `/etc/profile.d/ncs.sh`.
* The installer creates a `systemd` system service script in `/etc/systemd/system/ncs.service` and enables the NSO service to start at boot, but the service is *not* started immediately. See the steps below for starting NSO after installation and before rebooting.
* To allow package reload and to set other variables when starting NSO, an environment file called `/etc/ncs/ncs.systemd.conf` is created. This file is owned by the user that starts NSO.

For the `--system-install` option, you can also choose a user-defined (non-default) Installation Directory, Configuration Directory, Running Directory, and Log Directory with `--install-dir`, `--config-dir`, `--run-dir` and `--log-dir` parameters, and specify that NSO should run as a different user than root with the `--run-as-user` parameter.

If you choose a non-default Installation Directory by using `--install-dir`, you need to specify `--install-dir` for subsequent installs and also for backup and restore.

Use the `--ignore-init-scripts` option to disable provisioning the `systemd` system service.

If a legacy SysV service exists in `/etc/init.d/ncs` when installing in interactive mode, the user will be prompted to continue using the old SysV service behavior or prepare a `systemd` service. In non-interactive mode, a `systemd` service will be prepared where a `/etc/systemd/system/ncs.service.prepare` file is created. The service is not enabled to start at boot. To enable it, rename it to `/etc/systemd/system/ncs.service` and remove the old `/etc/init.d/ncs` SysV service. When using the `--non-interactive` option, the `/etc/systemd/system/ncs.service` file will be overwritten if it already exists.

For more information on the `ncs-installer`, see the [ncs-installer(1)](/guides/resources/man/ncs-installer.1) man page.

For an extensive guide to NSO deployment, refer to [Development to Production Deployment](/guides/administration/installation-and-deployment/development-to-production-deployment)*.*

</details>

<details>

<summary>Use NSO Memory Monitoring to Capture Debug Dumps Before an OOM Kill</summary>

NSO can monitor memory through the `/ncs-config/memory-management` section in `ncs.conf`. Configure it to trigger one or more debug dumps before memory pressure reaches the point where the Linux OOM-killer might terminate NSO without leaving useful diagnostics.

This feature can be used while leaving the host in Linux's default heuristic overcommit mode (`vm.overcommit_memory=0`); see [proc\_sys\_vm(5)](https://man7.org/linux/man-pages/man5/proc_sys_vm.5.html). In that mode, the kernel's allocation check is weak and there is still a risk that a process gets OOM-killed, so proactive debug dumps help preserve diagnostic information.

* On a regular host, NSO uses total memory and available memory (`MemAvailable`), which excludes caches.
* When NSO runs in a container, NSO uses the container cgroup memory limit and current usage instead of host-wide memory values.

Each `action` under `/ncs-config/memory-management/actions` defines:

* One threshold, either `used-memory-threshold-percentage` or `free-memory-threshold-bytes`.
* One compensating action, currently `debug-dump`.
* A required absolute dump `directory`.
* Optional rate limiting with `count` (default `5`) and `cooldown-period` (default `PT60M`).

When a threshold is crossed, NSO writes a timestamped debug dump such as `debug_dump_2026-04-17T09:12:34.567Z`, logs that it is creating the dump, and raises the `memory-management-action-triggered` alarm.

**Recommended Usage**

* Configure the feature in `/etc/ncs/ncs.conf` before starting NSO, or reload the configuration with `ncs --reload` after updating `ncs.conf`.
* Use a persistent directory that the NSO user can write to, typically under `NCS_RUN_DIR=/var/opt/ncs`, for example `/var/opt/ncs/debug-dumps`.
* For container-specific guidance, including mounted volumes, container memory limits, and dump locations, see [Use NSO Memory Monitoring to Capture Debug Dumps Before a Container OOM Kill](/guides/administration/installation-and-deployment/containerized-nso#d5e8605).
* Set the threshold early enough that NSO still has time to finish the dump. A good starting point is `90` for `used-memory-threshold-percentage`, or a `free-memory-threshold-bytes` value that leaves a few GiB of headroom on larger systems.
* Define more than one action if you want an early snapshot and then additional snapshots closer to the limit.
* Swapping may delay an OOM event, but it causes severe performance degradation and should not be part of the normal operating plan for NSO.
* Keep `NCS_DUMP` configured as well. If NSO runs as a non-root user, point it to a writable persistent location, typically under `NCS_RUN_DIR=/var/opt/ncs`. A proactive debug dump helps when the Linux OOM-killer would otherwise terminate NSO without producing a system dump.

**Example Configuration**

Add the following under the top-level `<ncs-config>` element in `ncs.conf`:

{% code title="ncs.conf memory-management example" %}

```xml
<memory-management>
  <actions>
    <action>
      <name>early-warning</name>
      <used-memory-threshold-percentage>90</used-memory-threshold-percentage>
      <debug-dump>
        <count>3</count>
        <cooldown-period>PT5M</cooldown-period>
        <directory>/var/opt/ncs/debug-dumps</directory>
      </debug-dump>
    </action>
    <action>
      <name>critical-free-memory</name>
      <free-memory-threshold-bytes>2147483648</free-memory-threshold-bytes>
      <debug-dump>
        <count>2</count>
        <cooldown-period>PT1M</cooldown-period>
        <directory>/var/opt/ncs/debug-dumps</directory>
      </debug-dump>
    </action>
  </actions>
</memory-management>
```

{% endcode %}

In the example above, the first action triggers when memory usage reaches 90% of the monitored limit. The second action triggers when less than 2 GiB remain available. Use either one threshold type or both, depending on how you size and operate the system.

**Verification**

After starting NSO or reloading the configuration:

* Check that the dump directory exists and is writable by the NSO user.
* When a threshold is crossed, look for `creating debug dump` in `ncs.log`.
* Confirm that a new file appears in the configured directory.
* Check `show alarms alarm-list` for the `memory-management-action-triggered` alarm.

{% hint style="info" %}
This feature does not stop the Linux OOM-killer by itself, but it prevents an OOM situation from leaving you without a debug dump from NSO.
{% endhint %}

{% hint style="warning" %}
Ensure that both the `/ncs-config/memory-management/actions/action/debug-dump/directory` path and the directory used for `NCS_DUMP` exist and are writable by the NSO user.

By default, NSO writes a system dump to the NSO run-time directory, typically `NCS_RUN_DIR=/var/opt/ncs` for a System Install. If that directory is not suitable, for example, if it is not writable by the NSO user or if you want unique dump names, set `NCS_DUMP` to a dump file path in a suitable directory, for example `NCS_DUMP="/path/to/dump-dir/ncs_crash.dump.$(date +%Y%m%d-%H%M%S)"`. For container-specific persistence guidance, see [Use NSO Memory Monitoring to Capture Debug Dumps Before a Container OOM Kill](/guides/administration/installation-and-deployment/containerized-nso#d5e8605).
{% endhint %}

</details>

{% hint style="info" %}
Some older NSO releases expect the `/etc/init.d/` folder to exist in the host operating system. If the folder does not exist, the installer may fail to successfully install NSO. A workaround that allows the installer to proceed is to create the folder manually, but the NSO process will not automatically start at boot.
{% endhint %}

### Step 5 - Set Up User Access <a href="#si.setup.user.access" id="si.setup.user.access"></a>

The installation is configured for PAM authentication, with group assignment based on the OS group database (e.g. `/etc/group` file). Users that need access to NSO must belong to either the `ncsadmin` group (for unlimited access rights) or the `ncsoper` group (for minimal access rights).

To set up user access:

1. To create the `ncsadmin` group, use the OS shell command:

   ```bash
   # groupadd ncsadmin
   ```
2. To create the `ncsoper` group, use the OS shell command:

   ```bash
   # groupadd ncsoper
   ```
3. To add an existing user to one of these groups, use the OS shell command:

   ```bash
   # usermod -a -G 'groupname' 'username'
   ```

### Step 6 - Review Server Configuration <a href="#si.set.env.variables" id="si.set.env.variables"></a>

Open the `/etc/ncs/ncs.conf` daemon configuration file in a text editor and review its contents. At the least, you should decide which northbound management interfaces to enable (if any). Refer to [System Management (ncs.conf)](https://nso-docs.cisco.com/guides/administration/installation-and-deployment/pages/hgWUBFw1TA0R6WyLxOgc#ncs.conf-file) for details.

For example, to enable the WebUI over HTTPS, find the `webui/transport/ssl` section and set `enabled` to `true`:

{% code title="Example: Enable Web UI Access" %}

```xml
  <webui>
    <!-- ... -->
    <transport>
      <!-- ... -->
      <ssl>
        <enabled>true</enabled>
```

{% endcode %}

Note that in the default configuration, all northbound interfaces are disabled unless enabled by a specific environment variable, such as `NCS_WEBUI_TRANSPORT_SSL`.

### Step 7 - Start Server <a href="#si.runtime.directory.creation" id="si.runtime.directory.creation"></a>

In a System Install, NSO runs as a system daemon that starts and stops with the operating system. However, right after installation, the server must be started manually.

{% hint style="info" %}
**Non-root Startup on SELinux**

If NSO was installed with `--run-as-user` and the host has SELinux enabled, starting the service may fail with an error similar to `/bin/su: Permission denied` due to missing SELinux permissions. Consider updating the SELinux policy for NSO or starting NSO unconfined. You may verify the SELinux mode with the command `getenforce`.

If the issue persists, you may consider disabling SELinux on the host before starting NSO as a non-root user. Set `SELINUX=disabled` in `/etc/sysconfig/selinux` and reboot the host before retrying the startup. Note that disabling SELinux significantly lowers the host security posture and should be done only with care.
{% endhint %}

1. Change to Super User privileges.

   ```bash
   $ sudo -s
   ```
2. Start NSO.

   ```bash
   # systemctl daemon-reload
   # systemctl start ncs
   ```

   The Runtime Directory for System Install is set up automatically; you do not need to create it. From now on, the NSO daemon `ncs` is automatically started at boot time.
3. The installation program creates a shell script file which sets the environment variables needed to run NSO. With the `--system-install` option, by default, these variables are set on shell startup. To explicitly set the variables, source `ncs.sh` or `ncs.csh` depending on your shell type.

   ```bash
   # source /etc/profile.d/ncs.sh
   ```
4. Once you log on with the user that belongs to the `ncsadmin` or `ncsoper` group, you can directly access the CLI as shown below:

   ```bash
   $ ncs_cli -C
   ```

### Step 8 - Generate License Registration Token <a href="#si.generate.license.token" id="si.generate.license.token"></a>

To conclude the NSO installation, a license registration token must be created using a (CSSM) account. This is because NSO uses [Cisco Smart Licensing](/guides/administration/management/system-management/cisco-smart-licensing) to make it easy to deploy and manage NSO license entitlements. Login credentials to the [Cisco Smart Software Manager](https://www.cisco.com/c/en/us/buy/smart-accounts/software-manager.html) (CSSM) account are provided by your Cisco contact and detailed instructions on how to [create a registration token](/guides/administration/management/system-management/cisco-smart-licensing#d5e2927) can be found in the Cisco Smart Licensing. General licensing information covering licensing models, how licensing works, usage compliance, etc., is covered in the [Cisco Software Licensing Guide](https://www.cisco.com/c/en/us/buy/licensing/licensing-guide.html).

To generate a license registration token:

1. When you have a token, start a Cisco CLI towards NSO and enter the token, for example:

   ```bash
   $ ncs_cli -Cu admin
   admin@ncs# license smart register idtoken
   YzIzMDM3MTgtZTRkNC00YjkxLTk2ODQtOGEzMTM3OTg5MG

   Registration process in progress.
   Use the 'show license status' command to check the progress and result.
   ```

   \
   Upon successful registration, NSO automatically requests a license entitlement for its own instance and for the number of devices it orchestrates and their NED types. If development mode has been enabled, only development entitlement for the NSO instance itself is requested.
2. Inspect the requested entitlements using the command `show license all` (or by inspecting the NSO daemon log). An example output is shown below.

   ```bash
   admin@ncs# show license all
   ...
   <INFO> 21-Apr-2016::11:29:18.022 miosaterm confd[8226]:
   Smart Licensing Global Notification:
   type = "notifyRegisterSuccess",
   agentID = "sa1",
   enforceMode = "notApplicable",
   allowRestricted = false,
   failReasonCode = "success",
   failMessage = "Successful."
   <INFO> 21-Apr-2016::11:29:23.029 miosaterm confd[8226]:
   Smart Licensing Entitlement Notification: type = "notifyEnforcementMode",
   agentID = "sa1",
   notificationTime = "Apr 21 11:29:20 2016",
   version = "1.0",
   displayName = "regid.2015-10.com.cisco.NSO-network-element",
   requestedDate = "Apr 21 11:26:19 2016",
   tag = "regid.2015-10.com.cisco.NSO-network-element",
   enforceMode = "inCompliance",
   daysLeft = 90,
   expiryDate = "Jul 20 11:26:19 2016",
   requestedCount = 8
   ...
   ```

<details>

<summary>Evaluation Period</summary>

If no registration token is provided, NSO enters a 90-day evaluation period and the remaining evaluation time is recorded hourly in the NSO daemon log:

```
      ...
<INFO> 13-Apr-2016::13:22:29.178 miosaterm confd[16260]:
Starting the NCS Smart Licensing Java VM
<INFO> 13-Apr-2016::13:22:34.737 miosaterm confd[16260]:
Smart Licensing evaluation time remaining: 90d 0h 0m 0s
...
<INFO> 13-Apr-2016::13:22:34.737 miosaterm confd[16260]:
Smart Licensing evaluation time remaining: 89d 23h 0m 0s
...
```

</details>

<details>

<summary>Communication Send Error</summary>

During upgrades, if you experience a 'Communication Send Error' during license registration, restart the Smart Agent.

</details>

<details>

<summary>If You are Unable to Access Cisco Smart Software Manager</summary>

In a situation where the NSO instance has no direct access to the Cisco Smart Software Manager, one option is the [Cisco Smart Software Manager Satellite](https://software.cisco.com/software/csws/ws/platform/home) which can be installed to manage software licenses on the premises. Install the satellite and use the command `call-home destination address http <url:port>` to point to the satellite.

Another option when direct access is not desired is to configure an HTTP or HTTPS proxy, e.g., `smart-license smart-agent proxy url https://127.0.0.1:8080`. If you plan to do this, take the note below regarding ignored CLI configurations into account:

If `ncs.conf` contains a configuration for any of the java-executable, java-options, override-url/url, or proxy/url under the configure path `/ncs-config/smart-license/smart-agent/`, then any corresponding configuration done via the CLI is ignored.

</details>

<details>

<summary>License Registration in HA Mode</summary>

When configuring NSO in High Availability (HA) mode, the license registration token must be provided to the CLI running on the primary node. Read more about HA and node types in [High Availability](/guides/administration/management/high-availability)*.*

</details>

<details>

<summary>Licensing Log</summary>

Licensing activities are also logged in the NSO daemon log as described in [Monitoring NSO](/guides/administration/management/system-management#d5e7876). For example, a successful token registration results in the following log entry:

```
<INFO> 21-Apr-2016::11:29:18.022 miosaterm confd[8226]:
Smart Licensing Global Notification:
type = "notifyRegisterSuccess"
```

</details>

<details>

<summary>Check Registration Status</summary>

To check the registration status, use the command `show license status`.

```bash
admin@ncs# show license status

Smart Licensing is ENABLED

Registration:
Status: REGISTERED
Smart Account: Network Services Orchestrator
Virtual Account: Default
Export-Controlled Functionality: Allowed
Initial Registration: SUCCEEDED on Apr 21 09:29:11 2016 UTC
Last Renewal Attempt: SUCCEEDED on Apr 21 09:29:16 2016 UTC
Next Renewal Attempt: Oct 18 09:29:16 2016 UTC
Registration Expires: Apr 21 09:26:13 2017 UTC
Export-Controlled Functionality: Allowed

License Authorization:

License Authorization:
Status: IN COMPLIANCE on Apr 21 09:29:18 2016 UTC
Last Communication Attempt: SUCCEEDED on Apr 21 09:26:30 2016 UTC
Next Communication Attempt: Apr 21 21:29:32 2016 UTC
Communication Deadline: Apr 21 09:26:13 2017 UTC
```

</details>

## System Install FAQs

Frequently Asked Questions (FAQs) about System Install.

<details>

<summary>Is there a dependency between the NSO Installation Directory and Runtime Directory?</summary>

No, there is no such dependency.

</details>

<details>

<summary>Do you need to source the <code>ncsrc</code> file before starting NSO?</summary>

No. By default, the environment variables are configured and set on the shell with System Install.

</details>

<details>

<summary>Can you start NSO from a directory that is not an NSO runtime directory?</summary>

Yes.

</details>

<details>

<summary>Can you stop NSO from a directory that is not an NSO runtime directory?</summary>

Yes.

</details>

<details>

<summary>For evaluation and development purposes, instead of a Local Install, you performed a System Install. Now you cannot build or run NSO examples as described in README files. How can you proceed further?</summary>

The easiest way is to uninstall the System install using `ncs-uninstall --all` and do a Local Install from scratch.

</details>

<details>

<summary>Can we move NSO Installation from one folder to another ?</summary>

No.

</details>

***

**Next Steps**

{% content-ref url="/pages/hgWUBFw1TA0R6WyLxOgc" %}
[System Management](/guides/administration/management/system-management)
{% endcontent-ref %}


# Post-Install Actions

Perform actions and activities possible after installing NSO.

The following actions are possible after installing NSO.

## After Local Install

{% content-ref url="/pages/0UTfvtVEqEJvhj44Jugi" %}
[Explore the Installation](/guides/administration/installation-and-deployment/post-install-actions/explore-the-installation)
{% endcontent-ref %}

{% content-ref url="/pages/4v0vE3z0KOlKWs1D1BnK" %}
[Start and Stop NSO](/guides/administration/installation-and-deployment/post-install-actions/start-stop-nso)
{% endcontent-ref %}

{% content-ref url="/pages/5YoUvbBwGoRxBonKbueS" %}
[Create NSO Instance](/guides/administration/installation-and-deployment/post-install-actions/create-nso-instance)
{% endcontent-ref %}

{% content-ref url="/pages/fggvogrFs6CfkCSMlJYz" %}
[Enable Development Mode](/guides/administration/installation-and-deployment/post-install-actions/enable-development-mode)
{% endcontent-ref %}

{% content-ref url="/pages/mTmkoQ4vdX8kNZEE3QKq" %}
[Running NSO Examples](/guides/administration/installation-and-deployment/post-install-actions/running-nso-examples)
{% endcontent-ref %}

{% content-ref url="/pages/HxbTVOTMq6KqD0BwV9q9" %}
[Migrate to System Install](/guides/administration/installation-and-deployment/post-install-actions/migrate-to-system-install)
{% endcontent-ref %}

{% content-ref url="/pages/vb7TJdQJv7gyHbWZmncS" %}
[Uninstall Local Install](/guides/administration/installation-and-deployment/post-install-actions/uninstall-local-install)
{% endcontent-ref %}

## After System Install

{% content-ref url="/pages/YfNefSSEIlHFDpSATnuA" %}
[Modify Examples for System Install](/guides/administration/installation-and-deployment/post-install-actions/modify-examples-for-system-install)
{% endcontent-ref %}

{% content-ref url="/pages/KsK0thxeGzqYKJdTbx8v" %}
[Uninstall System Install](/guides/administration/installation-and-deployment/post-install-actions/uninstall-system-install)
{% endcontent-ref %}


# Explore the Installation

Explore NSO contents after finishing the installation.

{% hint style="warning" %}
Applies to Local Install.
{% endhint %}

Before starting NSO, it is recommended to explore the installation contents.

Navigate to the newly created Installation Directory, for example:

```bash
cd ~/nso-6.0
```

## Contents of the Installation Directory

The installation directory includes the following contents:

* [Documentation](#d5e552)
* [Examples](#d5e560)
* [Network Element Drivers](#d5e564)
* [Shell scripts](#d5e604)

### Documentation <a href="#d5e552" id="d5e552"></a>

Along with the binaries, NSO installs a full set of documentation available in the `doc/` folder in the Installation Directory. There is also an [online version](https://nso-docs.cisco.com/guides).

```bash
ls -l doc/
drwxr-xr-x   5 user  staff   160B Nov 29 05:19 api/
drwxr-xr-x  14 user  staff   448B Nov 29 05:19 html/
-rw-r--r--   1 user  staff   202B Nov 29 05:19 index.html
drwxr-xr-x  17 user  staff   544B Nov 29 05:19 pdf/
```

Run `index.html` in your browser to explore further.

### Examples <a href="#d5e560" id="d5e560"></a>

Local Install comes with a rich set of [examples](https://github.com/NSO-developer/nso-examples/tree/6.7) to start using NSO.

```bash
$ ls -1 examples.ncs/
README.md
aaa
common
device-management
getting-started
high-availability
layered-services-architecture
misc
nano-services
northbound-interfaces
scaling-performance
sdk-api
service-management
```

### Network Element Drivers (NEDs) <a href="#d5e564" id="d5e564"></a>

In order to communicate with the network, NSO uses NEDs as device drivers for different device types. Cisco has NEDs for hundreds of different devices available for customers, and several are included in the installer in the `/packages/neds` directory.

In the example below, NEDs for Cisco ASA, IOS, IOS XR, and NX-OS are shown. Also included are NEDs for other vendors including Juniper JunOS, A10, ALU, and Dell.

```bash
$ ls -1 packages/neds
a10-acos-cli-3.0
alu-sr-cli-3.4
cisco-asa-cli-6.6
cisco-ios-cli-3.0
cisco-ios-cli-3.8
cisco-iosxr-cli-3.0
cisco-iosxr-cli-3.5
cisco-nx-cli-3.0
dell-ftos-cli-3.0
juniper-junos-nc-3.0
```

{% hint style="info" %}
The example NEDs included in the installer are intended for evaluation, demonstration, and use with the [examples.ncs](https://github.com/NSO-developer/nso-examples/tree/6.7) examples. These are not the latest versions available and often do not have all the features available in production NEDs.
{% endhint %}

#### **Install New NEDs**

A large number of pre-built supported NEDs are available which can be acquired and downloaded by the customers from [Cisco Software Download](https://software.cisco.com/). Note that the specific file names and versions that you download may be different from the ones in this guide. Therefore, remember to update the paths accordingly.

Like the NSO installer, the NEDs are `signed.bin` files that need to be run to validate the download and extract the new code.

To install new NEDs:

1. Change to the working directory where your downloads are. The filenames indicate which version of NSO the NEDs are pre-compiled for (in this case NSO 6.0), and the version of the NED. An example output is shown below.

   ```bash
   cd ~/Downloads/
   ls -l ncs*.bin

   # Output
   -rw-r--r--@ 1 user  staff   9708091 Dec 18 12:05 ncs-6.0-cisco-asa-6.7.7.signed.bin
   -rw-r--r--@ 1 user  staff  51233042 Dec 18 12:06 ncs-6.0-cisco-ios-6.42.1.signed.bin
   -rw-r--r--@ 1 user  staff   8400190 Dec 18 12:05 ncs-6.0-cisco-nx-5.13.1.1.signed.bin
   ```
2. Use the `sh` command to run `signed.bin` to verify the certificate and extract the NED tar.gz and other files. Repeat for all files. An example output is shown below.

   ```bash
   sh ncs-6.0-cisco-nx-5.13.1.1.signed.bin 
    
     Unpacking...  
     Verifying signature...
     Downloading CA certificate from http://www.cisco.com/security/pki/certs/crcam2.cer ...
     Successfully downloaded and verified crcam2.cer.
     Downloading SubCA certificate from http://www.cisco.com/security/pki/certs/innerspace.cer ...
     Successfully downloaded and verified innerspace.cer.
     Successfully verified root, subca and end-entity certificate chain.
     Successfully fetched a public key from tailf.cer.
     Successfully verified the signature of ncs-6.0-cisco-nx-5.13.1.1.tar.gz using tailf.cer
   ```
3. You now have three tar (.`tar.gz`) files. These are compressed versions of the NEDs. List the files to verify as shown in the example below.

   ```bash
   ls -l ncs*.tar.gz
   -rw-r--r--  1 user  staff   9704896 Dec 12 21:11 ncs-6.0-cisco-asa-6.7.7.tar.gz
   -rw-r--r--  1 user  staff  51260488 Dec 13 22:58 ncs-6.0-cisco-ios-6.42.1.tar.gz
   -rw-r--r--  1 user  staff   8409288 Dec 18 09:09 ncs-6.0-cisco-nx-5.13.1.1.tar.gz
   ```
4. Navigate to the `packages/neds` directory for your Local Install, for example:

   ```bash
   cd ~/nso-6.0/packages/neds
   ```
5. In the `/packages/neds` directory, extract the .tar files into this directory using the `tar` command with the path to where the compressed NED is located. An example is shown below.

   ```
   tar -zxvf ~/Downloads/ncs-6.0-cisco-nx-5.13.1.1.tar.gz
   tar -zxvf ~/Downloads/ncs-6.0-cisco-ios-6.42.1.tar.gz
   tar -zxvf ~/Downloads/ncs-6.0-cisco-asa-6.7.7.tar.gz
   ```

   \
   Here is a sample list of the newer NEDs extracted along with the ones bundled with the installation:

   ```
   drwxr-xr-x  13 user  staff   416 Nov 29 05:17 a10-acos-cli-3.0
   drwxr-xr-x  12 user  staff   384 Nov 29 05:17 alu-sr-cli-3.4
   drwxr-xr-x  13 user  staff   416 Nov 29 05:17 cisco-asa-cli-6.6
   drwxr-xr-x  13 user  staff   416 Dec 12 21:11 cisco-asa-cli-6.7
   drwxr-xr-x  12 user  staff   384 Nov 29 05:17 cisco-ios-cli-3.0
   drwxr-xr-x  12 user  staff   384 Nov 29 05:17 cisco-ios-cli-3.8
   drwxr-xr-x  13 user  staff   416 Dec 13 22:58 cisco-ios-cli-6.42
   drwxr-xr-x  13 user  staff   416 Nov 29 05:17 cisco-iosxr-cli-3.0
   drwxr-xr-x  13 user  staff   416 Nov 29 05:17 cisco-iosxr-cli-3.5
   drwxr-xr-x  13 user  staff   416 Nov 29 05:17 cisco-nx-cli-3.0
   drwxr-xr-x  14 user  staff   448 Dec 18 09:09 cisco-nx-cli-5.13
   drwxr-xr-x  13 user  staff   416 Nov 29 05:17 dell-ftos-cli-3.0
   drwxr-xr-x  10 user  staff   320 Nov 29 05:17 juniper-junos-nc-3.0
   ```

### Shell Scripts <a href="#d5e604" id="d5e604"></a>

The last thing to note is the files `ncsrc` and `ncsrc.tsch`. These are shell scripts for `bash` and `tsch` that set up your PATH and other environment variables for NSO. Depending on your shell, you need to source this file before starting NSO.

For more information on sourcing shell script, see the [Local Install steps](/guides/administration/installation-and-deployment/local-install).


# Start and Stop NSO

Start and stop the NSO daemon.

{% hint style="warning" %}
Applies to Local Install.
{% endhint %}

The command `ncs -h` shows various options when starting NSO. By default, NSO starts in the background without an associated terminal. It is recommended to add NSO to the `/etc/init` scripts of the deployment hosts. For more information, see the [ncs(1)](/guides/resources/man/ncs.1) in Manual Pages.

Whenever you start (or reload) the NSO daemon, it reads its configuration from `./ncs.conf` or `${NCS_DIR}/etc/ncs/ncs.conf` or from the file specified with the `-c` option. Parts of the configuration can also be placed in the `ncs.conf.d` directory that must be placed next to the actual `ncs.conf` file.

```bash
$ ncs
$ ncs --stop
$ ncs -h
...
```


# Create NSO Instance

Create a new NSO instance for Local Install.

{% hint style="warning" %}
Applies to Local Install.
{% endhint %}

One of the included scripts with an NSO installation is the `ncs-setup`, which makes it very easy to create instances of NSO from a Local Install. You can look at the `--help` or [ncs-setup(1)](/guides/resources/man/ncs-setup.1) in Manual Pages for more details, but the two options we need to know are:

* `--dest` defines the directory where you want to set up NSO. if the directory does not exist, it will be created.
* `--package` defines the NEDs that you want to have installed. You can specify this option multiple times.

{% hint style="info" %}
NCS is the original name of the NSO product. Therefore, many of the commands and application features are prefaced with `ncs`. You can think of NCS as another name for NSO.
{% endhint %}

To create an NSO instance:

1. Run the command to set up an NSO instance in the current directory with the IOS, NX-OS, IOS-XR and ASA NEDs. You only need one NED per platform that you want NSO to manage, even if you may have multiple versions in your installer `neds` directory.

   \
   Use the name of the NED folder in `${NCS_DIR}/packages/neds` for the latest NED version that you have installed for the target platform. Use the tab key to complete the path after you start typing (alternatively, copy and paste). Verify that the NED versions below match what is currently on the sandbox to avoid a syntax error. See the example below.

   ```bash
   ncs-setup --package ~/nso-6.0/packages/neds/cisco-ios-cli-6.44 \
   --package ~/nso-6.0/packages/neds/cisco-nx-cli-5.15 \
   --package ~/nso-6.0/packages/neds/cisco-iosxr-cli-7.20 \
   --package ~/nso-6.0/packages/neds/cisco-asa-cli-6.8 \
   --dest nso-instance
   ```
2. Check the `nso-instance` directory. Notice that several new files and folders are created.

   ```bash
   $ ls nso-instance/
   logs  ncs-cdb  ncs.conf  packages  README.ncs  scripts  state
   $ ls -l nso-instance/packages/
   total 0
   lrwxrwxrwx 1 user docker 51 Mar 19 12:44 cisco-asa-cli-6.8 ->
   /home/user/nso-6.0/packages/neds/cisco-asa-cli-6.8

   lrwxrwxrwx 1 user docker 52 Mar 19 12:44 cisco-ios-cli-6.44 ->
   /home/user/nso-6.0/packages/neds/cisco-ios-cli-6.44

   lrwxrwxrwx 1 user docker 54 Mar 19 12:44 cisco-iosxr-cli-7.20 ->
   /home/user/nso-6.0/packages/neds/cisco-iosxr-cli-7.20

   lrwxrwxrwx 1 user docker 51 Mar 19 12:44 cisco-nx-cli-5.15 ->
   /home/user/nso-6.0/packages/neds/cisco-nx-cli-5.15
   $
   ```

   Following is a description of the important files and folders:

   * `ncs.conf` is the NSO application configuration file and is used to customize aspects of the NSO instance (for example, to change ports, enable/disable features, and so on.) See [ncs.conf(5)](/guides/resources/man/ncs.conf.5) in Manual Pages for information.
   * `packages/` is the directory that has symlinks to the NEDs that we referenced in the `--package` arguments at the time of setup. See [NSO Packages](/guides/development/core-concepts/packages) in Development for more information.
   * `logs/` is the directory that contains all the logs from NSO. This directory is useful for troubleshooting.
3. Start the NSO instance by navigating to the `nso-instance` directory and typing the `ncs` command. You must be situated in the `nso-instance` directory each time you want to start or stop NSO. If you have multiple instances, you need to navigate to each one and use the `ncs` command to start or stop each one.
4. Verify that NSO is running by using the `ncs --status | grep status` command.

   ```bash
   $ ncs --status | grep status
   status: started
   db=running id=31 priority=1 path=/ncs:devices/device/live-status-protocol/device-type
   ```
5. Add Netsim or lab devices using the command `ncs-netsim -h`.


# Enable Development Mode

Enable your NSO instance for development purposes.

{% hint style="warning" %}
Applies to Local Install
{% endhint %}

If you intend to use your NSO instance for development purposes, enable the development mode using the command `license smart development enable`.


# Running NSO Examples

Run and interact with practice examples provided with the NSO installer.

{% hint style="warning" %}
Applies to Local Install.
{% endhint %}

This section provides an overview of how to run the examples provided with the NSO installer. By working through the examples, the reader should get a good overview of the various aspects of NSO and hands-on experience from interacting with it.

{% hint style="info" %}
This section references the examples located in [$NCS\_DIR/examples.ncs](https://github.com/NSO-developer/nso-examples/tree/6.7). The examples all have `README` files that include instructions related to the example.
{% endhint %}

## General Instructions <a href="#d5e1220" id="d5e1220"></a>

1. Make sure that NSO is installed with a Local Install according to the instructions in [Local Install](/guides/administration/installation-and-deployment/local-install).
2. Source the `ncsrc` file in the NSO installation directory to set up a local environment. For example:

   ```bash
   $ source ~/nso-6.0/ncsrc
   ```
3. Proceed to the example directory:

   ```bash
   $ cd $NCS_DIR/examples.ncs/device-management/simulated-devices
   ```
4. Follow the instructions in the `README` files that are located in the example directories.

Every example directory is a complete NSO run-time directory. The README file and the detailed instructions later in this guide show how to generate a simulated network and NSO configuration for running the specific examples. Basically, the following steps are done:

1. Create a simulated network using the `ncs-netsim --create-network` command:

   ```bash
   $ ncs-netsim create-network cisco-ios-cli-3.8 3 ios
   ```

   This creates 3 Cisco IOS devices called `ios0`, `ios1`, and `ios2`.
2. Create an NSO run-time environment using the `ncs-setup` command:

   ```bash
   $ ncs-setup --dest .
   ```

   This command uses the `--dest` option to create local directories for logs, database files, and the NSO configuration file to the current directory (note that `.` refers to the current directory).
3. Start NCS netsim:

   ```bash
   $ ncs-netsim start
   ```
4. Start NSO:

   ```bash
   $ ncs
   ```

{% hint style="warning" %}
It is important to make sure that you stop `ncs` and `ncs-netsim` when moving between examples using the `stop` option of the `netsim` and the `--stop` option of the `ncs`.

```bash
$ cd $NCS_DIR/examples.ncs/device-management/simulated-devices
$ ncs-netsim start
$ ncs
$ ncs-netsim stop
$ ncs --stop
```

{% endhint %}

## Common Mistakes <a href="#d5e1275" id="d5e1275"></a>

Some of the most common mistakes are:

<details>

<summary>Not Sourcing the <code>ncsrc</code> File</summary>

You have not sourced the `ncsrc` file:

```bash
$ ncs
-bash: ncs: command not found
```

</details>

<details>

<summary>Not Starting NSO from the Runtime Directory</summary>

You are trying to start NSO from a directory that is not set up as a runtime directory.

```bash
$ ncs
Bad configuration: /etc/ncs/ncs.conf:0: "./state/packages-in-use: \
   Failed to create symlink: no such file or directory"
Daemon died status=21
```

What happened above is that NSO did not find an `ncs.conf` in the local directory, so it uses the default in `/etc/ncs/ncs.conf`. That `ncs.conf` says there shall be directories at `./` such as `./state` which is not true. Make sure that you `cd` to the root of the example and check that there is a `ncs.conf` file and a `cdb-dir` directory.

</details>

<details>

<summary>Having Another Instance of NSO Running</summary>

You already have another instance of NSO running (or the same with netsim):

```bash
$ ncs
Cannot bind to internal socket /tmp/nso/nso-ipc : address already in use
Daemon died status=20
$ ncs-netsim start
DEVICE c0 Cannot bind to internal socket 127.0.0.1:5010 : \
  address already in use
Daemon died status=20
FAIL
```

To resolve the above, just stop the running instance of NSO and netsim. Remember that it does not matter where you started the "running" NSO and netsim; there is no need to `cd` back to the other example before stopping.

</details>

<details>

<summary>Not Having the NetSim Device Configuration Loaded into NSO</summary>

Another problem that users run into sometimes is where the NetSim device configuration is not loaded into NSO. This can happen if the order of commands is not followed. To resolve this, remove the database files in the `ncs_cdb` directory (keep any files with the `.xml` extension). In this way, NSO will reload the XML initialization files provided by **ncs-setup**.

```bash
$ ncs --stop
$ cd ncs-cdb/
$ ls
A.cdb
C.cdb
O.cdb
S.cdb
netsim_devices_init.xml
$ rm *.cdb
$ ncs
```

</details>


# Migrate to System Install

Convert your current Local Install setup to a System Install.

{% hint style="warning" %}
Applies to Local Install.
{% endhint %}

If you already have a Local Install with existing data that you would like to convert into a System Install, the following procedure allows you to do so. However, a reverse migration from System to Local Install is not supported.

{% hint style="info" %}
It is possible to perform the migration and upgrade simultaneously to a newer NSO version, however, doing so introduces additional complexity. If you run into issues, first migrate, and then perform the upgrade.
{% endhint %}

The following procedure assumes that NSO is installed as described in the NSO Local Install process and will perform an initial System Install of the same NSO version. After following these steps, consult the NSO System Install guide for additional steps that are required for a fully functional System Install.

The procedure also assumes you are using the `$HOME/ncs-run` folder as the run directory. If this is not the case, modify the following path accordingly.

To migrate to System Install:

1. Stop the current (local) NSO instance if it is running.

   ```bash
   $ ncs --stop
   ```
2. Take a complete backup of the Runtime Directory for potential disaster recovery.

   ```bash
   $ tar -czf $HOME/ncs-backup.tar.gz -C $HOME ncs-run
   ```
3. Change to Super User privileges.

   ```bash
   $ sudo -s
   ```
4. Start the NSO System Install.

   ```bash
   $ sh nso-VERSION.OS.ARCH.installer.bin --system-install
   ```
5. If you have multiple versions of NSO installed, verify that the symbolic link in `/opt/ncs` points to the correct version.
6. Copy the CDB files containing data to the central location.

   ```bash
   # cp $HOME/ncs-run/ncs-cdb/*.cdb /var/opt/ncs/cdb
   ```
7. Ensure that the `/var/opt/ncs/packages` directory includes all the necessary packages, appropriate for the NSO version. However, copying the packages directly could later on interfere with the operation of the `nct` command. It is better to only use symbolic links in that folder. Instead, copy the existing packages to the `/opt/ncs/packages` directory, either as directories or as tarball files. Make sure that each package includes the NSO version in its name and is not just a symlink, for example:

   ```bash
   # cd $HOME/ncs-run/packages
   # for pkg in *; do cp -RL $pkg /opt/ncs/packages/ncs-VERSION-$pkg; done
   ```
8. Link to these packages in the `/var/opt/ncs/packages` directory.

   ```bash
   # cd /var/opt/ncs/packages/
   # rm -f *
   # for pkg in /opt/ncs/packages/ncs-VERSION-*; do ln -s $pkg; done
   ```

   \
   The reason for prepending `ncs-VERSION` to the filename is to allow additional NSO commands, such as `nct upgrade` and `software packages` to work properly. These commands need to identify which NSO version a package was compiled for.
9. Edit the `/etc/ncs/ncs.conf` configuration file and make the necessary changes. If you wish to use the configuration from Local Install, disable the local authentication, unless you fully understand its security implications.

   ```xml
   <local-authentication>
     <enabled>false</enabled>
   </local-authentication>
   ```
10. When starting NSO at boot using `systemd`, make sure that you set the package reload option from the `/etc/ncs/ncs.systemd.conf` environment file to `true`. Or, for example, set `NCS_RELOAD_PACKAGES=true` before starting NSO if using the `ncs` command.

    ```bash
    # systemctl daemon-reload
    # systemctl start ncs
    ```
11. Review and complete the steps in NSO System Install, except running the installer, which you have done already. Once completed, you should have a running NSO instance with data from the Local Install.
12. Remove the package reload option if it was set.

    ```bash
    # unset NCS_RELOAD_PACKAGES
    ```
13. Update log file paths for Java and Python VM through the NSO CLI.

    ```bash
    $ ncs_cli -C -u admin
    admin@ncs# config
    Entering configuration mode terminal
    admin@ncs(config)# unhide debug
    admin@ncs(config)# show full-configuration java-vm stdout-capture file
    java-vm stdout-capture file ./logs/ncs-java-vm.log
    admin@ncs(config)# java-vm stdout-capture file /var/log/ncs/ncs-java-vm.log
    admin@ncs(config)# commit
    Commit complete.
    admin@ncs(config)# show full-configuration java-vm stdout-capture file
    java-vm stdout-capture file /var/log/ncs/ncs-java-vm.log
    admin@ncs(config)# show full-configuration python-vm logging log-file-prefix
    python-vm logging log-file-prefix ./logs/ncs-python-vm
    admin@ncs(config)# python-vm logging log-file-prefix /var/log/ncs/ncs-python-vm
    admin@ncs(config)# commit
    Commit complete.
    admin@ncs(config)# show full-configuration python-vm logging log-file-prefix
    python-vm logging log-file-prefix /var/log/ncs/ncs-python-vm
    admin@ncs(config)# exit
    admin@ncs#
    admin@ncs# exit
    ```
14. Verify that everything is working correctly.

At this point, you should have a complete copy of the previous Local Install running as a System Install. Should the migration fail at some point and you want to back out of it, the Local Install was not changed and you can easily go back to using it as before.

```bash
$ sudo systemctl stop ncs
$ source $HOME/ncs-VERSION/ncsrc
$ cd $HOME/ncs-run
$ ncs
```

In the unlikely event of Local Install becoming corrupted, you can restore it from the backup.

```bash
$ rm -rf $HOME/ncs-run
$ tar -xzf $HOME/ncs-backup.tar.gz -C $HOME
```


# Modify Examples for System Install

Alter your examples to work with System Install.

{% hint style="warning" %}
Applies to System Install.
{% endhint %}

Since all the NSO examples and README steps that come with the installer are primarily aimed at Local Install, you need to modify them to run them on a System Install.

To work with the System Install structure, this may require a little or bigger modification depending on the example.

For example, to port the [example.ncs/nano-services/basic-vrouter](https://github.com/NSO-developer/nso-examples/tree/6.7/nano-services/basic-vrouter) example to the System Install structure:

1. Make the following changes to the `basic-vrouter/ncs.conf` file:

   ```xml
   <enabled>false</enabled>
   <ip>0.0.0.0</ip>
   <port>8888</port>
   -<key-file>${NCS_DIR}/etc/ncs/ssl/cert/host.key</key-file>
   -<cert-file>${NCS_DIR}/etc/ncs/ssl/cert/host.cert</cert-file>
   +<key-file>${NCS_CONFIG_DIR}/etc/ncs/ssl/cert/host.key</key-file>
   +<cert-file>${NCS_CONFIG_DIR}/etc/ncs/ssl/cert/host.cert</cert-file>
   </ssl>
   </transport>
   ```
2. Copy the Local Install `$NCS_DIR/var/ncs/cdb/aaa_init.xml` file to the `basic-vrouter/` folder.

Other, more complex examples may require more `ncs.conf` file changes or require a copy of the Local Install default `$NCS_DIR/etc/ncs/ncs.conf` file together with the modification described above, or require the Local Install tool `$NCS_DIR/bin/ncs-setup` to be installed, as the `ncs-setup` command is usually not useful with a System Install. See [Migrate to System Install](/guides/administration/installation-and-deployment/post-install-actions/migrate-to-system-install) for more information.


# Uninstall Local Install

Remove Local Install.

{% hint style="warning" %}
Applies to Local Install.
{% endhint %}

To uninstall Local Install, simply delete the Install Directory.


# Uninstall System Install

Remove System Install.

{% hint style="warning" %}
Applies to System Install.
{% endhint %}

NSO can be uninstalled using the [ncs-installer(1)](/guides/resources/man/ncs-installer.1) option only if NSO is installed with `--system-install` option. Either part of the static files or the full installation can be removed using `ncs-uninstall` option. Ensure to stop NSO before uninstalling.

```bash
# ncs-uninstall --all
```

Executing the above command removes the Installation Directory `/opt/ncs` including symbolic links, Configuration Directory `/etc/ncs`, Run Directory `/var/opt/ncs`, Log Directory `/var/log/ncs`, `systemd` service file `/etc/systemd/system/ncs.service`, `systemd`environment file `/etc/ncs/ncs.systemd.conf`, and the user profile scripts from `/etc/profile.d`.

To make sure that no license entitlements are consumed after you have uninstalled NSO, be sure to perform the `deregister` command in the CLI:

```cli
admin@ncs# license smart deregister
```


# Containerized NSO

Deploy NSO in a containerized setup using Cisco-provided images.

NSO can be deployed in your environment using a container, such as Docker. Cisco offers two pre-built images for this purpose that you can use to run NSO and build packages (see [Overview of NSO Images](#d5e8294)).

***

**Migration Information**

If you are migrating from an existing NSO System Install to a container-based setup, follow the guidelines given below in [Migration to Containerized NSO](#sec.migrate-to-containerizednso).

***

## Use Cases for Containerized Approach

Running NSO in a container offers several benefits that you would generally expect from a containerized approach, such as ease of use and convenient distribution. More specifically, a containerized NSO approach allows you to:

* Run a container image of a specific version of NSO and your packages which can then be distributed as one unit.
* Deploy and distribute the same version across your production environment.
* Use the Build Image containing the necessary environment for compiling NSO packages.

## Overview of NSO Images <a href="#d5e8294" id="d5e8294"></a>

Cisco provides the following two NSO images based on Red Hat UBI.

* [Production Image](#production-image)
* [Build Image](#build-image)

<table data-full-width="false"><thead><tr><th valign="middle">Intended Use</th><th valign="middle">Develop NSO Packages</th><th>Build NSO Packages</th><th valign="middle">Run NSO</th><th valign="middle">NSO Install Type</th><th valign="middle">UBI Version</th></tr></thead><tbody><tr><td valign="middle">Development Host</td><td valign="middle"><img src="/files/4oQ1RJVzKC0m2c5ouEBZ" alt="" data-size="line"></td><td><img src="/files/sa0A8jerbRQ4lfDp8kVW" alt="" data-size="line"></td><td valign="middle"><img src="/files/sa0A8jerbRQ4lfDp8kVW" alt="" data-size="line"></td><td valign="middle">None or Local Install</td><td valign="middle">-</td></tr><tr><td valign="middle">Build Image</td><td valign="middle"><img src="/files/sa0A8jerbRQ4lfDp8kVW" alt="" data-size="line"></td><td><img src="/files/4oQ1RJVzKC0m2c5ouEBZ" alt="" data-size="line"></td><td valign="middle"><img src="/files/sa0A8jerbRQ4lfDp8kVW" alt="" data-size="line"></td><td valign="middle">System Install</td><td valign="middle">UBI 10</td></tr><tr><td valign="middle">Production Image</td><td valign="middle"><img src="/files/sa0A8jerbRQ4lfDp8kVW" alt="" data-size="line"></td><td><img src="/files/sa0A8jerbRQ4lfDp8kVW" alt="" data-size="line"></td><td valign="middle"><img src="/files/4oQ1RJVzKC0m2c5ouEBZ" alt="" data-size="line"></td><td valign="middle">System Install</td><td valign="middle">UBI 10</td></tr></tbody></table>

{% hint style="info" %}
The Red Hat UBI is an OCI-compliant image that is freely distributable and independent of platform and technical dependencies. You can read more about Red Hat UBI [here](https://www.redhat.com/en/blog/introducing-red-hat-universal-base-image), and about Open Container Initiative (OCI) [here](https://opencontainers.org/faq/).
{% endhint %}

### Production Image

The Production Image is a production-ready NSO image for system-wide deployment and use. It is based on NSO [System Install](/guides/administration/installation-and-deployment/system-install) and is available from the [Cisco Software Download](https://software.cisco.com/download/home) site.

Use the pre-built image as the base image in the container file (e.g., Dockerfile) and mount your own packages (such as NEDs and service packages) to run a final image for your production environment (see examples below).

{% hint style="info" %}
Consult the [Installation](/guides/administration/installation-and-deployment) documentation for information on installing NSO on a Docker host, building NSO packages, etc.
{% endhint %}

{% hint style="info" %}
See [Developing and Deploying a Nano Service](/guides/administration/installation-and-deployment/development-to-production-deployment/develop-and-deploy-a-nano-service) for an example that uses the container to deploy an SSH-key-provisioning nano service.

The README in the [examples.ncs/getting-started/netsim-sshkey](https://github.com/NSO-developer/nso-examples/tree/6.7/getting-started/netsim-sshkey) example provides a link to the container-based deployment variant of the example. See the `setup_ncip.sh` script and `README` in the `netsim-sshkey` deployment example for details.
{% endhint %}

### Build Image

The Build Image is a separate standalone NSO image with the necessary environment and software for building packages. It is provided specifically to address the developer needs of building packages.

The image is available as a signed package (e.g., `nso-VERSION.container-image-build.linux.ARCH.signed.bin`) from the Cisco [Software Download](https://software.cisco.com/download/home) site. You can run the Build Image in different ways, and a simple tool for defining and running multi-container Docker applications is [Docker Compose](https://docs.docker.com/compose/) (see examples below).

The container provides the necessary environment to build custom packages. The Build Image adds a few Linux packages that are useful for development, such as Ant, JDK, net-tools, pip, etc. Additional Linux packages can be added using, for example, the `dnf` command. The `dnf list installed` command lists all the installed packages.

## Downloading and Extracting the Images <a href="#sec.fetch-images" id="sec.fetch-images"></a>

To fetch and extract NSO images:

1. On Cisco's official [Software Download](https://software.cisco.com/download/home) site, search for "Network Services Orchestrator". Select the relevant NSO version in the drop-down list, e.g., "Crosswork Network Services Orchestrator 6"**,** and click "Network Services Orchestrator Software". Locate the binary, which is delivered as a signed package (e.g., `nso-6.4.container-image-prod.linux.x86_64.signed.bin`).
2. Extract the image and other files from the signed package, for example:

   ```bash
   sh nso-6.4.container-image-prod.linux.x86_64.signed.bin
   ```

{% hint style="info" %}
**Signed Archive File Pattern**

The signed archive file name has the following pattern:

`nso-VERSION.container-image-PROD_BUILD.linux.ARCH.signed.bin`, where:

* `VERSION` denotes the image's NSO version.
* `PROD_BUILD` denotes the type of the container (i.e., `prod` for Production, and `build` for Build).
* `ARCH` is the CPU architecture.
  {% endhint %}

## System Requirements <a href="#sec.system-reqs" id="sec.system-reqs"></a>

To run the images, make sure that your system meets the following requirements:

* A system running Linux `x86_64` or `ARM64`, or macOS `x86_64` or Apple Silicon. Linux for production.
* A container platform. Docker is the recommended platform and is used as an example in this guide for running NSO images. You may use another container runtime of your choice. Note that commands in this guide are Docker-specific. if you use another container runtime, make sure to use the respective commands.
* Ensure the system has sufficient resources. NSO containers require at least 4 CPU cores.
* To check the Java (JDK) and Python versions included in the container, use the following command, (where `cisco-nso-prod:6.5` is the image you want to check):

  <pre class="language-bash" data-title="Example: Check Java and Python Versions of Container"><code class="lang-bash">docker run --rm cisco-nso-prod:6.5 sh -c "java -version &#x26;&#x26; python --version"
  </code></pre>

{% hint style="info" %}
Docker on Mac uses a Linux VM to run the Docker engine, which is compatible with the normal Docker images built for Linux. You do not need to recompile your NSO-in-Docker images when moving between a Linux machine and Docker on Mac as they both essentially run Docker on Linux.
{% endhint %}

## Administrative Information <a href="#d5e8371" id="d5e8371"></a>

This section covers the necessary administrative information about the NSO Production Image.

### Migrate to Containerized NSO Setup <a href="#sec.migrate-to-containerizednso" id="sec.migrate-to-containerizednso"></a>

If you have NSO installed as a System Install, you can migrate to the Containerized NSO setup by following the instructions in this section. Migrating your Network Services Orchestrator (NSO) to a containerized setup can provide numerous benefits, including improved scalability, easier version management, and enhanced isolation of services.

The migration process is designed to ensure a smooth transition from a System-Installed NSO to a container-based deployment. Detailed steps guide you through preparing your existing environment, exporting the necessary configurations and state data, and importing them into your new containerized NSO instance. During the migration, consider the container runtime you plan to use, as this impacts the migration process.

**Before You Start**

* We recommend reading through this guide to understand better the expectations, requirements, and functioning aspects of a containerized deployment.
* Verify the compatibility of your current system configurations with the containerized NSO setup. See [System Requirements](#sec.system-reqs) for more information.
* Note that [NSO runs from a non-root user ](#nso-runs-from-a-non-root-user)with the containerized NSO setup[.](#nso-runs-from-a-non-root-user)
* Determine and install the container orchestration tool you plan to use (e.g., Docker, etc.).
* Ensure that your current NSO installation is fully operational and backed up and that you have a clear rollback strategy in case any issues arise. Pay special attention to customizations and integrations that your current NSO setup might have, and verify their compatibility with the containerized version of NSO.
* Have a contingency plan in place for quick recovery in case any issues are encountered during migration.

**Migration Steps**

Prepare:

1. Document your current NSO environment's specifics, including custom configurations and packages.
2. Perform a complete backup of your existing NSO instance, including configurations, packages, and data.
3. Set up the container environment and download/extract the NSO production image. See [Downloading and Extracting the Images](#sec.fetch-images) for details.

Migrate:

1. Stop the current NSO instance.
2. Save the run directory from the NSO instance in an appropriate place.
3. Use the same `ncs.conf` and High Availability (HA) setup previously used with your System Install. We assume that the `ncs.conf` follows the best practice and uses the `NCS_DIR`, `NCS_RUN_DIR`, `NCS_CONFIG_DIR`, and `NCS_LOG_DIR` variables for all paths. The `ncs.conf` can be added to a volume and mounted to `/nso/etc` in the container.

   ```bash
   docker container create --name temp -v NSO-evol:/nso/etc hello-world
   docker cp ncs.conf temp:/nso/etc
   docker rm temp
   ```
4. Add the run directory as a volume, mounted to `/nso/run` in the container and copy the CDB data, packages, etc., from the previous System Install instance.

   ```bash
   cd path-to-previous-run-dir
   docker container create --name temp -v NSO-rvol:/nso/run hello-world
   docker cp . temp:/nso/run
   docker rm temp
   ```
5. Create a volume for the log directory.

   ```bash
   docker volume create --name NSO-lvol
   ```
6. Start the container. Example:

   ```bash
   docker run -v NSO-rvol:/nso/run -v NSO-evol:/nso/etc -v NSO-lvol:/log -itd \
   --name cisco-nso -e EXTRA_ARGS=--with-package-reload -e ADMIN_USERNAME=admin \
   -e ADMIN_PASSWORD=admin cisco-nso-prod:6.4
   ```

Finalize:

1. Ensure that the containerized NSO instance functions as expected and validate system operations.
2. Plan and execute your cutover transition from the System-Installed NSO to the containerized version with minimal disruption.
3. Monitor the new setup thoroughly to ensure stability and performance.

### `ncs.conf` File Configuration and Preference <a href="#ug.admin_guide.containers.ncs" id="ug.admin_guide.containers.ncs"></a>

The `run-nso.sh` script runs a check at startup to determine which `ncs.conf` file to use. The order of preference is as below:

1. The `ncs.conf` file specified in the Dockerfile (i.e., `ENV $NCS_CONFIG_DIR /etc/ncs/`) is used as the first preference.
2. The second preference is to use the `ncs.conf` file mounted in the `/nso/etc/` run directory.
3. If no `ncs.conf` file is found at either `/etc/ncs` or `/nso/etc`, the default `ncs.conf` file provided with the NSO image in `/defaults` is used.

{% hint style="info" %}
If the `ncs.conf` file is edited after startup, it can be reloaded using MAAPI `reload_config()`. Example: `$ ncs_cmd -c "reload"`.
{% endhint %}

{% hint style="info" %}
The default `ncs.conf` file in `/defaults` has a set of environment variables that can be used to enable interfaces (all interfaces are disabled by default) which is useful when spinning up the Production container for quick testing. An interface can be enabled by setting the corresponding environment variable to `true`.

* `NCS_CLI_SSH`: Enables CLI over SSH on port `2024`.
* `NCS_WEBUI_TRANSPORT_TCP`: Enables JSON-RPC and RESTCONF over TCP on port `8080`.
* `NCS_WEBUI_TRANSPORT_SSL`: Enables JSON-RPC and RESTCONF over SSL/TLS on port `8888`.
* `NCS_NETCONF_TRANSPORT_SSH`: Enables NETCONF over SSH on port `2022`.
* `NCS_NETCONF_TRANSPORT_TCP`: Enables NETCONF over TCP on port `2023`.
  {% endhint %}

### Pre- and Post-Start Scripts <a href="#d5e8475" id="d5e8475"></a>

If you need to perform operations before or after the `ncs` process is started in the Production container, you can use Python and/or Bash scripts to achieve this. Add the scripts to the `$NCS_CONFIG_DIR/pre-ncs-start.d/` and `$NCS_CONFIG_DIR/post-ncs-start.d/` directories to have the `run-nso.sh` script run them.

### NSO Runs from a Non-Root User

NSO is installed with the `--run-as-user` option for build and production containers to run NSO from the non-root `nso` user that belongs to the `nso` user group.

When migrating from container versions where NSO has `root` privilege, ensure the `nso` user owns or has access rights to the required files and directories. Examples include application directories, SSH host keys, SSH keys used to authenticate with devices, etc. See the deployment example variant referenced by the [examples.ncs/getting-started/netsim-sshkey/README.md](https://github.com/NSO-developer/nso-examples/tree/6.7/getting-started/netsim-sshkey) for an example.

The NSO container runs a script called `take-ownership.sh` as part of its startup, which takes ownership of all the directories that NSO needs. The script will be one of the first things to run. The script can be overridden to take ownership of even more directories, such as mounted volumes or bind mounts.

### Admin User Creation <a href="#d5e8482" id="d5e8482"></a>

An admin user can be created on startup by the run script in the container. Three environment variables control the addition of an admin user:

* `ADMIN_USERNAME`: Username of the admin user to add, default is `admin`.
* `ADMIN_PASSWORD`: Password of the admin user to add.
* `ADMIN_SSHKEY`: Private SSH key of the admin user to add.

As `ADMIN_USERNAME` already has a default value, only `ADMIN_PASSWORD`, or `ADMIN_SSHKEY` need to be set in order to create an admin user. For example:

```bash
docker run -itd --name cisco-nso -e ADMIN_PASSWORD=admin cisco-nso-prod:6.4
```

This can be useful when starting up a container in CI for testing or development purposes. It is typically not required in a production environment where CDB already contains the required user accounts.

{% hint style="info" %}
When using a permanent volume for CDB, and restarting the NSO container multiple times with a different `ADMIN_USERNAME` or `ADMIN_PASSWORD`, the start script uses these environment variables to generate an XML file named `add_admin_user.xml`. The generated XML file is added to the CDB directory to be read at startup. But if the persisted CDB configuration file already exists in the CDB directory, NSO will not load any XML files at startup, instead the generated `add_admin_user.xml` in the CDB directory needs to be loaded manually.
{% endhint %}

{% hint style="info" %}
The default `ncs.conf` supplied with the NSO Production Image enables Linux PAM authentication and disables NSO local authentication. For `ADMIN_USERNAME`, `ADMIN_PASSWORD`, and `ADMIN_SSHKEY` to take effect, enable `/ncs-config/aaa/local-authentication`. Alternatively, you can create a local Linux admin user that is authenticated by NSO using Linux PAM.
{% endhint %}

### Exposing Ports <a href="#sec.exposed_ports" id="sec.exposed_ports"></a>

The default `ncs.conf` NSO configuration file does not enable any northbound interfaces, and no ports are exposed externally to the container. Ports can be exposed externally of the container when starting the container with the northbound interfaces and their ports correspondingly enabled in `ncs.conf`.

### Backup and Restore <a href="#d5e8524" id="d5e8524"></a>

The backup behavior of running NSO in vs. outside the container is largely the same, except that when running NSO in a container, the SSH and SSL certificates are not included in the backup produced by the `ncs-backup` script. This is different from running NSO outside a container where the default configuration path `/etc/ncs` is used to store the SSH and SSL certificates, i.e., `/etc/ncs/ssh` for SSH and `/etc/ncs/ssl` for SSL.

**Take a Backup**

Let's assume we start a production image container using:

```bash
docker run -d --name cisco-nso -v NSO-vol:/nso -v NSO-log-vol:/log cisco-nso-prod:6.4
```

To take a backup:

* Run the `ncs-backup` command. The backup file is written to `/nso/run/backups`.

  ```bash
  docker exec -it cisco-nso ncs-backup
  INFO  Backup /nso/run/backups/ncs-6.4@2024-11-03T11:31:07.backup.gz created successfully
  ```

**Restore a Backup**

To restore a backup, NSO must be stopped. As you likely only have access to the `ncs-backup` tool, the volume containing CDB and other run-time data from inside of the NSO container, this poses a slight challenge. Additionally, shutting down NSO will terminate the NSO container.

To restore a backup:

1. Shut down the NSO container:

   ```bash
   docker stop cisco-nso
   docker rm cisco-nso
   ```
2. Run the `ncs-backup --restore` command. Start a new container with the same persistent shared volumes mounted but with a different command. Instead of running the `/run-nso.sh`, which is the normal command of the NSO container, run the `restore` command.

   ```bash
   docker run -u root -it --rm -v NSO-vol:/nso -v NSO-log-vol:/log \
   --entrypoint ncs-backup cisco-nso-prod:6.4 \
   --restore /nso/run/backups/ncs-6.4@2024-11-03T11:31:07.backup.gz

   Restore /etc/ncs from the backup (y/n)? y
   Restore /nso/run from the backup (y/n)? y
   INFO  Restore completed successfully
   ```
3. Restoring an NSO backup should move the current run directory (`/nso/run` to `/nso/run.old`) and restore the run directory from the backup to the main run directory (`/nso/run`). After this is done, start the regular NSO container again as usual.\\

   ```bash
   docker run -d --name cisco-nso -v NSO-vol:/nso -v NSO-log-vol:/log cisco-nso-prod:6.4
   ```

### SSH Host Key <a href="#d5e8566" id="d5e8566"></a>

The NSO image `/run-nso.sh` script looks for an SSH host key named `ssh_host_ed25519_key` in the `/nso/etc/ssh` directory to be used by the NSO built-in SSH server for the CLI and NETCONF interfaces.

If an SSH host key exists, which is for a typical production setup stored in a persistent shared volume, it remains the same after restarts or upgrades of NSO. If no SSH host key exists, the script generates a private and public key.

In a high-availability (HA) setup, the host key is typically shared by all NSO nodes in the HA group and stored in a persistent shared volume. This is done to avoid fetching the public host key from the new primary after each failover.

### HTTPS TLS Certificate <a href="#d5e8574" id="d5e8574"></a>

NSO expects to find a TLS certificate and key at `/nso/ssl/cert/host.cert` and `/nso/ssl/cert/host.key` respectively. Since the `/nso` path is usually on persistent shared volume for production setups, the certificate remains the same across restarts or upgrades.

If no certificate is present, one will be generated. It is a self-signed certificate valid for 30 days making it possible to use both in development and staging environments. It is not meant for the production environment. You should replace it with a properly signed certificate for production and it is encouraged to do so even for test and staging environments. Simply generate one and place it at the provided path, for example using the following, which is the command used to generate the temporary self-signed certificate:

```
openssl req -new -newkey rsa:4096 -x509 -sha256 -days 30 -nodes \
-out /nso/ssl/cert/host.cert -keyout /nso/ssl/cert/host.key \
-subj "/C=SE/ST=NA/L=/O=NSO/OU=WebUI/CN=Mr. Self-Signed"
```

### YANG Model Changes (destructive) <a href="#d5e8584" id="d5e8584"></a>

The database in NSO, called CDB, uses YANG models as the schema for the database. It is only possible to store data in CDB according to the YANG models that define the schema.

If the YANG models are changed, particularly if the nodes are removed or renamed (rename is the removal of one leaf and an addition of another), any data in CDB for those leaves will also be removed. NSO normally warns about this when you attempt to load new packages, for example, `request packages reload` command refuses to reload the packages if the nodes in the YANG model have disappeared. You would then have to add the **force** argument, e.g., `request packages reload force`.

### Health Check <a href="#d5e8591" id="d5e8591"></a>

The base Production Image comes with a basic container health check. It uses `ncs_cmd` to get the state that NCS is currently in. Only the result status is observed to check if `ncs_cmd` was able to communicate with the `ncs` process. The result indicates if the `ncs` process is responding to IPC requests.

{% hint style="info" %}
The default `--health-start-period duration` in health check is set to 60 seconds. NSO will report an `unhealthy` state if it takes more than 60 seconds to start up. To resolve this, set the `--health-start-period duration` value to a relatively higher value, such as 600 seconds, or however long you expect NSO will take to start up.

To disable the health check, use the `--no-healthcheck` command.
{% endhint %}

### Use NSO Memory Monitoring to Capture Debug Dumps Before a Container OOM Kill <a href="#d5e8605" id="d5e8605"></a>

NSO can monitor memory through the `/ncs-config/memory-management` section in `ncs.conf`. In containerized deployments, configure it to trigger one or more debug dumps before memory pressure reaches the point where the container or the NSO process might be OOM-killed without leaving useful diagnostics.

This feature can be used while leaving the host in Linux's default heuristic overcommit mode (`vm.overcommit_memory=0`); see [proc\_sys\_vm(5)](https://man7.org/linux/man-pages/man5/proc_sys_vm.5.html). Host overcommit settings remain host-global and cannot be configured per container, but NSO can still use cgroup memory information to trigger debug dumps proactively.

* When NSO runs in a container with a configured memory limit, NSO uses the container cgroup memory limit and current usage instead of host-wide memory values.
* If you need host-based guidance instead, see [Use NSO Memory Monitoring to Capture Debug Dumps Before an OOM Kill](/guides/administration/installation-and-deployment/system-install#use-nso-memory-monitoring-to-capture-debug-dumps-before-an-oom-kill).

Each `action` under `/ncs-config/memory-management/actions` defines:

* One threshold, either `used-memory-threshold-percentage` or `free-memory-threshold-bytes`.
* One compensating action, currently `debug-dump`.
* A required absolute dump `directory`.
* Optional rate limiting with `count` (default `5`) and `cooldown-period` (default `PT60M`).

When a threshold is crossed, NSO writes a timestamped debug dump such as `debug_dump_2026-04-17T09:12:34.567Z`, logs that it is creating the dump, and raises the `memory-management-action-triggered` alarm.

**Recommended Usage**

* Configure the feature in the `ncs.conf` used by the container before starting NSO, or reload the configuration with `ncs --reload` after updating `ncs.conf`.
* Run the container with an explicit memory limit, for example `docker run --memory=<ram>` or an equivalent limit in your container platform.
* If swap should effectively be disabled for the container, set `--memory-swap=<ram>` equal to `--memory`.
* Use a mounted persistent volume that the NSO user can write to for the debug-dump directory, typically under `NCS_RUN_DIR=/nso/run`, for example `/nso/run/debug-dumps`.
* Set the threshold early enough that NSO still has time to finish the dump. A good starting point is `90` for `used-memory-threshold-percentage`, or a `free-memory-threshold-bytes` value that leaves a few GiB of headroom on larger systems.
* Define more than one action if you want an early snapshot and then additional snapshots closer to the limit.
* Keep `NCS_DUMP` configured as well. Point it to a writable persistent location, typically under `NCS_RUN_DIR=/nso/run`. A proactive debug dump helps when the Linux OOM-killer would otherwise terminate NSO without producing a system dump.

**Example Configuration**

Add the following under the top-level `<ncs-config>` element in the `ncs.conf` used by the container:

{% code title="ncs.conf memory-management example for a container" %}

```xml
<memory-management>
  <actions>
    <action>
      <name>early-warning</name>
      <used-memory-threshold-percentage>90</used-memory-threshold-percentage>
      <debug-dump>
        <count>3</count>
        <cooldown-period>PT5M</cooldown-period>
        <directory>/nso/run/debug-dumps</directory>
      </debug-dump>
    </action>
    <action>
      <name>critical-free-memory</name>
      <free-memory-threshold-bytes>2147483648</free-memory-threshold-bytes>
      <debug-dump>
        <count>2</count>
        <cooldown-period>PT1M</cooldown-period>
        <directory>/nso/run/debug-dumps</directory>
      </debug-dump>
    </action>
  </actions>
</memory-management>
```

{% endcode %}

In the example above, the first action triggers when memory usage reaches 90% of the configured container memory limit. The second action triggers when less than 2 GiB remain available within that limit. Use either one threshold type or both, depending on how you size and operate the container.

**Verification**

After starting NSO or reloading the configuration:

* Check that the dump directory exists and is writable by the NSO user inside the container.
* When a threshold is crossed, look for `creating debug dump` in `ncs.log`.
* Confirm that a new file appears in the configured directory.
* Check `show alarms alarm-list` for the `memory-management-action-triggered` alarm.

{% hint style="info" %}
This feature does not stop the Linux OOM-killer by itself, but it prevents an OOM situation from leaving you without a debug dump from NSO.
{% endhint %}

{% hint style="warning" %}
Ensure that both the `/ncs-config/memory-management/actions/action/debug-dump/directory` path and the directory used for `NCS_DUMP` exist and are writable by the NSO user.

By default, NSO writes a system dump to the NSO run-time directory, typically `NCS_RUN_DIR=/nso/run` in a container. If that directory is not backed by a persistent mounted volume or another suitable writable persistent location, dumps can be lost across container restarts. Set `NCS_DUMP` to a dump file path in a suitable mounted directory, for example `NCS_DUMP="/nso/run/ncs_crash.dump.$(date +%Y%m%d-%H%M%S)"`.
{% endhint %}

### Startup Arguments

The `/nso-run.sh` script that starts NSO is executed as an `ENTRYPOINT` instruction and the `CMD` instruction can be used to provide arguments to the entrypoint-script. Another alternative is to use the `EXTRA_ARGS` variable to provide arguments. The `/nso-run.sh` script will first check the `EXTRA_ARGS` variable before the `CMD` instruction.

An example using `docker run` with the `CMD` instruction:

```bash
docker run --name nso -itd cisco-nso-prod:6.4 --with-package-reload \
--ignore-initial-validation
```

With the `EXTRA_ARGS` variable:

```bash
docker run --name nso \
-e EXTRA_ARGS='--with-package-reload --ignore-initial-validation' \
-itd cisco-nso-prod:6.4
```

An example using a Docker Compose file, `compose.yaml`, with the `CMD` instruction:

```
services:
    nso:
        image: cisco-nso-prod:6.4
        container_name: nso
        command:
            - --with-package-reload
            - --ignore-initial-validation
```

With the `EXTRA_ARGS` variable:

```
services:
    nso:
        image: cisco-nso-prod:6.4
        container_name: nso
        environment:
            - EXTRA_ARGS=--with-package-reload --ignore-initial-validation
```

## Examples <a href="#d5e8625" id="d5e8625"></a>

This section provides examples to exhibit the use of NSO images.

### Running the Production Image using Docker CLI

This example shows how to run the standalone NSO Production Image using the Docker CLI.

The instructions and CLI examples used in this example are Docker-specific. If you are using a non-Docker container runtime, you will need to: fetch the NSO image from the Cisco software download site, then load and run the image with packages and networking, and finally log in to NSO CLI to run commands.

If you intend to run multiple images (i.e., both Production and Build), Docker Compose is a tool that simplifies defining and running multi-container Docker applications. See the example ([Running the NSO Images using Docker Compose](#sec.example-docker-compose)) below for detailed instructions.

**Steps**

Follow the steps below to run the Production Image using Docker CLI:

1. Start your container engine.
2. Next, load the image and run it. Navigate to the directory where you extracted the base image and load it. This will restore the image and its tag:

```bash
docker load -i nso-6.4.container-image-prod.linux.x86_64.tar.gz
```

3. Start a container from the image. Supply additional arguments to mount the packages and `ncs.conf` as separate volumes ([`-v` flag](https://docs.docker.com/engine/reference/commandline/run/)), and publish ports for networking ([`-p` flag](https://docs.docker.com/engine/reference/commandline/run/)) as needed. The container starts NSO using the `/run-nso.sh` script. To understand how the `ncs.conf` file is used, see [`ncs.conf` File Configuration and Preference](#ug.admin_guide.containers.ncs).

```bash
docker run -itd --name cisco-nso \
-v NSO-vol:/nso \
-v NSO-log-vol:/log \
--net=host \
-e ADMIN_USERNAME=admin \
-e ADMIN_PASSWORD=admin \
cisco-nso-prod:6.4
```

{% hint style="warning" %}
**Overriding Environment Variables**

Overriding basic environment variables (`NCS_CONFIG_DIR`, `NCS_LOG_DIR`, `NCS_RUN_DIR`, etc.) is not supported and therefore should be avoided. Using, for example, the `NCS_CONFIG_DIR` environment variable to mount a configuration directory will result in an error. Instead, to mount your configuration directory, do it appropriately in the correct place, which is under `/nso/etc`.
{% endhint %}

<details>

<summary>Examples: Running the Image with and without Named Volumes</summary>

The following examples show how to run the image with and without named volumes.

**Running without a named volume**: This is the minimal way of running the image but does not provide any persistence when the container is destroyed.

```bash
docker run -itd --name cisco-nso \
-p 8888:8888 \
-e ADMIN_USERNAME=admin\
-e ADMIN_PASSWORD=admin\
cisco-nso-prod
```

**Running with a single named volume**: This way provides persistence for the NSO mount point with a `NSO-vol` volume. Logs, however, are not persistent.

```bash

docker run -itd --name cisco-nso \
-v NSO-vol:/nso \
-p 8888:8888 \
-e ADMIN_USERNAME=admin\
-e ADMIN_PASSWORD=admin\
cisco-nso-prod
```

\
**Running with two named volumes**: This way provides full persistence for both the NSO and the log mount points.

```bash
docker run -itd --name cisco-nso \
-v NSO-vol:/nso \
-v NSO-log-vol:/log \
-p 8888:8888 \
-e ADMIN_USERNAME=admin\
-e ADMIN_PASSWORD=admin\
cisco-nso-prod
```

</details>

{% hint style="info" %}
**Loading the Packages**

* Loading the packages by mounting the default load path `/nso/run` as a volume is preferred. You can also load the packages by copying them manually into the `/nso/run/packages` directory in the container. During development, a bind mount of the package directory on the host machine makes it easy to update packages in NSO by simply changing the packages on the host.
* The default load path is configured in the `ncs.conf` file as `$NCS_RUN_DIR/packages`, where `$NCS_RUN_DIR` expands to `/nso/run` in the container. To find the load path, check the `ncs.conf` file in the `/etc/ncs/` directory.

  ```xml
  <load-path>
  <dir>${NCS_RUN_DIR}/packages</dir>
  <dir>${NCS_DIR}/etc/ncs</dir>
  ...
  </load-path>
  ```

{% endhint %}

{% hint style="info" %}
**Logging**

* With the Production Image, use a shared volume to persist data across restarts. If remote (Syslog) logging is used, there is little need to persist logs. If local logging is used, then persistent logging is recommended.
* NSO starts a cron job to handle logrotate of NSO logs by default. i.e., the `CRON_ENABLE` and `LOGROTATE_ENABLE` variables are set to `true` using the `/etc/logrotate.conf` configuration. See the `/etc/ncs/post-ncs-start.d/10-cron-logrotate.sh` script. To set how often the cron job runs, use the crontab file.
  {% endhint %}

4. Finally, log in to NSO CLI to run commands. Open an interactive shell on the running container and access the NSO CLI.

```bash
docker exec -it cisco-nso bash
#  ncs_cli -u admin
admin@ncs>
```

You can also use the `docker exec -it cisco-nso ncs_cli -u admin` command to access the CLI from the host's terminal.

### Upgrading NSO using Docker CLI <a href="#d5e8715" id="d5e8715"></a>

This example describes how to upgrade your NSO to run a newer NSO version in the container. The overall upgrade process is outlined in the steps below. In the example below, NSO is to be upgraded from version 6.3 to 6.4.

To upgrade your NSO version:

1. Start a container with the `docker run` command. In the example below, it mounts the `/nso` directory in the container to the `NSO-vol` named volume to persist the data. Another option is using a bind mount of the directory on the host machine. At this point, the `/cdb` directory is empty.

   ```bash
   docker run -itd -—name cisco-nso -v NSO-vol:/nso cisco-nso-prod:6.3
   ```
2. Perform a backup, either by running the `docker exec` command (make sure that the backup is placed somewhere we have mounted) or by creating a tarball of `/data/nso` on the host machine.

   ```bash
   docker exec -it cisco-nso ncs-backup
   ```
3. Stop the NSO by issuing the following command, or by stopping the container itself which will run the `ncs stop` command automatically.

   ```bash
   docker exec -it cisco-nso ncs --stop 
   ```
4. Remove the old NSO.

   ```bash
   docker rm -f cisco-nso
   ```
5. Start a new container and mount the `/nso` directory in the container to the `NSO-vol` named volume. This time the `/cdb` folder is not empty, so instead of starting a fresh NSO, an upgrade will be performed.

   ```bash
   docker run -itd --name cisco-nso -v NSO-vol:/nso cisco-nso-prod:6.4
   ```

At this point, you only have one container that is running the desired version 6.4 and you do not need to uninstall the old NSO.

### Running the NSO Images using Docker Compose <a href="#sec.example-docker-compose" id="sec.example-docker-compose"></a>

This example covers the necessary information to manifest the use of NSO images to compile packages and run NSO. Using Docker Compose is not a requirement, but a simple tool for defining and running a multi-container setup where you want to run both the Production and Build images in an efficient manner.

#### **Packages**

The packages used in this example are taken from the [examples.ncs/getting-started/netsim-sshkey](https://github.com/NSO-developer/nso-examples/tree/6.7/getting-started/netsim-sshkey) example:

* `distkey`: A simple Python + template service package that automates the setup of SSH public key authentication between netsim (ConfD) devices and NSO using a nano service.
* `ne`: A NETCONF NED package representing a netsim network element that implements a configuration subscriber Python application that adds or removes the configured public key, which the netsim (ConfD) network element checks when authenticating public key authentication clients.

#### **`docker-compose.yaml` - Docker Compose File Example**

A basic Docker Compose file is shown in the example below. It describes the containers running on a machine:

* The Production container runs NSO.
* The Build container builds the NSO packages.
* A third `example` container runs the netsim device.

Note that the packages use a shared volume in this simple example setup. In a more complex production environment, you may want to consider a dedicated redundant volume for your packages.

```
                version: '1.0'
                volumes:
                  NSO-1-rvol:

                networks:
                  NSO-1-net:

                services:
                  NSO-1:
                    image: cisco-nso-prod:6.4
                    container_name: nso1
                    profiles:
                      - prod
                    environment:
                      - EXTRA_ARGS=--with-package-reload
                      - ADMIN_USERNAME=admin
                      - ADMIN_PASSWORD=admin
                    networks:
                      - NSO-1-net
                    ports:
                      - "2024:2024"
                      - "8888:8888"
                    volumes:
                      - type: bind
                        source: /path/to/packages/NSO-1
                        target: /nso/run/packages
                      - type: bind
                        source: /path/to/log/NSO-1
                        target: /log
                      - type: volume
                        source: NSO-1-rvol
                        target: /nso
                    healthcheck:
                      test: ncs_cmd -c "wait-start 2"
                      interval: 5s
                      retries: 5
                      start_period: 10s
                      timeout: 10s

                  BUILD-NSO-PKGS:
                    image: cisco-nso-build:6.4
                    container_name: build-nso-pkgs
                    network_mode: none
                    profiles:
                      - build
                    volumes:
                      - type: bind
                        source: /path/to/packages/NSO-1
                        target: /nso/run/packages

                  EXAMPLE:
                    image: cisco-nso-prod:6.4
                    container_name: ex-netsim
                    profiles:
                      - example
                    networks:
                      - NSO-1-net
                    healthcheck:
                      test: test -f /nso-run-prod/etc/ncs.conf && ncs-netsim --dir /netsim is-alive ex0
                      interval: 5s
                      retries: 5
                      start_period: 10s
                      timeout: 10s
                      entrypoint: bash
                    command: -c 'rm -rf /netsim
                        && mkdir /netsim
                        && ncs-netsim --dir /netsim create-network /network-element 1 ex
                        && PYTHONPATH=/opt/ncs/current/src/ncs/pyapi ncs-netsim --dir
                            /netsim start
                        && mkdir -p /nso-run-prod/run/cdb
                        && echo "<devices xmlns=\"http://tail-f.com/ns/ncs\">
                            <authgroups><group><name>default</name>
                            <umap><local-user>admin</local-user>
                            <remote-name>admin</remote-name><remote-password>
                            admin</remote-password></umap></group>
                            </authgroups></devices>"
                            > /nso-run-prod/run/cdb/init1.xml
                        && ncs-netsim --dir /netsim ncs-xml-init >
                            /nso-run-prod/run/cdb/init2.xml
                        && sed -i.orig -e "s|127.0.0.1|ex-netsim|"
                            /nso-run-prod/run/cdb/init2.xml
                        && mkdir -p /nso-run-prod/etc
                        && sed -i.orig -e "s|</cli>|<style>c</style>
                            </cli>|" -e "/<ssh>/{n;s|<enabled>false
                                </enabled>|
                                <enabled>true</enabled>|}" defaults/ncs.conf
                        && sed -i.bak -e "/<local-authentication>/{n;s|
                            <enabled>false</enabled>|<enabled>true
                            </enabled>|}" defaults/ncs.conf
                        && sed "/<ssl>/{n;s|<enabled>false</enabled>|
                            <enabled>true</enabled>|}" defaults/ncs.conf
                            > /nso-run-prod/etc/ncs.conf
                        && mv defaults/ncs.conf.orig defaults/ncs.conf
                        && tail -f /dev/null'
                    volumes:
                      - type: bind
                        source: /path/to/packages/NSO-1/ne
                        target: /network-element
                      - type: volume
                        source: NSO-1-rvol
                        target: /nso-run-prod
```

<details>

<summary>Explanation of the Docker Compose File</summary>

A description of noteworthy Compose file items is given below.

* **`profiles`**: Profiles can be used to group containers in a Compose file, and they work perfectly for the Production, Build, and netsim containers. By adding multiple containers on the same machine (as a developer normally would), you can easily start the Production, Build, and netsim containers using their respective profiles (`prod`, `build`, and `example`).
* **The command used in the netsim example**: Creates a directory called `/netsim` where the netsims will be set up, then starts the netsims, followed by generating two `init.xml` files and editing the `ncs.conf` file for the Production container. Finally, it keeps the container running. If you want this to be more elegant, you need a netsim container image with a script in it that is well-documented.
* **`volumes`**: The Production and Build images are configured intentionally to have the same bind mount with `/path/to/packages/NSO-1` as the source and `/nso/run/packages` as the target. The Production Image mounts both the `/log` and `/nso` directories in the container. The `/log` directory is simply a bind mount, while the `/nso` directory is an actual volume.

  \
  Named volumes are recommended over bind mounts as described by the Docker Volumes documentation. The NSO `/run` directory should therefore be mounted as a named volume. However, you can make the `/run` directory a bind mount as well.

  The Compose file, typically named `docker-compose.yaml`, declares a volume called `NSO-1-rvol`. This is a named volume and will be created automatically by Compose. You can create this volume externally, at which point this volume must be declared as external. If the external volume doesn't exist, the container will not start.

  \
  The `example` netsim container will mount the network element NED in the packages directory. This package should be compiled. Note that the `NSO-1-rvol` volume is used by the `example` container to share the generated `init.xml` and `ncs.conf` files with the NSO Production container.
* **`healthcheck`**: The image comes with its own health check (similar to the one shown here in Compose), and this is how you configure it yourself. The health check for the netsim `example` container checks if the `ncs.conf` file has been generated, and the first Netsim instance started in the container. You could, in theory, start more netsims inside the container.

</details>

#### **Steps**

Follow the steps below to run the images using Docker Compose:

1. Start the Build container. This starts the services in the Compose file with the profile `build`.

   ```bash
   docker compose --profile build up -d
   ```
2. Copy the packages from the `netsim-sshkey` example and compile them in the NSO Build container. The easiest way to do this is by using the `docker exec` command, which gives more control over what to build and the order of it. You can also do this with a script to make it easier and less verbose. Normally you populate the package directory from the host. Here, we use the packages from an example.

   ```bash
   docker exec -it build-nso-pkgs sh -c 'cp -r ${NCS_DIR}/examples.ncs/getting-started \
       /netsim-sshkey/packages ${NCS_RUN_DIR}'

   docker exec -it build-nso-pkgs sh -c 'for f in ${NCS_RUN_DIR}/packages/*/src; \
       do make -C "$f" all || exit 1; done'
   ```
3. Start the netsim container. This outputs the generated `init.xml` and `ncs.conf` files to the NSO Production container. The `--wait` flag instructs to wait until the health check returns healthy.

   ```bash
   docker compose --profile example up --wait
   ```
4. Start the NSO Production container.

   ```bash
   docker compose --profile prod up --wait
   ```

   \
   At this point, NSO is ready to run the service example to configure the netsim device(s). A bash script (`demo.sh`) that runs the above steps and showcases the `netsim-sshkey` example is given below:

   ```
   #!/bin/bash
   set -eu # Abort the script if a command returns with a non-zero exit code or if
           # a variable name is dereferenced when the variable hasn't been set
   GREEN='\033[0;32m'
   PURPLE='\033[0;35m'
   NC='\033[0m' # No Color

   printf "${GREEN}##### Reset the container setup\n${NC}";
   docker compose --profile build down
   docker compose --profile example down -v
   docker compose --profile prod down -v
   rm -rf ./packages/NSO-1/* ./log/NSO-1/*

   printf "${GREEN}##### Start the build container used for building the NSO NED
       and service packages\n${NC}"
   docker compose --profile build up -d

   printf "${GREEN}##### Get the packages\n${NC}"
   printf "${PURPLE}##### NOTE: Normally you populate the package directory from the host.
   Here, we use packages from an NSO example\n${NC}"
   docker exec -it build-nso-pkgs sh -c 'cp -r
    ${NCS_DIR}/examples.ncs/getting-started/netsim-sshkey/packages ${NCS_RUN_DIR}'

   printf "${GREEN}##### Build the packages\n${NC}"
   docker exec -it build-nso-pkgs sh -c 'for f in ${NCS_RUN_DIR}/packages/*/src;
       do make -C "$f" all || exit 1; done'

   printf "${GREEN}##### Start the simulated device container and setup the example\n${NC}"
   docker compose --profile example up --wait

   printf "${GREEN}##### Start the NSO prod container\n${NC}"
   docker compose --profile prod up --wait

   printf "${GREEN}##### Showcase the netsim-sshkey example from NSO on the prod container\n${NC}"
   if [[ $# -eq 0 ]] ; then # Ask for input only if no argument was passed to this script
       printf "${PURPLE}##### Press any key to continue or ctrl-c to exit\n${NC}"
       read -n 1 -s -r
   fi
   docker exec -it nso1 sh -c 'sed -i.orig -e "s/make/#make/"
    ${NCS_DIR}/examples.ncs/getting-started/netsim-sshkey/showcase.sh'
   docker exec -it nso1 sh -c 'cd ${NCS_RUN_DIR};
    ${NCS_DIR}/examples.ncs/getting-started/netsim-sshkey/showcase.sh 1'
   ```

### Upgrading NSO using Docker Compose <a href="#d5e8861" id="d5e8861"></a>

This example describes how to upgrade NSO when using Docker Compose.

#### **Upgrade to a New Minor or Major Version**

To upgrade to a new minor or major version, for example, from 6.3 to 6.4, follow the steps below:

1. Change the image version in the Compose file to the new version, here 6.4.
2. Run the `docker compose up --profile build -d` command to start the Build container with the new image.
3. Compile the packages using the Build container.

   ```bash
   docker exec -it build-nso-pkgs sh -c 'for f in
   ${NCS_RUN_DIR}/packages/*/src;do make -C "$f" all || exit 1; done'
   ```
4. Run the `docker compose up --profile prod --wait` command to start the Production container with the new packages that were just compiled.

#### **Upgrade to a New Maintenance Version**

To upgrade to a new maintenance release version, for example, 6.4.1, follow the steps below:

1. Change the image version in the Compose file to the new version, here 6.4.1.
2. Run the `docker compose up --profile prod --wait` command.

   Upgrading in this way does not require a recompile. Docker detects changes and upgrades the image in the container to the new version.


# Development to Production Deployment

Deploy NSO from development to production.


# Develop and Deploy a Nano Service

Develop and deploy a nano service using a guided example.

This section shows how to develop and deploy a simple NSO nano service for managing the provisioning of SSH public keys for authentication. For more details on nano services, see [Nano Services for Staged Provisioning](/guides/development/core-concepts/nano-services) in Development. The example showcasing development is available under [examples.ncs/getting-started/netsim-sshkey](https://github.com/NSO-developer/nso-examples/tree/6.7/getting-started/netsim-sshkey). In addition, there is a reference from the `README` in the example's directory to the deployment version of the example.

## Development <a href="#d5e1323" id="d5e1323"></a>

<div data-with-frame="true"><figure><img src="/files/qkHJEc3CaYI1Zve85K82" alt="" width="375"><figcaption><p>The Development Host Topology</p></figcaption></figure></div>

After installing NSO with the [Local Install](/guides/administration/installation-and-deployment/local-install) option, development often begins with either retrieving an existing YANG model representing what the managed network element (a virtual or physical device, such as a router) can do or constructing a new YANG model that at least covers the configuration of interest to an NSO service. To enable NSO service development, the network element's YANG model can be used with NSO's netsim tool that uses ConfD (Configuration Daemon) to simulate the network elements and their management interfaces like NETCONF. Read more about netsim in [Network Simulator](/guides/operation-and-usage/operations/network-simulator-netsim).

The simple network element YANG model used for this example is available under `packages/ne/src/yang/ssh-authkey.yang`. The `ssh-authkey.yang` model implements a list of SSH public keys for identifying a user. The list of keys augments a list of users in the ConfD built-in `tailf-aaa.yang` module that ConfD uses to authenticate users.

```yang
module ssh-authkey {
  yang-version 1.1;
  namespace "http://example.com/ssh-authkey";
  prefix sa;

  import tailf-common {
    prefix tailf;
  }

  import tailf-aaa {
    prefix aaa;
  }

  description
    "List of SSH authorized public keys";

  revision 2023-02-02 {
    description
      "Initial revision.";
  }

  augment "/aaa:aaa/aaa:authentication/aaa:users/aaa:user" {
    list authkey {
      key pubkey-data;
      leaf pubkey-data {
        type string;
      }
    }
  }
}
```

On the network element, a Python application subscribes to ConfD to be notified of configuration changes to the user's public keys and updates the user's authorized\_keys file accordingly. See `packages/ne/netsim/ssh-authkey.py` for details.

The first step is to create an NSO package from the network element YANG model. Since NSO will use NETCONF over SSH to communicate with the device, the package will be a NETCONF NED. The package can be created using the `ncs-make-package` command or the NETCONF NED builder tool. The `ncs-make-package` command is typically used when the YANG models used by the network element are available. Hence, the packages/ne package for this example was generated using the `ncs-make-package` command.

As the `ssh-authkey.yang` model augments the users list in the ConfD built-in `tailf-aaa.yang` model, NSO needs a representation of that YANG model too to build the NED. However, the service will only configure the user's public keys, so only a subset of the `tailf-aaa.yang` model that only includes the user list is sufficient. To compare, see the `packages/ne/src/yang/tailf-aaa.yang` in the example vs. the network element's version under `$NCS_DIR/netsim/confd/src/confd/aaa/tailf-aaa.yang`.

Now that the network element package is defined, next up is the service package, beginning with finding out what steps are required for NSO to authenticate with the network element using SSH public key authentication:

1. First, generate private and public keys using, for example, the `ssh-keygen` OpenSSH authentication key utility.
2. Distribute the public keys to the ConfD-enabled network element's list of authorized keys.
3. Configure NSO to use public key authentication with the network element.
4. Finally, test the public key authentication by connecting NSO with the network element.

The outline above indicates that the service will benefit from implementing several smaller (nano) steps:

* The first step only generates private and public key files with no configuration. Thus, the first step should be implemented by an action before the second step runs, not as part of the second step transaction `create()` callback code configuring the network elements. The `create()` callback runs multiple times, for example, for service configuration changes, re-deploy, or commit dry-run. Therefore, generating keys should only happen when creating the service instance.
* The third step cannot be executed before the second step is complete, as NSO cannot use the public key for authenticating with the network element before the network element has it in its list of authorized keys.
* The fourth step uses the NSO built-in `connect()` action and should run after the third step finishes.

What configuration input do the above steps need?

* The name of the network element that will authenticate a user with an SSH public key.
* The name of the local NSO user that maps to the remote network element user the public key authenticates.
* The name of the remote network element user.
* A passphrase is used for encrypting the private key, guarding its privacy. The passphrase should be encrypted when storing it in the CDB, just like any other password.
* The name of the NSO authentication group to configure for public-key authentication with the NSO-managed network element.

A service YANG model that implements the above configuration:

```yang
  container pubkey-dist {
    list key-auth {
      key "ne-name local-user";

      uses ncs:nano-plan-data;
      uses ncs:service-data;
      ncs:servicepoint "distkey-servicepoint";

      leaf ne-name {
        type leafref {
          path "/ncs:devices/ncs:device/ncs:name";
        }
      }
      leaf local-user {
        type leafref {
          path "/ncs:devices/ncs:authgroups/ncs:group/ncs:umap/ncs:local-user";
          require-instance false;
        }
      }
      leaf remote-name {
        type leafref {
          path "/ncs:devices/ncs:authgroups/ncs:group/ncs:umap/ncs:remote-name";
          require-instance false;
        }
        mandatory true;
      }
      leaf authgroup-name {
        type leafref {
          path "/ncs:devices/ncs:authgroups/ncs:group/ncs:name";
          require-instance false;
        }
        mandatory true;
      }
      leaf passphrase {
        // Leave unset for no passphrase
        tailf:suppress-echo true;
        type tailf:aes-256-cfb-128-encrypted-string {
          length "10..max" {
            error-message "The passphrase must be at least 10 characters long";
          }
          pattern ".*[a-z]+.*" {
            error-message "The passphrase must have at least one lower case alpha";
          }
          pattern ".*[A-Z]+.*" {
            error-message "The passphrase must have at least one upper case alpha";
          }
          pattern ".*[0-9]+.*" {
            error-message "The passphrase must have at least one digit";
          }
          pattern ".*[<>~;:!@#/$%^&*=-]+.*" {
            error-message "The passphrase must have at least one of these" +
                          " symbols: [<>~;:!@#/$%^&*=-]+";
          }
          pattern ".* .*" {
            modifier invert-match;
            error-message "The passphrase must have no spaces";
          }
        }
      }
      ...
    }
  }
```

For details on the YANG statements used by the YANG model, such as `leaf`, `container`, `list`, `leafref`, `mandatory`, `length`, `pattern`, etc., see the [IETF RFC 7950](https://www.rfc-editor.org/rfc/rfc7950) that documents the YANG 1.1 Data Modeling Language. The `tailf:xyz` are YANG extension statements documented by [tailf\_yang\_extensions(5)](/guides/resources/man/tailf_yang_extensions.5) in Manual Pages.

The service configuration is implemented in YANG by a `key-auth` list where the network element and local user names are the list keys. In addition, the list has a `distkey-servicepoint` service point YANG extension statement to enable the list parameters used by the Python service callbacks that this example implements. Finally, the used `service-data` and `nano-plan-data` groupings add the common definitions for a service and the plan data needed when the service is a nano service.

For the nano service YANG part, an NSO YANG nano service behavior tree extension that references a plan outline extension implements the above steps for setting up SSH public key authentication with a network element:

```
    ncs:plan-outline distkey-plan {
    description "Plan for distributing a public key";
    ncs:component-type "dk:ne" {
      ncs:state "ncs:init";
      ncs:state "dk:generated" {
        ncs:create {
          // Request the generate-keys action
          ncs:post-action-node "$SERVICE" {
            ncs:action-name "generate-keys";
            ncs:result-expr "result = 'true'";
            ncs:sync;
          }
        }
        ncs:delete {
          // Request the delete-keys action
          ncs:post-action-node "$SERVICE" {
            ncs:action-name "delete-keys";
            ncs:result-expr "result = 'true'";
          }
        }
      }
      ncs:state "dk:distributed" {
        ncs:create {
          // Invoke a Python program to distribute the authorized public key to
          // the network element
          ncs:nano-callback;
          ncs:force-commit;
        }
      }
      ncs:state "dk:configured" {
        ncs:create {
          // Invoke a Python program that in turn invokes a service template to
          // configure NSO to use public key authentication with the network
          // element
          ncs:nano-callback;
          // Request the connect action to test the public key authentication
          ncs:post-action-node "/ncs:devices/device[name=$NE-NAME]" {
            ncs:action-name "connect";
            ncs:result-expr "result = 'true'";
          }
        }
      }
      ncs:state "ncs:ready";
    }
  }
  ncs:service-behavior-tree distkey-servicepoint {
    description "One component per distkey behavior tree";
    ncs:plan-outline-ref "dk:distkey-plan";
    ncs:selector {
      // The network element name used with this component
      ncs:variable "NE-NAME" {
        ncs:value-expr "current()/ne-name";
      }
      // The unique component name
      ncs:variable "NAME" {
        ncs:value-expr "concat(current()/ne-name, '-', current()/local-user)";
      }
      // Component for setting up public key authentication
      ncs:create-component "$NAME" {
        ncs:component-type-ref "dk:ne";
      }
    }
  }
```

The nano `service-behavior-tree` for the service point creates a nano service component for each list entry in the `key-auth` list. The last connection verification step of the nano service, the `connected` state, uses the `NE-NAME` variable. The `NAME` variable concatenates the `ne-name` and `local-user` keys from the `key-auth` list to create a unique nano service component name.

The only step that requires both a create and delete part is the `generated` state action that generates the SSH keys. If a user deletes a service instance and another network element does not currently use the generated keys, this deletes the keys too. NSO will revert the configuration automatically as part of the FASTMAP algorithm. Hence, the service list instances also need actions for generating and deleting keys.

```yang
  container pubkey-dist {
    list key-auth {
      key "ne-name local-user";
      ...
      action generate-keys {
        tailf:actionpoint generate-keys;
        output {
          leaf result {
            type boolean;
          }
        }
      }
      action delete-keys {
        tailf:actionpoint delete-keys;
        output {
          leaf result {
            type boolean;
          }
        }
      }
    }
  }
```

The actions have no input statements, as the input is the configuration in the service instance list entry.

The `generated` state uses the `ncs:sync` statement to ensure that the keys exist before the `distributed` state runs. Similarly, the `distributed` state uses the `force-commit` statement to commit the configuration to the NSO CDB and the network elements before the `configured` state runs.

See the `packages/distkey/src/yang/distkey.yang` YANG model for the nano service behavior tree, plan outline, and service configuration implementation.

Next, handling the key generation, distributing keys to the network element, and configuring NSO to authenticate using the keys with the network element requires some code, here written in Python, implemented by the `packages/distkey/python/distkey/distkey-app.py` script application.

The Python script application defines a Python `DistKeyApp` class specified in the `packages/distkey/package-meta-data.xml` file that NSO starts in a Python thread. This Python class inherits `ncs.application.Application` and implements the `setup()` and `teardown()` methods. The `setup()` method registers the nano service `create()` callbacks and the action handlers for generating and deleting the key files. Using the nano service state to separate the two nano service `create()` callbacks for the distribution and NSO configuration of keys, only one Python class, the `DistKeyServiceCallbacks` class, is needed to implement them.

```python
class DistKeyApp(ncs.application.Application):
    def setup(self):
        # Nano service callbacks require a registration for a service point,
        # component, and state, as specified in the corresponding data model
        # and plan outline.
        self.register_nano_service('distkey-servicepoint',  # Service point
                                    'dk:ne',                 # Component
                                    'dk:distributed',        # State
                                    DistKeyServiceCallbacks)
        self.register_nano_service('distkey-servicepoint',  # Service point
                                    'dk:ne',                 # Component
                                    'dk:configured',         # State
                                    DistKeyServiceCallbacks)

        # Side effect action that uses ssh-keygen to create the keyfiles
        self.register_action('generate-keys', GenerateActionHandler)
        # Action to delete the keys created by the generate keys action
        self.register_action('delete-keys', DeleteActionHandler)

    def teardown(self):
        self.log.info('DistKeyApp FINISHED')
```

The action for generating keys calls the OpenSSH `ssh-keygen` command to generate the private and public key files. Calling `ssh-keygen` is kept out of the service `create()` callback to avoid the key generation running multiple times, for example, for service changes, re-deploy, or dry-run commits. Also, NSO encrypts the passphrase used when generating the keys for added security, see the YANG model, so the Python code decrypts it before using it with the `ssh-keygen` command.

```python
class GenerateActionHandler(Action):
    @Action.action
    def cb_action(self, uinfo, name, keypath, ainput, aoutput, trans):
        '''Action callback'''
        service = ncs.maagic.get_node(trans, keypath)
        # Install the crypto keys used to decrypt the service passphrase leaf
        # as input to the key generation.
        with ncs.maapi.Maapi() as maapi:
            _maapi.install_crypto_keys(maapi.msock)
        # Decrypt the passphrase leaf for use when generating the keys
        encrypted_passphrase = service.passphrase
        decrypted_passphrase = _ncs.decrypt(str(encrypted_passphrase))
        aoutput = True
        # If it does not exist already, generate a private and public key
        if os.path.isfile(f'./{service.local_user}_ed25519') == False:
            result = subprocess.run(['ssh-keygen', '-N',
                                    f'{decrypted_passphrase}', '-t', 'ed25519',
                                    '-f', f'./{service.local_user}_ed25519'],
                                    stdout=subprocess.PIPE, check=True,
                                    encoding='utf-8')
            if "has been saved" not in result.stdout:
                aoutput = False
```

The `DeleteActionHandler` action deletes the key files if no more network elements use the user's keys:

```python
class DeleteActionHandler(Action):
    @Action.action
    def cb_action(self, uinfo, name, keypath, ainput, aoutput, trans):
        '''Action callback'''
        service = ncs.maagic.get_node(trans, keypath)
        # Only delete the key files if no more network elements use this
        # user's keys
        cur = trans.cursor('/pubkey-dist/key-auth')
        remove_key = True
        while True:
            try:
                value = next(cur)
                if value[0] != service.ne_name and value[1] == service.local_user:
                    remove_key = False
                    break
            except StopIteration:
                break
        aoutput = True
        if remove_key is True:
            try:
                os.remove(f'./{service.local_user}_ed25519.pub')
                os.remove(f'./{service.local_user}_ed25519')
            except OSError as e:
                if e.errno != errno.ENOENT:
                    aoutput = False
```

The Python class for the nano service `create()` callbacks handles both the distribution and NSO configuration of the keys. The `dk:distributed` state `create()` callback code adds the public key data to the network element's list of authorized keys. For the `create()` call for the `dk:configured` state, a template is used to configure NSO to use public key authentication with the network element. The template can be called directly from the nano service, but in this case, it needs to be called from the Python code to input the current working directory to the template:

```python
class DistKeyServiceCallbacks(NanoService):
    @NanoService.create
    def cb_nano_create(self, tctx, root, service, plan, component, state,
                        proplist, component_proplist):
        '''Nano service create callback'''
        if state == 'dk:distributed':
            # Distribute the public key to the network element's authorized
            # keys list
            with open(f'./{service.local_user}_ed25519.pub', 'r') as f:
                pubkey_data = f.read()
                config = root.devices.device[service.ne_name].config
                users = config.aaa.authentication.users
                users.user[service.local_user].authkey.create(pubkey_data)
        elif state == 'dk:configured':
            # Configure NSO to use a public key for authentication with
            # the network element
            template_vars = ncs.template.Variables()
            template_vars.add('CWD', os.getcwd())
            template = ncs.template.Template(service)
            template.apply('distkey-configured', template_vars)
```

The template to configure NSO to use public key authentication with the network element is available under `packages/distkey/templates/distkey-configured.xml`:

```xml
<config-template xmlns="http://tail-f.com/ns/config/1.0">
  <devices xmlns="http://tail-f.com/ns/ncs" tags="merge">
    <authgroups>
      <group>
        <name>{authgroup-name}</name>
        <umap>
          <local-user>{local-user}</local-user>
          <remote-name>{remote-name}</remote-name>
          <public-key>
            <private-key>
              <file>
                <name>{$CWD}/{local-user}_ed25519</name>
                <passphrase>{passphrase}</passphrase>
              </file>
            </private-key>
          </public-key>
        </umap>
      </group>
    </authgroups>
    <device>
      <name>{ne-name}</name>
      <authgroup>{authgroup-name}</authgroup>
    </device>
  </devices>
</config-template>}
```

The example uses three scripts to showcase the nano service:

* A shell script, `showcase.sh`, which uses the `ncs_cli` program to run CLI commands via the NSO IPC port.
* A Python script, `showcase-rc.sh`, which uses the `requests` package for RESTCONF edit operations and receiving event notifications.
* A Python script that uses NSO MAAPI, `showcase-maapi.sh`, via the NSO IPC port.

The `ncs_cli` program identifies itself with NSO as the `admin` user without authentication, and the RESTCONF client uses plain HTTP and basic user password authentication. All three scripts demonstrate the service by generating keys, distributing the public key, and configuring NSO for public key authentication with the network elements. To run the example, see the instructions in the `README` file of the example.

## Deployment <a href="#d5e1474" id="d5e1474"></a>

See the `README` in the `netsim-sshkey` example's directory for a reference to an NSO system installation in a container deployment variant.

<div data-with-frame="true"><figure><img src="/files/RrEQMfFimm8yWMv7NtrE" alt="" width="375"><figcaption><p>The Deployment Container Topology</p></figcaption></figure></div>

The deployment variant differs from the development example by:

* Installing NSO with a system installation for deployment instead of a local installation suitable for development
* Addressing NSO security by running NSO as the `admin` user and authenticating using a public key and token.
* Rotating NSO logs to avoid running out of disk space
* Installing the `distkey` service package and `ne` NED package at startup
* The NSO CLI showcase script uses SSH with public key authentication instead of the **ncs\_cli** program over unsecured IPC
* There is no Python MAAPI showcase script. Use RESTCONF over HTTPS with Python instead of Python MAAPI over unsecured IPC.
* Having NSO and the network elements (simulated by the ConfD subscriber application) run in separate containers
* NSO is either pre-installed in the NSO production container image or installed in a generic Linux container.

The deployment example sets up a minimal production installation where the NSO process runs as the `admin` OS user, relying on PAM authentication for the `admin` and `oper` NSO users. The `admin` user is authenticated over SSH using a public key for CLI and NETCONF access and over RESTCONF HTTPS using a token. The read-only `oper` user uses password authentication. The `oper` user can access the NSO WebUI over HTTPS port 443 from the container host.

A modified version of the NSO configuration file `ncs.conf` from the example running with a local install NSO is located in the `$NCS_CONFIG_DIR` (`/etc/ncs`) directory. The `packages`, `ncs-cdb`, `state`, and `scripts` directories are now under the `$NCS_RUN_DIR` (`/var/opt/ncs`) directory. The log directory is now the `$NCS_LOG_DIR` (`/var/log/ncs`) directory. Finally, the `$NCS_DIR` variable points to `/opt/ncs/current`.

Two scripts showcase the nano service:

* A shell script that runs NSO CLI commands over SSH.
* A Python script that uses the `requests` package to perform edit operations and receive event notifications.

As with the development version, both scripts will demo the service by generating keys, distributing the public key, and configuring NSO for public key authentication with the network elements.

To run the example and for more details, see the instructions in the `README` file of the [examples.ncs/getting-started/netsim-sshkey](https://github.com/NSO-developer/nso-examples/tree/6.7/getting-started/netsim-sshkey) deployment example.


# Secure Deployment

Security features to consider for NSO deployment.

When deploying NSO in production environments, security should be a primary consideration. This section guides the NSO features available for securing your NSO deployment.

## Development vs. Production Deployment

NSO installations can be configured for development or production use, with significantly different security implications.

### Production Installation

* Use the NSO Installer with the `--system-install` option for production deployments.
  * The `--local-install` option should only be used for development environments.
  * Use the NSO Installer `--run-as-user <User>` option to run NSO as a non-root user.
* Never use `ncs.conf` files from NSO distribution examples in production.
  * Evaluate and customize the default `ncs.conf` file provided with a system installation to meet your specific security requirements.

### Key Configuration Differences

The default `ncs.conf` for production installations differs from the development default `ncs.conf` in several critical security areas:

#### Encryption Keys

* Production (system) installations use external key management where `ncs.conf` points to `${NCS_CONFIG_DIR}/ncs.crypto_keys` using the `${NCS_DIR}/bin/ncs_crypto_keys` command to retrieve them.
* Development installations include the encryption keys directly in `ncs.conf`.

#### SSH Configuration

* Production restricts SSH host key algorithms to `ssh-ed25519` only.
* Development allows multiple algorithms for compatibility.

#### Authentication

* Production disables local authentication by default, using PAM with `system-auth`.
* Development enables local authentication and uses PAM with `common-auth`.
* Production includes password expiration warnings.

#### Network Interfaces

* Production disables CLI SSH, HTTP WebUI, and NETCONF SSH by default.
* Development enables these interfaces for convenience.
* Production enables restricted-file-access for CLI.

See [ncs.conf(5)](/guides/resources/man/ncs.conf.5) for all available options to configure the NSO daemon.

## Eliminating Root Access

Running NSO with minimal privileges is a fundamental security best practice:

* Use the NSO installer `--run-as-user User` option to run NSO as a non-root user.
* Some files are installed with elevated privileges - refer to the [ncs-installer(1)](/guides/resources/man/ncs-installer.1#system-installation) man page under the `--run-as-user User` option for details.
* The NSO production container runs NSO from a [non-root user](/guides/administration/installation-and-deployment/containerized-nso#nso-runs-from-a-non-root-user).
* If the CLI is used and we want to create CLI commands that run executables, we may want to modify the permissions of the `$NCS_DIR/lib/ncs/lib/confd-*/priv/cmdptywrapper` program.\
  To be able to run an executable as root or a specific user, we need to make `cmdptywrapper` `setuid` `root`, i.e.:

  1. `# chown root cmdptywrapper`
  2. `# chmod u+s cmdptywrapper`

  Failing that, all programs will be executed as the user running the `ncs` daemon. Consequently, if that user is the `root`, we do not have to perform the `chmod` operations above. The same applies to executables run via actions, but then we may want to modify the permissions of the `$NCS_DIR/lib/ncs/lib/confd-*/priv/cmdwrapper` program instead:

  1. `# chown root cmdwrapper`
  2. `# chmod u+s cmdwrapper`
* The deployment variant referenced in the README file of the [examples.ncs/getting-started/netsim-sshkey](https://github.com/NSO-developer/nso-examples/tree/6.7/getting-started/netsim-sshkey) example provides a native and NSO production container based example.

## Authentication, Authorization, and Accounting (AAA)

### PAM Authentication

PAM (Pluggable Authentication Modules) is the recommended authentication method for NSO:

* Group assignments based on the OS group database `/etc/group`.
* Default NACM (Network Configuration Access Control Module) settings provide two groups:
  * `ncsadmin`: unlimited access rights.
  * `ncsoper`: minimal access rights (read-only).

See [PAM](https://nso-docs.cisco.com/guides/administration/installation-and-deployment/development-to-production-deployment/pages/oUchYxSeOKkwhPOOefVU#ug.aaa.pam) for details.

### Customizing AAA Configuration

When customizing the default `aaa_init.xml` configuration:

* Exclude credentials unless local authentication is explicitly enabled.
* Never use default passwords.
* Carefully consider which groups can modify NACM rules.
* Tailor NACM settings for user groups based on the principle of least privilege.

See [AAA Infrastructure](/guides/administration/management/aaa-infrastructure) for details.

### Additional Authentication Methods

* CLI and NETCONF: SSH public key authentication.
* RESTCONF: Token, JWT, LDAP, or TACACS+ authentication.
* WebUI: HTTPS (TLS) with JSON-RPC SSO (Single Sign-On).

{% hint style="info" %}
Disable unused interfaces in `ncs.conf` to reduce the attack surface.
{% endhint %}

See [Authentication](https://nso-docs.cisco.com/guides/administration/installation-and-deployment/development-to-production-deployment/pages/oUchYxSeOKkwhPOOefVU#ug.aaa.authentication) for details.

## Securing IPC Access

Inter-Process Communication (IPC) security is crucial for safeguarding NSO's extensibility SDK API communications. Since the IPC socket allows full control of the system, it is important to ensure that only trusted or authorized clients can connect. See [Restricting Access to the IPC Socket](/guides/administration/advanced-topics/ipc-connection#restricting-access-to-the-ipc-socket).

Examples of programs that connect using IPC sockets:

* Remote commands, such as `ncs --reload`.
* MAAPI, CDB, DP, event notification API clients.
* The `ncs_cli` program.
* The `ncs_cmd` and `ncs_load` programs.

### Default Security

* Only local connections to IPC sockets are allowed by default.
* Unix domain sockets (Local IPC) are used by default, relying on filesystem permissions for access control.

### Best Practices

* Use Unix sockets for authenticating the client based on the UID of the other end of the socket connection.
  * Root and the user NSO runs from always have access.
  * If using TCP sockets, configure NSO to use access checks with a pre-shared key.
    * If enabling non-localhost IPC over TCP sockets, implement encryption.

See [Authenticating IPC Access](/guides/administration/management/aaa-infrastructure#authenticating-ipc-access) for details.

## Southbound Interface Security

Secure communication with managed devices:

* Use [Cisco-provided NEDs](/guides/administration/management/ned-administration) when possible.
* Refer to the [examples.ncs/getting-started/netsim-sshkey](https://github.com/NSO-developer/nso-examples/tree/6.7/getting-started/netsim-sshkey) README, which references a deployment variant of the example for SSH key update patterns using nano services.

## Cryptographic Key Management

### Hashing Algorithms

* Set the `ncs.conf` `/ncs-config/crypt-hash/algorithm` to SHA-512 for password hashing.
  * Used by the `ianach:crypt-hash` type for secure password storage.

### Encryption Keys

* Generate new encryption keys before or at startup.
* Replace or rotate keys generated by the NSO installer.
  * Rotate keys periodically.
* Store keys securely (default location: `/etc/ncs/ncs.crypto_keys`).
* The `ncs.crypto_keys` file contains the highly sensitive encryption keys for all encrypted CDB data.

See [Cryptographic Keys](/guides/administration/advanced-topics/cryptographic-keys) for details.

## Rate Limiting and Resource Protection

Implement various limiting mechanisms to prevent resource exhaustion:

### NSO Configuration Limits

NSO can be configured with some limits from `ncs.conf`:

* `/ncs-config/session-limits`: Limit concurrent sessions.
* `/ncs-config/transaction-limits`: Limit concurrent transactions.
* `/ncs-config/parser-limits`: Limit XML data parsing.
* `/ncs-config/webui/transport/unauthenticated-message-limit`: Limit unauthenticated message size.
* `/ncs-config/webui/rate-limiting`: Limit JSON-RPC requests per hour.

### External Rate Limiting

For additional protection, implement rate limiting at the network level using tools like Linux `iptables`.

## High Availability Security

When deploying NSO in [HA (High Availability)](/guides/administration/management/high-availability) configurations:

* RAFT HA:
  * Uses encrypted TLS with mutual X.509 authentication.
* Rule-based HA:
  * Unencrypted communication.
  * Shared token for authentication between HA group nodes.

{% hint style="info" %}
Encrypted strings for all encrypted CDB data, default stored in `/etc/ncs/ncs.crypto_keys`, must be identical across nodes
{% endhint %}

## Compliance Reporting

NSO provides comprehensive [compliance reporting](/guides/operation-and-usage/operations/compliance-reporting) capabilities:

* Track user actions - "Who has done what?"
* Verify network configuration compliance.
* Generate audit reports for regulatory requirements.

## FIPS Mode

For enhanced security and regulatory compliance:

* FIPS mode restricts NSO to use only FIPS 140-3 validated cryptographic modules.
* Enable with the `--fips-install` option during [installation](/guides/administration/installation-and-deployment/system-install).
* Required for certain government and regulated industry deployments.


# Deployment Example

Understand NSO deployment with an example setup.

This section shows examples of a typical deployment for a highly available (HA) setup. A reference to an example implementation of the `tailf-hcc` layer-2 upgrade deployment scenario described here, check the NSO example set under [examples.ncs/high-availability/hcc](https://github.com/NSO-developer/nso-examples/tree/6.7/high-availability/hcc). The example covers the following topics:

* Installation of NSO on all nodes in an HA setup
* Initial configuration of NSO on all nodes
* HA failover
* Upgrading NSO on all nodes in the HA cluster
* Upgrading NSO packages on all nodes in the HA cluster

The deployment examples use both the legacy rule-based and recommended HA Raft setup. See [High Availability](/guides/administration/management/high-availability) for HA details. The HA Raft deployment consists of three nodes running NSO and a node managing them, while the rule-based HA deployment uses only two nodes.

Based on the Raft consensus algorithm, the HA Raft version provides the best fault tolerance, performance, and security and is therefore recommended.

For the HA Raft setup, the NSO nodes `paris.fra`, `london.eng`, and `berlin.ger` nodes make up a cluster of one leader and two followers.

<div data-with-frame="true"><figure><img src="/files/KVCO3x2ayEZxSC37dbpw" alt="" width="375"><figcaption><p>The HA Raft Deployment Network</p></figcaption></figure></div>

For the rule-based HA setup, the NSO nodes `paris` and `london` make up one HA pair — one primary and one secondary.

<div data-with-frame="true"><figure><img src="/files/uMSQmO380RYUeVniSKfo" alt="" width="375"><figcaption><p>The Rule-Based HA Deployment Network</p></figcaption></figure></div>

HA is usually not optional for a deployment. Data resides in CDB, a RAM database with a disk-based journal for persistence. Both HA variants can be set up to avoid the need for manual intervention in a failure scenario, where HA Raft does the best job of keeping the cluster up. See [High Availability](/guides/administration/management/high-availability) for details.

## Initial NSO Installation <a href="#d5e7609" id="d5e7609"></a>

An NSO system installation on the NSO nodes is recommended for deployments. For System Installation details, see the [System Install](/guides/administration/installation-and-deployment/system-install) steps.

In this container-based example, Docker Compose uses a `Dockerfile` to build the container image and install NSO on multiple nodes, here containers. A shell script uses an SSH client to access the NSO nodes from the manager node to demonstrate HA failover and, as an alternative, a Python script that implements SSH and RESTCONF clients.

* An `admin` user is created on the NSO nodes. Password-less `sudo` access is set up to enable the `tailf-hcc` server to run the `ip` command. The manager's SSH client uses public key authentication, while the RESTCONF client uses a token to authenticate with the NSO nodes.

  The example creates two packages using the `ncs-make-package` command: `dummy` and `inert`. A third package, `tailf-hcc`, provides VIPs that point to the current HA leader/primary node.
* The packages are compressed into a `tar.gz` format for easier distribution, but that is not a requirement.

{% hint style="info" %}
While this deployment example uses containers, it is intended as a generic deployment guide. For details on running NSO in a container, such as Docker, see [Containerized NSO](/guides/administration/installation-and-deployment/containerized-nso).
{% endhint %}

This example uses a minimal Red Hat UBI distribution for hosting NSO with the following added packages:

* NSO's basic dependency requirements are fulfilled by adding the Java Runtime Environment (JRE), OpenSSH, and OpenSSL packages.
* The OpenSSH server is used for shell access and secure copy to the NSO Linux host for NSO version upgrade purposes. The NSO built-in SSH server provides CLI and NETCONF access to NSO.
* The NSO services require Python.
* To fulfill the `tailf-hcc` server dependencies, the `iproute2` utilities and `sudo` packages are installed. See [Dependencies](https://nso-docs.cisco.com/guides/administration/installation-and-deployment/development-to-production-deployment/pages/qG3CMifhI63daJ1BZfmB#ug.ha.hcc.deps) (in the section [Tailf HCC Package](https://nso-docs.cisco.com/guides/administration/installation-and-deployment/development-to-production-deployment/pages/qG3CMifhI63daJ1BZfmB#ug.ha.hcc)) for details on dependencies.
* The `rsyslog` package enables storing an NSO log file from several NSO logs locally and forwarding some logs to the manager.
* The `arp` command from the `net-tools` and `iputils` (`ping`) packages have been added for demonstration purposes.

The steps in the list below are performed as `root`. Docker Compose will build the container images, i.e., create the NSO installation as `root`.

The `admin` user will only need `root` access to run the `ip` command when `tailf-hcc` adds the Layer 2 VIP address to the leader/primary node interface.

The initialization steps are also performed as `root` for the nodes that make up the HA cluster:

* Create the `ncsadmin` and `ncsoper` Linux user groups.
* Create and add the `admin` and `oper` Linux users to their respective groups.
* Perform a system installation of NSO that runs NSO as the `admin` user.
* The `admin` user is granted access to run the `ip` command from the `vipctl` script as `root` using the `sudo` command as required by the `tailf-hcc` package.
* The `cmdwrapper` NSO program gets access to run the scripts executed by the `generate-token` action for generating RESTCONF authentication tokens as the current NSO user.
* Password authentication is set up for the read-only `oper` user for use with NSO only, which is intended for WebUI access.
* The `root` user is set up for Linux shell access only.
* The NSO installer, `tailf-hcc` package, application YANG modules, scripts for generating and authenticating RESTCONF tokens, and scripts for running the demo are all available to the NSO and manager containers.
* `admin` user permissions are set for the NSO directories and files created by the system install, as well as for the `root`, `admin`, and `oper` home directories.
* The `ncs.crypto_keys` are generated and distributed to all nodes.\
  \
  **Note**: The `ncs.crypto_keys` file is highly sensitive. It contains the encryption keys for all encrypted CDB data, which often includes passwords for various entities, such as login credentials to managed devices.\
  \
  **Note**: In an NSO System Install setup, not only the TLS certificates (HA Raft) or shared token (rule-based HA) need to match between the HA cluster nodes, but also the configuration for encrypted strings, by default stored in `/etc/ncs/ncs.crypto_keys`, needs to match between the nodes in the HA cluster. For rule-based HA, the tokens configured on the secondary nodes are overwritten with the encrypted token of type `aes-256-cfb-128-encrypted-string` from the primary node when the secondary connects to the primary. If there is a mismatch between the encrypted-string configuration on the nodes, NSO will not decrypt the HA token to match the token presented. As a result, the primary node denies the secondary node access the next time the HA connection needs to be re-established with a "Token mismatch, secondary is not allowed" error.
* For HA Raft, TLS certificates are generated for all nodes.
* The initial NSO configuration, `ncs.conf`, is updated and in sync (identical) on the nodes.
* The SSH servers are configured to allow only SSH public key authentication (no password). The `oper` user can use password authentication with the WebUI but has read-only NSO access.
* The `oper` user is denied access to the Linux shell.
* The `admin` user can access the Linux shell and NSO CLI using public key authentication.
* New keys for all users are distributed to the HA cluster nodes and the manager node when the HA cluster is initialized.
* The OpenSSH server and the NSO built-in SSH server use the same private and public key pairs located under `~/.ssh/id_ed25519`, while the manager public key is stored in the `~/.ssh/authorized_keys` file for both NSO nodes.
* Host keys are generated for all nodes to allow the NSO built-in SSH and OpenSSH servers to authenticate the server to the client.\
  \
  Each HA cluster node has its own unique SSH host keys stored under `${NCS_CONFIG_DIR}/ssh_host_ed25519_key`. The SSH client(s), here the manager, has the keys for all nodes in the cluster paired with the node's hostname and the VIP address in its `/root/.ssh/known_hosts` file.\
  \
  The host keys, like those used for client authentication, are generated each time the HA cluster nodes are initialized. The host keys are distributed to the manager and nodes in the HA cluster before the NSO built-in SSH and OpenSSH servers are started on the nodes.
* As NSO runs in containers, the environment variables are set to point to the system install directories in the Docker Compose `.env` file.
* NSO runs as the non-root `admin` user and, therefore, the NSO system installation is done using the `./nso-${VERSION}.linux.${ARCH}.installer.bin --system-install --run-as-user admin --ignore-init-scripts` options. By default, the NSO installation start script will create a `systemd` system service to run NSO as the `admin` user (default is the `root` user) when NSO is started using the `systemctl start ncs` command.\
  \
  However, this example uses the `--ignore-init-scripts` option to skip installing `systemd` scripts as it runs in a container that does not support `systemd`.\
  \
  The environment variables are copied to a `.pam_environment` file so the `root` and `admin` users can set the required environment variables when those users access the shell via SSH.\
  \
  The `/etc/systemd/system/ncs.service` `systemd` service script is installed as part of the NSO system install, if not using the `--ignore-init-scripts` option, and it can be customized if you would like to use it to start NSO. The script may provide what you need and can be a starting point.
* The OpenSSH `sshd` and `rsyslog` daemons are started.
* The packages from the package store are added to the `${NCS_RUN_DIR}/packages` directory before finishing the initialization part in the `root` context.
* The NSO smart licensing token is set.

## The `ncs.conf` Configuration <a href="#d5e7783" id="d5e7783"></a>

* The NSO IPC socket is configured in `ncs.conf` to only listen to localhost 127.0.0.1 connections, which is the default setting.\
  \
  By default, the clients connecting to the NSO IPC socket are considered trusted, i.e., no authentication is required, and the use of 127.0.0.1 with the `/ncs-config/ncs-ipc-address` IP address in `ncs.conf` to prevent remote access. See [Security Considerations](#ug.admin_guide.deployment.security) and [ncs.conf(5)](/guides/resources/man/ncs.conf.5) in Manual Pages for more details.
* `/ncs-config/aaa/pam` is set to enable PAM to authenticate users as recommended. All remote access to NSO must now be done using the NSO host's privileges. See [ncs.conf(5)](/guides/resources/man/ncs.conf.5) in Manual Pages for details.
* Depending on your Linux distribution, you may have to change the `/ncs-config/aaa/pam/service` setting. The default value is `common-auth`. Check the file `/etc/pam.d/common-auth` and make sure it fits your needs. See [ncs.conf(5)](/guides/resources/man/ncs.conf.5) in Manual Pages for details.\
  \
  Alternatively, or as a complement to the PAM authentication, users can be stored in the NSO CDB database or authenticated externally. See [Authentication](https://nso-docs.cisco.com/guides/administration/installation-and-deployment/development-to-production-deployment/pages/oUchYxSeOKkwhPOOefVU#ug.aaa.authentication) for details.
* RESTCONF token authentication under `/ncs-config/aaa/external-validation` is enabled using a `token_auth.sh` script that was added earlier together with a `generate_token.sh` script. See [ncs.conf(5)](/guides/resources/man/ncs.conf.5) in Manual Pages for details.\
  \
  The scripts allow users to generate a token for RESTCONF authentication through, for example, the NSO CLI and NETCONF interfaces that use SSH authentication or the Web interface.

  The token provided to the user is added to a simple YANG list of tokens where the list key is the username.
* The token list is stored in the NSO CDB operational data store and is only accessible from the node's local MAAPI and CDB APIs. See the HA Raft and rule-based HA `upgrade-l2/manager-etc/yang/token.yang` file in the examples.
* The NSO web server HTTPS interface should be enabled under `/ncs-config/webui`, along with `/ncs-config/webui/match-host-name = true` and `/ncs-config/webui/server-name` set to the hostname of the node, following security best practice. If the server needs to serve multiple domains or IP addresses, additional `server-alias` values can be configured. See [ncs.conf(5)](/guides/resources/man/ncs.conf.5) in Manual Pages for details.

  **Note**: The SSL certificates that NSO generates are self-signed:

  ```bash
        $ openssl x509 -in /etc/ncs/ssl/cert/host.cert -text -noout
        Certificate:
        Data:
        Version: 1 (0x0)
        Serial Number: 2 (0x2)
        Signature Algorithm: sha256WithRSAEncryption
        Issuer: C=US, ST=California, O=Internet Widgits Pty Ltd, CN=John Smith
        Validity
        Not Before: Dec 18 11:17:50 2015 GMT
        Not After : Dec 15 11:17:50 2025 GMT
        Subject: C=US, ST=California, O=Internet Widgits Pty Ltd
        Subject Public Key Info:
        .......
  ```

  Thus, if this is a production environment and the JSON-RPC and RESTCONF interfaces using the web server are not used solely for internal purposes, the self-signed certificate must be replaced with a properly signed certificate. See [ncs.conf(5)](/guides/resources/man/ncs.conf.5) in Manual Pages under `/ncs-config/webui/transport/ssl/cert-file` and `/ncs-config/restconf/transport/ssl/certFile` for more details.
* Disable `/ncs-config/webui/cgi` unless needed.
* The NSO SSH CLI login is enabled under `/ncs-config/cli/ssh/enabled`. See [ncs.conf(5)](/guides/resources/man/ncs.conf.5) in Manual Pages for details.
* The NSO CLI style is set to C-style, and the CLI prompt is modified to include the hostname under `/ncs-config/cli/prompt`. See [ncs.conf(5)](/guides/resources/man/ncs.conf.5) in Manual Pages for details.

  ```xml
      <prompt1>\u@nso-\H> </prompt1>
      <prompt2>\u@nso-\H% </prompt2>

      <c-prompt1>\u@nso-\H# </c-prompt1>
      <c-prompt2>\u@nso-\H(\m)# </c-prompt2>
  ```
* NSO HA Raft is enabled under `/ncs-config/ha-raft`, and the rule-based HA under `/ncs-config/ha`. See [ncs.conf(5)](/guides/resources/man/ncs.conf.5) in Manual Pages for details.
* Depending on your provisioned applications, you may want to turn `/ncs-config/rollback/enabled` off. Rollbacks do not work well with nano service reactive FASTMAP applications or if maximum transaction performance is a goal. If your application performs classical NSO provisioning, the recommendation is to enable rollbacks. Otherwise not. See [ncs.conf(5)](/guides/resources/man/ncs.conf.5) in Manual Pages for details.

## The `aaa_init.xml` Configuration <a href="#ug.admin_guide.deployment.aaa" id="ug.admin_guide.deployment.aaa"></a>

The NSO System Install places an AAA `aaa_init.xml` file in the `$NCS_RUN_DIR/cdb` directory. Compared to a Local Install for development, no users are defined for authentication in the `aaa_init.xml` file, and PAM is enabled for authentication. NACM rules for controlling NSO access are defined in the file for users belonging to a `ncsadmin` user group and read-only access for a `ncsoper` user group. As seen in the previous sections, this example creates Linux `root`, `admin`, and `oper` users, as well as the `ncsadmin` and `ncsoper` Linux user groups.

PAM authenticates the users using SSH public key authentication without a passphrase for NSO CLI and NETCONF login. Password authentication is used for the `oper` user intended for NSO WebUI login and token authentication for RESTCONF login.

Before the NSO daemon is running, and there are no existing CDB files, the default AAA configuration in the `aaa_init.xml` is used. It is restrictive and is used for this demo with only a minor addition to allow the oper user to generate a token for RESTCONF authentication.

The NSO authorization system is group-based; thus, for the rules to apply to a specific user, the user must be a member of the group to which the restrictions apply. PAM performs the authentication, while the NSO NACM rules do the authorization.

* Adding the `admin` user to the `ncsadmin` group and the `oper` user to the limited `ncsoper` group will ensure that the two users get properly authorized with NSO.
* Not adding the `root` user to any group matching the NACM groups results in zero access, as no NACM rule will match, and the default in the `aaa_init.xml` file is to deny all access.

The NSO NACM functionality is based on the [Network Configuration Access Control Model](https://datatracker.ietf.org/doc/html/rfc8341) IETF RFC 8341 with NSO extensions augmented by `tailf-acm.yang`. See [AAA infrastructure](/guides/administration/management/aaa-infrastructure), for more details.

The manager in this example logs into the different NSO hosts using the Linux user login credentials. This scheme has many advantages, mainly because all audit logs on the NSO hosts will show who did what and when. Therefore, the common bad practice of having a shared `admin` Linux user and NSO local user with a shared password is not recommended.

{% hint style="info" %}
The default `aaa_init.xml` file provided with the NSO system installation must not be used as-is in a deployment without reviewing and verifying that every NACM rule in the file matches the desired authorization level.
{% endhint %}

## The High Availability and VIP Configuration <a href="#d5e7892" id="d5e7892"></a>

This example sets up one HA cluster using HA Raft or rule-based HA with the `tailf-hcc` server to manage virtual IP addresses. See [NSO Rule-based HA](/guides/administration/management/high-availability) and [Tail-f HCC Package](https://nso-docs.cisco.com/guides/administration/installation-and-deployment/development-to-production-deployment/pages/qG3CMifhI63daJ1BZfmB#ug.ha.hcc) for details.

The NSO HA, together with the `tailf-hcc` package, provides three features:

* All CDB data is replicated from the leader/primary to the follower/secondary nodes.
* If the leader/primary fails, a follower/secondary takes over and starts to act as leader/primary. This is how HA Raft works and how the rule-based HA variant of this example is configured to handle failover automatically.
* At failover, `tailf-hcc` sets up a virtual alias IP address on the leader/primary node only and uses gratuitous ARP packets to update all nodes in the network with the new mapping to the leader/primary node.

Nodes in other networks can be updated using the `tailf-hcc` layer-3 BGP functionality or a load balancer. See the `load-balancer`and `hcc`examples in the NSO example set under [examples.ncs/high-availability](https://github.com/NSO-developer/nso-examples/tree/6.7/high-availability).

See the NSO example set under [examples.ncs/high-availability/hcc](https://github.com/NSO-developer/nso-examples/tree/6.7/high-availability/hcc) for a reference to an HA Raft and rule-based HA `tailf-hcc` Layer 3 BGP examples.

The HA Raft and rule-based HA upgrade-l2 examples also demonstrate HA failover, upgrading the NSO version on all nodes, and upgrading NSO packages on all nodes.

## Global Settings and Timeouts <a href="#d5e7915" id="d5e7915"></a>

Depending on your installation, e.g., the size and speed of the managed devices and the characteristics of your service applications, some default values of NSO may have to be tweaked, particularly some of the timeouts.

* Device timeouts. NSO has connect, read, and write timeouts for traffic between NSO and the managed devices. The default value may not be sufficient if devices/nodes are slow to commit, while some are sometimes slow to deliver their full configuration. Adjust timeouts under `/devices/global-settings` accordingly.
* Service code timeouts. Some service applications can sometimes be slow. Adjusting the `/services/global-settings/service-callback-timeout` configuration might be applicable depending on the applications. However, the best practice is to change the timeout per service from the service code using the Java `ServiceContext.setTimeout` function or the Python `data_set_timeout` function.

There are quite a few different global settings for NSO. The two mentioned above often need to be changed.

## Cisco Smart Licensing <a href="#d5e7928" id="d5e7928"></a>

NSO uses Cisco Smart Licensing, which is described in detail in [Cisco Smart Licensing](/guides/administration/management/system-management/cisco-smart-licensing). After registering your NSO instance(s), and receiving a token, following steps 1-6 as described in the [Create a License Registration Token](/guides/administration/management/system-management/cisco-smart-licensing#d5e2927) section of Cisco Smart Licensing, enter a token from your Cisco Smart Software Manager account on each host. Use the same token for all instances and script entering the token as part of the initial NSO configuration or from the management node:

```bash
admin@nso-paris# license smart register idtoken YzY2Yj...
admin@nso-london# license smart register idtoken YzY2Yj...
```

{% hint style="info" %}
The Cisco Smart Licensing CLI command is present only in the Cisco Style CLI, which is the default CLI for this setup.
{% endhint %}

## Log Management <a href="#d5e7939" id="d5e7939"></a>

### Log Rotate <a href="#d5e7941" id="d5e7941"></a>

The NSO system installations performed on the nodes in the HA cluster also install defaults for **logrotate**. Inspect `/etc/logrotate.d/ncs` and ensure that the settings are what you want. Note that the NSO error logs, i.e., the files `/var/log/ncs/ncserr.log*`, are internally rotated by NSO and must not be rotated by `logrotate`.

### Syslog <a href="#d5e7948" id="d5e7948"></a>

For the HA Raft and rule-based HA upgrade-l2 examples, see the reference from the `README` in the [examples.ncs/high-availability/hcc](https://github.com/NSO-developer/nso-examples/tree/6.7/high-availability/hcc) example directory; the examples integrate with `rsyslog` to log the `ncs`, `developer`, `upgrade`, `audit`, `netconf`, `snmp`, and `webui-access` logs to syslog with `facility` set to `daemon` in `ncs.conf`.

`rsyslogd` on the nodes in the HA cluster is configured to write the daemon facility logs to `/var/log/daemon.log`, and forward the daemon facility logs with the severity `info` or higher to the manager node's `/var/log/ha-cluster.log` syslog.

### Audit Network Log and NED Traces <a href="#d5e7968" id="d5e7968"></a>

Use the audit-network-log for recording southbound traffic towards devices. Enable by setting `/ncs-config/logs/audit-network-log/enabled` and `/ncs-config/logs/audit-network-log/file/enabled` to true in `$NCS_CONFIG_DIR/ncs.conf`, See [ncs.conf(5)](/guides/resources/man/ncs.conf.5) in Manual Pages for more information.

NED trace logs are a crucial tool for debugging NSO installations and not recommended for deployment. These logs are very verbose and for debugging only. Do not enable these logs in production.

Note that the NED logs include everything, even potentially sensitive data is logged. No filtering is done. The NED trace logs are controlled through the CLI under: `/device/global-settings/trace`. It is also possible to control the NED trace on a per-device basis under `/devices/device[name='x']/trace`.

There are three different settings for trace output. For various historical reasons, the setting that makes the most sense depends on the device type.

* For all CLI NEDs, use the `raw` setting.
* For all ConfD and netsim-based NETCONF devices, use the pretty setting. This is because ConfD sends the NETCONF XML unformatted, while `pretty` means that the XML is formatted.
* For Juniper devices, use the `raw` setting. Juniper devices sometimes send broken XML that cannot be formatted appropriately. However, their XML payload is already indented and formatted.
* For generic NED devices - depending on the level of trace support in the NED itself, use either `pretty` or `raw`.
* For SNMP-based devices, use the `pretty` setting.

Thus, it is usually not good enough to control the NED trace from `/devices/global-settings/trace`.

### Python Logs <a href="#d5e7999" id="d5e7999"></a>

While there is a global log for, for example, compilation errors in `/var/log/ncs/ncs-python-vm.log`, logs from user application packages are written to separate files for each package, and the log file naming is `ncs-python-vm-`*`pkg_name`*`.log`. The level of logging from Python code is controlled on a per package basis. See [Debugging of Python packages](/guides/development/core-concepts/nso-virtual-machines/nso-python-vm#debugging-of-python-packages) for more details.

### Java Logs <a href="#d5e8006" id="d5e8006"></a>

User application Java logs are written to `/var/log/ncs/ncs-java-vm.log`. The level of logging from Java code is controlled per Java package. See [Logging](/guides/development/core-concepts/nso-virtual-machines/nso-java-vm#logging) in Java VM for more details.

### Internal NSO Log <a href="#d5e8011" id="d5e8011"></a>

The internal NSO log resides at `/var/log/ncs/ncserr.*`. The log is written in a binary format. To view the internal error log, run the following command:

```bash
  $ ncs --printlog /var/log/ncs/ncserr.log.1
```

## Monitoring the Installation <a href="#d5e8018" id="d5e8018"></a>

All large-scale deployments employ monitoring systems. There are plenty of good tools to choose from, open source and commercial. All good monitoring tools can script (using various protocols) what should be monitored. It is recommended that a special read-only Linux user without shell access be set up like the `oper` user earlier in this chapter. A few commonly used checks include:

* At startup, check that NSO has been started using the `$NCS_DIR/bin/ncs_cmd -c "wait-start 2"` command.
* Use the `ssh` command to verify SSH access to the NSO host and NSO CLI.
* Check disk usage using, for example, the `df` utility.
* For example, use **curl** or the Python requests library to verify that the RESTCONF API is accessible.
* Check that the NETCONF API is accessible using, for example, the `$NCS_DIR/bin/netconf-console` tool with a `hello` message.
* Verify the NSO version using, for example, the `$NCS_DIR/bin/ncs --version` or RESTCONF `/restconf/data/tailf-ncs-monitoring:ncs-state/version`.
* Check if HA is enabled using, for example, RESTCONF `/restconf/data/tailf-ncs-monitoring:ncs-state/ha`.

### Alarms <a href="#d5e8046" id="d5e8046"></a>

RESTCONF can be used to view the NSO alarm table and subscribe to alarm notifications. NSO alarms are not events. Whenever an NSO alarm is created, a RESTCONF notification and SNMP trap are also sent, assuming that you have a RESTCONF client registered with the alarm stream or configured a proper SNMP target, unless the alarm type is filtered through `/alarms/control/filter-types`. Some alarms, like the rule-based HA `ha-secondary-down` alarm, require the intervention of an operator. Thus, a monitoring tool should also fetch the NSO alarm list.

```bash
$ curl -ik -H "X-Auth-Token: TsZTNwJZoYWBYhOPuOaMC6l41CyX1+oDaasYqQZqqok=" \
https://paris:8888/restconf/data/tailf-ncs-alarms:alarms
```

Or subscribe to the `ncs-alarms` RESTCONF notification stream.

### Metric - Counters, Gauges, and Rate of Change Gauges <a href="#d5e8053" id="d5e8053"></a>

NSO metric has different contexts all containing different counters, gauges, and rate of change gauges. There is a `sysadmin`, a `developer` and a `debug` context. Note that only the `sysadmin` context is enabled by default, as it is designed to be lightweight. Consult the YANG module `tailf-ncs-metric.yang` to learn the details of the different contexts.

### **Counters**

You may read counters by e.g. CLI, as in this example:

```bash
admin@ncs# show metric sysadmin counter session cli-total
metric sysadmin counter session cli-total 1
```

### **Gauges**

You may read gauges by e.g. CLI, as in this example:

```bash
admin@ncs# show metric sysadmin gauge session cli-open
metric sysadmin gauge session cli-open 1
```

### **Rate of Change Gauges**

You may read rate of change gauges by e.g. CLI, as in this example:

```bash
admin@ncs# show metric sysadmin gauge-rate session cli-open
NAME  RATE
-------------
1m    0.0
5m    0.2
15m   0.066
```

## Security Considerations <a href="#ug.admin_guide.deployment.security" id="ug.admin_guide.deployment.security"></a>

This section covers security considerations for this example. See [Secure Deployment Considerations](/guides/administration/installation-and-deployment/development-to-production-deployment/secure-deployment) for a general description.

The presented configuration enables the built-in web server for the WebUI and RESTCONF interfaces. It is paramount for security that you only enable HTTPS access with `/ncs-config/webui/match-host-name` and `/ncs-config/webui/server-name` properly set.

The AAA setup described so far in this deployment document is the recommended AAA setup. To reiterate:

* Have all users that need access to NSO authenticated through Linux PAM. This may then be through `/etc/passwd`. Avoid storing users in CDB.
* Given the default NACM authorization rules, you should have three different types of users on the system.
  * Users with shell access are members of the `ncsadmin` Linux group and are considered fully trusted because they have full access to the system.
  * Users without shell access who are members of the `ncsadmin` Linux group have full access to the network. They have access to the NSO SSH shell and can execute RESTCONF calls, access the NSO CLI, make configuration changes, etc. However, they cannot manipulate backups or perform system upgrades unless such actions are added to by NSO applications.
  * Users without shell access who are members of the `ncsoper` Linux group have read-only access. They can access the NSO SSH shell, read data using RESTCONF calls, etc. However, they cannot change the configuration, manipulate backups, and perform system upgrades.

If you have more fine-grained authorization requirements than read-write and read-only, additional Linux groups can be created, and the NACM rules can be updated accordingly. See [The `aaa_init.xml` Configuration](#ug.admin_guide.deployment.aaa) from earlier in this chapter on how the reference example implements users, groups, and NACM rules to achieve the above.

The default `aaa_init.xml` file must not be used as-is before reviewing and verifying that every NACM rule in the file matches the desired authorization level.

For a detailed discussion of the configuration of authorization rules through NACM, see [AAA infrastructure](/guides/administration/management/aaa-infrastructure), particularly the section [Authorization](https://nso-docs.cisco.com/guides/administration/installation-and-deployment/development-to-production-deployment/pages/oUchYxSeOKkwhPOOefVU#ug.aaa.authorization).

A considerably more complex scenario is when users require shell access to the host but are either untrusted or should not have any access to NSO at all. NSO listens on an IPC socket, which by default is a Unix domain socket (Local IPC) configured through `/ncs-config/ncs-local-ipc`. Alternatively, NSO can be configured to use a TCP socket through `/ncs-config/ncs-ipc-address`. In either case, the socket is typically limited to local connections for security. The socket multiplexes several different access methods to NSO.

The main security-related point is that no AAA checks are performed on this socket. If you have access to the socket, you also have complete access to all of NSO.

To drive this point home, when you invoke the `ncs_cli` command, a small C program that connects to the socket and tells NSO who you are, NSO assumes that authentication has already been performed. There is even a documented flag `--noaaa`, which tells NSO to skip all NACM rule checks for this session.

You must protect the socket to prevent untrusted Linux shell users from accessing the NSO instance using this method. This is done by using a file in the Linux file system. The file `/etc/ncs/ipc_access` gets created and populated with random data at install time. Enable `/ncs-config/ncs-ipc-access-check/enabled` in `ncs.conf` and ensure that trusted users can read the `/etc/ncs/ipc_access` file, for example, by changing group access to the file. See [ncs.conf(5)](/guides/resources/man/ncs.conf.5) in Manual Pages for details.

```bash
$ cat /etc/ncs/ipc_access
cat: /etc/ncs/ipc_access: Permission denied
$ sudo chown root:ncsadmin /etc/ncs/ipc_access
$ sudo chmod g+r /etc/ncs/ipc_access
$ ls -lat /etc/ncs/ipc_access
$ cat /etc/ncs/ipc_access
.......
```

For an HA setup, HA Raft is based on the Raft consensus algorithm and provides the best fault tolerance, performance, and security. It is therefore recommended over the legacy rule-based HA variant. The `raft-upgrade-l2` project, referenced from the NSO example set under [examples.ncs/high-availability/hcc](https://github.com/NSO-developer/nso-examples/tree/6.7/high-availability/hcc), together with this Deployment Example section, describes a reference implementation. See [NSO HA Raft](https://nso-docs.cisco.com/guides/administration/installation-and-deployment/development-to-production-deployment/pages/qG3CMifhI63daJ1BZfmB#ug.ha.raft) for more HA Raft details.


# Upgrade NSO

Upgrade NSO to a higher version.

Upgrading the NSO software gives you access to new features and product improvements. Every change carries a risk, and upgrades are no exception. To minimize the risk and make the upgrade process as painless as possible, this section describes the recommended procedures and practices to follow during an upgrade.

As usual, sufficient preparation avoids many pitfalls and makes the process more straightforward and less stressful.

## Preparing for Upgrade <a href="#d5e6830" id="d5e6830"></a>

There are multiple aspects that you should consider before starting with the actual upgrade procedure. While the development team tries to provide as much compatibility between software releases as possible, they cannot always avoid all incompatible changes. For example, when a deviation from an RFC standard is found and resolved, it may break clients that depend on the non-standard behavior. For this reason, a distinction is made between maintenance and a major NSO upgrade.

A maintenance NSO upgrade is within the same branch, i.e., when the first two version numbers stay the same (x.y in the x.y.z NSO version). An example is upgrading from version 6.2.1 to 6.2.2. In the case of a maintenance upgrade, the NSO release contains only corrections and minor enhancements, minimizing the changes. It includes binary compatibility for packages, so there is no need to recompile the .fxs files for a maintenance upgrade.

Correspondingly, when the first or second number in the version changes, that is called a full or major upgrade. For example, upgrading version 6.3.1 to 6.4 is a major, non-maintenance upgrade. Due to new features, packages must be recompiled, and some incompatibilities could manifest.

In addition to the above, a package upgrade is when you replace a package with a newer version, such as a NED or a service package. Sometimes, when package changes are not too big, it is possible to supply the new packages as part of the NSO upgrade, but this approach brings additional complexity. Instead, package upgrade and NSO upgrade should in general, be performed as separate actions and are covered as such.

To avoid surprises during any upgrade, first ensure the following:

* Hosts have sufficient disk space, as some additional space is required for an upgrade.
* The software is compatible with the target OS. However, sometimes a newer version of Java or system libraries, such as glibc, may be required.
* All the required NEDs and custom packages are compatible with the target NSO version. If you're planning to run the upgraded version in FIPS-compliant mode, make sure to upgrade the NEDs to the latest version.
* Existing packages have been compiled for the new version and are available to you during the upgrade.
* Check whether the existing `ncs.conf` file can be used as-is or needs updating. For example, stronger encryption algorithms may require you to configure additional keying material.
* Review the `CHANGES` file for information on what has changed.
* If upgrading from a no longer supported software version, verify that the upgrade can be performed directly. In situations where the currently installed version is very old, you may have to upgrade to one or more intermediate versions before upgrading to the target version.

In case it turns out that any of the packages are incompatible or cannot be recompiled, you will need to contact the package developers for an updated or recompiled version. For an official Cisco-supplied package, it is recommended that you always obtain a pre-compiled version if it is available for the target NSO release, instead of compiling the package yourself.

Additional preparation steps may be required based on the upgrade and the actual setup, such as when using the Layered Service Architecture (LSA) feature. In particular, for a major NSO upgrade in a multi-version LSA cluster, ensure that the new version supports the other cluster members and follow the additional steps outlined in [Deploying LSA](/guides/administration/advanced-topics/layered-service-architecture#deploying-lsa) in Layered Service Architecture.

If you use the High Availability (HA) feature, the upgrade consists of multiple steps on different nodes. To avoid mistakes, you are encouraged to script the process, for which you will need to set up and verify access to all NSO instances with either `ssh`, `nct`, or some other remote management command. For the reference example, we use in this chapter, see [examples.ncs/high-availability/hcc](https://github.com/NSO-developer/nso-examples/tree/6.7/high-availability/hcc). The management station uses shell and Python scripts that use `ssh` to access the Linux shell and NSO CLI and Python Requests for NSO RESTCONF interface access.

Likewise, NSO 5.3 added support for 256-bit AES encrypted strings, requiring the AES256CFB128 key in the `ncs.conf` configuration. You can generate one with the `openssl rand -hex 32` or a similar command. Alternatively, if you use an external command to provide keys, ensure that it includes a value for an `AES256CFB128_KEY` in the output.

With regard to init system, NSO 6.4 introduces `systemd` as the default option instead of SysV. In interactive mode, when upgrading to NSO 6.4 and later, the installer prompts the user to continue using the old SysV service or prepare a `systemd` service. In non-interactive mode, a `systemd` service is prepared by default. When using the `--non-interactive` option, the `/etc/systemd/system/ncs.service` file will be overwritten if it already exists.

Finally, regardless of the upgrade type, ensure that you have a working backup and can easily restore the previous configuration if needed, as described in [Backup and Restore](/guides/administration/management/system-management#backup-and-restore).

{% hint style="danger" %}
**Caution**

The `ncs-backup` (and consequently the `nct backup`) command does not back up the `/opt/ncs/packages` folder. If you make any file changes, back them up separately.

However, the best practice is not to modify packages in the `/opt/ncs/packages` folder. Instead, if an upgrade requires package recompilation, separate package folders (or files) should be used, one for each NSO version.
{% endhint %}

## Single Instance Upgrade <a href="#ug.admin_guide.manual_upgrade" id="ug.admin_guide.manual_upgrade"></a>

The upgrade of a single NSO instance requires the following steps:

1. Create a backup.
2. Perform a System Install of the new version.
3. Stop the old NSO server process.
4. Compact the CDB files write log.
5. Update the `/opt/ncs/current` symbolic link.
6. If required, update the `ncs.conf` configuration file.
7. Update the packages in `/var/opt/ncs/packages/` if recompilation is needed.
8. Start the NSO server process, instructing it to reload the packages.

{% hint style="info" %}
The following steps assume that you are upgrading to the 6.5 release. They pertain to a System Install of NSO, and you must perform them with Super User privileges.

If you're upgrading from a non-FIPS setup to a [FIPS](https://www.nist.gov/itl/publications-0/federal-information-processing-standards-fips)-compliant setup, ensure that the system requirements comply to FIPS mode install. This entails considering FIPS compliance at OS level as well as configuring NSO to use only FIPS-validated algorithms for keys and certificates.
{% endhint %}

{% stepper %}
{% step %}
As a best practice, always create a backup before trying to upgrade.

```bash
# ncs-backup
```

{% endstep %}

{% step %}
For the upgrade itself, you must first download to the host and install the new NSO release. At this point, you can choose to install NSO in standard mode or in FIPS mode.

{% tabs %}
{% tab title="Standard System Install" %}
The standard mode is the regular NSO install and is suitable for most installations. FIPS is disabled in this mode.

For standard NSO installation, run the installer as below:

```bash
# sh nso-6.5.linux.x86_64.installer.bin --system-install
```

{% endtab %}

{% tab title="FIPS System Install" %}
FIPS mode creates a FIPS-compliant NSO install.

FIPS mode should only be used for deployments that are subject to strict compliance regulations as the cryptographic functions are then confined to the CiscoSSL FIPS 140-3 module library.

For FIPS-compliant NSO install, run the installer with the additional `--fips-install` flag. Afterwards, if needed, enable FIPS in `ncs.conf` as described further below.

```bash
# sh nso-6.5.linux.x86_64.installer.bin --system-install --fips-install
```

{% endtab %}
{% endtabs %}
{% endstep %}

{% step %}
Stop the currently running server with the help of `systemd` or an equivalent command relevant to your system.

```bash
# systemctl stop ncs
Stopping ncs: .
```

{% endstep %}

{% step %}
Compact the CDB files write log using, for example, the `ncs --cdb-compact $NCS_RUN_DIR/cdb` command.
{% endstep %}

{% step %}
Next, you update the symbolic link for the currently selected version to point to the newly installed one, 6.5 in this case.

```bash
# cd /opt/ncs
# rm -f current
# ln -s ncs-6.5 current
```

{% endstep %}

{% step %}
While seldom necessary, at this point, you would also update the `/etc/ncs/ncs.conf` file. If you ran the installer with FIPS mode, update `ncs.conf` accordingly.

{% hint style="info" %}
**NSO Configuration for FIPS**

Note the following as part of FIPS-specific configuration:

1. If you're upgrading from a non-FIPS version (e.g., 6.4) to a FIPS-compliant version (e.g., 6.5), the following `ncs.conf` entry needs to be manually added to enable FIPS. Afterwards, upon upgrading between FIPS-compliant versions, the existing entry automatically updates, eliminating the need for any manual intervention.

```xml
<fips-mode>
    <enabled>true</enabled>
</fips-mode>
```

2. Additional environment variables (`NCS_OPENSSL_CONF_INCLUDE`, `NCS_OPENSSL_CONF`, `NCS_OPENSSL_MODULES`) are configured in `ncsrc` for FIPS compliance.
3. The default `crypto.so` is overwritten at install for FIPS compliance.

Additionally, note that:

* As certain algorithms typically available with CiscoSSL are not included in the FIPS 140-3 validated module (and therefore disabled in FIPS mode), you need to configure NSO to use only the algorithms and cryptographic suites available through the CiscoSSL FIPS 140-3 object module.
* With FIPS, NSO signals the NEDs to operate in FIPS mode using Bouncy Castle FIPS libraries for Java-based components, ensuring compliance with FIPS 140-3. To support this, NED packages may also require upgrading, as older versions — particularly SSH-based NEDs — often lack the necessary FIPS signaling or Bouncy Castle support required for cryptographic compliance.
* Configure SSH keys in `ncs.conf` and `init.xml`.
  {% endhint %}
  {% endstep %}

{% step %}
Now, ensure that the `/var/opt/ncs/packages/` directory has appropriate packages for the new version. It should be possible to continue using the same packages for a maintenance upgrade. But for a major upgrade, you must normally rebuild the packages or use pre-built ones for the new version. You must ensure this directory contains the exact same version of each existing package, compiled for the new release, and nothing else.

As a best practice, the available packages are kept in `/opt/ncs/packages/` and `/var/opt/ncs/packages/` only contains symbolic links. In this case, to identify the release for which they were compiled, the package file names all start with the corresponding NSO version. Then, you only need to rearrange the symbolic links in the `/var/opt/ncs/packages/` directory.

```bash
# cd /var/opt/ncs/packages/
# rm -f *
# for pkg in /opt/ncs/packages/ncs-6.5-*; do ln -s $pkg; done
```

{% hint style="warning" %}
Please note that the above package naming scheme is neither required nor enforced. If your package filesystem names differ from it, you will need to adjust the preceding command accordingly.
{% endhint %}
{% endstep %}

{% step %}
Finally, you start the new version of the NSO server with the `package reload` flag set. Set `NCS_RELOAD_PACKAGES=true` in `/etc/ncs/ncs.systemd.conf` and start NSO:

```bash
# systemctl start ncs
Starting ncs: ...
```

Set the `NCS_RELOAD_PACKAGES` variable in `/etc/ncs/ncs.systemd.conf` back to its previous value or the system would keep performing a packages reload at subsequent starts.

NSO will perform the necessary data upgrade automatically. However, this process may fail if you have changed or removed any packages. In that case, ensure that the correct versions of all packages are present in `/var/opt/ncs/packages/` and retry the preceding command.

Also, note that with many packages or data entries in the CDB, this process could take more than 90 seconds and result in the following error message:

```
Starting ncs (via systemctl): Job for ncs.service failed
because a timeout was exceeded. See "systemctl status
ncs.service" and "journalctl -xe" for details. [FAILED]
```

The above error does not imply that NSO failed to start, just that it took longer than 90 seconds. Therefore, it is recommended you wait some additional time before verifying.
{% endstep %}
{% endstepper %}

## Recover from a Failed Upgrade <a href="#d5e6931" id="d5e6931"></a>

It is imperative that you have a working copy of data available from which you can restore. That is why you must always create a backup before starting an upgrade. Only a backup guarantees that you can rerun the upgrade or back out of it, should it be necessary.

The same steps can also be used to restore data on a new, similar host if the OS of the initial host becomes corrupted beyond repair.

1. First, stop the NSO process if it is running.

   ```bash
   # systemctl stop ncs
   Stopping ncs: .
   ```
2. Verify and, if necessary, revert the symbolic link in `/opt/ncs/` to point to the initial NSO release.

   ```bash
   # cd /opt/ncs
   # ls -l current
   # ln -s ncs-VERSION current
   ```

   \
   In the exceptional case where the initial version installation was removed or damaged, you will need to re-install it first and redo the step above.
3. Verify if the correct (initial) version of NSO is being used.

   ```bash
   # ncs --version
   ```
4. Next, restore the backup.

   ```bash
   # ncs-backup --restore
   ```
5. Finally, start the NSO server and verify the restore was successful.

   ```bash
   # systemctl start ncs
   Starting ncs: .
   ```

## NSO HA Version Upgrade <a href="#ch_upgrade.ha" id="ch_upgrade.ha"></a>

Upgrading NSO in a highly available (HA) setup is a staged process. It entails running various commands across multiple NSO instances at different times.

The procedure described in this section is used with the rule-based built-in HA clusters. For HA Raft cluster instructions, refer to [Version Upgrade of Cluster Nodes](/guides/administration/management/high-availability) in the HA documentation.

The procedure is almost the same for a maintenance and major NSO upgrade. The difference is that a major upgrade requires the replacement of packages with recompiled ones. Still, a maintenance upgrade is often perceived as easier because there are fewer changes in the product.

The stages of the upgrade are:

1. First, enable read-only mode on the designated `primary`, and then on the `secondary` that is enabled for fail-over.
2. Take a full backup on all nodes.
3. If using a 3-node setup, disconnect the 3rd, non-fail-over `secondary` by disabling HA on this node.
4. Disconnect the HA pair by disabling HA on the designated `primary`, temporarily promoting the designated `secondary` to provide the read-only service (and advertise the shared virtual IP address if it is used).
5. Upgrade the designated `primary`.
6. Disable HA on the designated `secondary` node, to allow designated `primary` to become actual `primary` in the next step.
7. Activate HA on the designated `primary`, which will assume its assigned (`primary`) role to provide the full service (and again advertise the shared IP if used). However, at this point, the system is without HA.
8. Upgrade the designated `secondary` node.
9. Activate HA on the designated `secondary`, which will assume its assigned (`secondary`) role, connecting HA again.
10. Verify that HA is operational and has converged.
11. Upgrade the 3rd, non-fail-over `secondary` if it is used, and verify it successfully rejoins the HA cluster.

Enabling the read-only mode on both nodes is required to ensure the subsequent backup captures the full system state, as well as making sure the `failover-primary` does not start taking writes when it is promoted later on.

Disabling the non-fail-over `secondary` in a 3-node setup right after taking a backup is necessary when using the built-in HA rule-based algorithm (enabled by default in NSO 5.8 and later). Without it, the node might connect to the `failover-primary` when the failover happens, which disables read-only mode.

While not strictly necessary, explicitly promoting the designated `secondary` after disabling HA on the `primary` ensures a fast failover, avoiding the automatic reconnection attempts. If using a shared IP solution, such as the Tail-f HCC, this makes sure the shared VIP comes back up on the designated `secondary` as soon as possible. In addition, some older NSO versions do not reset the read-only mode upon disabling HA if they are not acting `primary`.

Another important thing to note is that all packages used in the upgrade must match the NSO release. If they do not, the upgrade will fail.

In the case of a major upgrade, you must recompile the packages for the new version. It is highly recommended that you use pre-compiled packages and do not compile them during this upgrade procedure since the compilation can prove nontrivial, and the production hosts may lack all the required (development) tooling. You should use a naming scheme to distinguish between packages compiled for different NSO versions. A good option is for package file names to start with the `ncs-MAJORVERSION-` prefix for a given major NSO version. This ensures multiple packages can co-exist in the `/opt/ncs/packages` folder, and the NSO version they can be used with becomes obvious.

The following is a transcript of a sample upgrade procedure, showing the commands for each step described above, in a 2-node HA setup, with nodes in their initial designated state. The procedure ensures that this is also the case in the end.

```xml
<switch to designated primary CLI>
admin@ncs# show high-availability status mode
high-availability status mode primary
admin@ncs# high-availability read-only mode true

<switch to designated secondary CLI>
admin@ncs# show high-availability status mode
high-availability status mode secondary
admin@ncs# high-availability read-only mode true

<switch to designated primary shell>
# ncs-backup

<switch to designated secondary shell>
# ncs-backup

<switch to designated primary CLI>
admin@ncs# high-availability disable

<switch to designated secondary CLI>
admin@ncs# high-availability be-primary

<switch to designated primary shell>
# <upgrade node>
# <set NCS_RELOAD_PACKAGES=true in `/etc/ncs/ncs.systemd.conf`>
# systemctl restart ncs
# <restore `/etc/ncs/ncs.systemd.conf`>

<switch to designated secondary CLI>
admin@ncs# high-availability disable

<switch to designated primary CLI>
admin@ncs# high-availability enable

<switch to designated secondary shell>
# <upgrade node>
# <set NCS_RELOAD_PACKAGES=true in `/etc/ncs/ncs.systemd.conf`>
# systemctl restart ncs
# <restore `/etc/ncs/ncs.systemd.conf`>

<switch to designated secondary CLI>
admin@ncs# high-availability enable
```

Scripting is a recommended way to upgrade the NSO version of an HA cluster. The following example script shows the required commands and can serve as a basis for your own customized upgrade script. In particular, the script requires a specific package naming convention above, and you may need to tailor it to your environment. In addition, it expects the new release version and the designated `primary` and `secondary` node addresses as the arguments. The recompiled packages are read from the `packages-MAJORVERSION/` directory.

For the below example script, we configured our `primary` and `secondary` nodes with their nominal roles that they assume at startup and when HA is enabled. Automatic failover is also enabled so that the `secondary` will assume the `primary` role if the `primary` node goes down.

{% code title="Configuration on Both Nodes" %}

```xml
<config xmlns="http://tail-f.com/ns/config/1.0">
  <high-availability xmlns="http://tail-f.com/ns/ncs">
    <ha-node>
      <id>n1</id>
      <nominal-role>primary</nominal-role>
    </ha-node>
    <ha-node>
      <id>n2</id>
      <nominal-role>secondary</nominal-role>
      <failover-primary>true</failover-primary>
    </ha-node>
    <settings>
      <enable-failover>true</enable-failover>
      <start-up>
        <assume-nominal-role>true</assume-nominal-role>
        <join-ha>true</join-ha>
      </start-up>
    </settings>
  </high-availability>
</config>
```

{% endcode %}

{% code title="Script for HA Major Upgrade (with Packages)" %}

```
#!/bin/bash
set -ex

vsn=$1
primary=$2
secondary=$3
installer_file=nso-${vsn}.linux.x86_64.installer.bin
pkg_vsn=$(echo $vsn | sed -e 's/^\([0-9]\+\.[0-9]\+\).*/\1/')
pkg_dir="packages-${pkg_vsn}"

function on_primary() { ssh $primary "$@" ; }
function on_secondary() { ssh $secondary "$@" ; }
function on_primary_cli() { ssh -p 2024 $primary "$@" ; }
function on_secondary_cli() { ssh -p 2024 $secondary "$@" ; }

function upgrade_nso() {
    target=$1
    scp $installer_file $target:
    ssh $target "sh $installer_file --system-install --non-interactive"
    ssh $target "rm -f /opt/ncs/current && \
                 ln -s /opt/ncs/ncs-${vsn} /opt/ncs/current"
}
function upgrade_packages() {
    target=$1
    do_pkgs=$(ls "${pkg_dir}/" || echo "")
    if [ -n "${do_pkgs}" ] ; then
        cd ${pkg_dir}
        ssh $target 'rm -rf /var/opt/ncs/packages/*'
        for p in ncs-${pkg_vsn}-*.gz; do
            scp $p $target:/opt/ncs/packages/
            ssh $target "ln -s /opt/ncs/packages/$p /var/opt/ncs/packages/"
        done
        cd -
    fi
}

# Perform the actual procedure

on_primary_cli 'request high-availability read-only mode true'
on_secondary_cli 'request high-availability read-only mode true'

on_primary 'ncs-backup'
on_secondary 'ncs-backup'

on_primary_cli 'request high-availability disable'
on_secondary_cli 'request high-availability be-primary'
upgrade_nso $primary
upgrade_packages $primary
on_primary `mv /etc/ncs/ncs.systemd.conf /etc/ncs/ncs.systemd.conf.bak'
on_primary 'echo "NCS_RELOAD_PACKAGES=true" > /etc/ncs/ncs.systemd.conf`
on_primary 'systemctl restart ncs'
on_primary `mv /etc/ncs/ncs.systemd.conf.bak /etc/ncs/ncs.systemd.conf'


on_secondary_cli 'request high-availability disable'
on_primary_cli 'request high-availability enable'
upgrade_nso $secondary
upgrade_packages $secondary
on_secondary `mv /etc/ncs/ncs.systemd.conf /etc/ncs/ncs.systemd.conf.bak'
on_secondary 'echo "NCS_RELOAD_PACKAGES=true" > /etc/ncs/ncs.systemd.conf`
on_secondary 'systemctl restart ncs'
on_secondary `mv /etc/ncs/ncs.systemd.conf.bak /etc/ncs/ncs.systemd.conf'

on_secondary_cli 'request high-availability enable'
```

{% endcode %}

Once the script is completed, it is paramount that you manually verify the outcome. First, check that the HA is enabled by using the `show high-availability` command on the CLI of each node. Then connect to the designated secondaries and ensure they have the complete latest copy of the data, synchronized from the primaries.

After the `primary` node is upgraded and restarted, the read-only mode is automatically disabled. This allows the `primary` node to start processing writes, minimizing downtime. However, there is no HA. Should the `primary` fail at this point or you need to revert to a pre-upgrade backup, the new writes would be lost. To avoid this scenario, again enable read-only mode on the `primary` after re-enabling HA. Then disable read-only mode only after successfully upgrading and reconnecting the `secondary`.

To further reduce time spent upgrading, you can customize the script to install the new NSO release and copy packages beforehand. Then, you only need to switch the symbolic links and restart the NSO process to use the new version.

You can use the same script for a maintenance upgrade as-is, with an empty `packages-MAJORVERSION` directory, or remove the `upgrade_packages` calls from the script.

Example implementations that use scripts to upgrade a 2- and 3-node setup using CLI/MAAPI or RESTCONF are available in the NSO example set under [examples.ncs/high-availability](https://github.com/NSO-developer/nso-examples/tree/6.7/high-availability).

We have been using a two-node HCC layer-2 upgrade reference example elsewhere in the documentation to demonstrate installing NSO and adding the initial configuration. The upgrade-l2 example referenced in [examples.ncs/high-availability/hcc](https://github.com/NSO-developer/nso-examples/tree/6.7/high-availability/hcc) implements shell and Python scripted steps to upgrade the NSO version using `ssh` to the Linux shell and the NSO CLI or Python Requests RESTCONF for accessing the `paris` and `london` nodes. See the example for details.

If you do not wish to automate the upgrade process, you will need to follow the instructions from [Single Instance Upgrade](#ug.admin_guide.manual_upgrade) and transfer the required files to each host manually. Additional information on HA is available in [High Availability](/guides/administration/management/high-availability). However, you can run the `high-availability` actions from the preceding script on the NSO CLI as-is. In this case, please take special care of which host you perform each command, as it can be easy to mix them up.

## Package Upgrade <a href="#d5e7083" id="d5e7083"></a>

Package upgrades are frequent and routine in development but require the same care as NSO upgrades in the production environment. The reason is that the new packages may contain an updated YANG model, resulting in a data upgrade process similar to a version upgrade. So, if a package is removed or uninstalled and a replacement is not provided, package-specific data, such as service instance data, will also be removed.

Another consideration are inbound requests on a live system. If it is likely such requests will arrive during an upgrade, consider using the in-service `packages reload optimistic` upgrade option. With it, you can also leverage the `backup` parameter to simplify the backup process. See [Package Management](/guides/administration/management/package-mgmt#open-transactions-during-upgrade) for the comparison of the two options.

In a single-node environment, the procedure is straightforward. Create a backup with the `ncs-backup` command and ensure the new package is compiled for the current NSO version and available under the `/opt/ncs/packages` directory. Then either manually rearrange the symbolic links in the `/var/opt/ncs/packages` directory or use the `software packages install` command in the NSO CLI. Finally, invoke the `packages reload` command. For example:

```bash
# ncs-backup
INFO  Backup /var/opt/ncs/backups/ncs-6.4@2024-04-21T10:34:42.backup.gz created
successfully
# ls /opt/ncs/packages
ncs-6.4-router-nc-1.0 ncs-6.4-router-nc-1.0.2
# ncs_cli -C
admin@ncs# software packages install package router-nc-1.0.2 replace-existing
installed ncs-6.4-router-nc-1.0.2
admin@ncs# packages reload

>>> System upgrade is starting.
>>> Sessions in configure mode must exit to operational mode.
>>> No configuration changes can be performed until upgrade has completed.
>>> System upgrade has completed successfully.
reload-result {
    package router-nc-1.0.2
    result true
}
```

On the other hand, upgrading packages in an HA setup is an error-prone process. Thus, NSO provides an action, `packages ha sync and-reload`to minimize such complexity. It is considerably faster and more efficient than upgrading one node at a time.

{% hint style="info" %}
If the only change in the packages is the addition of new NED packages, the `and-add` can replace `and-reload` command for an even more optimized and less intrusive update. See [Adding NED Packages](https://nso-docs.cisco.com/guides/administration/installation-and-deployment/pages/TpxOL8qGnNMkxkp0ZCeB#ug.package_mgmt.ned_package_add) for details.
{% endhint %}

The action executes on the `primary` node. First, it syncs the physical packages from the `primary` node to the `secondary` nodes as tar archive files, regardless if the packages were initially added as directories or tar archives. Then, it performs the upgrade on all nodes in one go. The action does not sync packages to or upgrade nodes with the `none` role.

The `packages ha sync` action only distributes new packages to the *secondary* nodes. If a package already exists on the `secondary` node, it will replace it with the one on the `primary` node. Deleting a package on the `primary` node will also delete it on the `secondary` node. Packages found in load paths under the installation destination (by default `/opt/ncs/current`) are not distributed as they belong to the system and should not differ between the `primary` and the `secondary` nodes.

It is crucial to ensure that the load path configuration is identical on both `primary` and `secondary` nodes. Otherwise, the distribution will not start, and the action output will contain detailed error information.

Using the `and-reload` parameter with the action starts the upgrade once packages are copied over. The action sets the `primary` node to read-only mode. After the upgrade is successfully completed, the node is set back to its previous mode.

If the parameter `and-reload` is also supplied with the `wait-commit-queue-empty` parameter, it will wait for the commit queue to become empty on the `primary` node and prevent other queue items from being added while the queue is being drained.

Using the `wait-commit-queue-empty` parameter is the recommended approach, as it minimizes the risk of the upgrade failing due to commit queue items still relying on the old schema.

{% code title="Package Upgrade Procedure" %}

```bash
primary@node1# software packages list
package {
  name dummy-1.0.tar.gz
  loaded
}
primary@node1# software packages fetch package-from-file \
$MY_PACKAGE_STORE/dummy-1.1.tar.gz
primary@node1# software packages install package dummy-1.1 replace-existing
primary@node1# packages ha sync and-reload { wait-commit-queue-empty }
```

{% endcode %}

The `packages ha sync and-reload` command has the following known limitations and side effects:

* The `primary` node is set to `read-only` mode before the upgrade starts, and it is set back to its previous mode if the upgrade is successfully upgraded. However, the node will always be in read-write mode if an error occurs during the upgrade. It is up to the user to set the node back to the desired mode by using the `high-availability read-only mode` command.
* As a best practice, you should create a backup of all nodes before upgrading. This action creates no backups, you must do that explicitly.

Example implementations that use scripts to upgrade a 2- and 3-node setup using CLI/MAAPI or RESTCONF are available in the NSO example set under [examples.ncs/high-availability](https://github.com/NSO-developer/nso-examples/tree/6.7/high-availability).

We have been using a two-node HCC layer 2 upgrade reference example elsewhere in the documentation to demonstrate installing NSO and adding the initial configuration. The `upgrade-l2` example referenced in [examples.ncs/high-availability/hcc](https://github.com/NSO-developer/nso-examples/tree/6.7/high-availability/hcc) implements shell and Python scripted steps to upgrade the `primary` `paris` package versions and sync the packages to the `secondary` `london` using `ssh` to the Linux shell and the NSO CLI or Python Requests RESTCONF for accessing the `paris` and `london` nodes. See the example for details.

In some cases, NSO may warn when the upgrade looks suspicious. For more information on this, see [Loading Packages](https://nso-docs.cisco.com/guides/administration/installation-and-deployment/pages/TpxOL8qGnNMkxkp0ZCeB#ug.package_mgmt.loading). If you understand the implications and are willing to risk losing data, use the `force` option with `packages reload` or set the `NCS_RELOAD_PACKAGES` environment variable to `force` when restarting NSO. It will force NSO to ignore warnings and proceed with the upgrade. In general, this is not recommended.

In addition, you must take special care of NED upgrades because services depend on them. For example, since NSO 5 introduced the CDM feature, which allows loading multiple versions of a NED, a major NED upgrade requires a procedure involving the `migrate` action.

When a NED contains nontrivial YANG model changes, that is called a major NED upgrade. The NED ID changes, and the first or second number in the NED version changes since NEDs follow the same versioning scheme as NSO. In this case, you cannot simply replace the package, as you would for a maintenance or patch NED release. Instead, you must load (add) the new NED package alongside the old one and perform the migration.

Migration uses the `/ncs:devices/device/migrate` action to change the ned-id of a single device or a group of devices. It does not affect the actual network device, except possibly reading from it. So, the migration does not have to be performed as part of the package upgrade procedure described above but can be done later, during normal operations. The details are described in [NED Migration](https://nso-docs.cisco.com/guides/administration/installation-and-deployment/pages/vGfK2qplvC1wu0670OgS#sec.ned_migration). Once the migration is complete, you can remove the old NED by performing another package upgrade, where you deinstall the old NED package. It can be done straight after the migration or as part of the next upgrade cycle.


# Management

Perform system management tasks on your NSO deployment.


# System Management

Perform NSO system management and configuration.

NSO consists of a number of modules and executable components. These executable components will be referred to by their command-line name, e.g. `ncs`, `ncs-netsim`, `ncs_cli`, etc. `ncs` is used to refer to the executable, the running daemon.

## Starting NSO <a href="#ug.sys_mgmt.starting_ncs" id="ug.sys_mgmt.starting_ncs"></a>

When NSO is started, it reads its configuration file and starts all subsystems configured to start (such as NETCONF, CLI, etc.).

By default, NSO starts in the background without an associated terminal. It is recommended to use a [System Install](/guides/administration/installation-and-deployment/system-install) when installing NSO for production deployment. This will create an `init` script that starts NSO when the system boots, and makes NSO start the service manager.

## Licensing NSO <a href="#ug.ncs_sys_mgmt.licensing" id="ug.ncs_sys_mgmt.licensing"></a>

NSO is licensed using Cisco Smart Licensing. To register your NSO instance, you need to enter a token from your Cisco Smart Software Manager account. For more information on this topic, see [Cisco Smart Licensing](/guides/administration/management/system-management/cisco-smart-licensing)*.*

## Configuring NSO

NSO is configured in the following two ways:

* Through its configuration file, `ncs.conf`.
* Through whatever data is configured at run-time over any northbound, for example, turning on trace using the CLI.

### `ncs.conf` File

The configuration file `ncs.conf` is read at startup and can be reloaded. Below is an example of the most common settings. It is included here as an example and should be self-explanatory. See [ncs.conf](/guides/resources/man/ncs.conf.5) in Manual Pages for more information. Important configuration settings are:

* `load-path`: where NSO should look for compiled YANG files, such as data models for NEDs or Services.
* `db-dir`: the directory on disk that CDB uses for its storage and any temporary files being used. It is also the directory where CDB searches for initialization files. This should be a local disk and not NFS mounted for performance reasons.
* Various log settings.
* AAA configuration.
* Rollback file directory and history length.
* Enabling north-bound interfaces like REST, and WebUI.
* Enabling of High-Availability mode.

The `ncs.conf` file is described in the [NSO Manual Pages](/guides/resources/man/ncs.conf.5). There is a large number of configuration items in `ncs.conf`, most of them have sane default values. The `ncs.conf` file is an XML file that must adhere to the `tailf-ncs-config.yang` model. If we start the NSO daemon directly, we must provide the path to the NCS configuration file as in:

```bash
# ncs -c /etc/ncs/ncs.conf
```

However, in a System Install, `systemd` is typically used to start NSO, and it will pass the appropriate options to the `ncs` command. Thus, NSO is started with the command:

```bash
# systemctl nso start
```

It is possible to edit the `ncs.conf` file, and then tell NSO to reload the edited file without restarting the daemon as in:

```bash
# ncs --reload
```

This command also tells NSO to close and reopen all log files, which makes it suitable to use from a system like `logrotate`.

In this section, some of the important configuration settings will be described and discussed.

### Exposed Interfaces

NSO allows access through a number of different interfaces, depending on the use case. In the default configuration, clients can access the system locally through an unauthenticated IPC socket (with the `ncs*` family of commands, using Local IPC over a Unix domain socket) and plain (non-HTTPS) HTTP web server (port 8080). Additionally, the system enables remote access through SSH-secured NETCONF and CLI (ports 2022 and 2024).

We strongly encourage you to review and customize the exposed interfaces to your needs in the `ncs.conf` configuration file. In particular, set:

* `/ncs-config/webui/match-host-name` to `true`.
* `/ncs-config/webui/server-name` to the hostname of the server.
* `/ncs-config/webui/server-alias` to additional domains or IP addresses used for serving HTTP(S).

If you decide to allow remote access to the web server, make sure you use TLS-secured HTTPS instead of HTTP and keep `match-host-name` enabled. Not doing so exposes you to security risks.

{% hint style="info" %}
Using `/ncs-config/webui/match-host-name = true` requires you to use the configured hostname when accessing the server. Web browsers do this automatically but you may need to set the `Host` header when performing requests programmatically using an IP address instead of the hostname.
{% endhint %}

To additionally secure IPC access, refer to [Restricting Access to the IPC Socket](/guides/administration/advanced-topics/ipc-connection#restricting-access-to-the-ipc-socket).

For more details on individual interfaces and their use, see [Northbound APIs](/guides/development/core-concepts/northbound-apis).

### Dynamic Configuration <a href="#d5e81" id="d5e81"></a>

Let's look at all the settings that can be manipulated through the NSO northbound interfaces. NSO itself has a number of built-in YANG modules. These YANG modules describe the structure that is stored in CDB. Whenever we change anything under, say `/devices/device`, it will change the CDB, but it will also change the configuration of NSO. We call this dynamic configuration since it can be changed at will through all northbound APIs.

We summarize the most relevant parts below:

```cli
ncs@ncs(config)#
Possible completions:
  aaa                        AAA management, users and groups
  cluster                    Cluster configuration
  devices                    Device communication settings
  java-vm                    Control of the NCS Java VM
  nacm                       Access control
  packages                   Installed packages
  python-vm                  Control of the NCS Python VM
  services                   Global settings for services, (the services themselves might be augmented somewhere else)
  session                    Global default CLI session parameters
  snmp                       Top-level container for SNMP related configuration and status objects.
  snmp-notification-receiver Configure reception of SNMP notifications
  software                   Software management
  ssh                        Global SSH connection configuration
```

#### **`tailf-ncs.yang` Module**

This is the most important YANG module that is used to control and configure NSO. The module can be found at: `$NCS_DIR/src/ncs/yang/tailf-ncs.yang` in the release. Everything in that module is available through the northbound APIs. The YANG module has descriptions for everything that can be configured.

`tailf-common-monitoring2.yang` and `tailf-ncs-monitoring2.yang` are two modules that are relevant to monitoring NSO.

### Built-in or External SSH Server <a href="#d5e95" id="d5e95"></a>

NSO has a built-in SSH server which makes it possible to SSH directly into the NSO daemon. Both the NSO northbound NETCONF agent and the CLI need SSH. To configure the built-in SSH server we need a directory with server SSH keys - it is specified via `/ncs-config/aaa/ssh-server-key-dir` in `ncs.conf`. We also need to enable `/ncs-config/netconf-north-bound/transport/ssh` and `/ncs-config/cli/ssh` in `ncs.conf`. In a System Install, `ncs.conf` is installed in the "config directory", by default `/etc/ncs`, with the SSH server keys in `/etc/ncs/ssh`.

### Run-time Configuration <a href="#ncsnwe.admin.runtime.cfg" id="ncsnwe.admin.runtime.cfg"></a>

There are also configuration parameters that are more related to how NSO behaves when talking to the devices. These reside in `devices global-settings`.

```cli
admin@ncs(config)# devices global-settings
Possible completions:
  backlog-auto-run               Auto-run the backlog at successful connection
  backlog-enabled                Backlog requests to non-responding devices
  commit-queue
  commit-retries                 Retry commits on transient errors
  connect-timeout                Timeout in seconds for new connections
  ned-settings                   Control which device capabilities NCS uses
  out-of-sync-commit-behaviour   Specifies the behaviour of a commit operation involving a device that is out of sync with NCS.
  read-timeout                   Timeout in seconds used when reading data
  report-multiple-errors         By default, when the NCS device manager commits data southbound and when there are errors, we only
                                 report the first error to the operator, this flag makes NCS report all errors reported by managed
                                 devices
  trace                          Trace the southbound communication to devices
  trace-dir                      The directory where trace files are stored
  write-timeout                  Timeout in seconds used when writing
  data
```

## User Management

Users are configured at the path `aaa authentication users`.

```cli
admin@ncs(config)# show full-configuration aaa authentication users user
aaa authentication users user admin
 uid        1000
 gid        1000
 password   $1$GNwimSPV$E82za8AaDxukAi8Ya8eSR.
 ssh_keydir /var/ncs/homes/admin/.ssh
 homedir    /var/ncs/homes/admin
!
aaa authentication users user oper
 uid        1000
 gid        1000
 password   $1$yOstEhXy$nYKOQgslCPyv9metoQALA.
 ssh_keydir /var/ncs/homes/oper/.ssh
 homedir    /var/ncs/homes/oper
!...
```

Access control, including group memberships, is managed using the NACM model (RFC 6536).

```cli
admin@ncs(config)# show full-configuration nacm
nacm write-default permit
nacm groups group admin
 user-name [ admin private ]
!
nacm groups group oper
 user-name [ oper public ]
!
nacm rule-list admin
 group [ admin ]
 rule any-access
  action permit
 !
!
nacm rule-list any-group
 group [ * ]
 rule tailf-aaa-authentication
  module-name       tailf-aaa
  path              /aaa/authentication/users/user[name='$USER']
  access-operations read,update
  action            permit
 !
```

### Adding a User

Adding a user includes the following steps:

1. Create the user: `admin@ncs(config)# aaa authentication users user <user-name>`.
2. Add the user to a NACM group: `admin@ncs(config)# nacm groups <group-name> admin user-name <user-name>`.
3. Verify/change access rules.

It is likely that the new user also needs access to work with device configuration. The mapping from NSO users and corresponding device authentication is configured in `authgroups`. So, the user needs to be added there as well.

```cli
admin@ncs(config)# show full-configuration devices authgroups
devices authgroups group default
 umap admin
  remote-name     admin
  remote-password $4$wIo7Yd068FRwhYYI0d4IDw==
 !
 umap oper
  remote-name     oper
  remote-password $4$zp4zerM68FRwhYYI0d4IDw==
 !
!
```

If the last step is forgotten, you will see the following error:

```cli
jim@ncs(config)# devices device c0 config ios:snmp-server community fee
jim@ncs(config-config)# commit
Aborted: Resource authgroup for jim doesn't exist
```

## Monitoring NSO <a href="#d5e7876" id="d5e7876"></a>

This section describes how to monitor NSO. See also [NSO Alarms](#nso-alarms).

Use the command `ncs --status` to get runtime information on NSO.

### NSO Status <a href="#d5e119" id="d5e119"></a>

Checking the overall status of NSO can be done using the shell:

```bash
$ ncs --status
```

Or, in the CLI:

```cli
ncs# show ncs-state
```

For details on the output see `$NCS_DIR/src/yang/tailf-common-monitoring2.yang`.

Below is an overview of the output:

<table data-header-hidden data-full-width="true"><thead><tr><th width="246" valign="top"></th><th valign="top"></th></tr></thead><tbody><tr><td valign="top"><code>daemon-status</code></td><td valign="top">You can see the NSO daemon mode, starting, phase0, phase1, started, stopping. The phase0 and phase1 modes are schema upgrade modes and will appear if you have upgraded any data models.</td></tr><tr><td valign="top"><code>version</code></td><td valign="top">The NSO version.</td></tr><tr><td valign="top"><code>smp</code></td><td valign="top">Number of threads used by the daemon.</td></tr><tr><td valign="top"><code>ha</code></td><td valign="top">The High-Availability mode of the NCS daemon will show up here: <code>secondary</code>, <code>primary</code>, <code>relay-secondary</code>.</td></tr><tr><td valign="top"><code>internal/callpoints</code></td><td valign="top"><p>The next section is callpoints. Make sure that any validation points, etc. are registered. (The <code>ncs-rfs-service-hook</code> is an obsolete callpoint, ignore this one).</p><ul><li><code>UNKNOWN</code> code tries to register a call-point that does not exist in a data model.</li><li><code>NOT-REGISTERED</code> a loaded data model has a call-point but no code has registered.</li></ul><p>Of special interest is of course the <code>servicepoints</code>. All your deployed service models should have a corresponding <code>service-point</code>. For example:</p><pre data-overflow="wrap"><code>servicepoints:
  id=l3vpn-servicepoint daemonId=10 daemonName=ncs-dp-6-l3vpn:L3VPN
  id=nsr-servicepoint daemonId=11 daemonName=ncs-dp-7-nsd:NSRService
  id=vm-esc-servicepoint daemonId=12 daemonName=ncs-dp-8-vm-manager-esc:ServiceforVMstarting
  id=vnf-catalogue-esc daemonId=13 daemonName=ncs-dp-9-vnf-catalogue-esc:ESCVNFCatalogueService
</code></pre></td></tr><tr><td valign="top"><code>internal/cdb</code></td><td valign="top">The <code>cdb</code> section is important. Look for any locks. This might be a sign that a developer has taken a CDB lock without releasing it. The subscriber section is also important. A design pattern is to register subscribers to wait for something to change in NSO and then trigger an action. Reactive FASTMAP is designed around that. Validate that all expected subscribers are OK.</td></tr><tr><td valign="top"><code>loaded-data-models</code></td><td valign="top">The next section shows all namespaces and YANG modules that are loaded. If you, for example, are missing a service model, make sure it is loaded.</td></tr><tr><td valign="top"><code>cli</code><em>,</em> <code>netconf</code><em>,</em> <code>rest</code><em>,</em> <code>snmp</code><em>,</em> <code>webui</code></td><td valign="top">All northbound agents like CLI, REST, NETCONF, SNMP, etc. are listed with their IP and port. So if you want to connect over REST, for example, you can see the port number here.</td></tr><tr><td valign="top"><code>patches</code></td><td valign="top">Lists any installed patches.</td></tr><tr><td valign="top"><code>upgrade-mode</code></td><td valign="top">If the node is in upgrade mode, it is not possible to get any information from the system over NETCONF. Existing CLI sessions can get system information.</td></tr></tbody></table>

It is also important to look at the packages that are loaded. This can be done in the CLI with:

```
admin> show packages
packages package cisco-asa
 package-version 3.4.0
 description     "NED package for Cisco ASA"
 ncs-min-version [ 3.2.2 3.3 3.4 4.0 ]
 directory       ./state/packages-in-use/1/cisco-asa
 component upgrade-ned-id
  upgrade java-class-name com.tailf.packages.ned.asa.UpgradeNedId
 component ASADp
  callback java-class-name [ com.tailf.packages.ned.asa.ASADp ]
 component cisco-asa
  ned cli ned-id  cisco-asa
  ned cli java-class-name com.tailf.packages.ned.asa.ASANedCli
  ned device vendor Cisco
```

### Monitoring the NSO Daemon <a href="#d5e174" id="d5e174"></a>

NSO runs the following processes:

* **The daemon**: `ncs.smp`: this is the NCS process running in the Erlang VM.
* **Java VM**: `com.tailf.ncs.NcsJVMLauncher`: service applications implemented in Java run in this VM. There are several options on how to start the Java VM, it can be monitored and started/restarted by NSO or by an external monitor. See the [ncs.conf(5)](/guides/resources/man/ncs.conf.5) Manual Page and the `java-vm` settings in the CLI.
* **Python VMs**: NSO packages can be implemented in Python. The individual packages can be configured to run a VM each or share a Python VM. Use the `show python-vm status current` to see current threads and `show python-vm status start` to see which threads were started at startup time.

### Logging <a href="#ug.ncs_sys_mgmt.logging" id="ug.ncs_sys_mgmt.logging"></a>

NSO has extensive logging functionality. Log settings are typically very different for a production system compared to a development system. Furthermore, the logging of the NSO daemon and the NSO Java VM/Python VM is controlled by different mechanisms. During development, we typically want to turn on the `developer-log`. The sample `ncs.conf` that comes with the NSO release has log settings suitable for development, while the `ncs.conf` created by a System Install are suitable for production deployment.

NSO logs in `/logs` in your running directory, (depends on your settings in `ncs.conf`). You might want the log files to be stored somewhere else. See man `ncs.conf` for details on how to configure the various logs. Below is a list of the most useful log files:

* `ncs.log` : NCS daemon log. See [Log Messages and Formats](/guides/administration/management/system-management/log-messages-and-formats). Can be configured to Syslog.
* `ncserr.log.1`*,* `ncserr.log.idx`*,* `ncserr.log.siz`: if the NSO daemon has a problem. this contains debug information relevant to support. The content can be displayed with `ncs --printlog ncserr.log`.
* `audit.log`: central audit log covering all northbound interfaces. See [Log Messages and Formats](/guides/administration/management/system-management/log-messages-and-formats). Can be configured to Syslog.
* `localhost:8080.access`: all HTTP requests to the daemon. This is an access log for the embedded Web server. This file adheres to the Common Log Format, as defined by Apache and others. This log is not enabled by default and is not rotated, i.e. use logrotate(8). Can be configured to Syslog.
* `devel.log`: developer-log is a debug log for troubleshooting user-written code. This log is enabled by default and is not rotated, i.e. use logrotate(8). This log shall be used in combination with the `java-vm` or `python-vm` logs. The user code logs in the VM logs and the corresponding library logs in `devel.log`. Disable this log in production systems. Can be configured to Syslog.\
  \
  You can manage this log and set its logging level in `ncs.conf`.

  ```xml
      <developer-log>
        <enabled>true</enabled>
        <file>
          <name>${NCS_LOG_DIR}/devel.log</name>
          <enabled>false</enabled>
        </file>
        <syslog>
          <enabled>true</enabled>
        </syslog>
      </developer-log>
      <developer-log-level>trace</developer-log-level>
  ```
* `ncs-java-vm`*.*`log`*,* `ncs-python-vm.log`: logger for code running in Java or Python VM, for example, service applications. Developers writing Java and Python code use this log (in combination with devel.log) for debugging. Both Java and Python log levels can be set from their respective VM settings in, for example, the CLI.

  ```cli
  admin@ncs(config)# python-vm logging level level-info
  admin@ncs(config)# java-vm java-logging logger com.tailf.maapi level level-info
  ```
* `netconf.log`*,* `snmp.log`: Log for northbound agents. Can be configured to Syslog.
* `rollbackNNNNN`: All NSO commits generate a corresponding rollback file. The maximum number of rollback files and file numbering can be configured in `ncs.conf`.
* `xpath.trace`: XPATH is used in many places, for example, XML templates. This log file shows the evaluation of all XPATH expressions and can be enabled in the `ncs.conf`.

  ```xml
      <xpathTraceLog>
        <enabled>true</enabled>
        <filename>${NCS_LOG_DIR}/xpath.trace</filename>
      </xpathTraceLog>
  ```

  To debug XPATH for a template, use the pipe target `debug` in the CLI instead.

  ```cli
  admin@ncs(config)# commit | debug template
  ```
* `ned-cisco-ios-xr-pe1.trace` (for example): if device trace is turned on a trace file will be created per device. The file location is not configured in `ncs.conf` but is configured when the device trace is turned on, for example in the CLI.

  ```cli
  admin@ncs(config)# devices device r0 trace pretty
  ```
* Progress trace log: When a transaction or action is applied, NSO emits specific progress events. These events can be displayed and recorded in a number of different ways, either in CLI with the pipe target `details` on a commit, or by writing it to a log file. You can read more about it in the [Progress Trace](/guides/development/advanced-development/progress-trace).
* Transaction error log: log for collecting information on failed transactions that lead to either a CDB boot error or a runtime transaction failure. The default is `false` (disabled). More information about the log is available in the Manual Pages under [Configuration Parameters](/guides/resources/man/ncs.conf.5#configuration-parameters) (see `logs/transaction-error-log`).
* Upgrade log: log containing information about CDB upgrade. The log is enabled by default and not rotated (i.e., use logrotate). With the NSO example set, the following examples populate the log in the `logs/upgrade.log` file: [examples.ncs/device-management/ned-yang-revision](https://github.com/NSO-developer/nso-examples/tree/6.7/device-management/ned-yang-revision), [examples.ncs/high-availability/upgrade-basic](https://github.com/NSO-developer/nso-examples/tree/6.7/high-availability/upgrade-basic), [examples.ncs/high-availability/upgrade-cluster](https://github.com/NSO-developer/nso-examples/tree/6.7/high-availability/upgrade-cluster), and [examples.ncs/service-management/upgrade-service](https://github.com/NSO-developer/nso-examples/tree/6.7/service-management/upgrade-service). More information about the log is available in the Manual Pages under [Configuration Parameters](/guides/resources/man/ncs.conf.5#configuration-parameters) (see `logs/upgrade-log)`.

### Syslog <a href="#d5e259" id="d5e259"></a>

NSO can syslog to a local Syslog. See `man ncs.conf` how to configure the Syslog settings. All Syslog messages are documented in Log Messages. The `ncs.conf` also lets you decide which of the logs should go into Syslog: `ncs.log, devel.log, netconf.log, snmp.log, audit.log, WebUI access log`. There is also a possibility to integrate with `rsyslog` to log the NCS, developer, audit, netconf, SNMP, and WebUI access logs to syslog with the facility set to daemon in `ncs.conf`. For reference, see the `upgrade-l2` example [examples.ncs/high-availability/hcc](https://github.com/NSO-developer/nso-examples/tree/6.7/high-availability/hcc) .

Below is an example of Syslog configuration:

```xml
    <syslog-config>
      <facility>daemon</facility>
    </syslog-config>

    <ncs-log>
      <enabled>true</enabled>
      <file>
        <name>./logs/ncs.log</name>
        <enabled>true</enabled>
      </file>
      <syslog>
        <enabled>true</enabled>
      </syslog>
    </ncs-log>
```

Log messages are described on the link below:

{% content-ref url="/pages/zu3mqwImBso718lveGTH" %}
[Log Messages and Formats](/guides/administration/management/system-management/log-messages-and-formats)
{% endcontent-ref %}

### NSO Alarms

NSO generates alarms for serious problems that must be remedied. Alarms are available over all the northbound interfaces and exist at the path `/alarms`. NSO alarms are managed as any other alarms by the general NSO Alarm Manager, see the specific section on the alarm manager in order to understand the general alarm mechanisms.

The NSO alarm manager also presents a northbound SNMP view, alarms can be retrieved as an alarm table, and alarm state changes are reported as SNMP Notifications. See the "NSO Northbound" documentation on how to configure the SNMP Agent.

This is also documented in the example [examples.ncs/northbound-interfaces/snmp-alarm](https://github.com/NSO-developer/nso-examples/tree/6.7/northbound-interfaces/snmp-alarm).

Alarms are described on the link below and information on alarm management is available under [Alarm Manager](/guides/operation-and-usage/operations/alarm-manager).

{% content-ref url="/pages/jNmCgNieU2LbzAWXSgC7" %}
[Alarm Types](/guides/administration/management/system-management/alarms)
{% endcontent-ref %}

### Tracing in NSO <a href="#d5e2587" id="d5e2587"></a>

Tracing enables observability across NSO operations by tagging requests with unique identifiers. NSO allows for using Trace Context on programmatic northbound interfaces, while the `label` commit parameter can be used to correlate events (e.g. notifications). These allow tracking of requests across service invocations, internal operations, and downstream device configurations.

#### **Trace Context**

NSO supports Trace Context based on the [W3C Trace Context specification](https://www.w3.org/TR/trace-context/), which is the recommended approach for distributed request tracing. This allows tracing information to flow between systems using standardized headers.

When using Trace Context:

* Trace information is carried in the `traceparent` and `tracestate` attributes.
* The trace ID is a UUID (RFC 4122) and is parsed from the `traceparent` attribute if Trace Context is provided with the request. Otherwise, a new one is generated.
* Trace Context is propagated automatically across NSO operations, including LSA setups and commit queues.
* It is supported across all major northbound protocols: NETCONF, RESTCONF, JSON-RPC, and MAAPI.
* Trace data appears in logs and trace files, enabling consistent request tracking across services and systems. For example, progress trace entries in text format include `trace-id=...` with the value of trace ID.

## Disaster Management <a href="#ug.ncs_sys_mgmt.disaster" id="ug.ncs_sys_mgmt.disaster"></a>

This section describes a number of disaster scenarios and recommends various actions to take in the different disaster variants.

### NSO Fails to Start <a href="#d5e2642" id="d5e2642"></a>

CDB keeps its data in a set of data files (`A` - configuration, `C` - schema, `O` - operational, and `S` - snapshot). If NSO is stopped, these four sets of files can be copied, and the copy is then a full backup of CDB.

Furthermore, if neither files exist in the configured CDB directory, CDB will attempt to initialize from all files in the CDB directory with the suffix `.xml`.

Thus, there exist two different ways to re-initiate CDB from a previously known good state, either from `.xml` files or from a CDB backup. The `.xml` files would typically be used to reinstall factory defaults whereas a CDB backup could be used in more complex scenarios.

If the `S` data files become inconsistent or have been removed, all commit queue items will be removed, and devices not yet processed out of sync. For such an event, appropriate alarms will be raised on the devices and any service instance that has unprocessed device changes will be set in the failed state.

When NSO starts and fails to initialize, the following exit codes can occur:

* Exit codes 1 and 19 mean that an internal error has occurred. A text message should be in the logs, or if the error occurred at startup before logging had been activated, on standard error (standard output if NSO was started with `--foreground --verbose`). Generally, the message will only be meaningful to the NSO developers, and an internal error should always be reported to support.
* Exit codes 2 and 3 are only used for the NCS control commands (see the section COMMUNICATING WITH NCS in the [ncs(1)](/guides/resources/man/ncs.1) in Manual Pages manual page) and mean that the command failed due to timeout. Code 2 is used when the initial connect to NSO didn't succeed within 5 seconds (or the `TryTime` if given), while code 3 means that the NSO daemon did not complete the command within the time given by the `--timeout` option.
* Exit code 10 means that one of the init files in the CDB directory was faulty in some way — further information in the log.
* Exit code 11 means that the CDB configuration was changed in an unsupported way. This will only happen when an existing database is detected, which was created with another configuration than the current in `ncs.conf`.
* Exit code 13 means that the schema change caused an upgrade, but for some reason, the upgrade failed. Details are in the log. The way to recover from this situation is either to correct the problem or to re-install the old schema (`fxs`) files.
* Exit code 14 means that the schema change caused an upgrade, but for some reason the upgrade failed, corrupting the database in the process. This is rare and usually caused by a bug. To recover, either start from an empty database with the new schema, or re-install the old schema files and apply a backup.
* Exit code 15 means that `A` or `C` data is corrupt in a non-recoverable way. Remove the files and re-start using a backup or init files.
* Exit code 16 means that CDB ran into an unrecoverable file error (such as running out of space on the device while performing journal compaction).
* Exit code 20 means that NSO failed to bind a socket.
* Exit code 21 means that some NSO configuration file is faulty. More information is in the logs.
* Exit code 22 indicates an NSO installation-related problem, e.g., that the user does not have read access to some library files, or that some file is missing.

If the NSO daemon starts normally, the exit code is 0.

If the AAA database is broken, NSO will start but with no authorization rules loaded. This means that all write access to the configuration is denied. The NSO CLI can be started with a flag `ncs_cli --noaaa` that will allow full unauthorized access to the configuration.

### NSO Failure After Startup <a href="#d5e2702" id="d5e2702"></a>

NSO attempts to handle all runtime problems without terminating, e.g., by restarting specific components. However, there are some cases where this is not possible, described below. When NSO is started the default way, i.e. as a daemon, the exit codes will of course not be available, but see the `--foreground` option in the [ncs(1)](/guides/resources/man/ncs.1) Manual Page.

* **Out of memory**: If NSO is unable to allocate memory, it will exit by calling abort(3). This will generate an exit code, as for reception of the SIGABRT signal - e.g. if NSO is started from a shell script, it will see 134, as the exit code (128 + the signal number).
* **Out of file descriptors for accept(2)**: If NSO fails to accept a TCP connection due to lack of file descriptors, it will log this and then exit with code 25. To avoid this problem, make sure that the process and system-wide file descriptor limits are set high enough, and if needed configure session limits in `ncs.conf`. The out-of-file descriptors issue may also manifest itself in that applications are no longer able to open new file descriptors.\
  \
  In many Linux systems, the default limit is 1024, but if we, for example, assume that there are four northbound interface ports, CLI, RESTCONF, SNMP, WebUI/JSON-RPC, or similar, plus a few hundred IPC ports, x 1024 == 5120. But one might as well use the next power of two, 8192, to be on the safe side.

  \
  Several application issues can contribute to consuming extra ports. In the scope of an NSO application that could, for example, be a script application that invokes CLI command or a callback daemon application that does not close the connection socket as it should.

  A commonly used command for changing the maximum number of open file descriptors is `ulimit -n [limit]`. Commands such as `netstat` and `lsof` can be useful to debug file descriptor-related issues.

### Transaction Commit Failure <a href="#d5e2722" id="d5e2722"></a>

When the system is updated, NSO executes a two-phase commit protocol towards the different participating databases including CDB. If a participant fails in the `commit()` phase although the participant succeeded in the preparation phase, the configuration is possibly in an inconsistent state.

When NSO considers the configuration to be in an inconsistent state, operations will continue. It is still possible to use NETCONF, the CLI, and all other northbound management agents. The CLI has a different prompt which reflects that the system is considered to be in an inconsistent state and also the Web UI shows this:

```
  -- WARNING ------------------------------------------------------
  Running db may be inconsistent. Enter private configuration mode and
  install a rollback configuration or load a saved configuration.
  ------------------------------------------------------------------
```

The MAAPI API has two interface functions that can be used to set and retrieve the consistency status, those are `maapi_set_running_db_status()` and `maapi_get_running_db_status()` corresponding. This API can thus be used to manually reset the consistency state. The only alternative to reset the state to a consistent state is by reloading the entire configuration.

## Backup and Restore

All parts of the NSO installation can be backed up and restored with standard file system backup procedures.

The most convenient way to do backup and restore is to use the `ncs-backup` command. In that case, the following procedure is used.

### Take a Backup <a href="#d5e7884" id="d5e7884"></a>

NSO Backup backs up the database (CDB) files, state files, config files, and rollback files from the installation directory. To take a complete backup (for disaster recovery), use:

```bash
# ncs-backup
```

The backup will be stored in the "run directory", by default `/var/opt/ncs`, as `/var/opt/ncs/backups/ncs-VERSION@DATETIME.backup`.

For more information on backup, refer to the [ncs-backup(1)](/guides/resources/man/ncs-backup.1) in Manual Pages.

### Restore a Backup <a href="#d5e7896" id="d5e7896"></a>

NSO Restore is performed if you would like to switch back to a previous good state or restore a backup.

It is always advisable to stop NSO before performing a restore.

1. First stop NSO if NSO is not stopped yet.

   ```
   systemctl stop ncs
   ```
2. Restore the backup.

   ```bash
   ncs-backup --restore
   ```

   \
   Select the backup to be restored from the available list of backups. The configuration and database with run-time state files are restored in `/etc/ncs` and `/var/opt/ncs`.
3. Start NSO.

   ```
   systemctl start ncs
   ```

## Rollbacks

NSO supports creating rollback files during the commit of a transaction that allows for rolling back the introduced changes. Rollbacks do not come without a cost and should be disabled if the functionality is not going to be used. Enabling rollbacks impacts both the time it takes to commit a change and requires sufficient storage on disk.

Rollback files contain a set of headers and the data required to restore the changes that were made when the rollback was created. One of the header fields includes a unique rollback ID that can be used to address the rollback file independent of the rollback numbering format.

The use of rollbacks from the supported APIs and the CLI is documented in the documentation for the given API.

### `ncs.conf` Config for Rollback <a href="#d5e5666" id="d5e5666"></a>

As described [earlier](#configuring-nso), NSO is configured through the configuration file, `ncs.conf`. In that file, we have the following items related to rollbacks:

* `/ncs-config/rollback/enabled`: If set to `true`, then a rollback file will be created whenever the running configuration is modified.
* `/ncs-config/rollback/directory`: Location where rollback files will be created.
* `/ncs-config/rollback/history-size`: The number of old rollback files to save.

## Troubleshooting <a href="#ug.sys_mgmt.tshoot" id="ug.sys_mgmt.tshoot"></a>

New users can face problems when they start to use NSO. If you face an issue, reach out to our support team regardless if your problem is listed here or not.

{% hint style="success" %}
A useful tool in this regard is the `ncs-collect-tech-report` tool, which is the Bash script that comes with the product. It collects all log files, CDB backup, and several debug dumps as a TAR file. Note that it works only with a System Install.

```bash
root@linux:/# ncs-collect-tech-report --full 
```

{% endhint %}

Some noteworthy issues are covered here.

<details>

<summary>Installation Problems: Error Messages During Installation</summary>

* **Error**

  ```
  tar: Skipping to next header
  gzip: stdin: invalid compressed data--format violated
  ```
* **Impact**\
  The resulting installation is incomplete.
* **Cause**\
  This happens if the installation program has been damaged, most likely because it has been downloaded in ASCII mode.
* **Resolution**\
  Remove the installation directory. Download a new copy of NSO from our servers. Make sure you use binary transfer mode every step of the way.

</details>

<details>

<summary>Problem Starting NSO: NSO Terminating with GLIBC Error</summary>

* **Error**

  ```
  Internal error: Open failed: /lib/tls/libc.so.6: version
  `GLIBC_2.3.4' not found (required by
  .../lib/ncs/priv/util/syst_drv.so)
  ```
* **Impact**\
  NSO terminates immediately with a message similar to the one above.
* **Cause**\
  This happens if you are running on a very old Linux version. The GNU libc (GLIBC) version is older than 2.3.4, which was released in 2004.
* **Resolution**\
  Use a newer Linux system, or upgrade the GLIBC installation.

</details>

<details>

<summary>Problem in Running Examples: The <code>netconf-console</code> Program Fails</summary>

* **Error**\
  You must install the Python SSH implementation Paramiko in order to use SSH.
* **Impact**\
  Sending NETCONF commands and queries with `netconf-console` fails, while it works using `netconf-console-tcp`.
* **Cause**\
  The `netconf-console` command is implemented using the Python programming language. It depends on the Python SSHv2 implementation Paramiko. Since you are seeing this message, your operating system doesn't have the Python module Paramiko installed.
* **Resolution**\
  Install Paramiko using the instructions from [https://www.paramiko.org](https://www.paramiko.org/).\
  \
  When properly installed, you will be able to import the Paramiko module without error messages.

  ```bash
  $ python
  ...
  >>> import paramiko
  >>>
  ```

  \
  Exit the Python interpreter with Ctrl+D.
* **Workaround**\
  A workaround is to use `netconf-console-tcp`. It uses TCP instead of SSH and doesn't require Paramiko. Note that TCP traffic is not encrypted.

</details>

<details>

<summary>Problems Using and Developing Services</summary>

If you encounter issues while loading service packages, creating service instances, or developing service models, templates, and code, you can consult the Troubleshooting section in [Implementing Services](/guides/development/core-concepts/implementing-services).

</details>

### General Troubleshooting Strategies <a href="#d5e2778" id="d5e2778"></a>

If you have trouble starting or running NSO, examples, or the clients you write, here are some troubleshooting tips.

<details>

<summary>Transcript</summary>

When contacting support, it often helps the support engineer to understand what you are trying to achieve if you copy-paste the commands, responses, and shell scripts that you used to trigger the problem, together with any CLI outputs and logs produced by NSO.

</details>

<details>

<summary>Source ENV Variables</summary>

If you have problems executing `ncs` commands, make sure you source the `ncsrc` script in your NSO directory (your path may be different than the one in the example if you are using a local install), which sets the required environmental variables.

```bash
$ source /etc/profile.d/ncs.sh
```

</details>

<details>

<summary>Log Files</summary>

To find out what NSO is/was doing, browsing NSO log files is often helpful. In the examples, they are called `devel.log`, `ncs.log`, `audit.log`. If you are working with your own system, make sure that the log files are enabled in `ncs.conf`. They are already enabled in all the examples. You can read more about how to enable and inspect various logs in the [Logging](#ug.ncs_sys_mgmt.logging) section.

</details>

<details>

<summary>Verify HW Resources</summary>

Both high CPU utilization and a lack of memory can negatively affect the performance of NSO. You can use commands such as `top` to examine resource utilization, and `free -mh` to see the amount of free and consumed memory. A common symptom of a lack of memory is NSO or Java-VM restarting. A sufficient amount of disk space is also required for CDB persistence and logs, so you can also check disk space with `df -h` command. In case there is enough space on the disk and you still encounter ENOSPC errors, check the inode usage with `df -i` command.

</details>

<details>

<summary>Status</summary>

NSO will give you a comprehensive status of daemon status, YANG modules, loaded packages, MIBs, active user sessions, CDB locks, and more if you run:

```bash
$ ncs --status
```

NSO status information is also available as operational data under `/ncs-state`.

</details>

<details>

<summary>Check Data Provider</summary>

If you are implementing a data provider (for operational or configuration data), you can verify that it works for all possible data items using:

```bash
$ ncs --check-callbacks
```

</details>

<details>

<summary>Debug Dump</summary>

If you suspect you have experienced a bug in NSO, or NSO told you so, you can give Support a debug dump to help us diagnose the problem. It contains a lot of status information (including a full `ncs --status report`) and some internal state information. This information is only readable and comprehensible to the NSO development team, so send the dump to your support contact. A debug dump is created using:

```bash
$ ncs --debug-dump mydump1
```

Just as in CSI on TV, the information must be collected as soon as possible after the event. Many interesting traces will wash away with time, or stay undetected if there are lots of irrelevant facts in the dump.

If NSO gets stuck while terminating, it can optionally create a debug dump after being stuck for 60 seconds. To enable this mechanism, set the environment variable `$NCS_DEBUG_DUMP_NAME` to a filename of your choice.

</details>

<details>

<summary>Error Log</summary>

Another thing you can do in case you suspect that you have experienced a bug in NSO is to collect the error log. The logged information is only readable and comprehensible to the NSO development team, so send the log to your support contact. The log actually consists of a number of files called `ncserr.log.*` - make sure to provide them all.

</details>

<details>

<summary>System Dump</summary>

If NSO aborts due to failure to allocate memory (see [Disaster Management](#ug.ncs_sys_mgmt.disaster)), and you believe that this is due to a memory leak in NSO, creating one or more debug dumps as described above (before NSO aborts) will produce the most useful information for Support. If this is not possible, NSO will produce a system dump by default before aborting, unless `DISABLE_NCS_DUMP` is set.

The default system dump file name is `ncs_crash.dump` and it could be changed by setting the environment variable `$NCS_DUMP` before starting NSO. The dumped information is only comprehensible to the NSO development team, so send the dump to your support contact.

</details>

<details>

<summary>System Call Trace</summary>

To catch certain types of problems, especially relating to system start and configuration, the operating system's system call trace can be invaluable. This tool is called `strace`/`ktrace`/`truss`. Please send the result to your support contact for a diagnosis.

By running the instructions below.

Linux:

```bash
# strace -f -o mylog1.strace -s 1024 ncs ...
```

BSD:

```bash
# ktrace -ad -f mylog1.ktrace ncs ...
# kdump -f mylog1.ktrace > mylog1.kdump
```

Solaris:

```bash
# truss -f -o mylog1.truss ncs ...
```

</details>


# Cisco Smart Licensing

Manage purchase and licensing of Cisco software.

[Cisco Smart Licensing](https://www.cisco.com/web/ordering/smart-software-licensing/index.html) is a cloud-based approach to licensing, and it simplifies the purchase, deployment, and management of Cisco software assets. Entitlements are purchased through a Cisco account via Cisco Commerce Workspace (CCW) and are immediately deposited into a Smart Account for usage. This eliminates the need to install license files on every device. Products that are smart-enabled communicate directly to Cisco to report consumption.

Cisco Smart Software Manager (CSSM) enables the management of software licenses and Smart Account from a single portal. The interface allows you to activate your product, manage entitlements, and renew and upgrade software.

A functioning Smart Account is required to complete the registration process. For detailed information about CSSM, see [Cisco Smart Software Manager](https://www.cisco.com/c/en/us/buy/smart-accounts/software-manager.html).

## Smart Accounts and Virtual Accounts <a href="#d5e2873" id="d5e2873"></a>

A virtual account exists as a sub-account within the Smart Account. Virtual accounts are a customer-defined structure based on organizational layout, business function, geography, or any defined hierarchy. They are created and maintained by the Smart Account administrator(s).

Visit [Cisco Cisco Software Central](https://software.cisco.com/) to learn about how to create and manage Smart Accounts.

### Request a Smart Account <a href="#d5e2878" id="d5e2878"></a>

The creation of a new Smart Account is a one-time event, and subsequent management of users is a capability provided through the tool. To request a Smart Account, visit [Cisco Cisco Software Central](https://software.cisco.com/) and take the following steps:

1. After logging in, select **Request a Smart Account** in the Administration section.

   <div data-with-frame="true"><figure><img src="/files/LkOyQq6HlXzJtCi46lUJ" alt="" width="375"><figcaption></figcaption></figure></div>
2. Select the type of Smart Account to create. There are two options: (a) Individual Smart Account requiring agreement to represent your company. By creating this Smart Account, you agree to authorization to create and manage product and service entitlements, users, and roles on behalf of your organization. (b) Create the account on behalf of someone else.

   <div data-with-frame="true"><figure><img src="/files/l39vnQ3KuiJfDJ5bnv4n" alt="" width="563"><figcaption></figcaption></figure></div>
3. Provide the required domain identifier and the preferred account name.

   <div data-with-frame="true"><figure><img src="/files/vFik9NTJOnXS4XjrPz9t" alt="" width="563"><figcaption></figcaption></figure></div>
4. The account request will be pending approval of the Account Domain Identifier. A subsequent email will be sent to the requester to complete the setup process.

   <div data-with-frame="true"><figure><img src="/files/Q2chXIX7lKiGLlLxoDjk" alt="" width="563"><figcaption></figcaption></figure></div>

### Adding Users to a Smart Account <a href="#d5e2905" id="d5e2905"></a>

Smart Account user management is available in the **Administration** section of [Cisco Cisco Software Central](https://software.cisco.com/). Take the following steps to add a new user to a Smart Account:

1. After logging in Select **Manage Smart Account** in the **Administration** section.

   <div data-with-frame="true"><figure><img src="/files/LkOyQq6HlXzJtCi46lUJ" alt="" width="375"><figcaption></figcaption></figure></div>
2. Choose the **Users** tab.

   <div data-with-frame="true"><figure><img src="/files/Adp7PEwSvJCT7QgLKxwC" alt="" width="375"><figcaption></figcaption></figure></div>
3. Select **New User** and follow the instructions in the wizard to add a new user.

   <div data-with-frame="true"><figure><img src="/files/iVR3xD3TV5uePygYbVJn" alt="" width="563"><figcaption></figcaption></figure></div>

### Create a License Registration Token <a href="#d5e2927" id="d5e2927"></a>

1. To create a new token, log into CSSM and select the appropriate Virtual Account.

   <div data-with-frame="true"><figure><img src="/files/riMwlTxg48VeJlVhyY7l" alt="" width="563"><figcaption></figcaption></figure></div>
2. Click on the **Smart Licenses** link to enter CSSM.

   <div data-with-frame="true"><figure><img src="/files/3vMhuJgeyWGXEawr32Js" alt="" width="563"><figcaption></figcaption></figure></div>
3. In CSSM click on **New Token**.

   <div data-with-frame="true"><figure><img src="/files/SvKsun6fwdMlHh7JwLhl" alt="" width="563"><figcaption></figcaption></figure></div>
4. Follow the dialog to provide a description, expiration, and export compliance applicability before accepting the terms and responsibilities. Click on **Create Token** to continue.

   <div data-with-frame="true"><figure><img src="/files/MAVdJszoXYCiV4wQN0F2" alt="" width="563"><figcaption></figcaption></figure></div>
5. Click on the new token.

   <div data-with-frame="true"><figure><img src="/files/fhePlfFqXL7WX7DMDMDL" alt="" width="563"><figcaption></figcaption></figure></div>
6. Copy the token from the dialogue window into your clipboard.

   <div data-with-frame="true"><figure><img src="/files/IQQAGZJwvof2pBUQLJUm" alt=""><figcaption></figcaption></figure></div>
7. Go to the NSO CLI and provide the token to the `license smart register idtoken` command:

   ```cli
   admin@ncs# license smart register idtoken YzY2YjFlOTYtOWYzZi00MDg1...
   Registration process in progress.
   Use the 'show license status' command to check the progress and result.
   ```

### Notes on Configuring Smart Licensing <a href="#d5e2966" id="d5e2966"></a>

* If `ncs.conf` contains configuration for any of java-executable, java-options, override-url/url, or proxy/url under the configure path `/ncs-config/smart-license/smart-agent/` any corresponding configuration done via the CLI is ignored.
* The smart licensing component of NSO runs its own Java virtual machine. Usually, the default Java options are sufficient:

  ```yang
            leaf java-options {
            tailf:info "Smart licensing Java VM start options";
            type string;
            default "-Xmx64M -Xms16M
            -Djava.security.egd=file:/dev/./urandom";
            description
            "Options which NCS will use when starting
            the Java VM.";}
  ```

  \
  If you, for some reason, need to modify the Java options, remember to include the default values as found in the YANG model.

### Validation and Troubleshooting <a href="#d5e2975" id="d5e2975"></a>

#### Available `show` and `debug` Commands <a href="#d5e2977" id="d5e2977"></a>

* `show license all`: Displays all information.
* `show license status`: Displays status information.
* `show license summary`: Displays summary.
* `show license tech`: Displays license tech support information.
* `show license usage`: Displays usage information.
* `debug smart_lic all`: All available Smart Licensing debug flags.


# Log Messages and Formats

<details>

<summary>AAA_LOAD_FAIL</summary>

`AAA_LOAD_FAIL`

* **Severity**\
  `CRIT`
* **Description**\
  Failed to load the AAA data, it could be that an external db is misbehaving or AAA is mounted/populated badly
* **Format String**\
  `"Failed to load AAA: ~s"`

</details>

<details>

<summary>ABORT_CAND_COMMIT</summary>

`ABORT_CAND_COMMIT`

* **Severity**\
  `INFO`
* **Description**\
  Aborting candidate commit, request from user, reverting configuration.
* **Format String**\
  `"Aborting candidate commit, request from user, reverting configuration."`

</details>

<details>

<summary>ABORT_CAND_COMMIT_REBOOT</summary>

`ABORT_CAND_COMMIT_REBOOT`

* **Severity**\
  `INFO`
* **Description**\
  ConfD restarted while having a ongoing candidate commit timer, reverting configuration.
* **Format String**\
  `"ConfD restarted while having a ongoing candidate commit timer, reverting configuration."`

</details>

<details>

<summary>ABORT_CAND_COMMIT_TERM</summary>

`ABORT_CAND_COMMIT_TERM`

* **Severity**\
  `INFO`
* **Description**\
  Candidate commit session terminated, reverting configuration.
* **Format String**\
  `"Candidate commit session terminated, reverting configuration."`

</details>

<details>

<summary>ABORT_CAND_COMMIT_TIMER</summary>

`ABORT_CAND_COMMIT_TIMER`

* **Severity**\
  `INFO`
* **Description**\
  Candidate commit timer expired, reverting configuration.
* **Format String**\
  `"Candidate commit timer expired, reverting configuration."`

</details>

<details>

<summary>ACCEPT_FATAL</summary>

`ACCEPT_FATAL`

* **Severity**\
  `CRIT`
* **Description**\
  ConfD encountered an OS-specific error indicating that networking support is unavailable.
* **Format String**\
  `"Fatal error for accept() - ~s"`

</details>

<details>

<summary>ACCEPT_FDLIMIT</summary>

`ACCEPT_FDLIMIT`

* **Severity**\
  `CRIT`
* **Description**\
  ConfD failed to accept a connection due to reaching the process or system-wide file descriptor limit.
* **Format String**\
  `"Out of file descriptors for accept() - ~s limit reached"`

</details>

<details>

<summary>AUTH_LOGIN_FAIL</summary>

`AUTH_LOGIN_FAIL`

* **Severity**\
  `INFO`
* **Description**\
  A user failed to log in to ConfD.
* **Format String**\
  `"login failed via ~s from ~s with ~s: ~s"`

</details>

<details>

<summary>AUTH_LOGIN_SUCCESS</summary>

`AUTH_LOGIN_SUCCESS`

* **Severity**\
  `INFO`
* **Description**\
  A user logged into ConfD.
* **Format String**\
  `"logged in to ~s via ~s from ~s with ~s using ~s authentication"`

</details>

<details>

<summary>AUTH_LOGOUT</summary>

`AUTH_LOGOUT`

* **Severity**\
  `INFO`
* **Description**\
  A user was logged out from ConfD.
* **Format String**\
  `"logged out <~s> user"`

</details>

<details>

<summary>BADCONFIG</summary>

`BADCONFIG`

* **Severity**\
  `CRIT`
* **Description**\
  confd.conf contained bad data.
* **Format String**\
  `"Bad configuration: ~s:~s: ~s"`

</details>

<details>

<summary>BAD_DEPENDENCY</summary>

`BAD_DEPENDENCY`

* **Severity**\
  `ERR`
* **Description**\
  A dependency was not found
* **Format String**\
  `"The dependency node '~s' for node '~s' in module '~s' does not exist"`

</details>

<details>

<summary>BAD_NS_HASH</summary>

`BAD_NS_HASH`

* **Severity**\
  `CRIT`
* **Description**\
  Two namespaces have the same hash value. The namespace hashvalue MUST be unique. You can pass the flag --nshash to confdc when linking the .xso files to force another value for the namespace hash.
* **Format String**\
  `"~s"`

</details>

<details>

<summary>BIND_ERR</summary>

`BIND_ERR`

* **Severity**\
  `CRIT`
* **Description**\
  ConfD failed to bind to one of the internally used listen sockets.
* **Format String**\
  `"~s"`

</details>

<details>

<summary>BRIDGE_DIED</summary>

`BRIDGE_DIED`

* **Severity**\
  `ERR`
* **Description**\
  ConfD is configured to start the confd\_aaa\_bridge and the C program died.
* **Format String**\
  `"confd_aaa_bridge died - ~s"`

</details>

<details>

<summary>CAND_COMMIT_ROLLBACK_DONE</summary>

`CAND_COMMIT_ROLLBACK_DONE`

* **Severity**\
  `INFO`
* **Description**\
  Candidate commit rollback done
* **Format String**\
  `"Candidate commit rollback done"`

</details>

<details>

<summary>CAND_COMMIT_ROLLBACK_FAILURE</summary>

`CAND_COMMIT_ROLLBACK_FAILURE`

* **Severity**\
  `ERR`
* **Description**\
  Failed to rollback candidate commit
* **Format String**\
  `"Failed to rollback candidate commit due to: ~s"`

</details>

<details>

<summary>CANDIDATE_BAD_FILE_FORMAT</summary>

`CANDIDATE_BAD_FILE_FORMAT`

* **Severity**\
  `WARNING`
* **Description**\
  The candidate database file has a bad format. The candidate database is reset to the empty database.
* **Format String**\
  `"Bad format found in candidate db file ~s; resetting candidate"`

</details>

<details>

<summary>CANDIDATE_CORRUPT_FILE</summary>

`CANDIDATE_CORRUPT_FILE`

* **Severity**\
  `WARNING`
* **Description**\
  The candidate database file is corrupt and cannot be read. The candidate database is reset to the empty database.
* **Format String**\
  `"Corrupt candidate db file ~s; resetting candidate"`

</details>

<details>

<summary>CDB_BACKUP</summary>

`CDB_BACKUP`

* **Severity**\
  `INFO`
* **Description**\
  CDB data backed up after migration to a new storage backend.
* **Format String**\
  `"CDB: ~s backed up to ~s"`

</details>

<details>

<summary>CDB_BOOT_ERR</summary>

`CDB_BOOT_ERR`

* **Severity**\
  `CRIT`
* **Description**\
  CDB failed to start. Some grave error in the cdb data files prevented CDB from starting - a recovery from backup is necessary.
* **Format String**\
  `"CDB boot error: ~s"`

</details>

<details>

<summary>CDB_CLIENT_TIMEOUT</summary>

`CDB_CLIENT_TIMEOUT`

* **Severity**\
  `ERR`
* **Description**\
  A CDB client failed to answer within the timeout period. The client will be disconnected.
* **Format String**\
  `"CDB client (~s) timed out, waiting for ~s"`

</details>

<details>

<summary>CDB_CONFIG_LOST</summary>

`CDB_CONFIG_LOST`

* **Severity**\
  `INFO`
* **Description**\
  CDB found it's data files but no schema file. CDB recovers by starting from an empty database.
* **Format String**\
  `"CDB: lost config, deleting DB"`

</details>

<details>

<summary>CDB_DB_LOST</summary>

`CDB_DB_LOST`

* **Severity**\
  `INFO`
* **Description**\
  CDB found it's data schema file but not it's data file. CDB recovers by starting from an empty database.
* **Format String**\
  `"CDB: lost DB, deleting old config"`

</details>

<details>

<summary>CDB_FATAL_ERROR</summary>

`CDB_FATAL_ERROR`

* **Severity**\
  `CRIT`
* **Description**\
  CDB encounterad an unrecoverable error
* **Format String**\
  `"fatal error in CDB: ~s"`

</details>

<details>

<summary>CDB_INIT_LOAD</summary>

`CDB_INIT_LOAD`

* **Severity**\
  `INFO`
* **Description**\
  CDB is processing an initialization file.
* **Format String**\
  `"CDB load: processing file: ~s"`

</details>

<details>

<summary>CDB_MIGRATE</summary>

`CDB_MIGRATE`

* **Severity**\
  `INFO`
* **Description**\
  CDB data migration to a new storage backend.
* **Format String**\
  `"CDB: migrate ~s to ~s"`

</details>

<details>

<summary>CDB_OFFLOAD</summary>

`CDB_OFFLOAD`

* **Severity**\
  `DEBUG`
* **Description**\
  CDB data offload started.
* **Format String**\
  `"CDB: offload ~s from memory"`

</details>

<details>

<summary>CDB_OP_INIT</summary>

`CDB_OP_INIT`

* **Severity**\
  `ERR`
* **Description**\
  The operational DB was deleted and re-initialized (because of upgrade or corrupt file)
* **Format String**\
  `"CDB: Operational DB re-initialized"`

</details>

<details>

<summary>CDB_STALE_BACKUP</summary>

`CDB_STALE_BACKUP`

* **Severity**\
  `INFO`
* **Description**\
  CDB backup data left on disk after migration that can be removed to free up disk space.
* **Format String**\
  `"CDB: ~s backup file(s) occupying ~sMiB, remove to free up disk space: ~s"`

</details>

<details>

<summary>CDB_UPGRADE_FAILED</summary>

`CDB_UPGRADE_FAILED`

* **Severity**\
  `ERR`
* **Description**\
  Automatic CDB upgrade failed. This means that the data model has been changed in a non-supported way.
* **Format String**\
  `"CDB: Upgrade failed: ~s"`

</details>

<details>

<summary>CGI_REQUEST</summary>

`CGI_REQUEST`

* **Severity**\
  `INFO`
* **Description**\
  CGI script requested.
* **Format String**\
  `"CGI: '~s' script with method ~s"`

</details>

<details>

<summary>CHANGE_USER</summary>

`CHANGE_USER`

* **Severity**\
  `INFO`
* **Description**\
  A NETCONF request to change user for authorization was succesfully done.
* **Format String**\
  `"changed user to ~s, groups ~s"`

</details>

<details>

<summary>CLI_CMD_ABORTED</summary>

`CLI_CMD_ABORTED`

* **Severity**\
  `INFO`
* **Description**\
  CLI command aborted.
* **Format String**\
  `"CLI aborted"`

</details>

<details>

<summary>CLI_CMD_DONE</summary>

`CLI_CMD_DONE`

* **Severity**\
  `INFO`
* **Description**\
  CLI command finished successfully.
* **Format String**\
  `"CLI done"`

</details>

<details>

<summary>CLI_CMD</summary>

`CLI_CMD`

* **Severity**\
  `INFO`
* **Description**\
  User executed a CLI command.
* **Format String**\
  `"CLI '~s'"`

</details>

<details>

<summary>CLI_DENIED</summary>

`CLI_DENIED`

* **Severity**\
  `INFO`
* **Description**\
  User was denied to execute a CLI command due to permissions.
* **Format String**\
  `"CLI denied '~s'"`

</details>

<details>

<summary>COMMIT_INFO</summary>

`COMMIT_INFO`

* **Severity**\
  `INFO`
* **Description**\
  Information about configuration changes committed to the running data store.
* **Format String**\
  `"commit ~s"`

</details>

<details>

<summary>COMMIT_QUEUE_CORRUPT</summary>

`COMMIT_QUEUE_CORRUPT`

* **Severity**\
  `ERR`
* **Description**\
  Failed to load commit queue. ConfD recovers by starting from an empty commit queue.
* **Format String**\
  `"Resetting commit queue due do inconsistent or corrupt data."`

</details>

<details>

<summary>CONFIG_CHANGE</summary>

`CONFIG_CHANGE`

* **Severity**\
  `INFO`
* **Description**\
  A change to ConfD configuration has taken place, e.g., by a reload of the configuration file
* **Format String**\
  `"ConfD configuration change: ~s"`

</details>

<details>

<summary>CONFIG_DEPRECATED</summary>

`CONFIG_DEPRECATED`

* **Severity**\
  `WARNING`
* **Description**\
  confd.conf contains a deprecated value
* **Format String**\
  `"Config value is deprecated: ~s"`

</details>

<details>

<summary>CONFIG_OBSOLETE</summary>

`CONFIG_OBSOLETE`

* **Severity**\
  `WARNING`
* **Description**\
  confd.conf contains an obsolete value
* **Format String**\
  `"Config value is obsolete: ~s"`

</details>

<details>

<summary>CONFIG_TRANSACTION_LIMIT</summary>

`CONFIG_TRANSACTION_LIMIT`

* **Severity**\
  `INFO`
* **Description**\
  Configuration transaction limit reached, rejected new transaction request.
* **Format String**\
  `"Configuration transaction limit of type '~s' reached, rejected new transaction request"`

</details>

<details>

<summary>CONSULT_FILE</summary>

`CONSULT_FILE`

* **Severity**\
  `INFO`
* **Description**\
  ConfD is reading its configuration file.
* **Format String**\
  `"Consulting daemon configuration file ~s"`

</details>

<details>

<summary>CRYPTO_KEYS_FAILED_LOADING</summary>

`CRYPTO_KEYS_FAILED_LOADING`

* **Severity**\
  `INFO`
* **Description**\
  Crypto keys failed to load because the old active generation is missing in the new configuration.
* **Format String**\
  `"Cannot reload crypto keys since the old active generation is missing in the new list of keys."`

</details>

<details>

<summary>DAEMON_DIED</summary>

`DAEMON_DIED`

* **Severity**\
  `CRIT`
* **Description**\
  An external database daemon closed its control socket.
* **Format String**\
  `"Daemon ~s died"`

</details>

<details>

<summary>DAEMON_TIMEOUT</summary>

`DAEMON_TIMEOUT`

* **Severity**\
  `CRIT`
* **Description**\
  An external database daemon did not respond to a query.
* **Format String**\
  `"Daemon ~s timed out"`

</details>

<details>

<summary>DEVEL_AAA</summary>

`DEVEL_AAA`

* **Severity**\
  `INFO`
* **Description**\
  Developer aaa log message
* **Format String**\
  `"~s"`

</details>

<details>

<summary>DEVEL_CAPI</summary>

`DEVEL_CAPI`

* **Severity**\
  `INFO`
* **Description**\
  Developer C api log message
* **Format String**\
  `"~s"`

</details>

<details>

<summary>DEVEL_CDB</summary>

`DEVEL_CDB`

* **Severity**\
  `INFO`
* **Description**\
  Developer CDB log message
* **Format String**\
  `"~s"`

</details>

<details>

<summary>DEVEL_CONFD</summary>

`DEVEL_CONFD`

* **Severity**\
  `INFO`
* **Description**\
  Developer ConfD log message
* **Format String**\
  `"~s"`

</details>

<details>

<summary>DEVEL_ECONFD</summary>

`DEVEL_ECONFD`

* **Severity**\
  `INFO`
* **Description**\
  Developer econfd api log message
* **Format String**\
  `"~s"`

</details>

<details>

<summary>DEVEL_SLS</summary>

`DEVEL_SLS`

* **Severity**\
  `INFO`
* **Description**\
  Developer smartlicensing api log message
* **Format String**\
  `"~s"`

</details>

<details>

<summary>DEVEL_SNMPA</summary>

`DEVEL_SNMPA`

* **Severity**\
  `INFO`
* **Description**\
  Developer snmp agent log message
* **Format String**\
  `"~s"`

</details>

<details>

<summary>DEVEL_SNMPGW</summary>

`DEVEL_SNMPGW`

* **Severity**\
  `INFO`
* **Description**\
  Developer snmp GW log message
* **Format String**\
  `"~s"`

</details>

<details>

<summary>DEVEL_WEBUI</summary>

`DEVEL_WEBUI`

* **Severity**\
  `INFO`
* **Description**\
  Developer webui log message
* **Format String**\
  `"~s"`

</details>

<details>

<summary>DUPLICATE_MODULE_NAME</summary>

`DUPLICATE_MODULE_NAME`

* **Severity**\
  `CRIT`
* **Description**\
  Duplicate module name found.
* **Format String**\
  `"The module name '~s' is both defined in '~s' and '~s'."`

</details>

<details>

<summary>DUPLICATE_NAMESPACE</summary>

`DUPLICATE_NAMESPACE`

* **Severity**\
  `CRIT`
* **Description**\
  Duplicate namespace found.
* **Format String**\
  `"The namespace ~s is defined in both module ~s and ~s."`

</details>

<details>

<summary>DUPLICATE_PREFIX</summary>

`DUPLICATE_PREFIX`

* **Severity**\
  `CRIT`
* **Description**\
  Duplicate prefix found.
* **Format String**\
  `"The prefix ~s is defined in both ~s and ~s."`

</details>

<details>

<summary>ERRLOG_SIZE_CHANGED</summary>

`ERRLOG_SIZE_CHANGED`

* **Severity**\
  `INFO`
* **Description**\
  Notify change of log size for error log
* **Format String**\
  `"Changing size of error log (~s) to ~s (was ~s)"`

</details>

<details>

<summary>EVENT_SOCKET_TIMEOUT</summary>

`EVENT_SOCKET_TIMEOUT`

* **Severity**\
  `CRIT`
* **Description**\
  An event notification subscriber did not reply within the configured timeout period
* **Format String**\
  `"Event notification subscriber with bitmask ~s timed out, waiting for ~s"`

</details>

<details>

<summary>EVENT_SOCKET_WRITE_BLOCK</summary>

`EVENT_SOCKET_WRITE_BLOCK`

* **Severity**\
  `CRIT`
* **Description**\
  Write on an event socket blocked for too long time
* **Format String**\
  `"~s"`

</details>

<details>

<summary>EXEC_WHEN_CIRCULAR_DEPENDENCY</summary>

`EXEC_WHEN_CIRCULAR_DEPENDENCY`

* **Severity**\
  `WARNING`
* **Description**\
  An error occurred while evaluating a when-expression.
* **Format String**\
  `"When-expression evaluation error: circular dependency in ~s"`

</details>

<details>

<summary>EXT_AUTH_2FA_FAIL</summary>

`EXT_AUTH_2FA_FAIL`

* **Severity**\
  `INFO`
* **Description**\
  External challenge authentication failed for a user.
* **Format String**\
  `"external challenge authentication failed via ~s from ~s with ~s: ~s"`

</details>

<details>

<summary>EXT_AUTH_2FA</summary>

`EXT_AUTH_2FA`

* **Severity**\
  `INFO`
* **Description**\
  External challenge sent to a user.
* **Format String**\
  `"external challenge sent to ~s from ~s with ~s"`

</details>

<details>

<summary>EXT_AUTH_2FA_SUCCESS</summary>

`EXT_AUTH_2FA_SUCCESS`

* **Severity**\
  `INFO`
* **Description**\
  An external challenge authenticated user logged in.
* **Format String**\
  `"external challenge authentication succeeded via ~s from ~s with ~s, member of groups: ~s~s"`

</details>

<details>

<summary>EXTAUTH_BAD_RET</summary>

`EXTAUTH_BAD_RET`

* **Severity**\
  `ERR`
* **Description**\
  Authentication is external and the external program returned badly formatted data.
* **Format String**\
  `"External auth program (user=~s) ret bad output: ~s"`

</details>

<details>

<summary>EXT_AUTH_FAIL</summary>

`EXT_AUTH_FAIL`

* **Severity**\
  `INFO`
* **Description**\
  External authentication failed for a user.
* **Format String**\
  `"external authentication failed via ~s from ~s with ~s: ~s"`

</details>

<details>

<summary>EXT_AUTH_SUCCESS</summary>

`EXT_AUTH_SUCCESS`

* **Severity**\
  `INFO`
* **Description**\
  An externally authenticated user logged in.
* **Format String**\
  `"external authentication succeeded via ~s from ~s with ~s, member of groups: ~s~s"`

</details>

<details>

<summary>EXT_AUTH_TOKEN_FAIL</summary>

`EXT_AUTH_TOKEN_FAIL`

* **Severity**\
  `INFO`
* **Description**\
  External token authentication failed for a user.
* **Format String**\
  `"external token authentication failed via ~s from ~s with ~s: ~s"`

</details>

<details>

<summary>EXT_AUTH_TOKEN_SUCCESS</summary>

`EXT_AUTH_TOKEN_SUCCESS`

* **Severity**\
  `INFO`
* **Description**\
  An externally token authenticated user logged in.
* **Format String**\
  `"external token authentication succeeded via ~s from ~s with ~s, member of groups: ~s~s"`

</details>

<details>

<summary>EXT_BIND_ERR</summary>

`EXT_BIND_ERR`

* **Severity**\
  `CRIT`
* **Description**\
  ConfD failed to bind to one of the externally visible listen sockets.
* **Format String**\
  `"~s"`

</details>

<details>

<summary>FILE_ERROR</summary>

`FILE_ERROR`

* **Severity**\
  `CRIT`
* **Description**\
  File error
* **Format String**\
  `"~s: ~s"`

</details>

<details>

<summary>FILE_LOAD</summary>

`FILE_LOAD`

* **Severity**\
  `DEBUG`
* **Description**\
  System loaded a file.
* **Format String**\
  `"Loaded file ~s"`

</details>

<details>

<summary>FILE_LOAD_ERR</summary>

`FILE_LOAD_ERR`

* **Severity**\
  `CRIT`
* **Description**\
  System tried to load a file in its load path and failed.
* **Format String**\
  `"Failed to load file ~s: ~s"`

</details>

<details>

<summary>FILE_LOADING</summary>

`FILE_LOADING`

* **Severity**\
  `DEBUG`
* **Description**\
  System starts to load a file.
* **Format String**\
  `"Loading file ~s"`

</details>

<details>

<summary>FXS_MISMATCH</summary>

`FXS_MISMATCH`

* **Severity**\
  `ERR`
* **Description**\
  A secondary connected to a primary where the fxs files are different
* **Format String**\
  `"Fxs mismatch, secondary is not allowed"`

</details>

<details>

<summary>GROUP_ASSIGN</summary>

`GROUP_ASSIGN`

* **Severity**\
  `INFO`
* **Description**\
  A user was assigned to a set of groups.
* **Format String**\
  `"assigned to groups: ~s"`

</details>

<details>

<summary>GROUP_NO_ASSIGN</summary>

`GROUP_NO_ASSIGN`

* **Severity**\
  `INFO`
* **Description**\
  A user was logged in but wasn't assigned to any groups at all.
* **Format String**\
  `"Not assigned to any groups - all access is denied"`

</details>

<details>

<summary>HA_BAD_VSN</summary>

`HA_BAD_VSN`

* **Severity**\
  `ERR`
* **Description**\
  A secondary connected to a primary with an incompatible HA protocol version
* **Format String**\
  `"Incompatible HA version (~s, expected ~s), secondary is not allowed"`

</details>

<details>

<summary>HA_DUPLICATE_NODEID</summary>

`HA_DUPLICATE_NODEID`

* **Severity**\
  `ERR`
* **Description**\
  A secondary arrived with a node id which already exists
* **Format String**\
  `"Nodeid ~s already exists"`

</details>

<details>

<summary>HA_FAILED_CONNECT</summary>

`HA_FAILED_CONNECT`

* **Severity**\
  `ERR`
* **Description**\
  An attempted library become secondary call failed because the secondary couldn't connect to the primary
* **Format String**\
  `"Failed to connect to primary: ~s"`

</details>

<details>

<summary>HA_SECONDARY_KILLED</summary>

`HA_SECONDARY_KILLED`

* **Severity**\
  `ERR`
* **Description**\
  A secondary node didn't produce its ticks
* **Format String**\
  `"Secondary ~s killed due to no ticks"`

</details>

<details>

<summary>INTERNAL_ERROR</summary>

`INTERNAL_ERROR`

* **Severity**\
  `CRIT`
* **Description**\
  A ConfD internal error - should be reported to <support@tail-f.com>.
* **Format String**\
  `"Internal error: ~s"`

</details>

<details>

<summary>IPC_CAPA_DBG_DUMP_DENIED</summary>

`IPC_CAPA_DBG_DUMP_DENIED`

* **Severity**\
  `INFO`
* **Description**\
  Debug dump denied for user - capability not enabled.
* **Format String**\
  `"Debug dump denied for user '~s' - capability not enabled."`

</details>

<details>

<summary>IPC_CAPA_DBG_DUMP_GRANTED</summary>

`IPC_CAPA_DBG_DUMP_GRANTED`

* **Severity**\
  `INFO`
* **Description**\
  Debug dump allowed for user.
* **Format String**\
  `"Debug dump allowed for user '~s'."`

</details>

<details>

<summary>JIT_ENABLED</summary>

`JIT_ENABLED`

* **Severity**\
  `INFO`
* **Description**\
  Show if JIT is enabled.
* **Format String**\
  `"JIT ~s"`

</details>

<details>

<summary>JSONRPC_LOG_MSG</summary>

`JSONRPC_LOG_MSG`

* **Severity**\
  `INFO`
* **Description**\
  JSON-RPC traffic log message
* **Format String**\
  `"JSON-RPC traffic log: ~s"`

</details>

<details>

<summary>JSONRPC_REQUEST_ABSOLUTE_TIMEOUT</summary>

`JSONRPC_REQUEST_ABSOLUTE_TIMEOUT`

* **Severity**\
  `INFO`
* **Description**\
  JSON-RPC absolute timeout.
* **Format String**\
  `"Stopping session due to absolute timeout: ~s"`

</details>

<details>

<summary>JSONRPC_REQUEST_IDLE_TIMEOUT</summary>

`JSONRPC_REQUEST_IDLE_TIMEOUT`

* **Severity**\
  `INFO`
* **Description**\
  JSON-RPC idle timeout.
* **Format String**\
  `"Stopping session due to idle timeout: ~s"`

</details>

<details>

<summary>JSONRPC_REQUEST</summary>

`JSONRPC_REQUEST`

* **Severity**\
  `INFO`
* **Description**\
  JSON-RPC method requested.
* **Format String**\
  `"JSON-RPC: '~s' with JSON params ~s"`

</details>

<details>

<summary>JSONRPC_WARN_MSG</summary>

`JSONRPC_WARN_MSG`

* **Severity**\
  `WARNING`
* **Description**\
  JSON-RPC warning message
* **Format String**\
  `"JSON-RPC warning: ~s"`

</details>

<details>

<summary>KICKER_MISSING_SCHEMA</summary>

`KICKER_MISSING_SCHEMA`

* **Severity**\
  `INFO`
* **Description**\
  Failed to load kicker schema
* **Format String**\
  `"Failed to load kicker schema"`

</details>

<details>

<summary>LIB_BAD_SIZES</summary>

`LIB_BAD_SIZES`

* **Severity**\
  `ERR`
* **Description**\
  An application connecting to ConfD used a library version that can't handle the depth and number of keys used by the data model.
* **Format String**\
  `"Got connect from library with insufficient keypath depth/keys support (~s/~s, needs ~s/~s)"`

</details>

<details>

<summary>LIB_BAD_VSN</summary>

`LIB_BAD_VSN`

* **Severity**\
  `ERR`
* **Description**\
  An application connecting to ConfD used a library version that doesn't match the ConfD version (e.g. old version of the client library).
* **Format String**\
  `"Got library connect from wrong version (~s, expected ~s)"`

</details>

<details>

<summary>LIB_NO_ACCESS</summary>

`LIB_NO_ACCESS`

* **Severity**\
  `ERR`
* **Description**\
  Access check failure occurred when an application connected to ConfD.
* **Format String**\
  `"Got library connect with failed access check: ~s"`

</details>

<details>

<summary>LISTENER_INFO</summary>

`LISTENER_INFO`

* **Severity**\
  `INFO`
* **Description**\
  ConfD starts or stops to listen for incoming connections.
* **Format String**\
  `"~s to listen for ~s on ~s:~s"`

</details>

<details>

<summary>LOCAL_AUTH_FAIL_BADPASS</summary>

`LOCAL_AUTH_FAIL_BADPASS`

* **Severity**\
  `INFO`
* **Description**\
  Authentication for a locally configured user failed due to providing bad password.
* **Format String**\
  `"local authentication failed via ~s from ~s with ~s: ~s"`

</details>

<details>

<summary>LOCAL_AUTH_FAIL</summary>

`LOCAL_AUTH_FAIL`

* **Severity**\
  `INFO`
* **Description**\
  Authentication for a locally configured user failed.
* **Format String**\
  `"local authentication failed via ~s from ~s with ~s: ~s"`

</details>

<details>

<summary>LOCAL_AUTH_FAIL_NOUSER</summary>

`LOCAL_AUTH_FAIL_NOUSER`

* **Severity**\
  `INFO`
* **Description**\
  Authentication for a locally configured user failed due to user not found.
* **Format String**\
  `"local authentication failed via ~s from ~s with ~s: ~s"`

</details>

<details>

<summary>LOCAL_AUTH_SUCCESS</summary>

`LOCAL_AUTH_SUCCESS`

* **Severity**\
  `INFO`
* **Description**\
  A locally authenticated user logged in.
* **Format String**\
  `"local authentication succeeded via ~s from ~s with ~s, member of groups: ~s"`

</details>

<details>

<summary>LOCAL_IPC_ACCESS_DENIED</summary>

`LOCAL_IPC_ACCESS_DENIED`

* **Severity**\
  `INFO`
* **Description**\
  Local IPC access denied for user.
* **Format String**\
  `"Local IPC access denied for user ~s connecting from ~s"`

</details>

<details>

<summary>LOGGING_DEST_CHANGED</summary>

`LOGGING_DEST_CHANGED`

* **Severity**\
  `INFO`
* **Description**\
  The target logfile will change to another file
* **Format String**\
  `"Changing destination of ~s log to ~s"`

</details>

<details>

<summary>LOGGING_SHUTDOWN</summary>

`LOGGING_SHUTDOWN`

* **Severity**\
  `INFO`
* **Description**\
  Logging subsystem terminating
* **Format String**\
  `"Daemon logging terminating, reason: ~s"`

</details>

<details>

<summary>LOGGING_STARTED</summary>

`LOGGING_STARTED`

* **Severity**\
  `INFO`
* **Description**\
  Logging subsystem started
* **Format String**\
  `"Daemon logging started"`

</details>

<details>

<summary>LOGGING_STARTED_TO</summary>

`LOGGING_STARTED_TO`

* **Severity**\
  `INFO`
* **Description**\
  Write logs for a subsystem to a specific file
* **Format String**\
  `"Writing ~s log to ~s"`

</details>

<details>

<summary>LOGGING_STATUS_CHANGED</summary>

`LOGGING_STATUS_CHANGED`

* **Severity**\
  `INFO`
* **Description**\
  Notify a change of logging status (enabled/disabled) for a subsystem
* **Format String**\
  `"~s ~s log"`

</details>

<details>

<summary>LOGIN_REJECTED</summary>

`LOGIN_REJECTED`

* **Severity**\
  `INFO`
* **Description**\
  Authentication for a user was rejected by application callback.
* **Format String**\
  `"~s"`

</details>

<details>

<summary>MAAPI_LOGOUT</summary>

`MAAPI_LOGOUT`

* **Severity**\
  `INFO`
* **Description**\
  A maapi user was logged out.
* **Format String**\
  `"Logged out from maapi ctx=~s (~s)"`

</details>

<details>

<summary>MAAPI_WRITE_TO_SOCKET_FAIL</summary>

`MAAPI_WRITE_TO_SOCKET_FAIL`

* **Severity**\
  `INFO`
* **Description**\
  maapi failed to write to a socket.
* **Format String**\
  `"maapi server failed to write to a socket. Op: ~s Ecode: ~s Error: ~s~s"`

</details>

<details>

<summary>MISSING_AES256CFB128_SETTINGS</summary>

`MISSING_AES256CFB128_SETTINGS`

* **Severity**\
  `ERR`
* **Description**\
  AES256CFB128 keys were not found in confd.conf
* **Format String**\
  `"AES256CFB128 keys were not found in confd.conf"`

</details>

<details>

<summary>MISSING_AESCFB128_SETTINGS</summary>

`MISSING_AESCFB128_SETTINGS`

* **Severity**\
  `ERR`
* **Description**\
  AESCFB128 keys were not found in confd.conf
* **Format String**\
  `"AESCFB128 keys were not found in confd.conf"`

</details>

<details>

<summary>MISSING_DES3CBC_SETTINGS</summary>

`MISSING_DES3CBC_SETTINGS`

* **Severity**\
  `ERR`
* **Description**\
  DES3CBC keys were not found in confd.conf
* **Format String**\
  `"DES3CBC keys were not found in confd.conf"`

</details>

<details>

<summary>MISSING_NS2</summary>

`MISSING_NS2`

* **Severity**\
  `CRIT`
* **Description**\
  While validating the consistency of the config - a required namespace was missing.
* **Format String**\
  `"The namespace ~s (referenced by ~s) could not be found in the loadPath."`

</details>

<details>

<summary>MISSING_NS</summary>

`MISSING_NS`

* **Severity**\
  `CRIT`
* **Description**\
  While validating the consistency of the config - a required namespace was missing.
* **Format String**\
  `"The namespace ~s could not be found in the loadPath."`

</details>

<details>

<summary>MMAP_SCHEMA_FAIL</summary>

`MMAP_SCHEMA_FAIL`

* **Severity**\
  `ERR`
* **Description**\
  Failed to setup the shared memory schema
* **Format String**\
  `"Failed to setup the shared memory schema"`

</details>

<details>

<summary>NETCONF_HDR_ERR</summary>

`NETCONF_HDR_ERR`

* **Severity**\
  `ERR`
* **Description**\
  The cleartext header indicating user and groups was badly formatted.
* **Format String**\
  `"Got bad NETCONF TCP header"`

</details>

<details>

<summary>NETCONF</summary>

`NETCONF`

* **Severity**\
  `INFO`
* **Description**\
  NETCONF traffic log message
* **Format String**\
  `"~s"`

</details>

<details>

<summary>NIF_LOG</summary>

`NIF_LOG`

* **Severity**\
  `INFO`
* **Description**\
  Log message from NIF code.
* **Format String**\
  `"~s: ~s"`

</details>

<details>

<summary>NOAAA_CLI_LOGIN</summary>

`NOAAA_CLI_LOGIN`

* **Severity**\
  `INFO`
* **Description**\
  A user used the --noaaa flag to confd\_cli
* **Format String**\
  `"logged in from the CLI with aaa disabled"`

</details>

<details>

<summary>NO_CALLPOINT</summary>

`NO_CALLPOINT`

* **Severity**\
  `CRIT`
* **Description**\
  ConfD tried to populate an XML tree but no code had registered under the relevant callpoint.
* **Format String**\
  `"no registration found for callpoint ~s of type=~s"`

</details>

<details>

<summary>NO_SUCH_IDENTITY</summary>

`NO_SUCH_IDENTITY`

* **Severity**\
  `CRIT`
* **Description**\
  The fxs file with the base identity is not loaded
* **Format String**\
  `"The identity ~s in namespace ~s refers to a non-existing base identity ~s in namespace ~s"`

</details>

<details>

<summary>NO_SUCH_NS</summary>

`NO_SUCH_NS`

* **Severity**\
  `CRIT`
* **Description**\
  A nonexistent namespace was referred to. Typically this means that a .fxs was missing from the loadPath.
* **Format String**\
  `"No such namespace ~s, used by ~s"`

</details>

<details>

<summary>NO_SUCH_TYPE</summary>

`NO_SUCH_TYPE`

* **Severity**\
  `CRIT`
* **Description**\
  A nonexistent type was referred to from a ns. Typically this means that a bad version of an .fxs file was found in the loadPath.
* **Format String**\
  `"No such simpleType '~s' in ~s, used by ~s"`

</details>

<details>

<summary>NOTIFICATION_REPLAY_STORE_FAILURE</summary>

`NOTIFICATION_REPLAY_STORE_FAILURE`

* **Severity**\
  `CRIT`
* **Description**\
  A failure occurred in the builtin notification replay store
* **Format String**\
  `"~s"`

</details>

<details>

<summary>NS_LOAD_ERR2</summary>

`NS_LOAD_ERR2`

* **Severity**\
  `CRIT`
* **Description**\
  System tried to process a loaded namespace and failed.
* **Format String**\
  `"Failed to process namespaces: ~s"`

</details>

<details>

<summary>NS_LOAD_ERR</summary>

`NS_LOAD_ERR`

* **Severity**\
  `CRIT`
* **Description**\
  System tried to process a loaded namespace and failed.
* **Format String**\
  `"Failed to process namespace ~s: ~s"`

</details>

<details>

<summary>OPEN_LOGFILE</summary>

`OPEN_LOGFILE`

* **Severity**\
  `INFO`
* **Description**\
  Indicate target file for certain type of logging
* **Format String**\
  `"Logging subsystem, opening log file '~s' for ~s"`

</details>

<details>

<summary>PAM_AUTH_FAIL</summary>

`PAM_AUTH_FAIL`

* **Severity**\
  `INFO`
* **Description**\
  A user failed to authenticate through PAM.
* **Format String**\
  `"PAM authentication failed via ~s from ~s with ~s: phase ~s, ~s"`

</details>

<details>

<summary>PAM_AUTH_SUCCESS</summary>

`PAM_AUTH_SUCCESS`

* **Severity**\
  `INFO`
* **Description**\
  A PAM authenticated user logged in.
* **Format String**\
  `"pam authentication succeeded via ~s from ~s with ~s"`

</details>

<details>

<summary>PHASE0_STARTED</summary>

`PHASE0_STARTED`

* **Severity**\
  `INFO`
* **Description**\
  ConfD has just started its start phase 0.
* **Format String**\
  `"ConfD phase0 started"`

</details>

<details>

<summary>PHASE1_STARTED</summary>

`PHASE1_STARTED`

* **Severity**\
  `INFO`
* **Description**\
  ConfD has just started its start phase 1.
* **Format String**\
  `"ConfD phase1 started"`

</details>

<details>

<summary>READ_STATE_FILE_FAILED</summary>

`READ_STATE_FILE_FAILED`

* **Severity**\
  `CRIT`
* **Description**\
  Reading of a state file failed
* **Format String**\
  `"Reading state file failed: ~s: ~s (~s)"`

</details>

<details>

<summary>RELOAD</summary>

`RELOAD`

* **Severity**\
  `INFO`
* **Description**\
  Reload of daemon configuration has been initiated.
* **Format String**\
  `"Reloading daemon configuration."`

</details>

<details>

<summary>REOPEN_LOGS</summary>

`REOPEN_LOGS`

* **Severity**\
  `INFO`
* **Description**\
  Logging subsystem, reopening log files
* **Format String**\
  `"Logging subsystem, reopening log files"`

</details>

<details>

<summary>REST_AUTH_FAIL</summary>

`REST_AUTH_FAIL`

* **Severity**\
  `INFO`
* **Description**\
  Rest authentication for a user failed.
* **Format String**\
  `"rest authentication failed from ~s"`

</details>

<details>

<summary>REST_AUTH_SUCCESS</summary>

`REST_AUTH_SUCCESS`

* **Severity**\
  `INFO`
* **Description**\
  A rest authenticated user logged in.
* **Format String**\
  `"rest authentication succeeded from ~s , member of groups: ~s"`

</details>

<details>

<summary>RESTCONF_REQUEST</summary>

`RESTCONF_REQUEST`

* **Severity**\
  `INFO`
* **Description**\
  RESTCONF request
* **Format String**\
  `"RESTCONF: request with ~s: ~s"`

</details>

<details>

<summary>RESTCONF_RESPONSE</summary>

`RESTCONF_RESPONSE`

* **Severity**\
  `INFO`
* **Description**\
  RESTCONF response
* **Format String**\
  `"RESTCONF: response with ~s: ~s duration ~s us"`

</details>

<details>

<summary>REST_REQUEST</summary>

`REST_REQUEST`

* **Severity**\
  `INFO`
* **Description**\
  REST request
* **Format String**\
  `"REST: request with ~s: ~s"`

</details>

<details>

<summary>REST_RESPONSE</summary>

`REST_RESPONSE`

* **Severity**\
  `INFO`
* **Description**\
  REST response
* **Format String**\
  `"REST: response with ~s: ~s duration ~s ms"`

</details>

<details>

<summary>ROLLBACK_FAIL_CREATE</summary>

`ROLLBACK_FAIL_CREATE`

* **Severity**\
  `ERR`
* **Description**\
  Error while creating rollback file.
* **Format String**\
  `"Error while creating rollback file: ~s: ~s"`

</details>

<details>

<summary>ROLLBACK_FAIL_DELETE</summary>

`ROLLBACK_FAIL_DELETE`

* **Severity**\
  `ERR`
* **Description**\
  Failed to delete rollback file.
* **Format String**\
  `"Failed to delete rollback file ~s: ~s"`

</details>

<details>

<summary>ROLLBACK_FAIL_RENAME</summary>

`ROLLBACK_FAIL_RENAME`

* **Severity**\
  `ERR`
* **Description**\
  Failed to rename rollback file.
* **Format String**\
  `"Failed to rename rollback file ~s to ~s: ~s"`

</details>

<details>

<summary>ROLLBACK_FAIL_REPAIR</summary>

`ROLLBACK_FAIL_REPAIR`

* **Severity**\
  `ERR`
* **Description**\
  Failed to repair rollback files.
* **Format String**\
  `"Failed to repair rollback files."`

</details>

<details>

<summary>ROLLBACK_REMOVE</summary>

`ROLLBACK_REMOVE`

* **Severity**\
  `INFO`
* **Description**\
  Found half created rollback0 file - removing and creating new.
* **Format String**\
  `"Found half created rollback0 file - removing and creating new"`

</details>

<details>

<summary>ROLLBACK_REPAIR</summary>

`ROLLBACK_REPAIR`

* **Severity**\
  `INFO`
* **Description**\
  Found half created rollback0 file - repairing.
* **Format String**\
  `"Found half created rollback0 file - repairing"`

</details>

<details>

<summary>SESSION_CREATE</summary>

`SESSION_CREATE`

* **Severity**\
  `INFO`
* **Description**\
  A new user session was created
* **Format String**\
  `"created new session via ~s from ~s with ~s"`

</details>

<details>

<summary>SESSION_LIMIT</summary>

`SESSION_LIMIT`

* **Severity**\
  `INFO`
* **Description**\
  Session limit reached, rejected new session request.
* **Format String**\
  `"Session limit of type '~s' reached, rejected new session request"`

</details>

<details>

<summary>SESSION_MAX_EXCEEDED</summary>

`SESSION_MAX_EXCEEDED`

* **Severity**\
  `INFO`
* **Description**\
  A user failed to create a new user sessions due to exceeding sessions limits
* **Format String**\
  `"could not create new session via ~s from ~s with ~s due to session limits"`

</details>

<details>

<summary>SESSION_TERMINATION</summary>

`SESSION_TERMINATION`

* **Severity**\
  `INFO`
* **Description**\
  A user session was terminated due to specified reason
* **Format String**\
  `"terminated session (reason: ~s)"`

</details>

<details>

<summary>SKIP_FILE_LOADING</summary>

`SKIP_FILE_LOADING`

* **Severity**\
  `DEBUG`
* **Description**\
  System skips a file.
* **Format String**\
  `"Skipping file ~s: ~s"`

</details>

<details>

<summary>SNMP_AUTHENTICATION_FAILED</summary>

`SNMP_AUTHENTICATION_FAILED`

* **Severity**\
  `INFO`
* **Description**\
  An SNMP authentication failed.
* **Format String**\
  `"SNMP authentication failed: ~s"`

</details>

<details>

<summary>SNMP_CANT_LOAD_MIB</summary>

`SNMP_CANT_LOAD_MIB`

* **Severity**\
  `CRIT`
* **Description**\
  The SNMP Agent failed to load a MIB file
* **Format String**\
  `"Can't load MIB file: ~s"`

</details>

<details>

<summary>SNMP_MIB_LOADING</summary>

`SNMP_MIB_LOADING`

* **Severity**\
  `DEBUG`
* **Description**\
  SNMP Agent loading a MIB file
* **Format String**\
  `"Loading MIB: ~s"`

</details>

<details>

<summary>SNMP_NOT_A_TRAP</summary>

`SNMP_NOT_A_TRAP`

* **Severity**\
  `INFO`
* **Description**\
  An UDP package was received on the trap receiving port, but it's not an SNMP trap.
* **Format String**\
  `"SNMP gateway: Non-trap received from ~s"`

</details>

<details>

<summary>SNMP_READ_STATE_FILE_FAILED</summary>

`SNMP_READ_STATE_FILE_FAILED`

* **Severity**\
  `CRIT`
* **Description**\
  Read SNMP agent state file failed
* **Format String**\
  `"Read state file failed: ~s: ~s"`

</details>

<details>

<summary>SNMP_REQUIRES_CDB</summary>

`SNMP_REQUIRES_CDB`

* **Severity**\
  `WARNING`
* **Description**\
  The SNMP agent requires CDB to be enabled in order to be started.
* **Format String**\
  `"Can't start SNMP. CDB is not enabled"`

</details>

<details>

<summary>SNMP_TRAP_NOT_FORWARDED</summary>

`SNMP_TRAP_NOT_FORWARDED`

* **Severity**\
  `INFO`
* **Description**\
  An SNMP trap was to be forwarded, but couldn't be.
* **Format String**\
  `"SNMP gateway: Can't forward trap from ~s; ~s"`

</details>

<details>

<summary>SNMP_TRAP_NOT_RECOGNIZED</summary>

`SNMP_TRAP_NOT_RECOGNIZED`

* **Severity**\
  `INFO`
* **Description**\
  An SNMP trap was received on the trap receiving port, but its definition is not known
* **Format String**\
  `"SNMP gateway: Can't forward trap with OID ~s from ~s; There is no notification with this OID in the loaded models."`

</details>

<details>

<summary>SNMP_TRAP_OPEN_PORT</summary>

`SNMP_TRAP_OPEN_PORT`

* **Severity**\
  `ERR`
* **Description**\
  The port for listening to SNMP traps could not be opened.
* **Format String**\
  `"SNMP gateway: Can't open trap listening port ~s: ~s"`

</details>

<details>

<summary>SNMP_TRAP_UNKNOWN_SENDER</summary>

`SNMP_TRAP_UNKNOWN_SENDER`

* **Severity**\
  `INFO`
* **Description**\
  An SNMP trap was to be forwarded, but the sender was not listed in confd.conf.
* **Format String**\
  `"SNMP gateway: Not forwarding trap from ~s; the sender is not recognized"`

</details>

<details>

<summary>SNMP_TRAP_V1</summary>

`SNMP_TRAP_V1`

* **Severity**\
  `INFO`
* **Description**\
  An SNMP v1 trap was received on the trap receiving port, but forwarding v1 traps is not supported.
* **Format String**\
  `"SNMP gateway: V1 trap received from ~s"`

</details>

<details>

<summary>SNMP_WRITE_STATE_FILE_FAILED</summary>

`SNMP_WRITE_STATE_FILE_FAILED`

* **Severity**\
  `WARNING`
* **Description**\
  Write SNMP agent state file failed
* **Format String**\
  `"Write state file failed: ~s: ~s"`

</details>

<details>

<summary>SSH_HOST_KEY_UNAVAILABLE</summary>

`SSH_HOST_KEY_UNAVAILABLE`

* **Severity**\
  `ERR`
* **Description**\
  No SSH host keys available.
* **Format String**\
  `"No SSH host keys available"`

</details>

<details>

<summary>SSH_SUBSYS_ERR</summary>

`SSH_SUBSYS_ERR`

* **Severity**\
  `INFO`
* **Description**\
  Typically errors where the client doesn't properly send the "subsystem" command.
* **Format String**\
  `"ssh protocol subsys - ~s"`

</details>

<details>

<summary>STARTED</summary>

`STARTED`

* **Severity**\
  `INFO`
* **Description**\
  ConfD has started.
* **Format String**\
  `"ConfD started vsn: ~s"`

</details>

<details>

<summary>STARTING</summary>

`STARTING`

* **Severity**\
  `INFO`
* **Description**\
  ConfD is starting.
* **Format String**\
  `"Starting ConfD vsn: ~s"`

</details>

<details>

<summary>STOPPING</summary>

`STOPPING`

* **Severity**\
  `INFO`
* **Description**\
  ConfD is stopping (due to e.g. confd --stop).
* **Format String**\
  `"ConfD stopping (~s)"`

</details>

<details>

<summary>TOKEN_MISMATCH</summary>

`TOKEN_MISMATCH`

* **Severity**\
  `ERR`
* **Description**\
  A secondary connected to a primary with a bad auth token
* **Format String**\
  `"Token mismatch, secondary is not allowed"`

</details>

<details>

<summary>UPGRADE_ABORTED</summary>

`UPGRADE_ABORTED`

* **Severity**\
  `INFO`
* **Description**\
  In-service upgrade was aborted.
* **Format String**\
  `"Upgrade aborted"`

</details>

<details>

<summary>UPGRADE_COMMITTED</summary>

`UPGRADE_COMMITTED`

* **Severity**\
  `INFO`
* **Description**\
  In-service upgrade was committed.
* **Format String**\
  `"Upgrade committed"`

</details>

<details>

<summary>UPGRADE_INIT_STARTED</summary>

`UPGRADE_INIT_STARTED`

* **Severity**\
  `INFO`
* **Description**\
  In-service upgrade initialization has started.
* **Format String**\
  `"Upgrade init started"`

</details>

<details>

<summary>UPGRADE_INIT_SUCCEEDED</summary>

`UPGRADE_INIT_SUCCEEDED`

* **Severity**\
  `INFO`
* **Description**\
  In-service upgrade initialization succeeded.
* **Format String**\
  `"Upgrade init succeeded"`

</details>

<details>

<summary>UPGRADE_PERFORMED</summary>

`UPGRADE_PERFORMED`

* **Severity**\
  `INFO`
* **Description**\
  In-service upgrade has been performed (not committed yet).
* **Format String**\
  `"Upgrade performed"`

</details>

<details>

<summary>WEB_ACTION</summary>

`WEB_ACTION`

* **Severity**\
  `INFO`
* **Description**\
  User executed a Web UI action.
* **Format String**\
  `"WebUI action '~s'"`

</details>

<details>

<summary>WEB_CMD</summary>

`WEB_CMD`

* **Severity**\
  `INFO`
* **Description**\
  User executed a Web UI command.
* **Format String**\
  `"WebUI cmd '~s'"`

</details>

<details>

<summary>WEB_COMMIT</summary>

`WEB_COMMIT`

* **Severity**\
  `INFO`
* **Description**\
  User performed Web UI commit.
* **Format String**\
  `"WebUI commit ~s"`

</details>

<details>

<summary>WEBUI_LOG_MSG</summary>

`WEBUI_LOG_MSG`

* **Severity**\
  `INFO`
* **Description**\
  WebUI access log message
* **Format String**\
  `"WebUI access log: ~s"`

</details>

<details>

<summary>WRITE_STATE_FILE_FAILED</summary>

`WRITE_STATE_FILE_FAILED`

* **Severity**\
  `CRIT`
* **Description**\
  Writing of a state file failed
* **Format String**\
  `"Writing state file failed: ~s: ~s (~s)"`

</details>

<details>

<summary>XPATH_EVAL_ERROR1</summary>

`XPATH_EVAL_ERROR1`

* **Severity**\
  `WARNING`
* **Description**\
  An error occurred while evaluating an XPath expression.
* **Format String**\
  `"XPath evaluation error: ~s for ~s"`

</details>

<details>

<summary>XPATH_EVAL_ERROR2</summary>

`XPATH_EVAL_ERROR2`

* **Severity**\
  `WARNING`
* **Description**\
  An error occurred while evaluating an XPath expression.
* **Format String**\
  `"XPath evaluation error: '~s' resulted in ~s for ~s"`

</details>

<details>

<summary>COMMIT_UN_SYNCED_DEV</summary>

`COMMIT_UN_SYNCED_DEV`

* **Severity**\
  `INFO`
* **Description**\
  Data was committed toward a device with bad or unknown sync state
* **Format String**\
  `"Committed data towards device ~s which is out of sync"`

</details>

<details>

<summary>NCS_DEVICE_OUT_OF_SYNC</summary>

`NCS_DEVICE_OUT_OF_SYNC`

* **Severity**\
  `INFO`
* **Description**\
  A check-sync action reported out-of-sync for a device
* **Format String**\
  `"NCS device-out-of-sync Device '~s' Info '~s'"`

</details>

<details>

<summary>NCS_JAVA_VM_FAIL</summary>

`NCS_JAVA_VM_FAIL`

* **Severity**\
  `ERR`
* **Description**\
  The NCS Java VM failure/timeout
* **Format String**\
  `"The NCS Java VM ~s"`

</details>

<details>

<summary>NCS_JAVA_VM_START</summary>

`NCS_JAVA_VM_START`

* **Severity**\
  `INFO`
* **Description**\
  Starting the NCS Java VM
* **Format String**\
  `"Starting the NCS Java VM"`

</details>

<details>

<summary>NCS_MEMORY_ACTION_TRIGGERED</summary>

`NCS_MEMORY_ACTION_TRIGGERED`

* **Severity**\
  `WARNING`
* **Description**\
  Memory-management action threshold reached
* **Format String**\
  `"Memory threshold reached in memory-management action '~s': ~s"`

</details>

<details>

<summary>NCS_PACKAGE_AUTH_BAD_RET</summary>

`NCS_PACKAGE_AUTH_BAD_RET`

* **Severity**\
  `ERR`
* **Description**\
  Package authentication program returned badly formatted data.
* **Format String**\
  `"package authentication using ~s program ret bad output: ~s"`

</details>

<details>

<summary>NCS_PACKAGE_AUTH_ERR</summary>

`NCS_PACKAGE_AUTH_ERR`

* **Severity**\
  `ERR`
* **Description**\
  Package authentication failed due to incorrect setup.
* **Format String**\
  `"package authentication using ~s failed with reason '~s'"`

</details>

<details>

<summary>NCS_PACKAGE_AUTH_FAIL</summary>

`NCS_PACKAGE_AUTH_FAIL`

* **Severity**\
  `INFO`
* **Description**\
  Package authentication failed.
* **Format String**\
  `"package authentication using ~s failed via ~s from ~s with ~s: ~s"`

</details>

<details>

<summary>NCS_PACKAGE_AUTH_SUCCESS</summary>

`NCS_PACKAGE_AUTH_SUCCESS`

* **Severity**\
  `INFO`
* **Description**\
  A package authenticated user logged in.
* **Format String**\
  `"package authentication using ~s succeeded via ~s from ~s with ~s, member of groups: ~s~s"`

</details>

<details>

<summary>NCS_PACKAGE_BAD_DEPENDENCY</summary>

`NCS_PACKAGE_BAD_DEPENDENCY`

* **Severity**\
  `CRIT`
* **Description**\
  Bad NCS package dependency
* **Format String**\
  `"Failed to load NCS package: ~s; required package ~s of version ~s is not present (found ~s)"`

</details>

<details>

<summary>NCS_PACKAGE_BAD_NCS_VERSION</summary>

`NCS_PACKAGE_BAD_NCS_VERSION`

* **Severity**\
  `CRIT`
* **Description**\
  Bad NCS version for package
* **Format String**\
  `"Failed to load NCS package: ~s; requires NCS version ~s"`

</details>

<details>

<summary>NCS_PACKAGE_CHAL_2FA</summary>

`NCS_PACKAGE_CHAL_2FA`

* **Severity**\
  `INFO`
* **Description**\
  Package authentication challenge sent to a user.
* **Format String**\
  `"package authentication challenge sent to ~s from ~s with ~s"`

</details>

<details>

<summary>NCS_PACKAGE_CHAL_FAIL</summary>

`NCS_PACKAGE_CHAL_FAIL`

* **Severity**\
  `INFO`
* **Description**\
  Package authentication challenge failed.
* **Format String**\
  `"package authentication challenge using ~s failed via ~s from ~s with ~s: ~s"`

</details>

<details>

<summary>NCS_PACKAGE_CIRCULAR_DEPENDENCY</summary>

`NCS_PACKAGE_CIRCULAR_DEPENDENCY`

* **Severity**\
  `CRIT`
* **Description**\
  Circular NCS package dependency
* **Format String**\
  `"Failed to load NCS package: ~s; circular dependency found"`

</details>

<details>

<summary>NCS_PACKAGE_COPYING</summary>

`NCS_PACKAGE_COPYING`

* **Severity**\
  `DEBUG`
* **Description**\
  A package is copied from the load path to private directory
* **Format String**\
  `"Copying NCS package from ~s to ~s"`

</details>

<details>

<summary>NCS_PACKAGE_DUPLICATE</summary>

`NCS_PACKAGE_DUPLICATE`

* **Severity**\
  `CRIT`
* **Description**\
  Duplicate package found
* **Format String**\
  `"Failed to load duplicate NCS package ~s: (~s)"`

</details>

<details>

<summary>NCS_PACKAGE_STATUS_CHANGE</summary>

`NCS_PACKAGE_STATUS_CHANGE`

* **Severity**\
  `DEBUG`
* **Description**\
  Status changed for the given package.
* **Format String**\
  `"package '~s' status changed to '~s'."`

</details>

<details>

<summary>NCS_PACKAGE_SYNTAX_ERROR</summary>

`NCS_PACKAGE_SYNTAX_ERROR`

* **Severity**\
  `CRIT`
* **Description**\
  Syntax error in package file
* **Format String**\
  `"Failed to load NCS package: ~s; syntax error in package file"`

</details>

<details>

<summary>NCS_PACKAGE_UPGRADE_ABORTED</summary>

`NCS_PACKAGE_UPGRADE_ABORTED`

* **Severity**\
  `CRIT`
* **Description**\
  The CDB upgrade was aborted implying that CDB is untouched. However the package state is changed
* **Format String**\
  `"NCS package upgrade failed with reason '~s'"`

</details>

<details>

<summary>NCS_PACKAGE_UPGRADE_UNSAFE</summary>

`NCS_PACKAGE_UPGRADE_UNSAFE`

* **Severity**\
  `CRIT`
* **Description**\
  Package upgrade has been aborted due to warnings.
* **Format String**\
  `"NCS package upgrade has been aborted due to warnings:\n~s"`

</details>

<details>

<summary>NCS_PYTHON_VM_FAIL</summary>

`NCS_PYTHON_VM_FAIL`

* **Severity**\
  `ERR`
* **Description**\
  The NCS Python VM failure/timeout
* **Format String**\
  `"The NCS Python VM ~s"`

</details>

<details>

<summary>NCS_PYTHON_VM_START</summary>

`NCS_PYTHON_VM_START`

* **Severity**\
  `INFO`
* **Description**\
  Starting the named NCS Python VM
* **Format String**\
  `"Starting the NCS Python VM ~s"`

</details>

<details>

<summary>NCS_PYTHON_VM_START_UPGRADE</summary>

`NCS_PYTHON_VM_START_UPGRADE`

* **Severity**\
  `INFO`
* **Description**\
  Starting a Python VM to run upgrade code
* **Format String**\
  `"Starting upgrade of NCS Python package ~s"`

</details>

<details>

<summary>NCS_SERVICE_OUT_OF_SYNC</summary>

`NCS_SERVICE_OUT_OF_SYNC`

* **Severity**\
  `INFO`
* **Description**\
  A check-sync action reported out-of-sync for a service
* **Format String**\
  `"NCS service-out-of-sync Service '~s' Info '~s'"`

</details>

<details>

<summary>NCS_SET_PLATFORM_DATA_ERROR</summary>

`NCS_SET_PLATFORM_DATA_ERROR`

* **Severity**\
  `ERR`
* **Description**\
  The device failed to set the platform operational data at connect
* **Format String**\
  `"NCS Device '~s' failed to set platform data Info '~s'"`

</details>

<details>

<summary>NCS_SMART_LICENSING_ENTITLEMENT_NOTIFICATION</summary>

`NCS_SMART_LICENSING_ENTITLEMENT_NOTIFICATION`

* **Severity**\
  `INFO`
* **Description**\
  Smart Licensing Entitlement Notification
* **Format String**\
  `"Smart Licensing Entitlement Notification: ~s"`

</details>

<details>

<summary>NCS_SMART_LICENSING_EVALUATION_COUNTDOWN</summary>

`NCS_SMART_LICENSING_EVALUATION_COUNTDOWN`

* **Severity**\
  `INFO`
* **Description**\
  Smart Licensing evaluation time remaining
* **Format String**\
  `"Smart Licensing evaluation time remaining: ~s"`

</details>

<details>

<summary>NCS_SMART_LICENSING_FAIL</summary>

`NCS_SMART_LICENSING_FAIL`

* **Severity**\
  `INFO`
* **Description**\
  The NCS Smart Licensing Java VM failure/timeout
* **Format String**\
  `"The NCS Smart Licensing Java VM ~s"`

</details>

<details>

<summary>NCS_SMART_LICENSING_GLOBAL_NOTIFICATION</summary>

`NCS_SMART_LICENSING_GLOBAL_NOTIFICATION`

* **Severity**\
  `INFO`
* **Description**\
  Smart Licensing Global Notification
* **Format String**\
  `"Smart Licensing Global Notification: ~s"`

</details>

<details>

<summary>NCS_SMART_LICENSING_START</summary>

`NCS_SMART_LICENSING_START`

* **Severity**\
  `INFO`
* **Description**\
  Starting the NCS Smart Licensing Java VM
* **Format String**\
  `"Starting the NCS Smart Licensing Java VM"`

</details>

<details>

<summary>NCS_SNMP_INIT_ERR</summary>

`NCS_SNMP_INIT_ERR`

* **Severity**\
  `INFO`
* **Description**\
  Failed to locate snmp\_init.xml in loadpath
* **Format String**\
  `"Failed to locate snmp_init.xml in loadpath ~s"`

</details>

<details>

<summary>NCS_SNMPM_START</summary>

`NCS_SNMPM_START`

* **Severity**\
  `INFO`
* **Description**\
  Starting the NCS SNMP manager component
* **Format String**\
  `"Starting the NCS SNMP manager component"`

</details>

<details>

<summary>NCS_SNMPM_STOP</summary>

`NCS_SNMPM_STOP`

* **Severity**\
  `INFO`
* **Description**\
  The NCS SNMP manager component has been stopped
* **Format String**\
  `"The NCS SNMP manager component has been stopped"`

</details>

<details>

<summary>NCS_TLS_CERT_LOAD_FR_DB_ERR</summary>

`NCS_TLS_CERT_LOAD_FR_DB_ERR`

* **Severity**\
  `CRIT`
* **Description**\
  Failed to load SSL/TLS certificate from database.
* **Format String**\
  `"Failed to load SSL/TLS certificate from db: ~s."`

</details>

<details>

<summary>NCS_TLS_CERT_LOAD_FR_FILE_ERR</summary>

`NCS_TLS_CERT_LOAD_FR_FILE_ERR`

* **Severity**\
  `CRIT`
* **Description**\
  Failed to load SSL/TLS certificate from file.
* **Format String**\
  `"Failed to load SSL/TLS certificate from file: ~s; Please check files specified at /ncs-config/webui/transport/ssl/cert-file or /ncs-config/webui/transport/ssl/ca-cert-file"`

</details>

<details>

<summary>NCS_UPGRADE_ABORTED_INTERNAL</summary>

`NCS_UPGRADE_ABORTED_INTERNAL`

* **Severity**\
  `CRIT`
* **Description**\
  The CDB upgrade was aborted due to some internal error. CDB is left untouched
* **Format String**\
  `"NCS upgrade failed with reason '~s'"`

</details>

<details>

<summary>BAD_LOCAL_PASS</summary>

`BAD_LOCAL_PASS`

* **Severity**\
  `INFO`
* **Description**\
  A locally configured user provided a bad password.
* **Format String**\
  `"Provided bad password"`

</details>

<details>

<summary>EXT_LOGIN</summary>

`EXT_LOGIN`

* **Severity**\
  `INFO`
* **Description**\
  An externally authenticated user logged in.
* **Format String**\
  `"Logged in over ~s using externalauth, member of groups: ~s~s"`

</details>

<details>

<summary>EXT_NO_LOGIN</summary>

`EXT_NO_LOGIN`

* **Severity**\
  `INFO`
* **Description**\
  External authentication failed for a user.
* **Format String**\
  `"failed to login using externalauth: ~s"`

</details>

<details>

<summary>NO_SUCH_LOCAL_USER</summary>

`NO_SUCH_LOCAL_USER`

* **Severity**\
  `INFO`
* **Description**\
  A non existing local user tried to login.
* **Format String**\
  `"no such local user"`

</details>

<details>

<summary>PAM_LOGIN_FAILED</summary>

`PAM_LOGIN_FAILED`

* **Severity**\
  `INFO`
* **Description**\
  A user failed to login through PAM.
* **Format String**\
  `"pam phase ~s failed to login through PAM: ~s"`

</details>

<details>

<summary>PAM_NO_LOGIN</summary>

`PAM_NO_LOGIN`

* **Severity**\
  `INFO`
* **Description**\
  A user failed to login through PAM
* **Format String**\
  `"failed to login through PAM: ~s"`

</details>

<details>

<summary>SSH_LOGIN</summary>

`SSH_LOGIN`

* **Severity**\
  `INFO`
* **Description**\
  A user logged into ConfD's builtin ssh server.
* **Format String**\
  `"logged in over ssh from ~s with authmeth:~s"`

</details>

<details>

<summary>SSH_LOGOUT</summary>

`SSH_LOGOUT`

* **Severity**\
  `INFO`
* **Description**\
  A user was logged out from ConfD's builtin ssh server.
* **Format String**\
  `"Logged out ssh <~s> user"`

</details>

<details>

<summary>SSH_NO_LOGIN</summary>

`SSH_NO_LOGIN`

* **Severity**\
  `INFO`
* **Description**\
  A user failed to login to ConfD's builtin SSH server.
* **Format String**\
  `"Failed to login over ssh: ~s"`

</details>

<details>

<summary>WEB_LOGIN</summary>

`WEB_LOGIN`

* **Severity**\
  `INFO`
* **Description**\
  A user logged in through the WebUI.
* **Format String**\
  `"logged in through Web UI from ~s"`

</details>

<details>

<summary>WEB_LOGOUT</summary>

`WEB_LOGOUT`

* **Severity**\
  `INFO`
* **Description**\
  A Web UI user logged out.
* **Format String**\
  `"logged out from Web UI"`

</details>


# Alarm Types

```
alarm-type
    cdb-offload-threshold-too-low
    certificate-expiration
    ha-alarm
        ha-node-down-alarm
            ha-primary-down
            ha-secondary-down
    ha-raft-quorum-lost
    memory-management-action-triggered
    ncs-cluster-alarm
        cluster-subscriber-failure
    ncs-dev-manager-alarm
        abort-error
        auto-configure-failed
        commit-through-queue-blocked
        commit-through-queue-failed
        commit-through-queue-failed-transiently
        commit-through-queue-rollback-failed
        configuration-error
        connection-failure
        final-commit-error
        missing-transaction-id
        ned-live-tree-connection-failure
        out-of-sync
        revision-error
    ncs-package-alarm
        package-load-failure
        package-operation-failure
    ncs-service-manager-alarm
        service-activation-failure
    ncs-snmp-notification-receiver-alarm
        receiver-configuration-error
    time-violation-alarm
        transaction-lock-time-violation
```

## Alarm Type Descriptions

<details>

<summary>abort-error</summary>

`abort-error`

* **Initial Perceived Severity**\
  major
* **Description**\
  An error happened while aborting or reverting a transaction. Device's configuration is likely to be inconsistent with the NCS CDB.
* **Recommended Action**\
  Inspect the configuration difference with compare-config, resolve conflicts with sync-from or sync-to if any.
* **Clear Condition(s)**\
  If NCS achieves sync with the device, or receives a transaction id for a netconf session towards the device, the alarm is cleared.
* **Alarm Message(s)**
  * `Device {dev} is locked`
  * `Device {dev} is southbound locked`
  * `abort error`

</details>

<details>

<summary>alarm-type</summary>

`alarm-type`

* **Description**\
  Base identity for alarm types. A unique identification of the fault, not including the managed object. Alarm types are used to identify if alarms indicate the same problem or not, for lookup into external alarm documentation, etc. Different managed object types and instances can share alarm types. If the same managed object reports the same alarm type, it is to be considered to be the same alarm. The alarm type is a simplification of the different X.733 and 3GPP alarm IRP alarm correlation mechanisms and it allows for hierarchical extensions.\
  A 'specific-problem' can be used in addition to the alarm type in order to have different alarm types based on information not known at design-time, such as values in textual SNMP Notification varbinds.

</details>

<details>

<summary>auto-configure-failed</summary>

`auto-configure-failed`

* **Initial Perceived Severity**\
  warning
* **Description**\
  Device auto-configure exhausted its retry attempts trying to connect and sync the device.
* **Recommended Action**\
  Make sure that NCS can connect to the device and then sync the configuration.
* **Clear Condition(s)**\
  If NCS achieves sync with the device, the alarm is cleared.
* **Alarm Message(s)**
  * `Auto-configure has exhausted its retry attempts`

</details>

<details>

<summary>cdb-offload-threshold-too-low</summary>

`cdb-offload-threshold-too-low`

* **Initial Perceived Severity**\
  warning
* **Description**\
  The CDB offload threshold configuration is set too low, causing the CDB memory footprint to reach the threshold even when there is no offloadable data present in the memory.
* **Recommended Action**\
  If system memory is sufficient, increase the threshold value, otherwise increase the system memory capacity.
* **Clear Condition(s)**\
  This alarm is cleared when CDB offload can lower the CDB memory footprint below the configured threshold value.
* **Alarm Message(s)**
  * `CDB offload threshold is too low`

</details>

<details>

<summary>certificate-expiration</summary>

`certificate-expiration`

* **Description**\
  The certificate is nearing its expiry or has already expired. The severity depends on the time left to expiry, it ranges from warning to critical.
* **Recommended Action**\
  Replace certificate.
* **Clear Condition(s)**\
  This alarm is cleared when the certificate is no longer loaded.
* **Alarm Message(s)**
  * `Certificate expires in less than {days} day(s)`
  * `Certificate has expired`

</details>

<details>

<summary>cluster-subscriber-failure</summary>

`cluster-subscriber-failure`

* **Initial Perceived Severity**\
  critical
* **Description**\
  Failure to establish a notification subscription towards a remote node.
* **Recommended Action**\
  Verify IP connectivity between cluster nodes.
* **Clear Condition(s)**\
  This alarm is cleared if NCS succeeds to establish a subscription towards the remote node, or when the subscription is explicitly stopped.
* **Alarm Message(s)**
  * `Failed to establish netconf notification subscription to node ~s, stream ~s`
  * `Commit queue items with remote nodes will not receive required event notifications.`

</details>

<details>

<summary>commit-through-queue-blocked</summary>

`commit-through-queue-blocked`

* **Initial Perceived Severity**\
  warning
* **Description**\
  A commit was queued behind a queue item waiting to be able to connect to one of its devices. This is potentially dangerous since one unreachable device can potentially fill up the commit queue indefinitely.
* **Clear Condition(s)**\
  An alarm raised due to a transient error will be cleared when NCS is able to reconnect to the device.
* **Alarm Message(s)**
  * `Commit queue item ~p is blocked because item ~p cannot connect to ~s`

</details>

<details>

<summary>commit-through-queue-failed</summary>

`commit-through-queue-failed`

* **Initial Perceived Severity**\
  critical
* **Description**\
  A queued commit failed.
* **Recommended Action**\
  Resolve with rollback if possible.
* **Clear Condition(s)**\
  This alarm is not cleared.
* **Alarm Message(s)**
  * `Failed to authenticate towards device {device}: {reason}`
  * `Device {dev} is locked`
  * `{Reason}`
  * `Device {dev} is southbound locked`
  * `Commit queue item {CqId} rollback invoked`
  * `Commit queue item {CqId} has failed: Operation failed because: inconsistent database`
  * `Remote commit queue item ~p cannot be unlocked: cluster node not configured correctly`

</details>

<details>

<summary>commit-through-queue-failed-transiently</summary>

`commit-through-queue-failed-transiently`

* **Initial Perceived Severity**\
  critical
* **Description**\
  A queued commit failed as it exhausted its retry attempts on transient errors.
* **Recommended Action**\
  Resolve with rollback if possible.
* **Clear Condition(s)**\
  This alarm is not cleared.
* **Alarm Message(s)**
  * `Failed to connect to device {dev}: {reason}`
  * `Connection to {dev} timed out`
  * `Failed to authenticate towards device {device}: {reason}`
  * `The configuration database is locked for device {dev}: {reason}`
  * `the configuration database is locked by session {id} {identification}`
  * `the configuration database is locked by session {id} {identification}`
  * `{Dev}: Device is locked in a {Op} operation by session {session-id}`
  * `resource denied`
  * `Commit queue item {CqId} rollback invoked`
  * `Commit queue item {CqId} has failed: Operation failed because: inconsistent database`
  * `Remote commit queue item ~p cannot be unlocked: cluster node not configured correctly`

</details>

<details>

<summary>commit-through-queue-rollback-failed</summary>

`commit-through-queue-rollback-failed`

* **Initial Perceived Severity**\
  critical
* **Description**\
  Rollback of a commit-queue item failed.
* **Recommended Action**\
  Investigate the status of the device and resolve the situation by issuing the appropriate action, i.e., service redeploy or a sync operation.
* **Clear Condition(s)**\
  This alarm is not cleared.
* **Alarm Message(s)**
  * `{Reason}`

</details>

<details>

<summary>configuration-error</summary>

`configuration-error`

* **Initial Perceived Severity**\
  critical
* **Description**\
  Invalid configuration of NCS managed device, NCS cannot recognize parameters needed to connect to device.
* **Recommended Action**\
  Verify that the configuration parameters defined in tailf-ncs-devices.yang submodule are consistent for this device.
* **Clear Condition(s)**\
  The alarm is cleared when NCS reads the configuration parameters for the device, and is raised again if the parameters are invalid.
* **Alarm Message(s)**
  * `Failed to resolve IP address for {dev}`
  * `the configuration database is locked by session {id} {identification}`
  * `{Reason}`
  * `Resource {resource} doesn't exist`

</details>

<details>

<summary>connection-failure</summary>

`connection-failure`

* **Initial Perceived Severity**\
  major
* **Description**\
  NCS failed to connect to a managed device before the timeout expired.
* **Recommended Action**\
  Verify address, port, authentication, check that the device is up and running. If the error occurs intermittently, increase connect-timeout.
* **Clear Condition(s)**\
  If NCS successfully reconnects to the device, the alarm is cleared.
* **Alarm Message(s)**
  * `The connection to {dev} was closed`
  * `Failed to connect to device {dev}: {reason}`

</details>

<details>

<summary>final-commit-error</summary>

`final-commit-error`

* **Initial Perceived Severity**\
  critical
* **Description**\
  A managed device validated a configuration change, but failed to commit. When this happens, NCS and the device are out of sync.
* **Recommended Action**\
  Reconcile by comparing and sync-from or sync-to.
* **Clear Condition(s)**\
  If NCS achieves sync with the device, the alarm is cleared.
* **Alarm Message(s)**
  * `The connection to {dev} was closed`
  * `External error in the NED implementation for device {dev}: {reason}`
  * `Internal error in the NED NCS framework affecting device {dev}: {reason}`

</details>

<details>

<summary>ha-alarm</summary>

`ha-alarm`

* **Description**\
  Base type for all alarms related to high availablity. This is never reported, sub-identities for the specific high availability alarms are used in the alarms.

</details>

<details>

<summary>ha-node-down-alarm</summary>

`ha-node-down-alarm`

* **Description**\
  Base type for all alarms related to nodes going down in high availablity. This is never reported, sub-identities for the specific node down alarms are used in the alarms.

</details>

<details>

<summary>ha-primary-down</summary>

`ha-primary-down`

* **Initial Perceived Severity**\
  critical
* **Description**\
  The node lost the connection to the primary node.
* **Recommended Action**\
  Make sure the HA cluster is operational, investigate why the primary went down and bring it up again.
* **Clear Condition(s)**\
  This alarm is automatically cleared when the node is reconnected to the HA cluster.
* **Alarm Message(s)**
  * `Lost connection to primary due to: Primary closed connection`
  * `Lost connection to primary due to: Tick timeout`
  * `Lost connection to primary due to: code {Code}`

</details>

<details>

<summary>ha-raft-quorum-lost</summary>

`ha-raft-quorum-lost`

* **Initial Perceived Severity**\
  critical
* **Description**\
  HA Raft leader has lost quorum and cannot reach a majority of cluster nodes. The HA Raft subsystem will be restarted, which will demote the leader and the node will be disabled until quorum can be established again.
* **Recommended Action**\
  Urgently investigate cluster connectivity. Check if cluster nodes are running and are reachable. Verify network connectivity between nodes. Review '/ha-raft/status' for details on which nodes are unreachable.
* **Clear Condition(s)**\
  This alarm is never automatically cleared and has to be cleared manually when the HA cluster has been restored.
* **Alarm Message(s)**
  * `HA Raft quorum lost: Leader cannot reach majority of nodes.`

</details>

<details>

<summary>ha-secondary-down</summary>

`ha-secondary-down`

* **Initial Perceived Severity**\
  critical
* **Description**\
  The node lost the connection to a secondary node.
* **Recommended Action**\
  Investigate why the secondary node went down, fix the connectivity issue and reconnect the secondary to the HA cluster.
* **Clear Condition(s)**\
  This alarm is cleared when the secondary node is reconnected to the HA cluster.
* **Alarm Message(s)**
  * `Lost connection to secondary`

</details>

<details>

<summary>memory-management-action-triggered</summary>

`memory-management-action-triggered`

* **Initial Perceived Severity**\
  critical
* **Description**\
  A memory management action has triggered. The memory conditions have reached the threshold set in the memory management action, triggering it.
* **Recommended Action**\
  A memory actions threshold being reached suggests high memory usage.
* **Clear Condition(s)**\
  This alarm is cleared when the memory conditions go back below the threshold set in the memory management action.
* **Alarm Message(s)**
  * `Memory Management action triggered.`

</details>

<details>

<summary>missing-transaction-id</summary>

`missing-transaction-id`

* **Initial Perceived Severity**\
  warning
* **Description**\
  A device announced in its NETCONF hello message that it supports the transaction-id as defined in <http://tail-f.com/yang/netconf-monitoring>. However when NCS tries to read the transaction-id no data is returned. The NCS check-sync feature will not work. This is usually a case of misconfigured NACM rules on the managed device.
* **Recommended Action**\
  Verify NACM rules on the concerned device.
* **Clear Condition(s)**\
  If NCS successfully reads a transaction id for which it had previously failed to do so, the alarm is cleared.
* **Alarm Message(s)**
  * `{Reason}`

</details>

<details>

<summary>ncs-cluster-alarm</summary>

`ncs-cluster-alarm`

* **Description**\
  Base type for all alarms related to cluster. This is never reported, sub-identities for the specific cluster alarms are used in the alarms.

</details>

<details>

<summary>ncs-dev-manager-alarm</summary>

`ncs-dev-manager-alarm`

* **Description**\
  Base type for all alarms related to the device manager This is never reported, sub-identities for the specific device alarms are used in the alarms.

</details>

<details>

<summary>ncs-package-alarm</summary>

`ncs-package-alarm`

* **Description**\
  Base type for all alarms related to packages. This is never reported, sub-identities for the specific package alarms are used in the alarms.

</details>

<details>

<summary>ncs-service-manager-alarm</summary>

`ncs-service-manager-alarm`

* **Description**\
  Base type for all alarms related to the service manager This is never reported, sub-identities for the specific service alarms are used in the alarms.

</details>

<details>

<summary>ncs-snmp-notification-receiver-alarm</summary>

`ncs-snmp-notification-receiver-alarm`

* **Description**\
  Base type for SNMP notification receiver Alarms. This is never reported, sub-identities for specific SNMP notification receiver alarms are used in the alarms.

</details>

<details>

<summary>ned-live-tree-connection-failure</summary>

`ned-live-tree-connection-failure`

* **Initial Perceived Severity**\
  major
* **Description**\
  NCS failed to connect to a managed device using one of the optional live-status-protocol NEDs.
* **Recommended Action**\
  Verify the configuration of the optional NEDs. If the error occurs intermittently, increase connect-timeout.
* **Clear Condition(s)**\
  If NCS successfully reconnects to the managed device, the alarm is cleared.
* **Alarm Message(s)**
  * `The connection to {dev} was closed`
  * `Failed to connect to device {dev}: {reason}`

</details>

<details>

<summary>out-of-sync</summary>

`out-of-sync`

* **Initial Perceived Severity**\
  major
* **Description**\
  A managed device is out of sync with NCS. Usually it means that the device has been configured out of band from NCS point of view.
* **Recommended Action**\
  Inspect the difference with compare-config, reconcile by invoking sync-from or sync-to.
* **Clear Condition(s)**\
  If NCS achieves sync with the device, the alarm is cleared.
* **Alarm Message(s)**
  * `Device {dev} is out of sync`
  * `Out of sync due to no-networking or failed commit-queue commits.`
  * `got: ~s expected: ~s.`

</details>

<details>

<summary>package-load-failure</summary>

`package-load-failure`

* **Initial Perceived Severity**\
  critical
* **Description**\
  NCS failed to load a package.
* **Recommended Action**\
  Check the package for the reason.
* **Clear Condition(s)**\
  If NCS successfully loads a package for which an alarm was previously raised, it will be cleared.
* **Alarm Message(s)**
  * `failed to open file {file}: {str}`
  * `Specific to the concerned package.`

</details>

<details>

<summary>package-operation-failure</summary>

`package-operation-failure`

* **Initial Perceived Severity**\
  critical
* **Description**\
  A package has some problem with its operation.
* **Recommended Action**\
  Check the package for the reason.
* **Clear Condition(s)**\
  This alarm is not cleared.

</details>

<details>

<summary>receiver-configuration-error</summary>

`receiver-configuration-error`

* **Initial Perceived Severity**\
  major
* **Description**\
  The snmp-notification-receiver could not setup its configuration, either at startup or when reconfigured. SNMP notifications will now be missed.
* **Recommended Action**\
  Check the error-message and change the configuration.
* **Clear Condition(s)**\
  This alarm will be cleared when the NCS is configured to successfully receive SNMP notifications
* **Alarm Message(s)**
  * `Configuration has errors.`

</details>

<details>

<summary>revision-error</summary>

`revision-error`

* **Initial Perceived Severity**\
  major
* **Description**\
  A managed device arrived with a known module, but too new revision.
* **Recommended Action**\
  Upgrade the Device NED using the new YANG revision in order to use the new features in the device.
* **Clear Condition(s)**\
  If all device yang modules are supported by NCS, the alarm is cleared.
* **Alarm Message(s)**
  * `The device has YANG module revisions not supported by NCS. Use the /devices/device/check-yang-modules action to check which modules that are not compatible.`

</details>

<details>

<summary>service-activation-failure</summary>

`service-activation-failure`

* **Initial Perceived Severity**\
  critical
* **Description**\
  A service failed during re-deploy.
* **Recommended Action**\
  Corrective action and another re-deploy is needed.
* **Clear Condition(s)**\
  If the service is successfully redeployed, the alarm is cleared.
* **Alarm Message(s)**
  * `Multiple device errors: {str}`

</details>

<details>

<summary>time-violation-alarm</summary>

`time-violation-alarm`

* **Description**\
  Base type for all alarms related to time violations. This is never reported, sub-identities for the specific time violation alarms are used in the alarms.

</details>

<details>

<summary>transaction-lock-time-violation</summary>

`transaction-lock-time-violation`

* **Initial Perceived Severity**\
  warning
* **Description**\
  The transaction lock time exceeded its threshold and might be stuck in the critical section. This threshold is configured in /ncs-config/transaction-lock-time-violation-alarm/timeout.
* **Recommended Action**\
  Investigate if the transaction is stuck and possibly interrupt it by closing the user session which it is attached to.
* **Clear Condition(s)**\
  This alarm is cleared when the transaction has finished.
* **Alarm Message(s)**
  * `Transaction lock time exceeded threshold.`

</details>


# Package Management

Perform package management tasks.

All user code that needs to run in NSO must be part of a package. A package is basically a directory of files with a fixed file structure or a tar archive with the same directory layout. A package consists of code, YANG modules, etc., that are needed to add an application or function to NSO. Packages are a controlled way to manage loading and versions of custom applications.

Network Element Drivers (NEDs) are also packages. Each NED allows NSO to manage a network device of a specific type. Except for third-party YANG NED packages which do not contain a YANG device model by default (and must be downloaded and fixed before adding to the package), a NED typically contains a device YANG model and the code, specifying how NSO should connect to the device. For NETCONF devices, NSO includes built-in tools to help you build a NED, as described in [NED Administration](/guides/administration/management/ned-administration), that you can use if needed. Otherwise, a third-party YANG NED, if available, should be used instead. Vendors, in some cases, provide the required YANG device models but not the entire NED. In practice, all NSO instances use at least one NED. The set of used NED packages depends on the number of different device types the NSO manages.

When NSO starts, it searches for packages to load. The `ncs.conf` parameter `/ncs-config/load-path` defines a list of directories. At initial startup, NSO searches these directories for packages and copies the packages to a private directory tree in the directory defined by the `/ncs-config/state-dir` parameter in `ncs.conf`, and loads and starts all the packages found. On subsequent startups, NSO will by default only load and start the copied packages. The purpose of this procedure is to make it possible to reliably load new or updated packages while NSO is running, with a fallback to the previously existing version of the packages if the reload should fail.

The package management workflow depends on the NSO installation type:

* In a [Local Install](/guides/administration/installation-and-deployment/local-install), used for development, evaluation, and examples, you manage packages directly in one of the configured load-path directories, often the runtime `packages` directory. In this environment it is normal to edit package directories in place, copy directories, or create and remove symbolic links, and then use `packages reload`, `packages add`, or `packages package <name> redeploy`.
* In a [System Install](/guides/administration/installation-and-deployment/system-install), used for production, you manage pre-built packages with the `software packages` actions. These actions stage packages under `/opt/ncs/packages` and install or deinstall the active package set in `/var/opt/ncs/packages`. Symbolic links may exist there, but they are managed by NSO as part of this workflow and should not be created or changed manually as the normal package management procedure.

In a System Install of NSO, the active packages are available in the `packages` subdirectory of the run directory, i.e. by default `/var/opt/ncs/packages`, and the private directory tree is created in the `state` subdirectory, i.e. by default `/var/opt/ncs/state`.

## Loading Packages <a href="#ug.package_mgmt.loading" id="ug.package_mgmt.loading"></a>

Loading of new or updated packages (as well as removal of packages that should no longer be used) can be requested via the `reload` action - from the NSO CLI:

```bash
admin@ncs# packages reload
reload-result {
    package cisco-ios
    result true
}
```

This request makes NSO copy all packages found in the load path to a temporary version of its private directory, and load the packages from this directory. If the loading is successful, this temporary directory will be made permanent, otherwise, the temporary directory is removed and NSO continues to use the previous version of the packages. Thus when updating packages, always update the version in the load path, and request that NSO does the reload via this action.

If the package changes include modified, added, or deleted `.fxs` files or `.ccl` files, NSO needs to run a data model upgrade procedure, also called a CDB upgrade. NSO provides a `dry-run` option to `packages reload` action to test the upgrade without committing the changes. Using a reload dry-run, you can tell if a CDB upgrade is needed or not.

The `report all-schema-changes` option of the reload action instructs NSO to produce a report of how the current data model schema is being changed. Combined with a dry run, the report allows you to verify the modifications introduced with the new versions of the packages before actually performing the upgrade.

When reloading packages, NSO will give a warning when the upgrade looks suspicious, i.e., may break some functionality. Note that this is not a strict upgrade validation, but only intended as a hint to the NSO administrator early in the upgrade process that something might be wrong. Currently, the following scenarios will trigger the warnings:

* One or more namespaces are removed by the upgrade. The consequence of this is all data belonging to this namespace is permanently deleted from CDB upon upgrade. This may be intended in some scenarios, in which case it is advised to proceed with overriding warnings as described below.
* There are source `.java` files found in the package, but no matching `.class` files in the jars loaded by NSO. This likely means that the package has not been compiled.
* There are matching `.class` files with modification time older than the source files, which hints that the source has been modified since the last time the package was compiled. This likely means that the package was not re-compiled the last time the source code was changed.

If a warning has been triggered it is a strong recommendation to fix the root cause. If all of the warnings are intended, it is possible to proceed with `packages reload force` command.

In some specific situations, upgrading a package with newly added custom validation points in the data model may produce an error similar to `no registration found for callpoint NEW-VALIDATION/validate` or simply `application communication failure`, resulting in an aborted upgrade. See [New Validation Points](https://nso-docs.cisco.com/guides/administration/management/pages/FxpCNgv5QKnfWrJw4nXf#cdb.upgrade-add-vp) on how to proceed.

In some cases, we may want NSO to do the same operation as the `reload` action at NSO startup, i.e. copy all packages from the load path before loading, even though the private directory copy already exists. This can be achieved in the following ways:

* Setting the shell environment variable `$NCS_RELOAD_PACKAGES` to `true`. This will make NSO do the copy from the load path on every startup, as long as the environment variable is set. In a System Install, NSO is typically started as a `systemd` system service, and `NCS_RELOAD_PACKAGES=true` can be set in `/etc/ncs/ncs.systemd.conf` temporarily to reload the packages.
* Giving the option `--with-package-reload` to the `ncs` command when starting NSO. This will make NSO do the copy from the load path on this particular startup, without affecting the behavior on subsequent startups.
* If warnings are encountered when reloading packages at startup using one of the options above, the recommended way forward is to fix the root cause as indicated by the warnings as mentioned before. If the intention is to proceed with the upgrade without fixing the underlying cause for the warnings, it is possible to force the upgrade using `NCS_RELOAD_PACKAGES`=`force` environment variable or `--with-package-reload-force` option.

Always use one of these methods when upgrading to a new version of NSO in an existing directory structure, to make sure that new packages are loaded together with the other parts of the new system.

### Open Transactions During Upgrade

For a data model upgrade, except for a dry run, special considerations must be made, depending on the chosen upgrade mode.

By default, all transactions are closed and new transactions are not allowed. Users having CLI sessions in configure mode must exit to operational mode. This means that starting a new management session, such as a CLI or SSH connection to the NSO, will also fail, producing an error that the node is in upgrade mode.

Here, NSO will wait up to 10 seconds for the existing transactions to close. If there are still open transactions at the end of this period, the upgrade will be canceled and the reload operation will fail. The `max-wait-time` and `timeout-action` parameters to the action can modify this behavior. For example, to wait for up to 30 seconds, and forcibly terminate any transactions that still remain open after this period, invoke the action as:

```cli
admin@ncs# packages reload max-wait-time 30 timeout-action kill
```

For in-service upgrades, when `packages reload` and `packages ha sync and-reload` are called with the `optimistic` parameter, NSO keeps accepting and processing requests on the northbound interfaces. The in-flight transactions are upgraded (rebased) to the new data model as part of the upgrade process. However, rebasing, and subsequently the transaction, may fail if the transaction is incompatible with the new schema.

The in-service upgrade may be combined with a `backup` switch, which takes an NSO backup before the upgrade starts.

Examples:

```cli
admin@ncs# packages reload optimistic backup
admin@ncs# packages ha sync and-reload { optimistic backup wait-commit-queue-empty }
```

Regardless of the mode, if there are ongoing commit queue items, and the `wait-commit-queue-empty` parameter is supplied, the upgrade will wait for the `max-wait-time` for items to finish before proceeding with the reload. If new transactions are not allowed and one of the queue items fails with `rollback-on-error` option set, the commit queue's rollback will also fail, and the queue item will be locked. In this case, the reload will be canceled. A manual investigation of the failure is needed in order to proceed with the reload.

In case there are no changes to `.fxs` or .`ccl` files, the reload can be carried out without the data model upgrade procedure, and these parameters are ignored.

## Redeploying Packages <a href="#ug.package_mgmt.redeploying" id="ug.package_mgmt.redeploying"></a>

If it is known in advance that there were no data model changes, i.e. none of the `.fxs` or `.ccl` files changed, and none of the shared JARs changed in a Java package, and the declaration of the components in the `package-meta-data.xml` is unchanged, then it is possible to do a lightweight package upgrade, called package redeploy. Package redeploy only loads the specified package, unlike packages reload which loads all of the packages found in the load-path.

```bash
admin@ncs# packages package mserv redeploy
result true
```

Redeploying a package allows you to reload updated or load new templates, reload private JARs for a Java package, or reload the Python code which is a part of this package. Only the changed part of the package will be reloaded, e.g. if there were no changes to Python code, but only templates, then the Python VM will not be restarted, but only templates reloaded. The upgrade is not seamless however as the old templates will be unloaded for a short while before the new ones are loaded, so any user of the template during this period of time will fail; the same applies to changed Java or Python code. It is hence the responsibility of the user to make sure that the services or other code provided by the package is unused while it is being redeployed.

The `package redeploy` will return `true` if the package's resulting status after the redeploy is `up`. Consequently, if the result of the action is `false`, then it is advised to check the operational status of the package in the package list.

```bash
admin@ncs# show packages package mserv oper-status
oper-status file-load-error
oper-status error-info "template3.xml:2 Unknown servicepoint: templ42-servicepoint"
```

## Adding NED Packages <a href="#ug.package_mgmt.ned_package_add" id="ug.package_mgmt.ned_package_add"></a>

Unlike a full `packages reload` operation, new NED packages can be loaded into the system without disrupting existing transactions. This is only possible for new packages, since these packages don't yet have any instance data.

{% hint style="warning" %}
Loading additional or larger NED schemas increases memory use in NSO and, when the Java VM is running, its heap use. Before adding NED packages, ensure that the host or container has sufficient memory and that the configured Java VM heap is large enough. See [Java VM Heap Size](https://nso-docs.cisco.com/guides/administration/management/pages/INyCuLO2h06pPPZrhcFY#ncs.development.scaling.memory.jvm) for sizing and configuration guidance.
{% endhint %}

The operation is performed through the `/packages/add` action. No additional input is necessary. The operation scans all the load-paths for any new NED packages and also verifies the existing packages are still present. If packages are modified or deleted, the operation will fail.

Each NED package defines `ned-id`, an identifier that is used in selecting the NED for each managed device. A new NED package is therefore a package with a ned-id value that is not already in use.

In addition, the system imposes some additional constraints, so it is not always possible to add just any arbitrary NED. In particular, NED packages can also contain one or more shared data models, such as NED settings or operational data for private use by the NED, that are not specific to each version of NED package but rather shared between all versions. These are typically placed outside any mount point (device-specific data model), extending the NSO schema directly. So, if a NED defines schema nodes outside any mount point, there must be no changes to these nodes if they already exist.

Adding a NED package with a modified shared data model is therefore not allowed and all shared data models are verified to be identical before a NED package can be added. If they are not, the `/packages/add` action will fail and you will have to use the `/packages/reload` command.

```bash
admin@ncs# packages add
add-result {
    package router-nc-1.1
    result true
}
```

The command returns `true` if the package's resulting status after deployment is `up`. Likewise, if the result for a package is `false`, then the package was added but its code has not started successfully and you should check the operational status of the package with the `show packages package <PKG> oper-status` command for additional information. You may then use the `/packages/package/redeploy` action to retry deploying the package's code, once you have corrected the error.

{% hint style="info" %}
In a high-availability setup, you can perform this same operation on all the nodes in the cluster with a single `packages ha sync and-add` command.
{% endhint %}

## Managing Packages on System Install <a href="#ug.package_mgmt.managing" id="ug.package_mgmt.managing"></a>

{% hint style="warning" %}
Applies to System Install. In production, manage pre-built packages with the `software packages` actions. Do not create or remove symbolic links manually in `/var/opt/ncs/packages`; that directory is managed by NSO as part of the `software packages` workflow.
{% endhint %}

In a System Install of NSO, management of pre-built packages is supported through a number of actions. This support is not available in a Local Install, since it is dependent on the directory structure created by the System Install. The supported workflow is to make the package available under `/opt/ncs/packages`, install or deinstall it with the `software packages` actions, and then activate the change with `packages reload` or, in a high-availability setup, `packages ha sync and-reload`. Please refer to the YANG submodule `$NCS_DIR/src/ncs/yang/tailf-ncs-software.yang` for the full details of the functionality described in this section.

Both package reload actions support `optimistic` mode for in-service upgrade, which can be combined with `backup` to create an NSO backup before the upgrade starts.

For the full production package upgrade procedure on System Install, including backup recommendations, single-node upgrades, and high-availability upgrades with `packages ha sync and-reload`, see [Package Upgrade](/guides/administration/installation-and-deployment/upgrade-nso#d5e7083).

### Actions

Actions are provided to list local packages, to fetch packages from the file system, and to install or deinstall packages:

* `software packages list [...]`: List local packages, categorized into loaded, installed, and installable. The listing can be restricted to only one of the categories - otherwise, each package listed will include the category for the package.
* `software packages fetch package-from-file <file>`: Fetch a package by copying it from the file system, making it installable.
* `software packages install package <package-name> [...]`: Install a package, making it available for loading via the `packages reload` action, or via a system restart with package reload. The action ensures that only one version of the package is installed - if any version of the package is installed already, the `replace-existing` option can be used to deinstall it before proceeding with the installation.
* `software packages deinstall package <package-name>`: Deinstall a package, i.e. remove it from the set of packages available for loading.

There is also an `upload` action that can be used via NETCONF or RESTCONF to upload a package from the local host to the NSO host, making it installable there. It is not feasible to use in the CLI or Web UI, since the actual package file contents is a parameter for the action. It is also not suitable for very large (more than a few megabytes) packages, since the processing of action parameters is not designed to deal with very large values, and there is a significant memory overhead in the processing of such values.

The states reported by `software packages list` are:

* `installable`: The package is present in `/opt/ncs/packages` and can be installed, but it is not yet part of the active package set.
* `installed`: The package is present in `/var/opt/ncs/packages`, but it is not currently loaded by the running NSO process.
* `loaded`: The package is currently loaded by NSO.

For a single-node System Install, the normal package workflow is:

1. Copy or build the package on the NSO host, typically as a `.tar.gz` file.
2. Make it installable with `software packages fetch`, or use `upload` over NETCONF or RESTCONF.
3. Install it with `software packages install`, using `replace-existing` when replacing an already installed version.
4. Activate the change with `packages reload`.

For example:

```bash
admin@ncs# software packages fetch package-from-file /tmp/package-store/router-nc-1.0.2.tar.gz
admin@ncs# software packages install package router-nc-1.0.2 replace-existing
admin@ncs# packages reload
```

In a high-availability System Install, first manage the package set on the primary node with the `software packages` actions and then synchronize and activate it on the other nodes with `packages ha sync and-reload`. For example:

```bash
primary@node1# software packages fetch package-from-file /tmp/package-store/dummy-1.1.tar.gz
primary@node1# software packages install package dummy-1.1 replace-existing
primary@node1# packages ha sync and-reload { wait-commit-queue-empty }
```

If the change only adds new NED packages, `packages ha sync and-add` can be used instead of `and-reload`, as described in [Adding NED Packages](#ug.package_mgmt.ned_package_add).

Example implementations of this System Install workflow are provided in [examples.ncs/high-availability/upgrade-basic/upgrade\_pkgs\_sys.sh](https://github.com/NSO-developer/nso-examples/tree/6.7/high-availability/upgrade-cluster/upgrade_pkgs_sys.sh) and [examples.ncs/high-availability/upgrade-cluster/upgrade\_pkgs\_sys.sh](https://github.com/NSO-developer/nso-examples/tree/6.7/high-availability/upgrade-cluster/upgrade_pkgs_sys.sh).

## Local Install Example <a href="#ncsnwe.admin.packages" id="ncsnwe.admin.packages"></a>

{% hint style="info" %}
Applies to Local Install. The following example manages packages directly in a runtime `packages` directory. That is appropriate for Local Install development and example environments. On System Install, use the `software packages` actions described above instead of manually managing symbolic links in `$NCS_RUN_DIR/packages`.
{% endhint %}

NSO Packages contain data models and code for a specific function. It might be NED for a specific device, a service application like MPLS VPN, a WebUI customization package, etc. Packages can be added, removed, and upgraded in run-time. A common task is to add a package to NSO to support a new device type or upgrade an existing package when the device is upgraded.

(We assume you have the example up and running from the previous section). Currently installed packages can be viewed with the following command:

```bash
admin@ncs# show packages
packages package cisco-ios
 package-version 1.0
 description     "NED package for Cisco IOS"
 ncs-min-version [ 6.4 ]
 directory       ./state/packages-in-use/1/cisco-ios-netsim-cli-1.0
 component upgrade-ned-id
  upgrade java-class-name com.tailf.packages.ned.ios.UpgradeNedId
 component cisco-ios
  ned cli ned-id  cisco-ios-netsim-cli-1.0
  ned cli java-class-name com.tailf.packages.ned.ios.IOSNedCli
  ned device vendor Cisco
NAME      VALUE
---------------------
show-tag  interface

 oper-status up
```

So the above command shows that NSO currently has one package, the NED for Cisco IOS.

NSO reads global configuration parameters from `ncs.conf`. More on NSO configuration later in this guide. In this Local Install example, the configuration tells NSO to look for packages in a `packages` directory where NSO was started. Using the [examples.ncs/device-management/simulated-devices](https://github.com/NSO-developer/nso-examples/tree/6.7/device-management/simulated-devices) example to demonstrate:

```bash
$ pwd
examples.ncs/device-management/simulated-devices
$ NONINTERACTIVE=1 ./demo.sh
$ ls packages/
cisco-ios-netsim-cli-1.0
$ ls packages/cisco-ios-netsim-cli-1.0
doc
load-dir
netsim
package-meta-data.xml
private-jar
shared-jar
src
```

As seen above a package is a defined file structure with data models, code, and documentation. NSO comes with a few ready-made example packages: `$NCS_DIR/packages/`. Also, there is a library of packages available from Tail-f, especially for supporting specific devices.

### Adding and Upgrading a Package on Local Install <a href="#d5e7809" id="d5e7809"></a>

Assume you would like to add support for Nexus devices to the example. Nexus devices have different data models and another CLI flavor. There is an example NED package for that in `$NCS_DIR/examples.ncs/common/packages/cisco-nx-netsim-cli-1.0`.

We can keep NSO running all the time, but we will stop the network simulator to add the Nexus devices to the simulator.

```bash
$ ncs-netsim stop
```

Because this is a Local Install example, add the Nexus package to the runtime `packages` directory by creating a symbolic link (or copy):

```bash
$ cd $NCS_DIR/examples.ncs/device-management/simulated-devices/packages
$ ln -s $NCS_DIR/examples.ncs/common/packages/cisco-nx-netsim-cli-1.0 cisco-nx-netsim-cli-1.0
$ ls -l
...
cisco-nx-netsim-cli-1.0 -> $NCS_DIR/examples.ncs/common/packages/cisco-nx-netsim-cli-1.0
```

The package is now in place, but until we tell NSO to look for package changes nothing happens:

```bash
  admin@ncs# show packages packages package
  cisco-ios ...  admin@ncs# packages reload

>>> System upgrade is starting.
>>> Sessions in configure mode must exit to operational mode.
>>> No configuration changes can be performed until upgrade has
completed.
>>> System upgrade has completed successfully.
reload-result {
    package cisco-ios
    result true
}
reload-result {
    package cisco-nx
    result true
}
```

So after the `packages reload` operation NSO also knows about Nexus devices. The reload operation also takes any changes to existing packages into account. The data store is automatically upgraded to cater to any changes like added attributes to existing configuration data.

### Simulating the New Device <a href="#d5e7826" id="d5e7826"></a>

```bash
$ ncs-netsim add-to-network cisco-nx-netsim-cli-1.0 2 n
$ ncs-netsim list
ncs-netsim list for examples.ncs/device-management/simulated-devices/netsim

name=c0 ...
name=c1 ...
name=c2 ...
name=n0 ...
name=n1 ...


$ ncs-netsim start
DEVICE c0 OK STARTED
DEVICE c1 OK STARTED
DEVICE c2 OK STARTED
DEVICE n0 OK STARTED
DEVICE n1 OK STARTED
$ ncs-netsim cli-c n0
n0#show running-config
no feature ssh
no feature telnet
fex 101
 pinning max-links 1
!
fex 102
 pinning max-links 1
!
nexus:vlan 1
!
...
```

### Adding the New Devices to NSO <a href="#d5e7835" id="d5e7835"></a>

We can now add these Nexus devices to NSO according to the below sequence:

```bash
admin@ncs(config)# devices device n0 device-type cli ned-id cisco-nx-netsim-cli-1.0
admin@ncs(config-device-n0)# port 10025
admin@ncs(config-device-n0)# address 127.0.0.1
admin@ncs(config-device-n0)# authgroup default
admin@ncs(config-device-n0)# state admin-state unlocked
admin@ncs(config-device-n0)# commit
admin@ncs(config-device-n0)# top
admin@ncs(config)# devices device n0 sync-from
result true
```


# High Availability

Implement redundancy in your deployment using High Availability (HA) setup.

As a single NSO node can fail or lose network connectivity, you can configure multiple nodes in a highly available (HA) setup, which replicates the CDB configuration and operational data across participating nodes. It allows the system to continue functioning even when some nodes are inoperable.

The replication architecture is that of one active primary and a number of secondaries. This means all configuration write operations must occur on the primary, which distributes the updates to the secondaries.

Operational data in the CDB may be replicated or not based on the `tailf:persistent` statement in the data model. If replicated, operational data writes can only be performed on the primary, whereas non-replicated operational data can also be written on the secondaries.

Replication is supported in several different architectural setups. For example, two-node active/standby designs as well as multi-node clusters with runtime software upgrade.

<div data-with-frame="true"><figure><img src="/files/548XD13ekrSH5LfsAhJk" alt="" width="375"><figcaption><p>Primary - Secondary Configuration</p></figcaption></figure></div>

<div data-with-frame="true"><figure><img src="/files/r2tL9igE5U5hlaIOr8lU" alt="" width="375"><figcaption><p>One Primary - Several Secondaries</p></figcaption></figure></div>

This feature is independent of but compatible with the [Layered Service Architecture (LSA)](/guides/administration/advanced-topics/layered-service-architecture), which also configures multiple NSO nodes to provide additional scalability. When the following text simply refers to a cluster, it identifies the set of NSO nodes participating in the same HA group, not an LSA cluster, which is a separate concept.

NSO supports the following options for implementing an HA setup to cater to the widest possible range of use cases (only one can be used at a time):

* [**HA Raft**](#ug.ha.raft): Using a modern, consensus-based algorithm, it offers a robust, hands-off solution that works best in the majority of cases.
* [**Rule-based HA**](#ug.ha.builtin): A less sophisticated solution that allows you to influence the primary selection but may require occasional manual operator action.
* [**External HA**](#read-only-state): NSO only provides data replication; all other functions, such as primary selection and group membership management, are performed by an external application, using the HA framework (HAFW).

All of these options rely on a secure transport for communication, which uses TLS and host certificates. See [Managing Certificates](#managing-certificates) for details.

In addition to data replication, having a fixed address to connect to the current primary in an HA group greatly simplifies access for operators, users, and other systems alike. Use [Tail-f HCC Package](#ug.ha.hcc) or an [external load balancer](#ug.ha.lb) to manage it.

## NSO HA Raft <a href="#ug.ha.raft" id="ug.ha.raft"></a>

[Raft](https://raft.github.io/) is a consensus algorithm that reliably distributes a set of changes to a group of nodes and robustly handles network and node failure. It can operate in the face of multiple, subsequent failures, while also allowing a previously failed or disconnected node to automatically rejoin the cluster without risk of data conflicts.

Compared to traditional fail-over HA solutions, Raft relies on the consensus of the participating nodes, which addresses the so-called “split-brain” problem, where multiple nodes assume a primary role. This problem is especially characteristic of two-node systems, where it is impossible for a single node on its own to distinguish between losing network connectivity itself versus the other node malfunctioning. For this reason, Raft requires at least three nodes in the cluster.

Raft achieves robustness by requiring at least three nodes in the HA cluster. Three is the recommended cluster size, allowing the cluster to operate in the face of a single node failure. In case you need to tolerate two nodes failing simultaneously, you can add two additional nodes, for a 5-node cluster. However, permanently having more than five nodes in a single cluster is currently not recommended since Raft requires the majority of the currently configured nodes in the cluster to reach consensus. Without the consensus, the cluster cannot function.

You can start a sample HA Raft cluster using the [examples.ncs/high-availability/raft-cluster](https://github.com/NSO-developer/nso-examples/tree/6.7/high-availability/raft-cluster) example to test it out. The scripts in the example show various aspects of cluster setup and operation, which are further described in the rest of this section.

Optionally, examples using separate containers for each HA Raft cluster member with NSO system installations are available and referenced in the [examples.ncs/high-availability/hcc](https://github.com/NSO-developer/nso-examples/tree/6.7/high-availability/hcc) example in the NSO example set.

### Overview of Raft Operation <a href="#d5e4526" id="d5e4526"></a>

The Raft algorithm works with the concept of (election) terms. In each term, nodes in the cluster vote for a leader. The leader is elected when it receives the majority of the votes. Since each node only votes for a single leader in a given term, there can only be one leader in the cluster for this term.

Once elected, the leader becomes responsible for distributing the changes and ensuring consensus in the cluster for that term. Consensus means that the majority of the participating nodes must confirm a change before it is accepted. This is required for the system to ensure no changes ever get overwritten and provide reliability guarantees. On the other hand, it also means more than half of the nodes must be available for normal operation.

Changes can only be performed on the leader, that will accept the change after the majority of the cluster nodes confirm it. This is the reason a typical Raft cluster has an odd number of nodes; exactly half of the nodes agreeing on a change is not sufficient. It also makes a two-node cluster (or any even number of nodes in a cluster) impractical; the system as a whole is no more available than it is with one fewer node.

If the connection to the leader is broken, such as during a network partition, the nodes start a new term and a new election. Another node can become a leader if it gets the majority of the votes of all nodes initially in the cluster. While gathering votes, the node has the status of a candidate. In case multiple nodes assume candidate status, a split-vote scenario may occur, which is resolved by starting a fresh election until a candidate secures the majority vote.

If it happens that there aren't enough reachable nodes to obtain a majority, a candidate can stay in the candidate state for an indefinite time. Otherwise, when a node votes for a candidate, it becomes a follower and stays a follower in this term, regardless if the candidate is elected or not.

Additionally, the NSO node can also be in the stalled state, if HA Raft is enabled but the node has not joined a cluster.

### Node Names and Certificates <a href="#ch_ha.raft_names" id="ch_ha.raft_names"></a>

Each node in an HA Raft cluster needs a unique name. Names are composed of multiple parts but are usually configured in the format of a simple `ADDRESS`, which identifies a network host where the NSO process is running, such as a fully qualified domain name (FQDN) or an IPv4 address.

Other nodes in the cluster must be able to resolve and reach the `ADDRESS`, which creates a dependency on the DNS if you use domain names instead of IP addresses. `ADDRESS` also cannot be a simple short name (without a dot), even if the system is able to resolve such a name using `hosts` file or a similar mechanism.

The full node name contains node id and port in the format of `ID@ADDRESS:PORT`, such as `ncsd@192.0.2.1:4570`.

You specify the node address in the `ncs.conf` file as the value for `node-address`, under the `listen` container. You can also use the node name (with the "@" character), however, that is usually unnecessary as the system prepends `ncsd@` as-needed. The non-default port value can also be specified in the `ncs.conf`.

Another aspect in which `ADDRESS` plays a role is authentication. The HA system uses mutual TLS to secure communication between cluster nodes. This requires you to configure a trusted Certificate Authority (CA) and a key/certificate pair for each node. When nodes connect, they check that the certificate of the peer validates against the CA and matches the `ADDRESS` of the peer.

The following is a HA Raft configuration snippet for `ncs.conf` that includes certificate settings and a sample `ADDRESS`:

```xml
  <ha-raft>
    <!-- ... -->
    <listen>
      <node-address>198.51.100.10</node-address>
    </listen>
    <ssl>
      <ca-cert-file>${NCS_CONFIG_DIR}/dist/ssl/cert/myca.crt</ca-cert-file>
      <cert-file>${NCS_CONFIG_DIR}/dist/ssl/cert/node-100-10.crt</cert-file>
      <key-file>${NCS_CONFIG_DIR}/dist/ssl/cert/node-100-10.key</key-file>
    </ssl>
  </ha-raft>
```

See [Managing Certificates](#managing-certificates) for more information on creating TLS X509 certificates.

### Actions <a href="#ch_ha.raft_actions" id="ch_ha.raft_actions"></a>

NSO HA Raft can be controlled through several actions. All actions are found under `/ha-raft/`. In the best-case scenario, you will only need the `create-cluster` action to initialize the cluster and the `read-only` and `create-cluster` actions when upgrading the NSO version.

The available actions are listed below:

<table><thead><tr><th width="240" valign="top">Action</th><th valign="top">Description</th></tr></thead><tbody><tr><td valign="top"><code>create-cluster</code></td><td valign="top">Initialize an HA Raft cluster. This action should only be invoked once to form a new cluster when no HA Raft log exists.<br>The members of the HA Raft cluster consist of the NCS node where the <code>/ha-raft/create-cluster</code>action is invoked, which will become the leader of the cluster; and the members specified by the <code>member</code> parameter.</td></tr><tr><td valign="top"><code>adjust-membership</code></td><td valign="top">Add or remove an HA node from the HA Raft cluster.</td></tr><tr><td valign="top"><code>disconnect</code></td><td valign="top">Disconnect an HA node from all remaining nodes. In the event of revoking a TLS certificate, invoke this action to disconnect the already established connections to the node with the revoked certificate. A disconnected node with a valid TLS certificate may re-establish the connection.</td></tr><tr><td valign="top"><code>reset</code></td><td valign="top">Reset the (disabled) local node to make the leader perform a full sync to this local node if an HA Raft cluster exists. If reset is performed on the leader node, the node will step down from leadership and it will be synced by the next leader node.<br>An HA Raft member will change role to <code>disabled</code> if <code>ncs.conf</code> has incompatible changes to the <code>ncs.conf</code> on the leader; a member will also change role to <code>disabled</code> if there are non-recoverable failures upon opening a snapshot.<br>See the <code>/ha-raft/status/disable-reason</code> leaf for the reason.<br>Set force to <code>true</code> to override reset when <code>/ha-raft/status/role</code> is not set to <code>disabled</code>.</td></tr><tr><td valign="top"><code>handover</code></td><td valign="top">Handover leadership to another member of the HA Raft cluster or step down from leadership and start a new election.</td></tr></tbody></table>

Additionally, the `/ncs-state/set-read-only` action can be used to toggle administrator-configured read-only mode. See [Read-only State](#read-only-state).

### Network and `ncs.conf` Prerequisites <a href="#ch_ha.raft_ports" id="ch_ha.raft_ports"></a>

In addition to the network connectivity required for the normal operation of a standalone NSO node, nodes in the HA Raft cluster must be able to initiate TCP connections from a random ephemeral client port to the following ports on other nodes:

* Port 4570 (the default HA communication port, configurable)

The Raft implementation does not impose any other hard limits on the network but you should keep in mind that consensus requires communication with other nodes in the cluster. A high round-trip latency between cluster nodes is likely to negatively impact the transaction throughput of the system.

The HA Raft cluster also requires compatible `ncs.conf` files among the member nodes. In particular, `/ncs-config/cdb/operational/enabled` and `/ncs-config/rollback/enabled` values affect replication behavior and must match. Likewise, each member must have the same set of encryption keys and the keys cannot be changed while the cluster is in operation.

To update the `ncs.conf` configuration, you must manually update the copy on each member node, making sure the new versions contain compatible values. Then perform the reload on the leader and the follower members will automatically reload their copies of the configuration file as well.

If a node is a cluster member but has been configured with a new, incompatible `ncs.conf` file, it gets automatically disabled. See the `/ha-raft/status/disabled-reason` for reason. You can re-enable the node with the `ha-raft reset` command, once you have reconciled the incompatibilities.

### Connected Nodes and Node Discovery <a href="#d5e4598" id="d5e4598"></a>

Raft has a notion of cluster configuration, in particular, how many and which members the cluster has. You define member nodes when you first initialize the cluster with the `create-cluster` command or use the `adjust-membership` command. The member nodes allow the cluster to know how many nodes are needed for consensus and similar.

However, not all cluster members may be reachable or alive all the time. Raft implementation in NSO uses TCP connections between nodes to transport data. The TCP connections are authenticated and encrypted using TLS by default (see [Security Considerations](#ch_ha.raft_security)). A working connection between nodes is essential for the cluster to function but a number of factors, such as firewall rules or expired/invalid certificates, can prevent the connection from establishing.

Therefore, NSO distinguishes between configured member nodes and nodes to which it has established a working transport connection. The latter are called connected nodes. In a normal, fully working, and properly configured cluster, the connected nodes will be the same as member nodes (except for the current node).

To help troubleshoot connectivity issues without affecting cluster operation, connected nodes will show even nodes that are not actively participating in the cluster but have established a transport connection to nodes in the cluster. The optional discovery mechanism, described next, relies on this functionality.

NSO includes a mechanism that simplifies the initial cluster setup by enumerating known nodes. This mechanism uses a set of seed nodes to discover all connectable nodes, which can then be used with the `create-cluster` command to form a Raft cluster.

When you specify one or more nodes with the `/ha-raft/seed-nodes/seed-node` setting in the `ncs.conf` file, the current node tries to establish a connection to these seed nodes, in order to discover the list of all nodes potentially participating in the cluster. For the discovery to work properly, all other nodes must also use seed nodes and the set of seed nodes must overlap. The recommended practice is to use the same set of seed nodes on every participating node.

Along with providing an autocompletion list for the `create-cluster` command, this feature streamlines the discovery of node names when using NSO in containerized or other dynamic environments, where node addresses are not known in advance.

### Initial Cluster Setup <a href="#ch_ha.raft_setup" id="ch_ha.raft_setup"></a>

Creating a new HA cluster consists of two parts: configuring the individual nodes and running the `create-cluster` action.

First, you must update the `ncs.conf` configuration file for each node. All HA Raft configuration comes under the `/ncs-config/ha-raft` element.

As part of the configuration, you must:

* Enable HA Raft functionality through the `enabled` leaf.
* Set `node-address` and the corresponding TLS parameters (see [Node Names and Certificates](#ch_ha.raft_names)).
* Identify the cluster this node belongs to with `cluster-name`.
* Reload or restart the NSO process (if already running).
* Repeat the preceding steps for every participating node.
* Enable read-only mode on designated leader to avoid potential sync issues in cluster formation.
* Invoke the `create-cluster` action.

The cluster name is simply a character string that uniquely identifies this HA cluster. The nodes in the cluster must use the same cluster name or they will refuse to establish a connection. This setting helps prevent mistakenly adding a node to the wrong cluster when multiple clusters are in operation, such as in an LSA setup.

{% code title="Sample HA Raft config for a cluster node" %}

```xml
  <ha-raft>
    <enabled>true</enabled>
    <cluster-name>sherwood</cluster-name>
    <listen>
      <node-address>ash.example.org</node-address>
    </listen>
    <ssl>
      <ca-cert-file>${NCS_CONFIG_DIR}/dist/ssl/cert/myca.crt</ca-cert-file>
      <cert-file>${NCS_CONFIG_DIR}/dist/ssl/cert/ash.crt</cert-file>
      <key-file>${NCS_CONFIG_DIR}/dist/ssl/cert/ash.key</key-file>
    </ssl>
    <seed-nodes>
      <seed-node>
        <address>birch.example.org</address>
        <port>4570</port>
      </seed-node>
    </seed-nodes>
  </ha-raft>
```

{% endcode %}

With all the nodes configured and running, connect to the node that you would like to become the initial leader and invoke the `ha-raft create-cluster` action. The action takes a list of nodes identified by their names. If you have configured `seed-nodes`, you will get auto-completion support, otherwise, you have to type in the names of the nodes yourself.

This action makes the current node a cluster leader and joins the other specified nodes to the newly created cluster. For example:

```bash
admin@ncs# ncs-state set-read-only mode true
admin@ncs# ha-raft create-cluster member [ birch.example.org cedar.example.org ]
admin@ncs# show ha-raft
ha-raft status role leader
ha-raft status leader ash.example.org
ha-raft status member [ ash.example.org birch.example.org cedar.example.org ]
ha-raft status connected-node [ birch.example.org cedar.example.org ]
ha-raft status local-node ash.example.org
...
admin@ncs# ncs-state set-read-only mode false
```

You can use the `show ha-raft` command on any node to inspect the status of the HA Raft cluster. The output includes the current cluster leader and members according to this node, as well as information about the local node, such as node name (`local-node`) and role. The `status/connected-node` list contains the names of the nodes with which this node has active network connections.

<details>

<summary><code>show ha-raft</code> Field Definitions</summary>

The command `show ha-raft` is used in NSO to display the current state of the HA Raft cluster. The output typically includes the following information:

* The role of the local node (for example, whether it is the `leader`, `follower`, `candidate`, or `stalled`).
* The leader of the cluster, if one has been elected.
* The list of member nodes that belong to the HA Raft cluster.
* The connected nodes, which are the nodes with which the local node currently has active RAFT communication.
* The local node information, detailing the node’s name and status.

This command is useful for both verifying that the HA Raft cluster is set up correctly and for troubleshooting issues by checking the connectivity and role assignments of the nodes. Some noteworthy terms of output are defined in the table below.

<table><thead><tr><th width="287.47265625" valign="top">Term</th><th valign="top">Definition</th></tr></thead><tbody><tr><td valign="top"><code>role</code></td><td valign="top">The current node’s Raft role (<code>leader</code>, <code>follower</code>, or <code>candidate</code>). Occasionally, in NSO, a node might appear as <code>stalled</code> if it has lost contact with the leader or quorum.</td></tr><tr><td valign="top"><code>leader</code></td><td valign="top">The current known leader of the cluster.</td></tr><tr><td valign="top"><code>member</code></td><td valign="top">A node that is part of the RAFT consensus group (i.e., a voting participant, not an observer). Leaders, followers, and candidates are members; observers are not.</td></tr><tr><td valign="top"><code>connected-node</code></td><td valign="top">The nodes this instance is connected to.</td></tr><tr><td valign="top"><code>local-node</code></td><td valign="top">The name of the current node.</td></tr><tr><td valign="top"><code>lag</code></td><td valign="top">The number of indices the replicated log is behind the leader node. A value of <code>0</code> means no lag — the node's RAFT log is fully up-to-date with the leader. The larger the value, the more out-of-sync the node is, which may indicate a replication or connectivity issue.</td></tr><tr><td valign="top"><code>index</code></td><td valign="top">The last replicated HA Raft log index, i.e., this is the last log entry replicated to a node.</td></tr><tr><td valign="top"><code>state</code></td><td valign="top"><p>The synchronization status of the node’s RAFT log. Common values include:</p><ul><li><code>in-sync</code>: The node is up-to-date with the leader.</li><li><code>behind</code>: The node is lagging behind in log replication.</li><li><code>unreachable</code>: The node is not communicating with one or more RAFT peers, i.e., the node cannot reach the leader or other RAFT peers, preventing synchronization.</li><li><code>requires-snapshot</code>: The node has fallen too far behind to catch up using logs and needs a full snapshot from the leader.</li></ul></td></tr><tr><td valign="top"><code>current-index</code></td><td valign="top">The latest log index on this node.</td></tr><tr><td valign="top"><code>applied-index</code></td><td valign="top">The last index applied to CDB.</td></tr><tr><td valign="top"><code>serial-number</code></td><td valign="top">The certificate serial number. Used to uniquely identify the node.</td></tr></tbody></table>

</details>

In case you get an error, such as the `Error: NSO can't reach member node 'ncsd@ADDRESS'.`, verify all of the following:

* The node at the `ADDRESS` is reachable. You can use the `ping` `ADDRESS` command, for example.
* The problematic node has the correct `ncs.conf` configuration, especially `cluster-name` and `node-address`. The latter should match the `ADDRESS` and should contain at least one dot.
* Nodes use compatible configuration. For example, make sure that the `ncs.crypto_keys` file (if used) or the `encrypted-strings` configuration in `ncs.conf` is identical across all nodes in the cluster. In case the encrypted keys (e.g., the `ncs.crypto_keys` file or the key-generation in `ncs.conf`) become mismatched and are subsequently corrected, you must issue a `ha-raft reset` on the affected node(s) to fully re-establish RAFT state. Without a reset, the node may remain disabled or out of sync due to inconsistencies in its persisted RAFT state. Additionally, when a node enters the `disabled` state for any reason (for example, configuration mismatch, unrecoverable RAFT state corruption, or safety-rule violations), RAFT will not automatically recover. Because `disabled` is the terminal state of the RAFT state machine, the operator must manually issue `ha-raft reset` after resolving the underlying cause to return the node to normal operation.
* HA Raft is enabled, using the `show ha-raft` command on the unreachable node.
* The firewall configuration on the OS and on the network level permits traffic on the required ports (see [Network and `ncs.conf` Prerequisites](#ch_ha.raft_ports)).
* The node uses a certificate that the CA can validate. For example, copy the certificates to the same location and run `openssl verify -CAfile CA_CERT NODE_CERT` to verify this.
* Verify the `epmd -names` command on each node shows the ncsd process. If not, stop NSO, run `epmd -kill`, and then start NSO again.

In addition to the above, you may also examine the `logs/raft.log` file for detailed information on the error message and overall operation of the Raft algorithm. The amount of information in the file is controlled by the `/ncs-config/logs/raft-log` configuration in the `ncs.conf`.

### Cluster Management <a href="#d5e4696" id="d5e4696"></a>

After the initial cluster setup, you can add new nodes or remove existing nodes from the cluster with the help of the `ha-raft adjust-membership` action. For example:

```bash
admin@ncs# show ha-raft status member
ha-raft status member [ ash.example.org birch.example.org cedar.example.org ]
admin@ncs# ha-raft adjust-membership remove-node birch.example.org
admin@ncs# show ha-raft status member
ha-raft status member [ ash.example.org cedar.example.org ]
admin@ncs# ha-raft adjust-membership add-node dollartree.example.org
admin@ncs# show ha-raft status member
ha-raft status member [ ash.example.org cedar.example.org dollartree.example.org ]
```

When removing nodes using the `ha-raft adjust-membership remove-node` command, the removed node is not made aware that it is removed from the cluster and continues signaling the other nodes. This is a limitation in the algorithm, as it must also handle situations, where the removed node is down or unreachable. To prevent further communication with the cluster, it is important you ensure the removed node is shut down. You should shut down the to-be-removed node prior to removal from the cluster, or immediately after it. The former is recommended but the latter is required if there are only two nodes left in the cluster and shutting down prior to removal would prevent the cluster from reaching consensus.

Additionally, you can force an existing follower node to perform a full re-sync from the leader by invoking the `ha-raft reset` action with the `force` option. Using this action on the leader will make the node give up the leader role and perform a sync with the newly elected leader.

As leader selection during the Raft election is not deterministic, NSO provides the `ha-raft handover` action, which allows you to either trigger a new election if called with no arguments or transfer leadership to a specific node. The latter is especially useful when, for example, one of the nodes resides in a different location and more traffic between locations may incur extra costs or additional latency, so you prefer this node is not the leader under normal conditions.

#### Passive Follower

In certain situations, it may be advantageous to have a follower node that cannot be promoted to leader role. Consider a scenario with three Raft-enabled nodes distributed across two different data centers.

In this case, a node located without a peer in the same data center might experience increased latency due to the requirement for acknowledgments from at least one node in the other data center.

To address this, HA Raft provides the `/ncs-config/ha-raft/passive` setting. When this setting is enabled (set to `true`), it prevents the node from assuming the candidate or leader role. A passive follower still participates by voting in leader elections.

Note that the `passive` parameter is local to the node, meaning other nodes in the cluster are unaware that a particular follower is passive. Consequently, it is possible to initiate a handover action targeting the passive node, but the handover will ultimately fail at a later stage, allowing the current leader to retain its position.

### Three-Node Example with HCC VIP <a href="#ch_ha.raft.threenode" id="ch_ha.raft.threenode"></a>

The following example summarizes a common three-node HA Raft deployment together with HCC layer-2 VIP management. In this setup, the cluster consists of one current leader and two followers:

<table><thead><tr><th width="133" valign="top">Node</th><th width="195" valign="top">Address</th><th valign="top">Notes</th></tr></thead><tbody><tr><td valign="top"><code>nso-a</code></td><td valign="top"><code>nso-a.example.org</code></td><td valign="top">Initial leader when <code>create-cluster</code> is invoked on this node.</td></tr><tr><td valign="top"><code>nso-b</code></td><td valign="top"><code>nso-b.example.org</code></td><td valign="top">Follower.</td></tr><tr><td valign="top"><code>nso-c</code></td><td valign="top"><code>nso-c.example.org</code></td><td valign="top">Follower.</td></tr></tbody></table>

Each node uses the same `cluster-name` and seed-node list, but a different `node-address` and certificate/key pair. After the nodes are started, initialize the cluster on the node that should become the initial leader and then configure HCC on the leader:

```bash
admin@ncs# ha-raft create-cluster member [ nso-b.example.org nso-c.example.org ]
admin@ncs# show ha-raft
ha-raft status role leader
ha-raft status leader nso-a.example.org
ha-raft status member [ nso-a.example.org nso-b.example.org nso-c.example.org ]
ha-raft status connected-node [ nso-b.example.org nso-c.example.org ]
...
admin@ncs(config)# hcc enabled
admin@ncs(config)# hcc vip 192.0.2.100
admin@ncs(config)# commit
```

In steady state, the expected status is:

<table><thead><tr><th width="127" valign="top">Node</th><th width="286" valign="top">State</th><th valign="top">VIP state</th></tr></thead><tbody><tr><td valign="top"><code>nso-a</code></td><td valign="top"><code>role leader</code></td><td valign="top">HCC binds the VIP on the leader.</td></tr><tr><td valign="top"><code>nso-b</code></td><td valign="top"><code>role follower</code></td><td valign="top">No VIP is bound on the follower.</td></tr><tr><td valign="top"><code>nso-c</code></td><td valign="top"><code>role follower</code></td><td valign="top">No VIP is bound on the follower.</td></tr></tbody></table>

#### **HA Events in a Three-Node HA Raft Cluster**

<table><thead><tr><th width="184" valign="top">HA event</th><th width="286" valign="top">State change</th><th width="147" valign="top">VIP state</th><th valign="top">Manual action</th></tr></thead><tbody><tr><td valign="top">One follower is shut down or loses connectivity</td><td valign="top">The cluster still has quorum.<br>The leader raises <code>ha-secondary-down</code> for the lost follower and remains <code>leader</code> while the remaining peer stays <code>follower</code>.<br>The down node is absent from <code>connected-node</code>.</td><td valign="top">The VIP remains bound on the current leader.</td><td valign="top">No action required.</td></tr><tr><td valign="top">The stopped or disconnected follower returns</td><td valign="top">The previous <code>ha-secondary-down</code> alarm clears when connectivity is restored.<br>The returning node rejoins as <code>follower</code> and catches up automatically from the leader.</td><td valign="top">The VIP remains bound on the current leader.</td><td valign="top">No action required.</td></tr><tr><td valign="top">The leader is shut down or loses connectivity, but the two remaining nodes can still reach each other</td><td valign="top">The remaining quorum elects a new <code>leader</code> and the other surviving node becomes or remains <code>follower</code>.<br>One of the surviving nodes raises <code>ha-primary-down</code> for the lost leader.<br>Writes continue on the new leader.</td><td valign="top">The VIP moves to the newly elected leader.</td><td valign="top">No action required.</td></tr><tr><td valign="top">The former leader returns after failover</td><td valign="top">Any earlier <code>ha-primary-down</code> alarm clears when connectivity to the former leader is restored.<br>The old leader rejoins as <code>follower</code> and catches up automatically.<br>The current leader remains <code>leader</code>.</td><td valign="top">The VIP remains bound on the current leader.</td><td valign="top">No action required.</td></tr><tr><td valign="top">One node is isolated from the other two by a network partition</td><td valign="top">The two-node side still has quorum and keeps or elects a leader.<br>If the isolated node was a follower, the leader on the majority side raises <code>ha-secondary-down</code>.<br>If the isolated node was the leader, one of the surviving nodes raises <code>ha-primary-down</code> before or during leader re-election.<br>The isolated node cannot make durable progress and rejoins as a <code>follower</code> once it learns of the higher term after connectivity is restored.</td><td valign="top">The VIP follows the leader on the majority side.</td><td valign="top">No action required if connectivity is restored cleanly.</td></tr><tr><td valign="top">The leader loses quorum because both other members are unavailable or unreachable</td><td valign="top">The cluster cannot commit new writes without a majority.<br>On the isolated leader, write attempts will block or fail, <code>ha-raft-quorum-lost</code> is raised, and the node transition to <code>disabled</code>.<br>Once enough members reconnect, one node can become leader again.</td><td valign="top">Do not rely on the VIP for write traffic until quorum is restored and a stable leader exists again.</td><td valign="top">Restore quorum first. If a node remains <code>disabled</code>, inspect <code>/ha-raft/status/disable-reason</code> and invoke <code>/ha-raft/reset</code> after resolving the underlying cause.</td></tr></tbody></table>

As in HA Raft generally, leader election among the surviving quorum members is not deterministic. If you want the cluster to return to a preferred leader after recovery, use `ha-raft handover` once the cluster is healthy again.

### Migrating From Existing Rule-based HA <a href="#d5e4714" id="d5e4714"></a>

If you have an existing HA cluster using the rule-based built-in HA, you can migrate it to use HA Raft instead. This procedure is performed in four distinct high-level steps:

* Ensuring the existing cluster meets migration prerequisites.
* Preparing the required HA Raft configuration files.
* Switching to HA Raft.
* Adding additional nodes to the cluster.

The procedure does not perform an NSO version upgrade, so the cluster remains on the same version. It also does not perform any schema upgrades, it only changes the type of the HA cluster.

The migration procedure is in place, that is, the existing nodes are disconnected from the old cluster and connected to the new one. This results in a temporary disruption of the service, so it should be performed during a service window.

First, you should ensure the cluster meets migration prerequisites. The cluster must use:

* NSO 6.1.2 or later
* tailf-hcc 6.0 or later (if used)

In case these prerequisites are not met, follow the standard upgrade procedures to upgrade the existing cluster to supported versions first.

Additionally, ensure that all used packages are compatible with HA Raft, as NSO uses some new or updated notifications about HA state changes. Also, verify the network supports the new cluster communications (see [Network and `ncs.conf` Prerequisites](#ch_ha.raft_ports)).

Secondly, prepare all the `ncs.conf` and related files for each node, such as certificates and keys. Create a copy of all the `ncs.conf` files and disable or remove the existing `>ha<` section in the copies. Then add the required configuration items to the copies, as described in [Initial Cluster Setup](#ch_ha.raft_setup) and [Node Names and Certificates](#ch_ha.raft_names). Do not update the `ncs.conf` files used by the nodes yet.

It is recommended but not necessary that you set the seed nodes in `ncs.conf` to the designated primary and fail-over primary. Do this for all `ncs.conf` files for all nodes.

#### Procedure 1. Migration to HA Raft

1. With the new configurations at hand and verified, start the switch to HA Raft. The cluster nodes should be in their nominal, designated roles. If not, perform a failover first.
2. On the designated (actual) primary, called `node1`, enable read-only mode.

   ```bash
   admin@node1# ncs-state set-read-only mode true
   ```
3. Then take a backup of all nodes.
4. Once the backup successfully completes, stop the designated fail-over primary (actual secondary) NSO process, update its `ncs.conf` and the related (certificate) files for HA Raft, and then start it again. Connect to this node's CLI, here called node2, and verify HA Raft is enabled with the `show` `ha-raft` command.

   ```bash
   admin@node2# show ha-raft
   ha-raft status role stalled
   ha-raft status local-node node2.example.org
   > ... output omitted ... <
   ```
5. Now repeat the same for the designated primary (`node1`). If you have set the seed nodes, you should see the fail-over primary show under `connected-node`.

   ```bash
   admin@node1# show ha-raft
   ha-raft status role stalled
   ha-raft status connected-node [ node2.example.org ]
   ha-raft status local-node node1.example.org
   > ... output omitted ... <
   ```
6. On the old designated primary (node1) invoke the `ha-raft create-cluster` action and create a two-node Raft cluster with the old fail-over primary (`node2`, actual secondary). The action takes a list of nodes identified by their names. If you have configured `seed-nodes`, you will get auto-completion support, otherwise you have to type in the name of the node yourself.

   ```bash
   admin@node1# ha-raft create-cluster member [ node2.example.org ]
   admin@node1# show ha-raft
   ha-raft status role leader
   ha-raft status leader node1.example.org
   ha-raft status member [ node1.example.org node2.example.org ]
   ha-raft status connected-node [ node2.example.org ]
   ha-raft status local-node node1.example.org
   > ... output omitted ... <
   ```

   In case of errors running the action, refer to [Initial Cluster Setup](#ch_ha.raft_setup) for possible causes and troubleshooting steps.
7. Raft requires at least three nodes to operate effectively (as described in [NSO HA Raft](#ug.ha.raft)) and currently, there are only two in the cluster. If the initial cluster had only two nodes, you must provision an additional node and set it up for HA Raft. If the cluster initially had three nodes, there is the remaining secondary node, `node3`, which you must stop, update its configuration as you did with the other two nodes, and start it up again.
8. Finally, on the old designated primary and current HA Raft leader, use the `ha-raft adjust-membership add-node` action to add this third node to the cluster.

   ```bash
   admin@node1# ha-raft adjust-membership add-node node3.example.org
   admin@node1# show ha-raft status member
   ha-raft status member [ node1.example.org node2.example.org node3.example.org ]
   ```

### Security Considerations <a href="#ch_ha.raft_security" id="ch_ha.raft_security"></a>

Communication between the NSO nodes in an HA Raft cluster relies on an RPC protocol transported over TLS (unless explicitly disabled by setting `/ncs-config/ha-raft/ssl/enabled` to 'false').

TLS (Transport Layer Security) provides Authentication and Privacy by only allowing NSO nodes to connect using certificates and keys issued from the same Certificate Authority (CA). It is *paramount* to security that TLS is always used.

Access to a host can be revoked by the CA through the means of a CRL (Certificate Revocation List). To enforce certificate revocation within an HA Raft cluster, invoke the action /ha-raft/disconnect to terminate the pre-existing connection. A connection to the node can re-establish once the node's certificate is valid.

Also ensure the node keys are kept safe by the correct filesystem permissions and keep the CA key in a safe place since it can be used to generate new certificates and key pairs for peers.

### Packages Upgrades in Raft Cluster

NSO contains a mechanism for distributing packages to nodes in a Raft cluster, greatly simplifying package management in a highly-available setup.

You perform all package management operations on the current leader node. To identify the leader node, you can use the `show ha-raft status leader` command on a running cluster.

Invoking the `packages reload` command makes the leader node update its currently loaded packages, identical to a non-HA, single-node setup. At the same time, the leader also distributes these packages to the followers to load. However, the load paths on the follower nodes, such as `/var/opt/ncs/packages/`, are not updated. This means, that if a leader election took place, a different leader was elected, and you performed another `packages reload`, the system would try to load the versions of the packages on this other leader, which may be out of date or not even present.

The recommended approach is, therefore, to use the `packages ha sync and-reload` command instead, unless a load path is shared between NSO nodes, such as the same network drive. This command distributes and updates packages in the load paths on the follower nodes, as well as loading them.

For the full procedure, first, ensure all cluster nodes are up and operational, then follow these steps on the leader node:

* Perform a full backup of the NSO instance, such as running `ncs-backup`.
* Add, replace, or remove packages on the filesystem. The exact location depends on the type of NSO deployment, for example `/var/opt/ncs/packages/`.
* Invoke the `packages ha sync and-reload` or `packages ha sync and-add` command to start the upgrade process.

Note that while the upgrade is in progress, writes to the CDB are not allowed and will be rejected.

For a `packages ha sync and-reload` example see the `raft-upgrade-l2` NSO system installation-based example referenced by the [examples.ncs/high-availability/hcc](https://github.com/NSO-developer/nso-examples/tree/6.7/high-availability/hcc) example in the NSO example set.

For more details, troubleshooting, and general upgrade recommendations, see [NSO Packages](/guides/administration/management/package-mgmt) and [Upgrade](/guides/administration/installation-and-deployment/upgrade-nso).

### Version Upgrade of Cluster Nodes <a href="#ch_ha.raft_upgrade" id="ch_ha.raft_upgrade"></a>

Currently, the only supported and safe way of upgrading the Raft HA cluster NSO version requires that the cluster be taken offline since the nodes must, at all times, run the same software version.

Do not attempt an upgrade unless all cluster member nodes are up and actively participating in the cluster. Verify the current cluster state with the `show ha-raft status` command. All member nodes must also be present in the connected-node list.

The procedure differentiates between the current leader node versus followers. To identify the leader, you can use the `show ha-raft status leader` command on a running cluster.

**Procedure 2. Cluster Version Upgrade**

1. On the leader, first enable read-only mode using the `ha-raft read-only mode true` command and then verify that all cluster nodes are in sync with the `show ha-raft status log replications state` command.
2. Before embarking on the upgrade procedure, it's imperative to backup each node. This ensures that you have a safety net in case of any unforeseen issues. For example, you can use the `$NCS_DIR/bin/ncs-backup` command.
3. Delete the `$NCS_RUN_DIR/cdb/compact.lock` file and compact the CDB write log on all nodes using, for example, the `$NCS_DIR/bin/ncs --cdb-compact $NCS_RUN_DIR/cdb` command.
4. On all nodes, delete the `$NCS_RUN_DIR/state/raft/` directory with a command such as `rm -rf $NCS_RUN_DIR/state/raft/`.
5. Stop NSO on all the follower nodes, for example, invoking the `$NCS_DIR/bin/ncs --stop` or `systemctl stop ncs` command on each node.
6. Stop NSO on the leader node only after you have stopped all the follower nodes in the previous step. Alternatively NSO can be stopped on the nodes before deleting the HA Raft state and compacting the CDB write log without needing to delete the `compact.lock` file.
7. Upgrade the NSO packages on the leader to support the new NSO version.
8. Install the new NSO version on all nodes.
9. Start NSO on all nodes.
10. Re-initialize the HA cluster using the `ha-raft create-cluster` action on the node to become the leader.
11. Finally, verify the cluster's state through the `show ha-raft status` command. Ensure that all data has been correctly synchronized across all cluster nodes and that the leader is no longer read-only. The latter happens automatically after re-initializing the HA cluster.

For a standard System Install, the single-node procedure is described in [Single Instance Upgrade](https://nso-docs.cisco.com/guides/administration/management/pages/mKh0oXFZkQsBaoCMMyEo#ug.admin_guide.manual_upgrade), but in general depends on the NSO deployment type. For example, it will be different for containerized environments. For specifics, please refer to the documentation for the deployment type.

For an example see the `raft-upgrade-l2` NSO system installation-based example referenced by the [examples.ncs/high-availability/hcc](https://github.com/NSO-developer/nso-examples/tree/6.7/high-availability/hcc) example in the NSO example set.

If the upgrade fails before or during the upgrade of the original leader, start up the original followers to restore service and then restore the original leader, using backup as necessary.

However, if the upgrade fails after the original leader was successfully upgraded, you should still be able to complete the cluster upgrade. If you are unable to upgrade a follower node, you may provision a (fresh) replacement and the data and packages in use will be copied from the leader.

## NSO Rule-based HA <a href="#ug.ha.builtin" id="ug.ha.builtin"></a>

NSO can manage the HA groups based on a set of predefined rules. This functionality was added in NSO 5.4 and is sometimes referred to simply as the built-in HA. However, since NSO 6.1, HA Raft (which is also built-in) is available as well, and is likely a better choice in most situations.

Rule-based HA allows administrators to:

* Configure HA group members with IP addresses and default roles
* Configure failover behavior
* Configure start-up behavior
* Configure HA group members with IP addresses and default roles
* Assign roles, join HA group, enable/disable rule-based HA through actions
* View the state of the current HA setup

NSO rule-based HA is defined in `tailf-ncs-high-availability.yang`, with data residing under the `/high-availability/` container. Since NSO 6.7, the HA transport uses TLS and host certificates to secure node communication. See [Managing Certificates](#managing-certificates) for details and recipes.

{% hint style="info" %}
In environments with high NETCONF traffic, particularly when using `ncs_device_notifs`, it's recommended to enable read-only mode on the designated primary node before performing HA activation or sync. This prevents `app_sync` from being blocked by notification processing.

Use the following command prior to enabling HA or assigning roles:

```bash
admin@ncs# ncs-state set-read-only mode true
```

After successful sync and HA establishment, disable read-only mode:

```bash
admin@ncs# ncs-state set-read-only mode false
```

{% endhint %}

NSO rule-based HA does not manage any virtual IP addresses, or advertise any BGP routes or similar. This must be handled by an external package. Tail-f HCC 5.x and greater has this functionality compatible with NSO rule-based HA. You can read more about the HCC package in the [following chapter](#ug.ha.hcc).

### Prerequisites <a href="#d5e4824" id="d5e4824"></a>

To use NSO rule-based HA, HA must first be enabled in `ncs.conf` - See [Mode of Operation](#ha.moo).

{% hint style="info" %}
If the package tailf-hcc with a version less than 5.0 is loaded, NSO rule-based HA will not function. These HCC versions may still be used but NSO built-in HA will not function in parallel.
{% endhint %}

### HA Member Configuration <a href="#d5e4830" id="d5e4830"></a>

All HA group members are defined under `/high-availability/ha-node`. Each configured node must have a unique `/high-availability/ha-node/address` (IP address or host name) configured and a unique HA ID. Additionally, nominal roles and fail-over settings may be configured on a per-node basis.

The HA Node ID is a unique identifier used to identify NSO instances in an HA group. The HA ID of the local node - relevant amongst others when an action is called - is determined by matching configured HA node IP addresses against IP addresses assigned to the host machine of the NSO instance. As the HA ID is crucial to NSO HA, NSO rule-based HA will not function if the local node cannot be identified.

To join a HA group, a shared secret must be configured on the active primary and any prospective secondary. This is used for a CHAP-2-like authentication and is specified under `/high-availability/token/`.

{% hint style="info" %}
In an NSO System Install setup, not only does the shared token need to match between the HA group nodes but the configuration for encrypted strings, default stored in `/etc/ncs/ncs.crypto_keys`, need also to match between the nodes in the HA group.
{% endhint %}

The token configured on the secondary node is overwritten with the encrypted token of type `aes-256-cfb-128-encrypted-string` from the primary node when the secondary node connects to the primary. If there is a mismatch between the encrypted-string configuration on the nodes, NSO will not decrypt the HA token to match the token presented. As a result, the primary node denies the secondary node access the next time the HA connection needs to reestablish with a "Token mismatch, secondary is not allowed" error.

See the `upgrade-l2` example, referenced from [examples.ncs/high-availability/hcc](https://github.com/NSO-developer/nso-examples/tree/6.7/high-availability/hcc), for an example setup and the [Deployment Example](/guides/administration/installation-and-deployment/development-to-production-deployment/deployment-example) for a description of the example.

Also, note that the `ncs.crypto_keys` file is highly sensitive. The file contains the encryption keys for all CDB data that is encrypted on disk. Besides the HA token, this often includes passwords for various entities, such as login credentials to managed devices.

### HA Roles <a href="#d5e4846" id="d5e4846"></a>

NSO can assume HA roles `primary`, `secondary` and `none`. Roles can be assigned directly through actions, or at startup or failover. See [HA Framework Requirements](#ha-framework-requirements) for the definition of these roles.

{% hint style="info" %}
NSO rule-based HA does not support relay-secondaries.
{% endhint %}

NSO rule-based HA distinguishes between the concepts of nominal role and assigned role. Nominal-role is configuration data that applies when an NSO instance starts up and at failover. The assigned role is the role that the NSO instance has been ordered to assume either by an action or as a result of startup or failover.

### Failover <a href="#d5e4857" id="d5e4857"></a>

Failover may occur when a secondary node loses the connection to the primary node. A secondary may then take over the primary role. Failover behavior is configurable and controlled by the parameters:

* `/high-availability/ha-node{id}/failover-primary`
* `/high-availability/settings/enable-failover`

For automatic failover to function, `/high-availability/settings/enable-failover` must be se to `true`. It is then possible to enable at most one node with a nominal role secondary as failover-primary, by setting the parameter `/high-availability/ha-node{id}/failover-primary`. The failover works in both directions; if a nominal primary is currently connected to the failover-primary as a secondary and loses the connection, then it will attempt to take over as a primary.

Before failover happens, a failover-primary-enabled secondary node may attempt to reconnect to the previous primary before assuming the primary role. This behavior is configured by the parameters denoting how many reconnect attempts will be made, and with which interval, respectively.

* `/high-availability/settings/reconnect-attempts`
* `/high-availability/settings/reconnect-interval`

HA Members that are assigned as secondaries, but are neither failover-primaries nor set with a nominal-role primary, may attempt to rejoin the HA group after losing connection to the primary.

This is controlled by `/high-availability/settings/reconnect-secondaries`. If this is true, secondary nodes will query the nodes configured under `/high-availability/ha-node` for an NSO instance that currently has the primary role. Any configured nominal roles will not be considered. If no primary node is found, subsequent attempts to rejoin the HA setup will be issued with an interval defined by `/high-availability/settings/reconnect-interval`.

In case a net-split provokes a failover it is possible to end up in a situation with two primaries, both nodes accepting writes. The primaries are then not synchronized and will end up in a split brain. Once one of the primaries joins the other as a secondary, the HA cluster is once again consistent but any out-of-sync changes will be overwritten.

To prevent split-brain from occurring, NSO 5.7 or later comes with a rule-based algorithm. The algorithm is enabled by default, it can be disabled or changed from the parameters:

* `/high-availability/settings/consensus/enabled [true]`
* `/high-availability/settings/consensus/algorithm [ncs:rule-based]`

The rule-based algorithm can be used in either of the two HA constellations:

* Two nodes: one nominal primary and one nominal secondary configured as failover-primary.
* Three nodes: one nominal primary, one nominal secondary configured as failover-primary, and one perpetual secondary.

On failover:

* Failover-primary: become primary but enable read-only mode. Once the secondary joins, disable read-only.
* Nominal primary: on loss of all secondaries, change role to none. If one secondary node is connected, stay primary.

{% hint style="info" %}
In certain cases, the rule-based consensus algorithm results in nodes being disconnected and will not automatically rejoin the HA cluster, such as in the example above when the nominal primary becomes none on the loss of all secondaries.
{% endhint %}

To restore the HA cluster one may need to manually invoke the `/high-availability/be-secondary-to` action.

{% hint style="info" %}
In the case where the failover-primary takes over as primary, it will enable administrator-configured read-only mode; if no secondary connects it will remain read-only. This is done to guarantee consistency.
{% endhint %}

{% hint style="info" %}
In a three-node cluster, when the nominal primary takes over as actual primary, it first enables administrator-configured read-only mode until a secondary connects. This is done to guarantee consistency.
{% endhint %}

The administrator-configured read-only mode can be manually controlled using the `/ncs-state/set-read-only` action. Note that disabling the administrator-configured read-only mode does not affect read-only mode imposed by the HA operational state (e.g. on secondary nodes).

When any node loses connection, this can also be observed in high-availability alarms as either a `ha-primary-down` or a `ha-secondary-down` alarm.

```bash
alarms alarm-list alarm ncs ha-primary-down /high-availability/ha-node[id='paris']
 is-cleared              false
 last-status-change      2022-05-30T10:02:45.706947+00:00
 last-perceived-severity critical
 last-alarm-text         "Lost connection to primary due to: Primary closed connection"
 status-change 2022-05-30T10:02:45.706947+00:00
  received-time      2022-05-30T10:02:45.706947+00:00
  perceived-severity critical
  alarm-text         "Lost connection to primary due to: Primary closed connection"
```

```bash
alarms alarm-list alarm ncs ha-secondary-down /high-availability/ha-node[id='london'] ""
 is-cleared              false
 last-status-change      2022-05-30T10:04:33.231808+00:00
 last-perceived-severity critical
 last-alarm-text         "Lost connection to secondary"
 status-change 2022-05-30T10:04:33.231808+00:00
  received-time      2022-05-30T10:04:33.231808+00:00
  perceived-severity critical
  alarm-text         "Lost connection to secondary"
```

### Startup <a href="#ug.ha.startup" id="ug.ha.startup"></a>

Startup behavior is defined by a combination of the parameters `/high-availability/settings/start-up/assume-nominal-role` and `/high-availability/settings/start-up/join-ha` as well as the node's nominal role:

<table><thead><tr><th width="188" valign="top">assume-nominal-role</th><th width="137" valign="top">join-ha</th><th width="154" valign="top">nominal-role</th><th valign="top">Behaviour</th></tr></thead><tbody><tr><td valign="top"><code>true</code></td><td valign="top"><code>false</code></td><td valign="top"><code>primary</code></td><td valign="top">Assume primary role.</td></tr><tr><td valign="top"><code>true</code></td><td valign="top"><code>false</code></td><td valign="top"><code>secondary</code></td><td valign="top">Attempt to connect as secondary to the node (if any), which has nominal-role primary. If this fails, make no retry attempts and assume none role.</td></tr><tr><td valign="top"><code>true</code></td><td valign="top"><code>false</code></td><td valign="top"><code>none</code></td><td valign="top">Assume none role</td></tr><tr><td valign="top"><code>false</code></td><td valign="top"><code>true</code></td><td valign="top"><code>primary</code></td><td valign="top">Attempt to join HA setup as secondary by querying for the current primary. Retries will be attempted. Retry attempt interval is defined by <code>/high-availability/settings/reconnect-interval</code>.</td></tr><tr><td valign="top"><code>false</code></td><td valign="top"><code>true</code></td><td valign="top"><code>secondary</code></td><td valign="top">Attempt to join HA setup as secondary by querying for the current primary. Retries will be attempted. Retry attempt interval is defined by <code>/high-availability/settings/reconnect-interval</code>. If all retry attempts fail, assume none role.</td></tr><tr><td valign="top"><code>false</code></td><td valign="top"><code>true</code></td><td valign="top"><code>none</code></td><td valign="top">Assume none role.</td></tr><tr><td valign="top"><code>true</code></td><td valign="top"><code>true</code></td><td valign="top"><code>primary</code></td><td valign="top">Query HA setup once for a node with primary role. If found, attempt to connect as secondary to that node. If no current primary is found, assume primary role.</td></tr><tr><td valign="top"><code>true</code></td><td valign="top"><code>true</code></td><td valign="top"><code>secondary</code></td><td valign="top">Attempt to join HA setup as secondary by querying for the current primary. Retries will be attempted. Retry attempt interval is defined by <code>/high-availability/settings/reconnect-interval</code>. If all retry attempts fail, assume none role.</td></tr><tr><td valign="top"><code>true</code></td><td valign="top"><code>true</code></td><td valign="top"><code>none</code></td><td valign="top">Assume none role.</td></tr><tr><td valign="top"><code>false</code></td><td valign="top"><code>false</code></td><td valign="top"><code>-</code></td><td valign="top">Assume none role.</td></tr></tbody></table>

### Actions <a href="#d5e5031" id="d5e5031"></a>

NSO rule-based HA can be controlled through several actions. All actions are found under `/high-availability/`. The available actions are listed below:

<table><thead><tr><th width="240" valign="top">Action</th><th valign="top">Description</th></tr></thead><tbody><tr><td valign="top"><code>be-primary</code></td><td valign="top">Order the local node to assume the HA role primary.</td></tr><tr><td valign="top"><code>be-none</code></td><td valign="top">Order the local node to assume the HA role none.</td></tr><tr><td valign="top"><code>be-secondary-to</code></td><td valign="top">Order the local node to connect as secondary to the provided HA node. This is an asynchronous operation; the result can be found under <code>/high-availability/status/be-secondary-result</code>.</td></tr><tr><td valign="top"><code>local-node-id</code></td><td valign="top">Identify which of the nodes in <code>/high-availability/ha-node</code> (if any) corresponds to the local NSO instance.</td></tr><tr><td valign="top"><code>enable</code></td><td valign="top">If disabled, enable NSO rule-based HA and optionally assume an HA role according to /high-availability/settings/start-up/ parameters.</td></tr><tr><td valign="top"><code>disable</code></td><td valign="top">Disable NSO rule-based HA and assume a HA role none.</td></tr></tbody></table>

### Status Check <a href="#d5e5077" id="d5e5077"></a>

The current state of NSO rule-based HA can be monitored by observing `/high-availability/status/`. Information can be found about the current active HA mode and the current assigned role. For nodes with active mode primary, a list of connected nodes and their source IP addresses is shown. For nodes with assigned role secondary the latest result of the be-secondary operation is listed. All NSO rule-based HA status information is non-replicated operational data - the result here will differ between nodes connected in an HA setup.

## Managing Certificates

Secure communication between NSO nodes in an HA setup (rule-based or Raft) relies on Transport Layer Security (TLS). TLS uses public key cryptography, where each node requires a separate public/private key pair and a corresponding certificate. Key and certificate management is a broad topic and is critical to the overall security of the system.

Certificates are used to authenticate each node in the cluster. The TLS protocol not only verifies that the certificate/key pair comes from a trusted source (certificate is signed by a trusted CA), it also checks that the certificate matches the host you are connecting to. This means host A may have a valid certificate and key, signed by a trusted CA; however, if the certificate is for another host, say host B, the authentication will fail. Typically, this means the host address must appear in the node certificate's Subject Alternative Name (SAN) extension, as `dNSName` or `iPAddress` (see [RFC2459](https://datatracker.ietf.org/doc/html/rfc2459)).

It is important that you create and use a self-signed CA to provision the node certificates and that the CA is used only to sign the certificates of the member nodes in one NSO HA cluster. Do not sign any other certificate with this CA - any certificate signed by the CA can be used to authenticate with and control the cluster.

{% hint style="danger" %}
Never disable TLS. Without TLS, anyone with access to the HA port can gain complete control of the NSO cluster.
{% endhint %}

NSO uses standard PEM-encoded X509 certificates for TLS, which you can generate with any compatible tool, such as the `openssl` command-line utility. Note that certificates have an expiry date and should be rotated regularly.

For your convenience, NSO contains a set of scripts for certificate management as part of the example set, located in [examples.ncs/high-availability/ca](https://github.com/NSO-developer/nso-examples/tree/6.7/high-availability/ca). These scripts call out to `openssl` and produce certificates with algorithms and key strength that are considered production-grade in many environments. However, always vet the scripts and the produced certificates before use.

The example set also contains examples for setting up Raft- and rule-based HA clusters using certificate management scripts, such as [examples.ncs/high-availability/raft-cluster](https://github.com/NSO-developer/nso-examples/tree/6.7/high-availability/raft-cluster) and [examples.ncs/high-availability/rule-based-basic](https://github.com/NSO-developer/nso-examples/tree/6.7/high-availability/rule-based-basic).

### Using Certificate Management Scripts

The scripts provided in `examples.ncs/high-availability/ca` are designed to simplify common certificate operations. You can use them to quickly provision node certificates for test and development clusters or as a base for your own scripts. Before using them for production, audit the scripts and ensure they match your requirements, such as generated certificate lifetime.

The scripts are part of the example set, which may need to be installed separately for an NSO system install. They depend on the system `openssl` command and are tested with versions 1.1 and 3.0 of `openssl`, but newer versions likely work as well. You can verify the version you have installed with the `openssl version` command. You should also verify the system time is set correctly, e.g. by running `date`, since certificates are time-bound.

Before you start, it is highly recommended that you create a new, separate directory to hold certificate data for each cluster. Including the cluster name in the directory name helps distinguish certificates of one HA cluster from another, such as when using an LSA deployment in an HA configuration. You may also want to create a symlink to the `$NCS_DIR/examples.ncs/high-availability/ca` to save you some typing.

```bash
$ mkdir ~/nso-cluster1-certs
$ cd ~/nso-cluster1-certs
$ ln -s "$NCS_DIR/examples.ncs/high-availability/ca" ./bin
```

To create a new certificate for a cluster node (here called `n1`):

```bash
$ bin/create-cert n1 192.0.2.1 n1.example.com
```

where `192.0.2.1` and `n1.example.com` are the IP address and hostname of the node; you can specify only one if you wish. The use of the first argument, `n1`, is not strictly necessary but gives the certificate a short name to make it easy to refer to.

Since the CA certificate does not exist yet, one is created and you must enter the CA key passphrase multiple times (to set it, then for verification and signing). You will need this passphrase for practically all certificate operations.

You can create more certificates, say for another node `n2`:

```bash
$ bin/create-cert n2 192.0.2.2 n2.example.com
```

Naturally, the certificates for one cluster must be created (signed) by the same CA, so create as many certificates as there are cluster nodes. You can check the existing certificates with the `list-certs` command.

To use the certificates with NSO, invoke `export-cert` for each certificate, specifying the folder where you want to save the files. If the filesystem of a node is directly accessible (via NFS or container bind mount for example), you can save the files directly next to the node's `ncs.conf`, in a folder named `tls` or similar. Otherwise, export to a temporary folder and securely transfer it to the node, e.g. using `scp`.

```bash
$ bin/export-cert n1 n1/tls
$ scp -r n1/tls n1.example.com:/etc/ncs/
```

Since certificates are valid for a limited time, you can check the validity with the `verify-cert` command, which takes into account the CRL if present:

```bash
$ bin/verify-cert n2
```

Executing these commands creates a directory named `ca-...` that holds all the certificate and auxiliary data, allowing you to renew or export certificates again. There are a number of other operations available too, see [$NCS\_DIR/examples.ncs/high-availability/ca/README.md](https://github.com/NSO-developer/nso-examples/tree/6.7/high-availability/ca/README.md) for details.

### Two-Node Example with HCC VIP <a href="#ug.ha.builtin.twonode" id="ug.ha.builtin.twonode"></a>

The following example summarizes a common two-node rule-based HA deployment together with HCC layer-2 VIP management. It applies to one nominal primary node and one nominal secondary node configured as `failover-primary`.

The two nodes must also share the same HA token. The example below uses generic addresses and a single VIP:

```bash
high-availability token <same-token-on-both-nodes>
high-availability ha-node nso-a
 address 192.0.2.10
 nominal-role primary
!
high-availability ha-node nso-b
 address 192.0.2.11
 nominal-role secondary
 failover-primary true
!
high-availability settings enable-failover true
high-availability settings reconnect-secondaries true
high-availability settings start-up assume-nominal-role true
high-availability settings start-up join-ha true
high-availability settings reconnect-interval 5
high-availability settings reconnect-attempts 3
hcc enabled
hcc vip-address [ 192.0.2.100 ]
```

For the consensus-enabled cases in the tables below, set `/high-availability/settings/consensus/enabled` to `true` and `/high-availability/settings/consensus/algorithm` to `ncs:rule-based`. For the consensus-disabled cases, set `/high-availability/settings/consensus/enabled` to `false`.

In steady state, the expected status is:

<table><thead><tr><th width="127" valign="top">Node</th><th width="276" valign="top">State</th><th valign="top">VIP state</th></tr></thead><tbody><tr><td valign="top"><code>nso-a</code></td><td valign="top"><code>mode primary</code>, <code>assigned-role primary</code>, <code>read-only-mode false</code></td><td valign="top">HCC binds the VIP on <code>nso-a</code>.</td></tr><tr><td valign="top"><code>nso-b</code></td><td valign="top"><code>mode secondary</code>, <code>assigned-role secondary</code>, <code>read-only-mode false</code></td><td valign="top">No VIP is bound on <code>nso-b</code>.</td></tr></tbody></table>

#### **HA Events with Consensus Enabled**

With consensus enabled, some two-node failure cases deliberately reduce availability in order to avoid split brain and protect data consistency.

<table><thead><tr><th width="175" valign="top">HA event</th><th width="281" valign="top">State change</th><th width="154" valign="top">VIP state</th><th valign="top">Manual action</th></tr></thead><tbody><tr><td valign="top">Secondary node is shut down or lost</td><td valign="top"><code>nso-a</code> raises <code>ha-secondary-down</code> and changes from <code>primary</code> to <code>none</code> with <code>read-only-mode false</code>.<br><code>nso-b</code> is unavailable.</td><td valign="top">No node has the VIP bound because there is no current primary.</td><td valign="top">If service must continue before <code>nso-b</code> returns, promote <code>nso-a</code> with <code>/high-availability/be-primary</code>.</td></tr><tr><td valign="top">Secondary node returns while no node is primary</td><td valign="top">The previous <code>ha-secondary-down</code> alarm clears when connectivity is restored.<br><code>nso-a</code> stays in <code>none</code>.<br><code>nso-b</code> retries according to the reconnect settings and then also settles in <code>none</code>.</td><td valign="top">The VIP remains unbound.</td><td valign="top">Select one node to become primary with <code>/high-availability/be-primary</code>, then use <code>/high-availability/be-secondary-to</code> on the other node.</td></tr><tr><td valign="top">Secondary node returns after the surviving node has been promoted back to primary</td><td valign="top">The previous <code>ha-secondary-down</code> alarm clears.<br><code>nso-a</code> remains <code>primary</code> with <code>read-only-mode false</code>.<br><code>nso-b</code> joins as <code>secondary</code>.</td><td valign="top">The VIP is bound on <code>nso-a</code>.</td><td valign="top">No action required.</td></tr><tr><td valign="top">Primary node is shut down or lost</td><td valign="top"><code>nso-a</code> is unavailable.<br><code>nso-b</code> raises <code>ha-primary-down</code>, retries, and then becomes <code>primary</code> with <code>read-only-mode true</code>.</td><td valign="top">The VIP moves to <code>nso-b</code>.</td><td valign="top">If write access must resume before another secondary joins, clear the administrator-configured read-only mode with <code>/ncs-state/set-read-only mode false</code>.</td></tr><tr><td valign="top">Network partition between the nodes</td><td valign="top">No split brain occurs.<br><code>nso-a</code> raises <code>ha-secondary-down</code> and changes to <code>none</code> with <code>read-only-mode false</code>.<br><code>nso-b</code> raises <code>ha-primary-down</code> and becomes <code>primary</code> with <code>read-only-mode true</code>.</td><td valign="top">The VIP is bound on <code>nso-b</code>.</td><td valign="top">Pick the node that should remain primary, promote it if needed, clear administrator-configured read-only mode if needed, and rejoin the other node with <code>/high-availability/be-secondary-to</code>.</td></tr></tbody></table>

For a three-node rule-based cluster, the consensus-enabled cases use one nominal primary (`n1`), one nominal secondary configured as `failover-primary` (`n2`), and one perpetual secondary (`n3`). This table describes only NSO rule-based HA state; NSO rule-based HA does not manage a VIP by itself.

<table><thead><tr><th width="209" valign="top">HA event</th><th width="401" valign="top">State change</th><th valign="top">Manual action</th></tr></thead><tbody><tr><td valign="top">One secondary node is shut down or lost while the primary can still reach the other secondary</td><td valign="top"><code>n1</code> raises <code>ha-secondary-down</code> for the lost secondary and remains <code>primary</code> with <code>read-only-mode false</code>.<br>The remaining secondary stays <code>secondary</code> to <code>n1</code> with <code>read-only-mode false</code>.</td><td valign="top">No action required.</td></tr><tr><td valign="top">The stopped or lost secondary returns while a primary exists</td><td valign="top">The previous <code>ha-secondary-down</code> alarm clears when connectivity is restored.<br>The returning node joins the current primary as <code>secondary</code> with <code>read-only-mode false</code>.</td><td valign="top">No action required.</td></tr><tr><td valign="top">Both secondary nodes are shut down or lost while the nominal primary remains available</td><td valign="top"><code>n1</code> raises <code>ha-secondary-down</code> for both secondaries and changes from <code>primary</code> to <code>none</code> with <code>read-only-mode false</code>.</td><td valign="top">If service must continue, promote the intended primary with <code>/high-availability/be-primary</code> and connect at least one secondary with <code>/high-availability/be-secondary-to</code>.</td></tr><tr><td valign="top">Secondary nodes return while no node is primary</td><td valign="top"><code>n1</code>, <code>n2</code>, and <code>n3</code> settle in <code>none</code>; no automatic primary is selected.</td><td valign="top">Promote one node with <code>/high-availability/be-primary</code>, then connect the other nodes to it with <code>/high-availability/be-secondary-to</code>.</td></tr><tr><td valign="top">Nominal primary is shut down or lost while both secondaries are available</td><td valign="top"><code>n1</code> is unavailable.<br><code>n2</code> raises <code>ha-primary-down</code>, retries, and becomes <code>primary</code>. It may enter administrator-configured read-only mode during takeover; once <code>n3</code> joins as secondary, <code>n2</code> has <code>read-only-mode false</code>.<br><code>n3</code> reconnects as <code>secondary</code> to <code>n2</code>.</td><td valign="top">No action required.</td></tr><tr><td valign="top">Former nominal primary returns after failover</td><td valign="top"><code>n1</code> finds the current primary, <code>n2</code>, and joins as <code>secondary</code> with <code>read-only-mode false</code>.<br><code>n2</code> remains <code>primary</code> and <code>n3</code> remains <code>secondary</code>.</td><td valign="top">No action required. To return to the nominal roles, use <code>/high-availability/be-primary</code> on <code>n1</code> and <code>/high-availability/be-secondary-to</code> on <code>n2</code> and <code>n3</code>.</td></tr><tr><td valign="top">Nominal primary and perpetual secondary are shut down or lost, leaving only the failover-primary node</td><td valign="top"><code>n2</code> raises <code>ha-primary-down</code>, retries, and becomes <code>primary</code> with <code>read-only-mode true</code> because no secondary is connected.</td><td valign="top">If write access must resume before any secondary returns, clear the administrator-configured read-only mode with <code>/ncs-state/set-read-only mode false</code>. Otherwise, wait for <code>n1</code> or <code>n3</code> to return and join as secondary.</td></tr><tr><td valign="top">A secondary joins a read-only failover-primary</td><td valign="top">When <code>n1</code> or <code>n3</code> joins <code>n2</code> as secondary, <code>n2</code> remains <code>primary</code> and changes to <code>read-only-mode false</code>.</td><td valign="top">No action required.</td></tr><tr><td valign="top">Network partition splits the cluster</td><td valign="top">No writable split brain occurs.<br>The side that has a primary with at least one secondary remains or becomes the writable HA group.<br>A nominal primary with no secondaries changes to <code>none</code>; a failover-primary with no secondary can only take over with <code>read-only-mode true</code>.</td><td valign="top">After connectivity returns, verify the intended primary and reconnect other nodes with <code>/high-availability/be-secondary-to</code> as needed.</td></tr></tbody></table>

#### **HA Events with Consensus Disabled**

With consensus disabled, the two-node setup favors availability during ordinary failover, but a network partition can lead to split brain.

<table><thead><tr><th width="175" valign="top">HA event</th><th width="281" valign="top">State change</th><th width="154" valign="top">VIP state</th><th valign="top">Manual action</th></tr></thead><tbody><tr><td valign="top">Primary node is shut down or lost</td><td valign="top"><code>nso-a</code> is unavailable.<br><code>nso-b</code> raises <code>ha-primary-down</code>, retries, and then becomes <code>primary</code> with <code>read-only-mode false</code>.</td><td valign="top">The VIP moves to <code>nso-b</code>.</td><td valign="top">No action required.</td></tr><tr><td valign="top">Former primary returns after failover</td><td valign="top">The previous <code>ha-primary-down</code> alarm clears.<br><code>nso-a</code> looks for the current primary and rejoins as <code>secondary</code> with <code>read-only-mode false</code>.<br><code>nso-b</code> remains <code>primary</code>.</td><td valign="top">The VIP remains bound on <code>nso-b</code>.</td><td valign="top">No action required.</td></tr><tr><td valign="top">Secondary node is shut down or lost while the primary remains available</td><td valign="top"><code>nso-a</code> raises <code>ha-secondary-down</code> and stays <code>primary</code> with <code>read-only-mode false</code>.<br><code>nso-b</code> is unavailable.</td><td valign="top">The VIP remains bound on <code>nso-a</code>.</td><td valign="top">No action required.</td></tr><tr><td valign="top">Secondary node returns</td><td valign="top">The previous <code>ha-secondary-down</code> alarm clears.<br><code>nso-a</code> stays <code>primary</code> with <code>read-only-mode false</code>.<br><code>nso-b</code> reconnects and becomes <code>secondary</code> with <code>read-only-mode false</code>.</td><td valign="top">The VIP remains bound on <code>nso-a</code>.</td><td valign="top">No action required.</td></tr><tr><td valign="top">Network partition between the nodes</td><td valign="top">Split brain occurs.<br><code>nso-a</code> raises <code>ha-secondary-down</code> and remains <code>primary</code> with <code>read-only-mode false</code>.<br><code>nso-b</code> raises <code>ha-primary-down</code> and also becomes <code>primary</code> with <code>read-only-mode false</code>.</td><td valign="top">Both nodes can bind the same VIP until the split brain is resolved.</td><td valign="top">Manual recovery is required. Choose the authoritative primary, avoid further writes on the other node, and reconnect that node as a secondary with <code>/high-availability/be-secondary-to</code>.</td></tr></tbody></table>

## Tail-f HCC Package <a href="#ug.ha.hcc" id="ug.ha.hcc"></a>

The Tail-f HCC package extends the built-in HA functionality by providing virtual IP addresses (VIPs) that can be used to connect to the NSO HA group primary node. HCC ensures that the VIP addresses are always bound by the HA group primary and never bound by a secondary. Each time a node transitions between primary and secondary states HCC reacts by binding (primary) or unbinding (secondary) the VIP addresses.

HCC manages IP addresses at the link layer (OSI layer 2) for Ethernet interfaces, and optionally, also at the network layer (OSI layer 3) using BGP router advertisements. The layer-2 and layer-3 functions are mostly independent and this document describes the details of each one separately. However, the layer-3 function builds on top of the layer-2 function. The layer-2 function is always necessary, otherwise, the Linux kernel on the primary node would not recognize the VIP address or accept traffic directed to it.

{% hint style="info" %}
Tail-f HCC version 5.x is non-backward compatible with previous versions of Tail-f HCC and requires functionality provided by NSO version 5.4 and greater. For more details, see the [following chapter](#ug.ha.hcc.compared).
{% endhint %}

### Dependencies <a href="#ug.ha.hcc.deps" id="ug.ha.hcc.deps"></a>

Both the HCC layer-2 VIP and layer-3 BGP functionality depend on `iproute2` utilities and `awk`. An optional dependency is `arping` (either from `iputils` or Thomas Habets `arping` implementation) which allows HCC to announce the VIP to MAC mapping to all nodes in the network by sending Gratuitous ARP (GARP) requests for IPv4 addresses. For IPv6 VIP addresses, the optional `ndsend` tool (from the `ndisc6` package) enables HCC to send unsolicited Neighbor Advertisements to update NDP caches on the local link, providing equivalent functionality to GARP for IPv6.

The HCC layer-3 BGP functionality depends on the [`GoBGP`](https://osrg.github.io/gobgp/) daemon version 2.x being installed on each NSO host that is configured to run HCC in BGP mode.

GoBGP is open-source software originally developed by NTT Communications and released under the Apache License 2.0. GoBGP can be obtained directly from <https://osrg.github.io/gobgp/> and is also packaged for mainstream Linux distributions.

The HCC layer-3 DNS Update functionality depends on the command line utility `nsupdate`.

Tools Dependencies are listed below:

<table><thead><tr><th width="192">Tool</th><th width="164">Package</th><th width="128">Required</th><th>Description</th></tr></thead><tbody><tr><td><code>ip</code></td><td><code>iproute2</code></td><td>yes</td><td>Adds and deletes the virtual IP from the network interface.</td></tr><tr><td><code>awk</code></td><td><code>mawk</code> or <code>gawk</code></td><td>yes</td><td>Installed with most Linux distributions.</td></tr><tr><td><code>sed</code></td><td><code>sed</code></td><td>yes</td><td>Installed with most Linux distributions.</td></tr><tr><td><code>arping</code></td><td><code>iputils</code> or <code>arping</code></td><td>optional</td><td>Recommended for IPv4 VIP configurations. Sends Gratuitous ARP requests to update ARP caches on the local network when a VIP moves between nodes.</td></tr><tr><td><code>ndsend</code></td><td><code>ndisc6</code></td><td>optional</td><td>Recommended for IPv6 VIP configurations. Sends unsolicited Neighbor Advertisements to update NDP caches on the local link when a VIP moves between nodes (IPv6 equivalent of Gratuitous ARP).</td></tr><tr><td><code>gobgpd</code> and <code>gobgp</code></td><td><code>GoBGP 2.x</code></td><td>optional</td><td>Required for layer-3 configurations. gobgpd is started by the HCC package and advertises the virtual IP using BGP. gobgp is used to get advertised routes.</td></tr><tr><td><code>nsupdate</code></td><td><code>bind-tools</code> or <code>knot-dnsutils</code></td><td>optional</td><td>Required for layer-3 DNS update functionality and is used to submit Dynamic DNS Update requests to a name server.</td></tr></tbody></table>

Same as with built-in HA functionality, all NSO instances must be configured to run in HA mode. See the [following instructions](#ha.moo) on how to enable HA on NSO instances.

### Running the HCC Package with NSO as a Non-Root User <a href="#ug.ha.hcc.nonroot" id="ug.ha.hcc.nonroot"></a>

By default, GoBGP binds to TCP port 179 at startup. Since port 179 is a privileged port (below 1024), it requires root privileges. When NSO runs as a non-root user, GoBGP runs as the same user and cannot bind to port 179.

There are several ways to handle this:

1. **Configure an alternative port or disable the listener** (recommended when applicable). The `port` parameter under `/hcc/bgp/node{id}` controls which TCP port GoBGP listens on. Disabling the listener can work when HCC/GoBGP is configured to actively initiate sessions to its neighbors and those peers are listening for and allow inbound TCP connections from HCC. In deployments where the peer is expected to initiate the session, or where policy/firewall rules do not allow the peer to accept the connection, keep a listener enabled.

   To disable the listener entirely:

   ```cli
   admin@ncs(config)# hcc bgp node paris port disabled
   admin@ncs(config)# commit
   ```

   To use a non-privileged port (e.g., 1790):

   ```cli
   admin@ncs(config)# hcc bgp node paris port 1790
   admin@ncs(config)# commit
   ```
2. Set capability `CAP_NET_BIND_SERVICE` on the `gobgpd` file. This allows GoBGP to bind port 179 without full root privileges. May not be supported by all Linux distributions.

   ```bash
   $ sudo setcap 'cap_net_bind_service=+ep' /usr/bin/gobgpd
   ```
3. Set the owner to `root` and the `setuid` bit of the `gobgpd` file. Works on all Linux distributions but grants broader privileges than option 2.

   ```bash
   $ sudo chown root /usr/bin/gobgpd
   $ sudo chmod u+s /usr/bin/gobgpd
   ```
4. The `vipctl` script, included in the HCC package, uses `sudo` to run the `ip`, `arping`, and `ndsend` commands when NSO is not running as root. If `sudo` is used, you must ensure it does not require password input. For example, if NSO runs as `admin` user, the `sudoers` file can be edited similarly to the following:

   ```bash
   $ sudo echo "admin ALL = (root) NOPASSWD: /bin/ip" >> /etc/sudoers
   $ sudo echo "admin ALL = (root) NOPASSWD: /path/to/arping" >> /etc/sudoers
   $ sudo echo "admin ALL = (root) NOPASSWD: /path/to/ndsend" >> /etc/sudoers
   ```

{% hint style="info" %}
Options 2 and 3 are only needed if GoBGP must accept incoming BGP connections on port 179. Option 4 applies to the VIP management functionality and is independent of the BGP port setting.
{% endhint %}

### Tail-f HCC Compared with HCC Version 4.x and Older <a href="#ug.ha.hcc.compared" id="ug.ha.hcc.compared"></a>

#### **HA Group Management Decisions**

Tail-f HCC 5.x or later does not participate in decisions on which NSO node is primary or secondary. These decisions are taken by NSO's built-in HA and then pushed as notifications to HCC. The NSO built-in HA functionality is available in NSO starting with version 5.4, where older NSO versions are not compatible with the HCC 5.x or later.

#### **Embedded BGP Daemon**

HCC 5.x or later operates a GoBGP daemon as a subprocess completely managed by NSO. The old HCC function pack interacted with an external Quagga BGP daemon using a NED interface.

#### **Automatic Interface Assignment**

HCC 5.x or later automatically associates VIP addresses with Linux network interfaces using the `ip` utility from the iproute2 package. VIP addresses are also treated as `/32` without defining a new subnet. The old HCC function pack used explicit configuration to associate VIPs with existing addresses on each NSO host and define IP subnets for VIP addresses.

### Upgrading <a href="#ug.ha.hcc.upgrade" id="ug.ha.hcc.upgrade"></a>

Since version 5.0, HCC relies on the NSO built-in HA for cluster management and only performs address or route management in reaction to cluster changes. Therefore, no special measures are necessary if using HCC when performing an NSO version upgrade or a package upgrade. Instead, you should follow the standard best practice HA upgrade procedure from [NSO HA Version Upgrade](https://nso-docs.cisco.com/guides/administration/management/pages/mKh0oXFZkQsBaoCMMyEo#ch_upgrade.ha).

A reference to upgrade examples can be found in the README under [examples.ncs/high-availability/hcc](https://github.com/NSO-developer/nso-examples/tree/6.7/high-availability/hcc).

### Layer-2 <a href="#ug.ha.hcc.layer2" id="ug.ha.hcc.layer2"></a>

The purpose of the HCC layer-2 functionality is to ensure that the configured VIP addresses are bound in the Linux kernel of the NSO primary node only. This ensures that the primary node (and only the primary node) will accept traffic directed toward the VIP addresses.

HCC also notifies the local layer-2 network when VIP addresses are bound. For IPv4 addresses, HCC sends Gratuitous ARP (GARP) packets via `arping`. For IPv6 addresses, HCC sends unsolicited Neighbor Advertisements via `ndsend` (if installed). In both cases, nodes on the network update their address resolution caches (ARP for IPv4, NDP for IPv6) with the new mapping so they can continue to send traffic to the non-failed, now primary node.

#### **Operational Details**

HCC binds the VIP addresses as additional (alias) addresses on existing Linux network interfaces (e.g. `eth0`). The network interface for each VIP is chosen automatically by performing a kernel routing lookup on the VIP address. That is, the VIP will automatically be associated with the same network interface that the Linux kernel chooses to send traffic to the VIP.

This means that you can map each VIP onto a particular interface by defining a route for a subnet that includes the VIP. If no such specific route exists the VIP will automatically be mapped onto the interface of the default gateway.

{% hint style="info" %}
To check which interface HCC will choose for a particular VIP address, simply run for example and look at the device `dev` in the output, for example `eth0`:

```bash
admin@paris:~$ ip route get 192.168.123.22
```

{% endhint %}

#### **Configuration**

The layer-2 functionality is configured by providing a list of IPv4 and/or IPv6 VIP addresses and enabling HCC. The VIP configuration parameters are found under `/hcc:hcc`.

Global Layer-2 Configuration:

<table><thead><tr><th width="168" valign="top">Parameters</th><th width="199" valign="top">Type</th><th valign="top">Description</th></tr></thead><tbody><tr><td valign="top"><code>enabled</code></td><td valign="top">boolean</td><td valign="top">If set to 'true', the primary node in an HA group automatically binds the set of Virtual IPv[46] addresses.</td></tr><tr><td valign="top"><code>vip-address</code></td><td valign="top">list of inet:ip-address</td><td valign="top">The list of virtual IPv[46] addresses to bind on the primary node. The addresses are automatically unbound when a node becomes secondary. The addresses can therefore be used externally to reliably connect to the HA group primary node.</td></tr></tbody></table>

#### **Example Configuration**

```bash
admin@ncs(config)# hcc enabled
admin@ncs(config)# hcc vip 192.168.123.22
admin@ncs(config)# hcc vip 2001:db8::10
admin@ncs(config)# commit
```

#### **VIP Behavior in Three-Node Raft-based HA**

In the three-node HA Raft example in [NSO HA Raft](#ch_ha.raft.threenode), HCC layer-2 VIP ownership follows the node that currently has `role leader`:

* Steady state: the VIP is bound on the current leader only.
* One follower down or returning: the VIP stays on the same leader.
* Leader failover with the remaining two nodes still connected: the VIP moves to the newly elected leader.
* One node isolated from the other two: the VIP follows the leader on the majority side, while the isolated node eventually rejoins as a follower.
* Complete quorum loss: do not rely on the VIP for writes until quorum is restored and the cluster has a stable leader again.
* If a follower briefly disconnects and the same node remains leader, HCC keeps the VIP on that same leader.

If a failover leaves leadership on a different node than you prefer, use `ha-raft handover` after the cluster has recovered to move the VIP back together with leadership.

#### **VIP Behavior in Two-Node Rule-based HA**

In the two-node rule-based HA example in [NSO Rule-based HA](#ug.ha.builtin.twonode), HCC layer-2 VIP ownership follows the node that currently has `mode primary`:

* Steady state: the VIP is bound on the nominal primary only.
* Consensus enabled, nominal primary changes to `none`: no node owns the VIP, so HCC unbinds it.
* Consensus enabled, failover-primary takes over: the VIP moves to the failover-primary node, which may still be administrator-configured read-only.
* Consensus disabled, ordinary failover: the VIP moves to the surviving primary and stays there when the failed node returns as a secondary.
* Consensus disabled, network partition: both nodes may end up in `primary` mode and both may bind the same VIP until the split brain is manually resolved.

### Layer-3 BGP <a href="#ug.ha.hcc.layer3" id="ug.ha.hcc.layer3"></a>

The purpose of the HCC layer-3 BGP functionality is to operate a BGP daemon on each NSO node and to ensure that routes for the VIP addresses are advertised by the BGP daemon on the primary node only.

The layer-3 functionality is an optional add-on to the layer-2 functionality. When enabled, the set of BGP neighbors must be configured separately for each NSO node. Each NSO node operates an embedded BGP daemon and maintains connections to peers but only the primary node announces the VIP addresses.

The layer-3 functionality relies on the layer-2 functionality to assign the virtual IP addresses to one of the host's interfaces. One notable difference in assigning virtual IP addresses when operating in Layer-3 mode is that the virtual IP addresses are assigned to the loopback interface `lo` rather than to a specific physical interface.

#### **Operational Details**

HCC operates a [`GoBGP`](https://osrg.github.io/gobgp/) subprocess as an embedded BGP daemon. The BGP daemon is started, configured, and monitored by HCC. The HCC YANG model includes basic BGP configuration data and state data.

Operational data in the YANG model includes the state of the BGP daemon subprocess and the state of each BGP neighbor connection. The BGP daemon writes log messages directly to NSO where the HCC module extracts updated operational data and then repeats the BGP daemon log messages into the HCC log verbatim. You can find these log messages in the developer log (`devel.log`).

```bash
admin@ncs# show hcc
NODE    BGPD  BGPD
ID      PID   STATUS   ADDRESS       STATE        CONNECTED
-------------------------------------------------------------
london  -     -        192.168.30.2  -            -
paris   827   running  192.168.31.2  ESTABLISHED  true
```

{% hint style="info" %}
GoBGP must be installed separately. The `gobgp` and `gobgpd` binaries must be found in paths specified by the `$PATH` environment variable. For system install, NSO reads `$PATH` in the `systemd` service file `/etc/systemd/system/ncs.service`. Since tailf-hcc 6.0.2, the path to `gobgp`/`gobgpd` is no longer possible to specify from the configuration data leaf `/hcc/bgp/node/gobgp-bin-dir`. The leaf has been removed from the `tailf-hcc/src/yang/tailf-hcc.yang` module.

Upgrades: If BGP is enabled and the `gobgp` or `gobgpd` binaries are not found, the tailf-hcc package will fail to load. The user must then install GoBGP and invoke the `packages reload` action or restart NSO with `NCS_RELOAD_PACKAGES=true` in `/etc/ncs/ncs.systemd.conf` and `systemctl restart ncs`.
{% endhint %}

#### **Configuration**

The layer-3 BGP functionality is configured as a list of BGP configurations with one list entry per node. Configurations are separate because each NSO node usually has different BGP neighbors with their own IP addresses, authentication parameters, etc.

The BGP configuration parameters are found under `/hcc:hcc/bgp/node{id}`.

Per-Node Layer-3 Configuration:

<table><thead><tr><th width="194">Parameters</th><th width="190">Type</th><th>Description</th></tr></thead><tbody><tr><td><code>node-id</code></td><td><code>string</code></td><td>Unique node ID. A reference to <code>/ncs:high-availability/ha-node/id</code>.</td></tr><tr><td><code>enabled</code></td><td><code>boolean</code></td><td>If set to <code>true</code>, this node uses BGP to announce VIP addresses when in the HA primary state.</td></tr><tr><td><code>as</code></td><td><code>inet:as-number</code></td><td>The BGP Autonomous System Number for the local BGP daemon.</td></tr><tr><td><code>router-id</code></td><td><code>inet:ip-address</code></td><td>The router ID for the local BGP daemon.</td></tr><tr><td><code>port</code></td><td><code>union (enumeration | int32)</code></td><td>The TCP port GoBGP listens on. Default is <code>179</code>. Set to <code>disabled</code> to turn off the listener entirely, which avoids binding to privileged port <code>179</code> and allows running GoBGP without root privileges when peers accept outbound-only connections. Alternatively, set to a non-privileged port (e.g., <code>1790</code>). See <a href="#ug.ha.hcc.nonroot">Running as Non-Root User</a>.</td></tr></tbody></table>

Each NSO node can connect to a different set of BGP neighbors. For each node, the BGP neighbor list configuration parameters are found under `/hcc:hcc/bgp/node{id}/neighbor{address}`.

Per-Neighbor BGP Configuration:

<table><thead><tr><th width="178" valign="top">Parameters</th><th width="201" valign="top">Type</th><th valign="top">Description</th></tr></thead><tbody><tr><td valign="top"><code>address</code></td><td valign="top"><code>inet:ip-address</code></td><td valign="top">BGP neighbor IP address.</td></tr><tr><td valign="top"><code>as</code></td><td valign="top"><code>inet:as-number</code></td><td valign="top">BGP neighbor Autonomous System Number.</td></tr><tr><td valign="top"><code>ttl-min</code></td><td valign="top"><code>uint8</code></td><td valign="top">Optional minimum TTL value for BGP packets. When configured, enables BGP Generalized TTL Security Mechanism (GTSM).</td></tr><tr><td valign="top"><code>password</code></td><td valign="top"><code>string</code></td><td valign="top">Optional password to use for BGP authentication with this neighbor.</td></tr><tr><td valign="top"><code>enabled</code></td><td valign="top"><code>boolean</code></td><td valign="top">If set to <code>true</code>, then an outgoing BGP connection to this neighbor is established by the HA group primary node.</td></tr></tbody></table>

#### **Example**

```bash
admin@ncs(config)# hcc bgp node paris enabled
admin@ncs(config)# hcc bgp node paris as 64512
admin@ncs(config)# hcc bgp node paris router-id 192.168.31.99
admin@ncs(config)# hcc bgp node paris neighbor 192.168.31.2 as 64514
admin@ncs(config)# ... repeated for each neighbor if more than one ...
            ... repeated for each node ...
admin@ncs(config)# commit
```

### Layer-3 DNS Update

The purpose of the HCC layer-3 DNS Update functionality is to notify a DNS server of the IP address change of the active primary NSO server, allowing the DNS server to update the DNS record for the given domain name.

Geographically redundant NSO setup typically relies on DNS support. To enable this use case, tailf-hcc can dynamically update DNS with the `nsupdate` utility on HA status change notification.

The DNS server used should support updates through `nsupdate` command (RFC 2136).

#### Operational Details

HCC listens on the underlying NSO HA notifications stream. When HCC receives a notification about an NSO node being Primary, it updates the DNS Server with the IP address of the Primary NSO for the given hostname. The HCC YANG model includes basic DNS configuration data and operational status data.

Operational data in the YANG model includes the result of the latest DNS update operation.

```bash
admin@ncs# show hcc dns
hcc dns status time 2023-10-20T23:16:33.472522+00:00
hcc dns status exit-code 0
```

If the DNS Update is unsuccessful, an error message will be populated in operational data, for example:

```bash
admin@ncs# show hcc dns
hcc dns status time 2023-10-20T23:36:33.372631+00:00
hcc dns status exit-code 2
hcc dns status error-message "; Communication with 10.0.0.10#53 failed: timed out"
```

{% hint style="info" %}
The DNS Server must be installed and configured separately, and details are provided to HCC as configuration data. The DNS Server must be configured to update the reverse DNS record.
{% endhint %}

#### Configuration

The layer-3 DNS Update functionality needs DNS-related information like DNS server IP address, port, zone, etc, and information about NSO nodes involved in HA - node, ip, and location.

The DNS configuration parameters are found under `/hcc:hcc/dns`.

Layer-3 DNS Configuration:

<table><thead><tr><th valign="top">Parameters</th><th valign="top">Type</th><th valign="top">Description</th></tr></thead><tbody><tr><td valign="top"><code>enabled</code></td><td valign="top">boolean</td><td valign="top">If set to <code>true</code>, DNS updates will be enabled.</td></tr><tr><td valign="top"><code>fqdn</code></td><td valign="top">inet:domain-name</td><td valign="top">DNS domain-name for the HA primary.</td></tr><tr><td valign="top"><code>ttl</code></td><td valign="top">uint32</td><td valign="top">Time to live for DNS record, default 86400.</td></tr><tr><td valign="top"><code>key-file</code></td><td valign="top">string</td><td valign="top">Specifies the file path for <code>nsupdate</code> keyfile.</td></tr><tr><td valign="top"><code>server</code></td><td valign="top">inet:ip-address</td><td valign="top">DNS Server IP Address.</td></tr><tr><td valign="top"><code>port</code></td><td valign="top">uint32</td><td valign="top">DNS Server port, default 53.</td></tr><tr><td valign="top"><code>zone</code></td><td valign="top">inet:host</td><td valign="top">DNS Zone to update on the server.</td></tr><tr><td valign="top"><code>timeout</code></td><td valign="top">uint32</td><td valign="top">Timeout for <code>nsupdate</code> command, default 300.</td></tr></tbody></table>

Each NSO node can be placed in a separate Location/Site/Availability-Zone. This is configured as a list member configuration, with one list entry per node ID. The member list configuration parameters are found under `/hcc:hcc/dns/member{node-id}`.

<table><thead><tr><th valign="top">Parameter</th><th valign="top">Type</th><th valign="top">Description</th></tr></thead><tbody><tr><td valign="top"><code>node-id</code></td><td valign="top">string</td><td valign="top">Unique NSO HA node ID. Valid values are: /high-availability/ha-node when built-in HA is used or /ha-raft/status/member for HA Raft.</td></tr><tr><td valign="top"><code>ip-address</code></td><td valign="top">inet:ip-address</td><td valign="top">IP where NSO listens for incoming requests to any northbound interfaces.</td></tr><tr><td valign="top"><code>location</code></td><td valign="top">string</td><td valign="top">Name of the Location/Site/Availability-Zone where node is placed.</td></tr></tbody></table>

#### Example

Here is an example configuration for a setup of two dual-stack NSO nodes, node-1 and node-2, that have an IPv4 and an IPv6 address configured. The configuration also sets up an update signing with the specified key.

```bash
admin@ncs(config)#  hcc dns enabled
admin@ncs(config)#  hcc dns fqdn example.com
admin@ncs(config)#  hcc dns ttl 120
admin@ncs(config)#  hcc dns key-file /home/cisco/DNS-testing/good.key
admin@ncs(config)#  hcc dns server 10.0.0.10
admin@ncs(config)#  hcc dns port 53
admin@ncs(config)#  hcc dns zone zone1.nso
admin@ncs(config)#  hcc dns member node-1 ip-address [ 10.0.0.20 ::10 ]
admin@ncs(config)#  hcc dns member node-1 location SanJose
admin@ncs(config)#  hcc dns member node-2 ip-address [ 10.0.0.30 ::20 ]
admin@ncs(config)#  hcc dns member node-2 location NewYork
admin@ncs(config)# commit
```

### Usage

This section describes basic deployment scenarios for HCC. Layer-2 mode is demonstrated first and then the layer-3 BGP functionality is configured in addition:

* [Layer-2 Deployment](#layer-2-deployment)
* [Enabling Layer-3 BGP](#enabling-layer-3-bgp)
* [Enabling Layer-3 DNS](#enabling-layer-3-dns)

A reference to container-based examples for the layer-2 and layer-3 deployment scenarios described here can be found in the NSO example set under [examples.ncs/high-availability/hcc](https://github.com/NSO-developer/nso-examples/tree/6.7/high-availability/hcc).

Both scenarios consist of two test nodes: `london` and `paris` with a single IPv4 VIP address. For the layer-2 scenario, the nodes are on the same network. The layer-3 scenario also involves a BGP-enabled `router` node as the `london` and `paris` nodes are on two different networks.

#### **Layer-2 Deployment**

The layer-2 operation is configured by simply defining the VIP addresses and enabling HCC. The HCC configuration on both nodes should match, otherwise, the primary node's configuration will overwrite the secondary node configuration when the secondary connects to the primary node.

Addresses:

<table><thead><tr><th valign="top">Hostname</th><th valign="top">Address</th><th valign="top">Role</th></tr></thead><tbody><tr><td valign="top"><code>paris</code></td><td valign="top">192.168.23.99</td><td valign="top">Paris service node.</td></tr><tr><td valign="top"><code>london</code></td><td valign="top">192.168.23.98</td><td valign="top">London service node.</td></tr><tr><td valign="top"><code>vip4</code></td><td valign="top">192.168.23.122</td><td valign="top">NSO primary node IPv4 VIP address.</td></tr></tbody></table>

Configuring VIPs:

```bash
admin@ncs(config)# hcc enabled
admin@ncs(config)# hcc vip 192.168.23.122
admin@ncs(config)# commit
```

Verifying VIP Availability:

Once enabled, HCC on the HA group primary node will automatically assign the VIP addresses to corresponding Linux network interfaces.

```bash
root@paris:/var/log/ncs# ip address list
1: lo: <LOOPBACK,UP,LOWER_UP> mtu 65536 qdisc noqueue state UNKNOWN group default qlen 1
    link/loopback 00:00:00:00:00:00 brd 00:00:00:00:00:00
    inet 127.0.0.1/8 scope host lo
    valid_lft forever preferred_lft forever
    inet6 ::1/128 scope host
    valid_lft forever preferred_lft forever
2: enp0s3: <BROADCAST,MULTICAST,UP,LOWER_UP> mtu 1500 qdisc pfifo_fast state UP group default qlen 1000
    link/ether 52:54:00:fa:61:99 brd ff:ff:ff:ff:ff:ff
    inet 192.168.23.99/24 brd 192.168.23.255 scope global enp0s3
    valid_lft forever preferred_lft forever
    inet 192.168.23.122/32 scope global enp0s3
    valid_lft forever preferred_lft forever
    inet6 fe80::5054:ff:fefa:6199/64 scope link
    valid_lft forever preferred_lft forever
```

On the secondary node, HCC will not configure these addresses.

```bash
root@london:~# ip address list
1: lo: <LOOPBACK,UP,LOWER_UP> mtu 65536 ...
    link/loopback 00:00:00:00:00:00 brd 00:00:00:00:00:00
    inet 127.0.0.1/8 scope host lo
    valid_lft forever preferred_lft forever
    inet6 ::1/128 scope host
    valid_lft forever preferred_lft forever
2: enp0s3: <BROADCAST,MULTICAST,UP,LOWER_UP> mtu 1500 ...
    link/ether 52:54:00:fa:61:98 brd ff:ff:ff:ff:ff:ff
    inet 192.168.23.98/24 brd 192.168.23.255 scope global enp0s3
    valid_lft forever preferred_lft forever
    inet6 fe80::5054:ff:fefa:6198/64 scope link
    valid_lft forever preferred_lft forever
```

Layer-2 Example Implementation:

A reference to a container-based example of the layer-2 scenario can be found in the NSO example set. See the [examples.ncs/high-availability/hcc](https://github.com/NSO-developer/nso-examples/tree/6.7/high-availability/hcc) `README`.

#### **Enabling Layer-3 BGP**

Layer-3 operation is configured for each NSO HA group node separately. The HCC configuration on both nodes should match, otherwise, the primary node's configuration will overwrite the configuration on the secondary node.

Addresses:

<table><thead><tr><th valign="top">Hostname</th><th valign="top">Address</th><th width="101" valign="top">AS</th><th valign="top">Role</th></tr></thead><tbody><tr><td valign="top"><code>paris</code></td><td valign="top">192.168.31.99</td><td valign="top">64512</td><td valign="top">Paris node</td></tr><tr><td valign="top"><code>london</code></td><td valign="top">192.168.30.98</td><td valign="top">64513</td><td valign="top">London node</td></tr><tr><td valign="top"><code>router</code></td><td valign="top"><p>192.168.30.2</p><p>192.168.31.2</p></td><td valign="top">64514</td><td valign="top">BGP-enabled router</td></tr><tr><td valign="top"><code>vip4</code></td><td valign="top">192.168.23.122</td><td valign="top"></td><td valign="top">Primary node IPv4 VIP address</td></tr></tbody></table>

Configuring BGP for Paris Node:

```bash
admin@ncs(config)# hcc bgp node paris enabled
admin@ncs(config)# hcc bgp node paris as 64512
admin@ncs(config)# hcc bgp node paris router-id 192.168.31.99
admin@ncs(config)# hcc bgp node paris neighbor 192.168.31.2 as 64514
admin@ncs(config)# commit
```

Configuring BGP for London Node:

```bash
admin@ncs(config)# hcc bgp node london enabled
admin@ncs(config)# hcc bgp node london as 64513
admin@ncs(config)# hcc bgp node london router-id 192.168.30.98
admin@ncs(config)# hcc bgp node london neighbor 192.168.30.2 as 64514
admin@ncs(config)# commit
```

Check BGP Neighbor Connectivity:

Check neighbor connectivity on the `paris` primary node. Note that its connection to neighbor 192.168.31.2 (`router`) is `ESTABLISHED`.

```bash
admin@ncs# show hcc
      BGPD  BGPD
NODE ID PID   STATUS   ADDRESS       STATE        CONNECTED
----------------------------------------------------------------
london  -     -        192.168.30.2  -            -
paris   2486  running  192.168.31.2  ESTABLISHED  true
```

Check neighbor connectivity on the `london` secondary node. Note that the primary node also has an `ESTABLISHED` connection to its neighbor 192.168.30.2 (`router`). The primary and secondary nodes both maintain their BGP neighbor connections at all times when BGP is enabled, but only the primary node announces routes for the VIPs.

```bash
admin@ncs# show hcc
      BGPD  BGPD
NODE ID PID   STATUS   ADDRESS       STATE        CONNECTED
----------------------------------------------------------------
london  494   running  192.168.30.2  ESTABLISHED  true
paris   -     -        192.168.31.2  -            -
```

Check Advertised BGP Routes Neighbors:

Check the BGP routes received by the `router`.

```bash
admin@ncs# show ip bgp
...
Network          Next Hop            Metric LocPrf Weight Path
*> 192.168.23.122/32
                  192.168.31.99                          0 64513 ?
```

The VIP subnet is routed to the `paris` host, which is the primary node.

Layer-3 BGP Example Implementation:

A reference to a container-based example of the combined layer-2 and layer-3 BGP scenario can be found in the NSO example set. See the [examples.ncs/high-availability/hcc](https://github.com/NSO-developer/nso-examples/tree/6.7/high-availability/hcc) `README`.

#### **Enabling Layer-3 DNS**

If enabled prior to the HA being established, HCC will update the DNS server with the IP address of the Primary node once a primary is selected.

If an HA is already operational, and Layer-3 DNS is enabled and configured afterward, HCC will not update the DNS server automatically. An automatic DNS server update will only happen if a HA switchover happens. HCC exposes an update action to manually trigger an update to the DNS server with the IP address of the primary node.

DNS Update Action:

The user can explicitly update DNS from the specific NSO node by running the update action.

```bash
admin@ncs# hcc dns update
```

Check the result of invoking the DNS update utility using the operational data in `/hcc/dns`:

```bash
admin@ncs# show hcc dns
hcc dns status time 2023-10-10T20:47:31.733661+00:00
hcc dns status exit-code 0
hcc dns status error-message ""
```

One way to verify DNS server updates is through the `nslookup` program. However, be mindful of the DNS caching mechanism, which may cache the old value for the amount of time controlled by the TTL setting.

```bash
cisco@node-2:~$ nslookup example.com
Server:   10.0.0.10
Address:  10.0.0.10#53

Name: example.com
Address: 10.0.0.20
Name: example.com
Address: ::10
```

DNS get-node-location Action:

/hcc/dns/member holds the information about all members involved in HA. The `get-node-location` action provides information on the location of an NSO node.

```bash
admin@ncs(config)# hcc dns get-node-location
location SanJose
```

### Data Model <a href="#ug.ha.hcc.data_models" id="ug.ha.hcc.data_models"></a>

The HCC data model can be found in the HCC package (`tailf-hcc.yang`).

## Setup with an External Load Balancer <a href="#ug.ha.lb" id="ug.ha.lb"></a>

As an alternative to the HCC package, NSO built-in HA, either rule-based or HA Raft, can also be used in conjunction with a load balancer device in a reverse proxy configuration. Instead of managing the virtual IP address directly as HCC does, this setup relies on an external load balancer to route traffic to the currently active primary node.

<div data-with-frame="true"><figure><img src="/files/lwd8IYxK6kupFpOZPKXY" alt="" width="375"><figcaption><p>Load Balancer Routes Connections to the Appropriate NSO Node</p></figcaption></figure></div>

The load balancer uses HTTP health checks to determine which node is currently the active primary. The example, found in the [examples.ncs/high-availability/load-balancer](https://github.com/NSO-developer/nso-examples/tree/6.7/high-availability/load-balancer) directory uses HTTP status codes on the health check endpoint to easily distinguish whether the node is currently primary or not.

In the example, freely available HAProxy software is used as a load balancer to demonstrate the functionality. It is configured to steer connections on localhost to either of the TCP port 2024 (SSH CLI) and TCP port 8080 (web UI and RESTCONF) to the active node in a 2-node HA cluster. The HAProxy software is required if you wish to run this example yourself.

<div data-with-frame="true"><figure><img src="/files/0MQhY5J9qMFvhr7tg5vd" alt="" width="375"><figcaption><p>Load Balancer Uses Health Checks to Determine the Currently Active Primary Node</p></figcaption></figure></div>

You can start all the components in the example by running the `make build start` command. At the beginning, the first node `n1` is the active primary. Connecting to the localhost port 2024 will establish a connection to this node:

```bash
$ make build start
Setting up run directory for nso-node1
 ... make output omitted ...
Waiting for n2 to connect: .
$ ssh -p 2024 admin@localhost
admin@localhost's password: admin

admin connected from 127.0.0.1 using ssh on localhost
admin@n1> switch cli
admin@n1# show high-availability
high-availability enabled
high-availability status mode primary
high-availability status current-id n1
high-availability status assigned-role primary
high-availability status read-only-mode false
ID  ADDRESS
---------------
n2  127.0.0.1
```

Then, you can disable the high availability subsystem on `n1` to simulate a node failure.

```bash
admin@n1# high-availability disable
result NSO Built-in HA disabled
admin@n1# exit
Connection to localhost closed.
```

Disconnect and wait a few seconds for the built-in HA to perform the failover to node `n2`. The time depends on the `high-availability/settings/reconnect-interval` and is set quite aggressively in this example to make the failover in about 6 seconds. Reconnect with the SSH client and observe the connection is now made to the fail-over node which has become the active primary:

```bash
$ ssh -p 2024 admin@localhost
admin@localhost's password: admin

admin connected from 127.0.0.1 using ssh on localhost
admin@n2> switch cli
admin@n2# show high-availability
high-availability enabled
high-availability status mode primary
high-availability status current-id n2
high-availability status assigned-role primary
high-availability status read-only-mode false
```

Finally, shut down the example with the `make stop clean` command.

## NB Listens to Addresses on HA Primary for Load Balancers <a href="#d5e5523" id="d5e5523"></a>

NSO can be configured for the HA primary to listen on additional ports for the northbound interfaces NETCONF, RESTCONF, the web server (including JSON-RPC), and the CLI over SSH. Once a different node transitions to role primary the configured listen addresses are brought up on that node instead.

When the following configuration is added to `ncs.conf`, then the primary HA node will listen(2) and bind(2) port 1830 on the wildcard IPv4 and IPv6 addresses.

```xml
<netconf-north-bound>
  <transport>
    <ssh>
      <enabled>true</enabled>
      <ip>0.0.0.0</ip>
      <port>830</port>
      <ha-primary-listen>
        <ip>0.0.0.0</ip>
        <port>1830</port>
      </ha-primary-listen>
      <ha-primary-listen>
        <ip>::</ip>
        <port>1830</port>
      </ha-primary-listen>
    </ssh>
  </transport>
</netconf-north-bound>
```

A similar configuration can be added for other NB interfaces, see the ha-primary-listen list under `/ncs-config/{restconf,webui,cli}`.

## Read-only State

Administrators can explicitly configure read-only mode on any NSO node (primary or secondary) using either `/ncs-state/set-read-only` action or `maapi_set_readonly_mode()` MAAPI function, which disables further configuration changes and may be required, for example, during cluster maintenance. This is called *administrator-configured* read-only mode and can be verified by querying the `/ncs-state/admin-configured-read-only-mode` leaf.

However, an NSO node may also become read-only due to its HA state. The overall read-only state is therefore governed by:

* **Administrator-configured read-only mode**: Explicitly set by an administrator as described above (or invoked by [#ug.ha.builtin](#ug.ha.builtin "mention") in certain situations).
* **Read-only mode imposed by HA operational state**: Automatically set based on the node's HA role, e.g. when secondary or follower.

The node will enter a read-only state, as shown in `/ncs-state/read-only-mode`, and reject write transactions when either of these conditions is true. This separation allows administrators to place a node in read-only mode for maintenance or upgrade scenarios, independent of the HA configuration.

## HA Framework Requirements

If an external HAFW is used, NSO only replicates the CDB data. NSO must be told by the HAFW which node should be primary and which nodes should be secondaries.

The HA framework must also detect when nodes fail and instruct NSO accordingly. If the primary node fails, the HAFW must elect one of the remaining secondaries and appoint it the new primary. The remaining secondaries must also be informed by the HAFW about the new primary situation.

### Mode of Operation <a href="#ha.moo" id="ha.moo"></a>

NSO must be instructed through the `ncs.conf` configuration file that it should run in HA mode. The following configuration snippet enables HA mode:

```xml
<ha>
  <enabled>true</enabled>
  <ip>0.0.0.0</ip>
  <port>4570</port>
  <extra-listen>
    <ip>::</ip>
    <port>4569</port>
  </extra-listen>
  <tick-timeout>PT20S</tick-timeout>
</ha>
```

Make sure to restart the `ncs` process for the changes to take effect.

The IP address and the port above indicate which IP and which port should be used for the communication between the HA nodes. `extra-listen` is an optional list of `ip:port` pairs that a HA primary also listens to for secondary connections. For IPv6 addresses, the syntax `[ip]:port` may be used. If the `:port` is omitted, `port` is used. The `tick-timeout` is a duration indicating how often each secondary must send a tick message to the primary indicating liveness. If the primary has not received a tick from a secondary within 3 times the configured tick time, the secondary is considered to be dead. Similarly, the primary sends tick messages to all the secondaries. If a secondary has not received any tick messages from the primary within the 3 times the timeout, the secondary will consider the primary dead and report accordingly.

A HA node can be in one of three states: `NONE`, `SECONDARY` or `PRIMARY`. Initially, a node is in the `NONE` state. This implies that the node will read its configuration from CDB, stored locally on file. Once the HA framework has decided whether the node should be a secondary or a primary the HAFW must invoke either the methods `Ha.beSecondary(primary)` or `Ha.bePrimary()`

When an NSO HA node starts, it always starts up in mode `NONE`. At this point, there are no other nodes connected. Each NSO node reads its configuration data from the locally stored CDB and applications on or off the node may connect to NSO and read the data they need. Although write operations are allowed in the `NONE` state it is highly discouraged to initiate southbound communication unless necessary. A node in `NONE` state should only be used to configure NSO itself or to do maintenance such as upgrades. When in `NONE` state, some features are disabled, including but not limited to:

* commit queue
* NSO scheduler
* nano-service side effect queue

This is to avoid situations where multiple NSO nodes are trying to perform the same southbound operation simultaneously.

At some point, the HAFW will command some nodes to become secondary nodes of a named primary node. When this happens, each secondary node tracks changes and (logically or physically) copies all the data from the primary. Previous data at the secondary node is overwritten.

Note that the HAFW, by using NSO's start phases, can make sure that NSO does not start its northbound interfaces (NETCONF, CLI, ...) until the HAFW has decided what type of node it is. Furthermore once a node has been set to the `SECONDARY` state, it is not possible to initiate new write transactions towards the node. It is thus never possible for an agent to write directly into a secondary node. Once a node is returned either to the `NONE` state or to the `PRIMARY` state, write transactions can once again be initiated towards the node.

The HAFW may command a secondary node to become primary at any time. The secondary node already has up-to-date data, so it simply stops receiving updates from the previous primary. Presumably, the HAFW also commands the primary node to become a secondary node or takes it down, or handles the situation somehow. If it has crashed, the HAFW tells the secondary to become primary, restarts the necessary services on the previous primary node, and gives it an appropriate role, such as secondary. This is outside the scope of NSO.

Each of the primary and secondary nodes has the same set of all callpoints and validation points locally on each node. The start sequence has to make sure the corresponding daemons are started before the HAFW starts directing secondary nodes to the primary, and before replication starts. The associated callbacks will however only be executed at the primary. If e.g. the validation executing at the primary needs to read data that is not stored in the configuration and only available on another node, the validation code must perform any needed RPC calls.

If the order from the HAFW is to become primary, the node will start to listen for incoming secondaries at the `ip:port` configured under `/ncs-config/ha`. The secondaries TCP connect to the primary and this socket is used by NSO to distribute the replicated data.

If the order is to be a secondary, the node will contact the primary and possibly copy the entire configuration from the primary. This copy is not performed if the primary and secondary decide that they have the same version of the CDB database loaded, in which case nothing needs to be copied. This mechanism is implemented by use of a unique token, the `transaction id` - it contains the node id of the node that generated it and a time stamp, but is effectively "opaque".

This transaction ID is generated by the cluster primary each time a configuration change is committed, and all nodes write the same transaction ID into their copy of the committed configuration. If the primary dies and one of the remaining secondaries is appointed the new primary, the other secondaries must be told to connect to the new primary. They will compare their last transaction ID to the one from the newly appointed primary. If they are the same, no CDB copy occurs. This will be the case unless a configuration change has sneaked in since both the new primary and the remaining secondaries will still have the last transaction ID generated by the old primary - the new primary will not generate a new transaction ID until a new configuration change is committed. The same mechanism works if a secondary node is simply restarted. No cluster reconfiguration will lead to a CDB copy unless the configuration has been changed in between.

Northbound agents should run on the primary, an agent can't commit write operations at a secondary node.

When an agent commits its CDB data, CDB will stream the committed data out to all registered secondaries. If a secondary dies during the commit, nothing will happen, the commit will succeed anyway. When and if the secondary reconnects to the cluster, the secondary will have to copy the entire configuration. All data on the HA sockets between NSO nodes only go in the direction from the primary to the secondaries. A secondary that isn't reading its data will eventually lead to a situation with full TCP buffers at the primary. In principle, it is the responsibility of HAFW to discover this situation and notify the primary NSO about the hanging secondary. However, if 3 times the tick timeout is exceeded, NSO will itself consider the node dead and notify the HAFW. The default value for tick timeout is 20 seconds.

The primary node holds the active copy of the entire configuration data in CDB. All configuration data has to be stored in CDB for replication to work. At a secondary node, any request to read will be serviced while write requests will be refused. Thus, the CDB subscription code works the same regardless of whether the CDB client is running at the primary or at any of the secondaries. Once a secondary has received the updates associated to a commit at the primary, all CDB subscribers at the secondary will be duly notified about any changes using the normal CDB subscription mechanism.

If the system has been set up to subscribe for NETCONF notifications, the secondaries will have all subscriptions as configured in the system, but the subscription will be idle. All NETCONF notifications are handled by the primary, and once the notifications get written into stable storage (CDB) at the primary, the list of received notifications will be replicated to all secondaries.

## Security Aspects <a href="#d5e5586" id="d5e5586"></a>

We specify in `ncs.conf` which IP address the primary should bind for incoming secondaries. If we choose the default value `0.0.0.0` it is the responsibility of the application to ensure that connection requests only arrive from acceptable trusted sources through some means of firewalling.

A cluster is also protected by a token, a secret string only known to the application. The `Ha.connect()` method must be given the token. A secondary node that connects to a primary node negotiates with the primary using a CHAP-2-like protocol, thus both the primary and the secondary are ensured that the other end has the same token without ever revealing their own token. The token is never sent in clear text over the network. This mechanism ensures that a connection from an NSO secondary to a primary can only succeed if they both have the same token.

It is indeed possible to store the token itself in CDB, thus an application can initially read the token from the local CDB data, and then use that token in . the constructor for the `Ha` class. In this case, it may very well be a good idea to have the token stored in CDB be of type tailf:aes-256-cfb-128-encrypted-string.

If the actual CDB data that is sent on the wire between cluster nodes is sensitive, and the network is untrusted, the recommendation is to use IPSec between the nodes. An alternative option is to decide exactly which configuration data is sensitive and then use the tailf:aes-256-cfb-128-encrypted-string type for that data. If the configuration data is of type tailf:aes-256-cfb-128-encrypted-string the encrypted data will be sent on the wire in update messages from the primary to the secondaries.

## API <a href="#d5e5602" id="d5e5602"></a>

There are two APIs used by the HA framework to control the replication aspects of NSO. First, there exists a synchronous API used to tell NSO what to do, secondly, the application may create a notifications socket and subscribe to HA-related events where NSO notifies the application on certain HA-related events such as the loss of the primary, etc. The HA-related notifications sent by NSO are crucial to how to program the HA framework.

The HA-related classes reside in the `com.tailf.ha` package. See Javadocs for reference. The HA notifications-related classes reside in the `com.tailf.notif` package, See Javadocs for reference.

## Ticks <a href="#d5e5608" id="d5e5608"></a>

The configuration parameter `/ncs-cfg/ha/tick-timeout` is by default set to 20 seconds. This means that every 20 seconds each secondary will send a tick message on the socket leading to the primary. Similarly, the primary will send a tick message every 20 seconds on every secondary socket.

This aliveness detection mechanism is necessary for NSO. If a socket gets closed all is well, NSO will clean up and notify the application accordingly using the notifications API. However, if a remote node freezes, the socket will not get properly closed at the other end. NSO will distribute update data from the primary to the secondaries. If a remote node is not reading the data, TCP buffer will get full and NSO will have to start to buffer the data. NSO will buffer data for at most `tickTime` times 3 time units. If a `tick` has not been received from a remote node within that time, the node will be considered dead. NSO will report accordingly over the notifications socket and either remove the hanging secondary or, if it is a secondary that loses contact with the primary, go into the initial `NONE` state.

If the HAFW can be really trusted, it is possible to set this timeout to `PT0S`, i.e zero, in which case the entire dead-node-detection mechanism in NSO is disabled.

## CDB Replication <a href="#d5e5651" id="d5e5651"></a>

When HA is enabled in `ncs.conf`, CDB automatically replicates data written on the primary to the connected secondary nodes. Replication is done on a per-transaction basis to all the secondaries in parallel and is synchronous. When NSO is in secondary mode the northbound APIs are in read-only mode, that is the configuration can not be changed on a secondary other than through replication updates from the primary. It is still possible to read from for example NETCONF or CLI (if they are enabled) on a secondary. CDB subscriptions work as usual. When NSO is in the `NONE` state CDB is unlocked and it behaves as when NSO is not in HA mode at all.

Unlike configuration data, operational data is replicated only if it is defined as persistent in the data model (using the `tailf:persistent` extension).


# AAA Infrastructure

Manage user authentication, authorization, and audit using NSO's AAA mechanism.

## The Problem

Users log into NSO through the CLI, NETCONF, RESTCONF, SNMP, or via the Web UI. In either case, users need to be authenticated. That is, a user needs to present credentials, such as a password or a public key to gain access. As an alternative, for RESTCONF, users can be authenticated via token validation.

Once a user is authenticated, all operations performed by that user need to be authorized. That is, certain users may be allowed to perform certain tasks, whereas others are not. This is called authorization. We differentiate between the authorization of commands and the authorization of data access.

## Structure - Data Models <a href="#d5e5697" id="d5e5697"></a>

The NSO daemon manages device configuration including AAA information. NSO manages AAA information as well as uses it. The AAA information describes which users may log in, what passwords they have, and what they are allowed to do. This is solved in NSO by requiring a data model to be both loaded and populated with data. NSO uses the YANG module `tailf-aaa.yang` for authentication, while `ietf-netconf-acm.yang` (NETCONF Access Control Model (NACM), [RFC 8341](https://tools.ietf.org/html/rfc8341)) as augmented by `tailf-acm.yang` is used for group assignment and authorization.

### Data Model Contents <a href="#d5e5706" id="d5e5706"></a>

The NACM data model is targeted specifically towards access control for NETCONF operations and thus lacks some functionality that is needed in NSO, in particular, support for the authorization of CLI commands and the possibility to specify the context (NETCONF, CLI, etc.) that a given authorization rule should apply to. This functionality is modeled by augmentation of the NACM model, as defined in the `tailf-acm.yang` YANG module.

The `ietf-netconf-acm.yang` and `tailf-acm.yang` modules can be found in `$NCS_DIR/src/ncs/yang` directory in the release, while `tailf-aaa.yang` can be found in the `$NCS_DIR/src/ncs/aaa` directory.

NACM options related to services are modeled by augmentation of the NACM model, as defined in the `tailf-ncs-acm.yang` YANG module. The `tailf-ncs-acm.yang` can be found in `$NCS_DIR/src/ncs/yang` directory in the release.

The complete AAA data model defines a set of users, a set of groups, and a set of rules. The data model must be populated with data that is subsequently used by by NSO itself when it authenticates users and authorizes user data access. These YANG modules work exactly like all other `fxs` files loaded into the system with the exception that NSO itself uses them. The data belongs to the application, but NSO itself is the user of the data.

Since NSO requires a data model for the AAA information for its operation, it will report an error and fail to start if these data models cannot be found.

## AAA-related Items in `ncs.conf` <a href="#d5e5722" id="d5e5722"></a>

NSO itself is configured through a configuration file - `ncs.conf`. In that file, we have the following items related to authentication and authorization:

* `/ncs-config/aaa/ssh-server-key-dir`: If SSH termination is enabled for NETCONF or the CLI, the NSO built-in SSH server needs to have server keys. These keys are generated by the NSO install script and by default end up in `$NCS_DIR/etc/ncs/ssh`.\
  \
  It is also possible to use OpenSSH to terminate NETCONF or the CLI. If OpenSSH is used to terminate SSH traffic, this setting has no effect.
* `/ncs-config/aaa/ssh-pubkey-authentication`: If SSH termination is enabled for NETCONF or the CLI, this item controls how the NSO SSH daemon locates the user keys for public key authentication. See [Public Key Login](#ug.aaa.public_key_login) for details.
* `/ncs-config/aaa/local-authentication/enabled`: The term 'local user' refers to a user stored under `/aaa/authentication/users`. The alternative is a user unknown to NSO, typically authenticated by PAM. By default, NSO first checks local users before trying PAM or external authentication.\
  \
  Local authentication is practical in test environments. It is also useful when we want to have one set of users that are allowed to log in to the host with normal shell access and another set of users that are only allowed to access the system using the normal encrypted, fully authenticated, northbound interfaces of NSO.\
  \
  If we always authenticate users through PAM, it may make sense to set this configurable to `false`. If we disable local authentication, it implicitly means that we must use either PAM authentication or external authentication. It also means that we can leave the entire data trees under `/aaa/authentication/users` and, in the case of external authentication, also `/nacm/groups` (for NACM) or `/aaa/authentication/groups` (for legacy tailf-aaa) empty.
* `/ncs-config/aaa/pam`: NSO can authenticate users using PAM (Pluggable Authentication Modules). PAM is an integral part of most Unix-like systems.\
  \
  PAM is a complicated - albeit powerful - subsystem. It may be easier to have all users stored locally on the host, However, if we want to store users in a central location, PAM can be used to access the remote information. PAM can be configured to perform most login scenarios including RADIUS and LDAP. One major drawback with PAM authentication is that there is no easy way to extract the group information from PAM. PAM authenticates users, it does not also assign a user to a set of groups. PAM authentication is thoroughly described later in this chapter.
* `/ncs-config/aaa/default-group`: If this configuration parameter is defined and if the group of a user cannot be determined, a logged-in user ends up in the given default group.
* `/ncs-config/aaa/external-authentication`: NSO can authenticate users using an external executable. This is further described later in [External Authentication](#ug.aaa.external_authentication). As an alternative, you may consider using package authentication.
* `/ncs-config/aaa/external-validation`: NSO can authenticate users by validation of tokens using an external executable. This is further described later in [External Token Validation](#ug.aaa.external_validation). Where external authentication uses a username and password to authenticate a user, external validation uses a token. The validation script should use the token to authenticate a user and can, optionally, also return a new token to be returned with the result of the request. It is currently only supported for RESTCONF.
* `/ncs-config/aaa/external-challenge`: NSO has support for multi-factor authentication by sending challenges to a user. Challenges may be sent from any of the external authentication mechanisms but are currently only supported by JSON-RPC and CLI over SSH. This is further described later in [External Multi-factor Authentication](#ug.aaa.external_challenge).
* `/ncs-config/aaa/package-authentication`: NSO can authenticate users using package authentication. It extends the concept of external authentication by allowing multiple packages to be used for authentication instead of a single executable. This is further described in [Package Authentication](#ug.aaa.packageauth).
* `/ncs-config/aaa/single-sign-on`: With this setting enabled, NSO invokes Package Authentication on all requests to HTTP endpoints with the `/sso` prefix. This way, Package Authentication packages that require custom endpoints can expose them under the `/sso` base route.\
  \
  For example, a SAMLv2 Single Sign-On (SSO) package needs to process requests to an AssertionConsumerService endpoint, such as `/sso/saml/acs`, and therefore requires enabling this setting.\
  \
  This is a valid authentication method for WEB UI and JSON-RPC interfaces and needs Package Authentication to be enabled as well.
* `/ncs-config/aaa/single-sign-on/enable-automatic-redirect`: If only one Single Sign-On package is configured (a package with `single-sign-on-url` set in `package-meta-data.xml`) and also this setting is enabled, NSO automatically redirects all unauthenticated access attempts to the configured `single-sign-on-url`.

## Authentication <a href="#ug.aaa.authentication" id="ug.aaa.authentication"></a>

Depending on the northbound management protocol, when a user session is created in NSO, it may or may not be authenticated. If the session is not yet authenticated, NSO's AAA subsystem is used to perform authentication and authorization, as described below. If the session already has been authenticated, NSO's AAA assigns groups to the user as described in [Group Membership](#ug.aaa.groups), and performs authorization, as described in [Authorization](#ug.aaa.authorization).

The authentication part of the data model can be found in `tailf-aaa.yang`:

```yang
    container authentication {
      tailf:info "User management";
      container users {
        tailf:info "List of local users";
        list user {
          key name;
          leaf name {
            type string;
            tailf:info "Login name of the user";
          }
          leaf uid {
            type int32;
            mandatory true;
            tailf:info "User Identifier";
          }
          leaf gid {
            type int32;
            mandatory true;
            tailf:info "Group Identifier";
          }
          leaf password {
            type passwdStr;
            mandatory true;
          }
          leaf ssh_keydir {
            type string;
            mandatory true;
            tailf:info "Absolute path to directory where user's ssh keys
                        may be found";
          }
          leaf homedir {
            type string;
            mandatory true;
            tailf:info "Absolute path to user's home directory";
          }
        }
      }
    }
```

AAA authentication is used in the following cases:

* When the built-in SSH server is used for NETCONF and CLI sessions.
* For Web UI sessions and REST access.
* When the method `Maapi.Authenticate()` is used.

NSO's AAA authentication is not used in the following cases:

* When NETCONF uses an external SSH daemon, such as OpenSSH.

  \
  In this case, the NETCONF session is initiated using the program `netconf-subsys`, as described in [NETCONF Transport Protocols](https://nso-docs.cisco.com/guides/administration/management/pages/TItWBhukD9D6FJkD3eWB#ug.netconf_agent.transport) in Northbound APIs.
* When NETCONF uses TCP, as described in [NETCONF Transport Protocols](https://nso-docs.cisco.com/guides/administration/management/pages/TItWBhukD9D6FJkD3eWB#ug.netconf_agent.transport) in Northbound APIs, e.g. through the command `netconf-console`.
* When accessing the CLI by invoking the `ncs_cli`, e.g. through an external SSH daemon, such as OpenSSH, or a telnet daemon.\
  \
  An important special case here is when a user has shell access to the host and runs **ncs\_cli** from the shell. This command, as well as direct access to the IPC socket, allows for authentication bypass. It is crucial to consider this case for your deployment. If non-trusted users have shell access to the host, IPC access must be restricted. See [Authenticating IPC Access](#authenticating-ipc-access).
* When SNMP is used, SNMP has its own authentication mechanisms. See [NSO SNMP Agent](/guides/development/core-concepts/northbound-apis#the-nso-snmp-agent) in Northbound APIs.
* When the method `Maapi.startUserSession()` is used without a preceding call of `Maapi.authenticate()`.

### Public Key Login <a href="#ug.aaa.public_key_login" id="ug.aaa.public_key_login"></a>

When a user logs in over NETCONF or the CLI using the built-in SSH server, with a public key login, the procedure is as follows.

The user presents a username in accordance with the SSH protocol. The SSH server consults the settings for `/ncs-config/aaa/ssh-pubkey-authentication` and `/ncs-config/aaa/local-authentication/enabled` .

1. If `ssh-pubkey-authentication` is set to `local`, and the SSH keys in `/aaa/authentication/users/user{$USER}/ssh_keydir` match the keys presented by the user, authentication succeeds.
2. Otherwise, if `ssh-pubkey-authentication` is set to `system`, `local-authentication` is enabled, and the SSH keys in `/aaa/authentication/users/user{$USER}/ssh_keydir` match the keys presented by the user, authentication succeeds.
3. Otherwise, if `ssh-pubkey-authentication` is set to `system` and the user `/aaa/authentication/users/user{$USER}` does not exist, but the user does exist in the OS password database, the keys in the user's `$HOME/.ssh` directory are checked. If these keys match the keys presented by the user, authentication succeeds.
4. Otherwise, authentication fails.

In all cases the keys are expected to be stored in a file called `authorized_keys` (or `authorized_keys2` if `authorized_keys` does not exist), and in the native OpenSSH format (i.e. as generated by the OpenSSH `ssh-keygen` command). If authentication succeeds, the user's group membership is established as described in [Group Membership](#ug.aaa.groups).

This is exactly the same procedure that is used by the OpenSSH server with the exception that the built-in SSH server also may locate the directory containing the public keys for a specific user by consulting the `/aaa/authentication/users` tree.

### **Setting up Public Key Login**

We need to provide a directory where SSH keys are kept for a specific user and give the absolute path to this directory for the `/aaa/authentication/users/user/ssh_keydir` leaf. If a public key login is not desired at all for a user, the value of the `ssh_keydir` leaf should be set to `""`, i.e. the empty string. Similarly, if the directory does not contain any SSH keys, public key logins for that user will be disabled.

The built-in SSH daemon supports DSA, RSA, and ED25519 keys. To generate and enable RSA keys of size 4096 bits for, say, user "bob", the following steps are required.

On the client machine, as user "bob", generate a private/public key pair as:

```bash
# ssh-keygen -b 4096 -t rsa
Generating public/private rsa key pair.
Enter file in which to save the key (/home/bob/.ssh/id_rsa):
Created directory '/home/bob/.ssh'.
Enter passphrase (empty for no passphrase):
Enter same passphrase again:
Your identification has been saved in /home/bob/.ssh/id_rsa.
Your public key has been saved in /home/bob/.ssh/id_rsa.pub.
The key fingerprint is:
ce:1b:63:0a:f9:d4:1d:04:7a:1d:98:0c:99:66:57:65 bob@buzz
# ls -lt ~/.ssh
total 8
-rw-------  1 bob users 3247 Apr  4 12:28 id_rsa
-rw-r--r--  1 bob users  738 Apr  4 12:28 id_rsa.pub
```

Now we need to copy the public key to the target machine where the NETCONF or CLI SSH client runs.

Assume we have the following user entry:

```xml
<user>
  <name>bob</name>
  <uid>100</uid>
  <gid>10</gid>
  <password>$1$feedbabe$nGlMYlZpQ0bzenyFOQI3L1</password>
  <ssh_keydir>/var/system/users/bob/.ssh</ssh_keydir>
  <homedir>/var/system/users/bob</homedir>
</user>
```

We need to copy the newly generated file `id_rsa.pub`, which is the public key, to a file on the target machine called `/var/system/users/bob/.ssh/authorized_keys`.

{% hint style="info" %}
Since the release of [OpenSSH 7.0](https://www.openssh.com/txt/release-7.0), support of `ssh-dss` host and user keys is disabled by default. If you want to continue using these, you may re-enable it using the following options for OpenSSH client:

```
HostKeyAlgorithms=+ssh-dss
PubkeyAcceptedKeyTypes=+ssh-dss
```

You can find full instructions at [OpenSSH Legacy Options](https://www.openssh.com/legacy.html) webpage.
{% endhint %}

### Password Login <a href="#d5e5901" id="d5e5901"></a>

Password login is triggered in the following cases:

* When a user logs in over NETCONF or the CLI using the built-in SSH server, with a password. The user presents a username and a password in accordance with the SSH protocol.
* When a user logs in using the Web UI. The Web UI asks for a username and password.
* When the method `Maapi.authenticate()` is used.

In this case, NSO will by default try local authentication, PAM, external authentication, and package authentication in that order, as described below. It is possible to change the order in which these are tried, by modifying the `ncs.conf`. parameter `/ncs-config/aaa/auth-order`. See [ncs.conf(5)](/guides/resources/man/ncs.conf.5) in Manual Pages for details.

1. If `/aaa/authentication/users/user{$USER}` exists and the presented password matches the encrypted password in `/aaa/authentication/users/user{$USER}/password`, the user is authenticated.
2. If the password does not match or if the user does not exist in `/aaa/authentication/users`, PAM login is attempted, if enabled. See [PAM](#ug.aaa.pam) for details.
3. If all of the above fails and external authentication is enabled, the configured executable is invoked. See [External Authentication](#ug.aaa.external_authentication) for details.

If authentication succeeds, the user's group membership is established as described in [Group Membership](#ug.aaa.groups).

### PAM <a href="#ug.aaa.pam" id="ug.aaa.pam"></a>

On operating systems supporting PAM, NSO also supports PAM authentication. Using PAM, authentication with NSO can be very convenient since it allows us to have the same set of users and groups having access to NSO as those that have access to the UNIX/Linux host itself.

{% hint style="info" %}
PAM is the recommended way to authenticate NSO users.
{% endhint %}

If we use PAM, we do not have to have any users or any groups configured in the NSO aaa namespace at all.

To configure PAM we typically need to do the following:

1. Remove all users and groups from the AAA initialization XML file.
2. Enable PAM in `ncs.conf` by adding the following to the AAA section in `ncs.conf`. The `service` name specifies the PAM service, typically a file in the directory `/etc/pam.d`, but may alternatively, be an entry in a file `/etc/pam.conf` depending on OS and version. Thus, it is possible to have a different login procedure for NSO than for the host itself.

   ```xml
   <pam>
     <enabled>true</enabled>
     <service>common-auth</service>
   </pam>
   ```
3. If PAM is enabled and we want to use PAM for login, the system may have to run as `root`. This depends on how PAM is configured locally. However, the default system authentication will typically require `root`, since the PAM libraries then read `/etc/shadow`. If we don't want to run NSO as root, the solution here is to change the owner of a helper program called `$NCS_DIR/lib/ncs/lib/pam-*/priv/epam` and also set the `setuid` bit.

   ```bash
   # cd $NCS_DIR/lib/ncs/lib/pam-*/priv/
   # chown root:root epam
   # chmod u+s epam
   ```

As an example, say that we have a user test in `/etc/passwd`, and furthermore:

```bash
# grep test /etc/group
operator:x:37:test
admin:x:1001:test
```

Thus, the `test` user is part of the `admin` and the `operator` groups and logging in to NSO as the `test` user through CLI SSH, Web UI, or NETCONF, renders the following in the audit log.

```
<INFO> 28-Jan-2009::16:05:55.663 buzz ncs[14658]: audit user: test/0 logged
    in over ssh from 127.0.0.1 with authmeth:password
<INFO> 28-Jan-2009::16:05:55.670 buzz ncs[14658]: audit user: test/5 assigned
    to groups: operator,admin
<INFO> 28-Jan-2009::16:05:57.655 buzz ncs[14658]: audit user: test/5 CLI 'exit'
```

Thus, the `test` user was found and authenticated from `/etc/passwd`, and the crucial group assignment of the test user was done from `/etc/group`.

If we wish to be able to also manipulate the users, their passwords, etc on the device, we can write a private YANG model for that data, store that data in CDB, set up a normal CDB subscriber for that data, and finally when our private user data is manipulated, our CDB subscriber picks up the changes and changes the contents of the relevant `/etc` files.

### External Authentication <a href="#ug.aaa.external_authentication" id="ug.aaa.external_authentication"></a>

A common situation is when we wish to have all authentication data stored remotely, not locally, for example on a remote RADIUS or LDAP server. This remote authentication server typically not only stores the users and their passwords but also the group information.

If we wish to have not only the users but also the group information stored on a remote server, the best option for NSO authentication is to use external authentication.

If this feature is configured, NSO will invoke the executable configured in `/ncs-config/aaa/external-authentication/executable` in `ncs.conf` , and pass the username and the clear text password on `stdin` using the string notation: `"[user;password;]\n"`.

For example, if the user `bob` attempts to log in over SSH using the password 'secret', and external authentication is enabled, NSO will invoke the configured executable and write `"[bob;secret;]\n"` on the `stdin` stream for the executable. The task of the executable is then to authenticate the user and also establish the username-to-groups mapping.

For example, the executable could be a RADIUS client which utilizes some proprietary vendor attributes to retrieve the groups of the user from the RADIUS server. If authentication is successful, the program should write `accept` followed by a space-separated list of groups that the user is a member of, and additional information as described below. Again, assuming that bob's password indeed was 'secret', and that bob is a member of the `admin` and the `lamers` groups, the program should write `accept admin lamers $uid $gid $supplementary_gids $HOME` on its standard output and then exit.

{% hint style="info" %}
There is a general limit of 16000 bytes of output from the `externalauth` program.
{% endhint %}

Thus, the format of the output from an `externalauth` program when authentication is successful should be:

**`"accept $groups $uid $gid $supplementary_gids $HOME\n"`**

Where:

* `$groups` is a space-separated list of the group names the user is a member of.
* `$uid` is the UNIX integer user ID that NSO should use as a default when executing commands for this user.
* `$gid` is the UNIX integer group ID that NSO should use as a default when executing commands for this user.
* `$supplementary_gids` is a (possibly empty) space-separated list of additional UNIX group IDs the user is also a member of.
* `$HOME` is the directory that should be used as HOME for this user when NSO executes commands on behalf of this user.

It is further possible for the program to return a token on successful authentication, by using `"accept_token"` instead of `"accept"`:

**`"accept_token $groups $uid $gid $supplementary_gids $HOME $token\n"`**

Where:

* `$token` is an arbitrary string. NSO will then, for some northbound interfaces, include this token in responses.

It is also possible for the program to return additional information on successful authentication, by using `"accept_info"` instead of `"accept"`:

**`"accept_info $groups $uid $gid $supplementary_gids $HOME $info\n"`**

Where:

* `$info` is some arbitrary text. NSO will then just append this text to the generated audit log message (CONFD\_EXT\_LOGIN).

Yet another possibility is for the program to return a warning that the user's password is about to expire, by using `"accept_warning"` instead of `"accept"`:

**`"accept_warning $groups $uid $gid $supplementary_gids $HOME $warning\n"`**

Where:

* `$warning` is an appropriate warning message. The message will be processed by NSO according to the setting of `/ncs-config/aaa/expiration-warning` in `ncs.conf`.

There is also support for token variations of `"accept_info"` and `"accept_warning"` namely `"accept_token_info"` and `"accept_token_warning"`. Both `"accept_token_info"` and `"accept_token_warning"` expect the external program to output exactly the same as described above with the addition of a token after `$HOME`:

* `"accept_token_info $groups $uid $gid $supplementary_gids $HOME $token $info\n"`
* `"accept_token_warning $groups $uid $gid $supplementary_gids $HOME $token $warning\n"`

If authentication failed, the program should write `"reject"` or `"abort"`, possibly followed by a reason for the rejection, and a trailing newline. For example, `"reject Bad password\n"` or just `"abort\n"`. The difference between `"reject"` and `"abort"` is that with `"reject"`, NSO will try subsequent mechanisms configured for `/ncs-config/aaa/auth-order` in `ncs.conf` (if any), while with `"abort"`, the authentication fails immediately. Thus `"abort"` can prevent subsequent mechanisms from being tried, but when external authentication is the last mechanism (as in the default order), it has the same effect as `"reject"`.

Supported by some northbound APIs, such as JSON-RPC and CLI over SSH, the external authentication may also choose to issue a challenge:

`"challenge $challenge-id $challenge-prompt\n"`

{% hint style="info" %}
The challenge-prompt may be multi-line, why it must be base64 encoded.
{% endhint %}

For more information on multi-factor authentication, see [External Multi-Factor Authentication](#ug.aaa.external_challenge).

When external authentication is used, the group list returned by the external program is prepended by any possible group information stored locally under the `/aaa` tree. Hence when we use external authentication it is indeed possible to have the entire `/aaa/authentication` tree empty. The group assignment performed by the external program will still be valid and the relevant groups will be used by NSO when the authorization rules are checked.

### External Token Validation <a href="#ug.aaa.external_validation" id="ug.aaa.external_validation"></a>

When username and password authentication is not feasible, authentication by token validation is possible. Currently, only RESTCONF supports this mode of authentication. It shares all properties of external authentication, but instead of a username and password, it takes a token as input. The output is also almost the same, the only difference is that it is also expected to output a username.

If this feature is configured, NSO will invoke the executable configured in `/ncs-config/aaa/external-validation/executable` in `ncs.conf` , and pass the token on `stdin` using the string notation: `"[token;]\n"`.

For example if the user `bob` attempts to log over RESTCONF using the token `topsecret`, and external validation is enabled, NSO will invoke the configured executable and write `"[topsecret;]\n"` on the `stdin` stream for the executable.

The task of the executable is then to validate the token, thereby authenticating the user and also establishing the username and username-to-groups mapping.

For example, the executable could be a FUSION client that utilizes some proprietary vendor attributes to retrieve the username and groups of the user from the FUSION server. If token validation is successful, the program should write `accept` followed by a space-separated list of groups that the user is a member of, and additional information as described below. Again, assuming that `bob`'s token indeed was `topsecret`, and that `bob` is a member of the `admin` and the `lamers` groups, the program should write `accept admin lamers $uid $gid $supplementary_gids $HOME $USER` on its standard output and then exit.

{% hint style="info" %}
There is a general limit of 16000 bytes of output from the `externalvalidation` program.
{% endhint %}

Thus the format of the output from an `externalvalidation` program when token validation authentication is successful should be:

`"accept $groups $uid $gid $supplementary_gids $HOME $USER\n"`

Where:

* `$groups` is a space-separated list of the group names the user is a member of.
* `$uid` is the UNIX integer user ID NSO should use as a default when executing commands for this user.
* `$gid` is the UNIX integer group ID NSO should use as a default when executing commands for this user.
* `$supplementary_gids` is a (possibly empty) space-separated list of additional UNIX group IDs the user is also a member of.
* `$HOME` is the directory that should be used as HOME for this user when NSO executes commands on behalf of this user.
* `$USER` is the user derived from mapping the token.

It is further possible for the program to return a new token on successful token validation authentication, by using `"accept_token"` instead of `"accept"`:

`"accept_token $groups $uid $gid $supplementary_gids $HOME $USER $token\n"`

Where:

* `$token` is an arbitrary string. NSO will then, for some northbound interfaces, include this token in responses.

It is also possible for the program to return additional information on successful token validation authentication, by using `"accept_info"` instead of `"accept"`:

`"accept_info $groups $uid $gid $supplementary_gids $HOME $USER $info\n"`

Where:

* `$info` is some arbitrary text. NSO will then just append this text to the generated audit log message (CONFD\_EXT\_LOGIN).

Yet another possibility is for the program to return a warning that the user's password is about to expire, by using `"accept_warning"` instead of `"accept"`:

`"accept_warning $groups $uid $gid $supplementary_gids $HOME $USER $warning\n"`

Where:

* `$warning` is an appropriate warning message. The message will be processed by NSO according to the setting of `/ncs-config/aaa/expiration-warning` in `ncs.conf`.

There is also support for token variations of `"accept_info"` and `"accept_warning"` namely `"accept_token_info"` and `"accept_token_warning"`. Both `"accept_token_info"` and `"accept_token_warning"` expect the external program to output exactly the same as described above with the addition of a token after `$USER`:

`"accept_token_info $groups $uid $gid $supplementary_gids $HOME $USER $token $info\n"`

`"accept_token_warning $groups $uid $gid $supplementary_gids $HOME $USER $token $warning\n"`

If token validation authentication fails, the program should write `"reject"` or `"abort"`, possibly followed by a reason for the rejection and a trailing newline. For example `"reject Bad password\n"` or just `"abort\n"`. The difference between `"reject"` and `"abort"` is that with `"reject"`, NSO will try subsequent mechanisms configured for `/ncs-config/aaa/validation-order` in `ncs.conf` (if any), while with `"abort"`, the token validation authentication fails immediately. Thus `"abort"` can prevent subsequent mechanisms from being tried. Currently, the only available token validation authentication mechanism is the external one.

Supported by some northbound APIs, such as JSON-RPC and CLI over SSH, the external validation may also choose to issue a challenge:

`"challenge $challenge-id $challenge-prompt\n"`

{% hint style="info" %}
The challenge prompt may be multi-line, why it must be base64 encoded.
{% endhint %}

For more information on multi-factor authentication, see [External Multi-Factor Authentication](#ug.aaa.external_challenge).

### Multi-Factor Authentication <a href="#ug.aaa.external_challenge" id="ug.aaa.external_challenge"></a>

When authentication requires MFA, NSO issues a challenge as part of that authentication method. A challenge consists of a challenge ID and a base64 encoded challenge prompt, and a user is supposed to send a response to the challenge. Currently, only JSONRPC and CLI over SSH support multi-factor authentication. Responses to challenges of multi-factor authentication have the same output as the token authentication mechanism.

MFA is supported by External and Package Authentication (preferred). To configure Package Multi-Factor Authentication, refer to the section [Package Challenges](#package-challenges).

When an authentication method responds with a challenge and user provides a response, NSO invokes the challenge script associated with that authentication method.

{% hint style="info" %}
**Deprecated Configuration**

The configuration option `/ncs-config/aaa/challenge-order` is deprecated and, if present, will be ignored at runtime. Challenge handling is instead tied directly to the authentication method being attempted. NSO does not select or order challenge mechanisms independently of authentication method.
{% endhint %}

The challenge script must write one of the following responses to standard output:

* `accept`: The challenge succeeds and authentication completes successfully.
* `reject`: The challenge fails for the current authentication method. NSO proceeds with the next authentication method configured in `/ncs-config/aaa/auth-order`, if any. In case of package authentication with multiple packages configured, a `reject` will lead to invoking challenge of the next package.
* `abort`: The challenge fails and authentication terminates immediately. No further authentication methods are attempted.

Note that `reject` and `abort` affect authentication flow only. They do not influence which challenge mechanism is invoked.

**Example Authentication Flow**

Suppose both authentication methods are configured:

* Package authentication invokes the package-provided challenge handler.
* External authentication invokes the external challenge executable.

{% code title="Example: Authentication and Challenge Flow" overflow="wrap" expandable="true" %}

```
Assume the following authentication order:

<auth-order>package-authentication external-authentication</auth-order>

1. Package authentication is attempted.
   - The package authentication script returns "challenge".
   - NSO invokes the package challenge handler.
   - The package challenge handler returns "reject".

   Result: NSO proceeds to the next authentication method.

2. External authentication is attempted.
   - The external authentication script returns "challenge".
   - NSO invokes the external challenge executable configured under 
     /ncs-config/aaa/external-challenge/executable.
   - The external challenge returns "accept".

   Result: Authentication completes successfully.

This example demonstrates that the challenge mechanism invoked is determined by the authentication method currently being attempted.

In particular, when external authentication is used, the package challenge handler is not invoked. Conversely, when package authentication is used, the package challenge handler is invoked as expected.

The deprecated /ncs-config/aaa/challenge-order configuration has no effect on this behavior.
```

{% endcode %}

{% hint style="info" %}
While not common practice, it is possible to set up multiple authentication packages. If multiple authentication packages are configured, and the challenge of the first package fails, the challenge handler of the second package will be tried.
{% endhint %}

Supported by some northbound APIs, such as JSON-RPC and CLI over SSH, the challenge script may also choose to issue a new challenge:

`"challenge $challenge-id $challenge-prompt\n"`

{% hint style="info" %}
The challenge prompt may be multi-line, so it must be base64 encoded.
{% endhint %}

{% hint style="info" %}
Note that when using challenges with the CLI over SSH, the `/ncs-config/cli/ssh/use-keyboard-interactive>` need to be set to true for the challenges to be sent correctly to the client.
{% endhint %}

{% hint style="info" %}
The configuration of the SSH client used may need to be given the option to allow a higher number of allowed number of password prompts, e.g. `-o NumberOfPasswordPrompts`, else the default number may introduce an unexpected behavior when the client is presented with multiple challenges.
{% endhint %}

#### External Multi-Factor Authentication

If external authentication is configured, NSO will invoke the executable configured in `/ncs-config/aaa/external-challenge/executable` in `ncs.conf` , and pass the challenge ID and response on `stdin` using the string notation: `"[challenge-id;response;]\n"`.

For example, a user `bob` has received a challenge during external authentication and attempts to log in over JSON-RPC with a response to the challenge using challenge ID `"22efa"` with response `"ae457b"`. The external challenge mechanism is enabled and NSO invokes the configured executable, writing `"[22efa;ae457b;]\n"` on the `stdin` stream for the executable.

The task of the executable is then to validate the challenge ID, and response combination, thereby authenticating the user and also establishing the username and username-to-groups mapping.

For example, the executable could be a RADIUS client which utilizes some proprietary vendor attributes to retrieve the username and groups of the user from the RADIUS server. If challenge ID+response validation is successful, the program should write `"accept "` followed by a space-separated list of groups the user is a member of, and additional information as described below. Again, assuming that `bob`'s challenge ID+response combination indeed was `"22efa", "ae457b"`, and that `bob` is a member of the `admin` and the `lamers` groups, the program should write `"accept admin lamers $uid $gid $supplementary_gids $HOME $USER\n"` on its standard output and then exit.

{% hint style="info" %}
There is a general limit of 16000 bytes of output from the `externalchallenge` program.
{% endhint %}

Thus the format of the output from an `externalchallenge` program when challenge-based authentication is successful should be:

`"accept $groups $uid $gid $supplementary_gids $HOME $USER\n"`

Where:

* `$groups` is a space-separated list of the group names the user is a member of.
* `$uid` is the UNIX integer user ID NSO should use as a default when executing commands for this user.
* `$gid` is the UNIX integer group ID NSO should use as a default when executing commands for this user.
* `$supplementary_gids` is a (possibly empty) space-separated list of additional UNIX group IDs the user is also a member of.
* `$HOME` is the directory that should be used as HOME for this user when NSO executes commands on behalf of this user.
* `$USER` is the user derived from mapping the challenge ID, response.

It is further possible for the program to return a token on successful authentication, by using `"accept_token"` instead of `"accept"`:

`"accept_token $groups $uid $gid $supplementary_gids $HOME $USER $token\n"`

Where:

* `$token` is an arbitrary string. NSO will then, for some northbound interfaces, include this token in responses.

It is also possible for the program to return additional information on successful authentication, by using `"accept_info"` instead of `"accept"`:

`"accept_info $groups $uid $gid $supplementary_gids $HOME $USER $info\n"`

Where:

* `$info` is some arbitrary text. NSO will then just append this text to the generated audit log message (CONFD\_EXT\_LOGIN).

Yet another possibility is for the program to return a warning that the user's password is about to expire, by using `"accept_warning"` instead of `"accept"`:

`"accept_warning $groups $uid $gid $supplementary_gids $HOME $USER $warning\n"`

Where:

* `$warning` is an appropriate warning message. The message will be processed by NSO according to the setting of `/ncs-config/aaa/expiration-warning` in `ncs.conf`.

There is also support for token variations of `"accept_info"` and `"accept_warning"` namely `"accept_token_info"` and `"accept_token_warning"`. Both `"accept_token_info"` and `"accept_token_warning"` expect the external program to output exactly the same as described above with the addition of a token after `$USER`:

`"accept_token_info $groups $uid $gid $supplementary_gids $HOME $USER $token $info\n"`

`"accept_token_warning $groups $uid $gid $supplementary_gids $HOME $USER $token $warning\n"`

If authentication fails, the program should write `"reject"` or `"abort"`, possibly followed by a reason for the rejection and a trailing newline. For example `"reject Bad challenge response\n"` or just `"abort\n"`.

The difference between `"reject"` and `"abort"` is that with `"reject"`, NSO proceeds with the next authentication method configured in `/ncs-config/aaa/auth-order` (if any), while with `"abort"`, the challenge-response authentication fails immediately. Thus, `"abort"` can prevent subsequent mechanisms from being tried.

### Package Authentication <a href="#ug.aaa.packageauth" id="ug.aaa.packageauth"></a>

The Package Authentication functionality allows for packages to handle the NSO authentication in a customized fashion. Authentication data can e.g. be stored remotely, and a script in the package is used to communicate with the remote system.

Compared to external authentication, the Package Authentication mechanism allows specifying multiple packages to be invoked in the order they appear in the configuration. NSO provides implementations for LDAP, SAMLv2, and TACACS+ protocols with packages available in `$NCS_DIR/packages/auth/`. Additionally, you can implement your own authentication packages as detailed below.

Authentication packages are NSO packages with the required content of an executable file `scripts/authenticate`. This executable basically follows the same API, and limitations, as the external auth script, but with a different input format and some additional functionality. Other than these requirements, it is possible to customize the package arbitrarily.

{% hint style="info" %}
Package authentication is supported for Single Sign-On (see [Single Sign-On](/guides/development/advanced-development/web-ui-development#single-sign-on-sso) in Web UI), JSON-RPC, and RESTCONF. Note that Single Sign-On and (non-batch) JSON-RPC allow all functionality while the RESTCONF interface will treat anything other than a "`accept_username`" reply from the package as if authentication failed!
{% endhint %}

Package authentication is enabled by setting the `ncs.conf` options `/ncs-config/aaa/package-authentication/enabled` to true, and adding the package by name in the `/ncs-config/aaa/package-authentication/packages` list. The order of the configured packages is the order that the packages will be used when attempting to authenticate a user. See [ncs.conf(5)](/guides/resources/man/ncs.conf.5) in Manual Pages for details.

{% hint style="info" %}
When package authentication is used for RESTCONF, each authentication attempt requires NSO to invoke an external program through an Erlang port. This adds non-trivial processing overhead and can become a bottleneck in high-frequency request scenarios.

Northbound applications that send bursts of RESTCONF requests may therefore experience increased latency or reduced throughput when package authentication is enabled.
{% endhint %}

If this feature is configured in `ncs.conf`, NSO will for each configured package invoke `script/authenticate`, and pass username, password, and original HTTP request (i.e. the user-supplied `next` query parameter), HTTP request, HTTP headers, HTTP body, client source IP, client source port, northbound API context, and protocol on `stdin` using the string notation: `"[user;password;orig_request;request;headers;body;src-ip;src-port;ctx;proto;]\n"`.

{% hint style="info" %}
The fields user, password, orig\_request, request, headers, and body are all base64 encoded.
{% endhint %}

{% hint style="info" %}
If the body length exceeds the `partial_post_size` of the RESTCONF server, the body passed to the authenticate script will only contain the string `'==nso_package_authentication_partial_body==`'.
{% endhint %}

{% hint style="info" %}
The original request will be prefixed with the string `==nso_package_authentication_next==` before the base64 encoded part. This means supplying the `next` query parameter value `/my-location` will pass the following string to the authentication script: `==nso_package_authentication_next==L215LWxvY2F0aW9u`.
{% endhint %}

For example, if an unauthenticated user attempts to start a single sign-on process over northbound HTTP-based APIs with the cisco-nso-saml2-auth package, package authentication is enabled and configured with packages, and also single sign-on is enabled, NSO will, for each configured package, invoke the executable `scripts/authenticate` and write `"[;;;R0VUIC9zc28vc2FtbC9sb2dpbi8gSFRUUC8xLjE=;;;127.0.0.1;59226;webui;https;]\n"`. on the `stdin` stream for the executable.

For clarity, the base64 decoded contents sent to `stdin` presented: `"[;;;GET /sso/saml/login/ HTTP/1.1;;;127.0.0.1;54321;webui;https;]\n"`.

The task of the package is then to authenticate the user and also establish the username-to-groups mapping.

For example, the package could support a SAMLv2 authentication protocol which communicates with an Identity Provider (IdP) for authentication. If authentication is successful, the program should write either `"accept"`, or `"accept_username"`, depending on whether the authentication is started with a username or if an external entity handles the entire authentication and supplies the username for a successful authentication. (SAMLv2 uses `accept_username`, since the IdP handles the entire authentication.) The "accept\_username " is followed by a username and then followed by a space-separated list of groups the user is a member of, and additional information as described below. If authentication is successful and the authenticated user `bob` is a member of the groups `admin` and `wheel`, the program should write `"accept_username bob admin wheel 1000 1000 100 /home/bob\n"` on its standard output and then exit.

{% hint style="info" %}
There is a general limit of 16000 bytes of output from the "packageauth" program.
{% endhint %}

Thus the format of the output from a `packageauth` program when authentication is successful should be either the same as from `externalauth` (see [External Authentication](#ug.aaa.external_authentication)) or the following:

`"accept_username $USER $groups $uid $gid $supplementary_gids $HOME\n"`

Where:

* `$USER` is the user derived during the execution of the "packageauth" program. Base64 encoded.
* `$groups` is a space-separated list of the group names the user is a member of.
* `$uid` is the UNIX integer user ID NSO should use as a default when executing commands for this user.
* `$gid` is the UNIX integer group ID NSO should use as a default when executing commands for this user.
* `$supplementary_gids` is a (possibly empty) space-separated list of additional UNIX group IDs the user is also a member of.
* `$HOME` is the directory that should be used as HOME for this user when NSO executes commands on behalf of this user.

In addition to the `externalauth` API, the authentication packages can also return the following responses:

* `unknown '`*`reason`*`'` - (*`reason`* being plain-text) if they can't handle authentication for the supplied input.
* `redirect '`*`url`*`'` - (*`url`* being base64 encoded) for an HTTP redirect.
* `content '`*`content-type`*`' '`*`content`*`'` - (*`content-type`* being plain-text mime-type and *`content`* being base64 encoded) to relay supplied content.
* `accept_username_redirect url $USER $groups $uid $gid $supplementary_gids $HOME` - which combines the `accept_username` and `redirect`.

It is also possible for the program to return additional information on successful authentication, by using `"accept_info"` instead of `"accept"`:

`"accept_info $groups $uid $gid $supplementary_gids $HOME $info\n"`

Where:

* `$info` is some arbitrary text. NSO will then just append this text to the generated audit log message (NCS\_PACKAGE\_AUTH\_SUCCESS).

Yet another possibility is for the program to return a warning that the user's password is about to expire, by using `"accept_warning"` instead of `"accept"`:

`"accept_warning $groups $uid $gid $supplementary_gids $HOME $warning\n"`

Where:

* `$warning` is an appropriate warning message. The message will be processed by NSO according to the setting of `/ncs-config/aaa/expiration-warning` in `ncs.conf`.

If authentication fails, the program should write `"reject"` or `"abort"`, possibly followed by a reason for the rejection and a trailing newline. For example `"reject 'Bad password'\n"` or just `"abort\n"`. The difference between `"reject"` and `"abort"` is that with `"reject"`, NSO will try subsequent mechanisms configured for `/ncs-config/aaa/auth-order`, and packages configured for `/ncs-config/aaa/package-authentication/packages` in `ncs.conf` (if any), while with `"abort"`, the authentication fails immediately. Thus `"abort"` can prevent subsequent mechanisms from being tried, but when external authentication is the last mechanism (as in the default order), it has the same effect as `"reject"`.

When package authentication is used, the group list returned by the package executable is prepended by any possible group information stored locally under the `/aaa` tree. Hence when package authentication is used, it is indeed possible to have the entire `/aaa/authentication` tree empty. The group assignment performed by the external program will still be valid and the relevant groups will be used by NSO when the authorization rules are checked.

### **Username/Password Package Authentication for CLI**

Package authentication will invoke the `scripts/authenticate` when a user tries to authenticate using CLI. In this case, only the username, password, client source IP, client source port, northbound API context, and protocol will be passed to the script.

{% hint style="info" %}
When serving a username/password request, script output other than accept, challenge or abort will be treated as if authentication failed.
{% endhint %}

### **Package Challenges**

When `/ncs-config/aaa/package-authentication/package-challenge/enabled` is set to true, packages will also be used to try to resolve challenges sent to the server and are only supported by CLI over SSH. The script `script/challenge` will be invoked passing challenge ID, response, client source IP, client source port, northbound API context, and protocol on `stdin` using the string notation: `"[challengeid;response;src-ip;src-port;ctx;proto;]\n"` . The output should follow that of the authenticate script.

{% hint style="info" %}
The fields `challengeid` and `response` are base64 encoded when passed to the script.
{% endhint %}

## Authenticating IPC Access

NSO communicates with clients (Python and Java client libraries, `ncs_cli`, `netconf-subsys`, and others) using the NSO IPC socket. By default, only local connections to the IPC socket are allowed and clients are authenticated based on their Unix UID.

If the client is trusted (the UID-based authentication is not used), the IPC protocol allows the client to supply user and group information to use for authorization in NSO. Effectively, this delegates the authentication to the client and enables the `ncs_cli` and similar commands to support the `--user` and `--groups` options. For example, the following command connects to NSO as user `admin` when it is run as the same Unix user as NSO:

```bash
ncs_cli --user admin
```

In general, authenticating access to the IPC socket is a security best practice and should always be used. NSO uses Unix domain sockets for IPC by default, and authenticates the client based on the UID of the other end of the socket connection. Alternatively, the system can be configured to use TCP sockets. In this case, the system should be configured to use an access check, where every IPC client must prove that it has access to a pre-shared key. See [Restricting Access to the IPC Socket](/guides/administration/advanced-topics/ipc-connection#restricting-access-to-the-ipc-socket) on how to enable it.

### UID-based Authentication for Unix Sockets

NSO uses Unix domain sockets for IPC communications by default (the `ncs-local-ipc/enabled` configuration in `ncs.conf` defaults to `true`). The main benefit of this communication method is that it is generally more secure than TCP sockets. It also provides additional information on the communicating peer, such as the user ID of the calling process. NSO can then use this information to authenticate the peer.

As part of the initial handshake, NSO reads the effective UID (euid) of the process initiating the Unix socket connection. The system then finds an `/aaa/authentication/users/user` entry with the corresponding `uid` value. Access is permitted or denied based on the `local_ipc_access` value. If access is permitted, the user connects as the user, found in the `/aaa/authentication/users/user` list. The following is an example of such a user list entry:

```bash
aaa authentication users user admin
 uid              500
 gid              500
 password         $6$...
 ssh_keydir       /var/ncs/homes/admin/.ssh
 homedir          /var/ncs/homes/admin
 local_ipc_access true
!
```

NSO will skip this access check in case the euid of the connecting process is 0 (root user) or same as the user NSO is running as. (In both these cases, the connecting user could access NSO data directly, bypassing the access check.)

If the default Local IPC socket path has been changed, clients and client libraries must specify the path that identifies the socket. The path must match the one set under `ncs-local-ipc/path` in `ncs.conf`. Clients may expose a client-specific way to set it, such as the `-S` option of the `ncs_cli` command. Alternatively, you can use the `NCS_IPC_PATH` environment variable to specify the socket path independently of the used client.

See [examples.ncs/aaa/ipc](https://github.com/NSO-developer/nso-examples/tree/6.7/aaa/ipc) for a working example.

## Group Membership <a href="#ug.aaa.groups" id="ug.aaa.groups"></a>

Once a user is authenticated, group membership must be established. A single user can be a member of several groups. Group membership is used by the authorization rules to decide which operations a certain user is allowed to perform. Thus, the NSO AAA authorization model is entirely group-based. This is also sometimes referred to as role-based authorization.

All groups are stored under `/nacm/groups`, and each group contains a number of usernames. The `ietf-netconf-acm.yang` model defines a group entry:

```yang
list group {
  key name;

  description
    "One NACM Group Entry.  This list will only contain
     configured entries, not any entries learned from
     any transport protocols.";

  leaf name {
    type group-name-type;
    description
      "Group name associated with this entry.";
  }

  leaf-list user-name {
    type user-name-type;
    description
      "Each entry identifies the username of
       a member of the group associated with
       this entry.";
  }
}
```

The `tailf-acm.yang` model augments this with a `gid` leaf:

```yang
augment /nacm:nacm/nacm:groups/nacm:group {
  leaf gid {
    type int32;
    description
      "This leaf associates a numerical group ID with the group.
       When a OS command is executed on behalf of a user,
       supplementary group IDs are assigned based on 'gid' values
       for the groups that the use is a member of.";
  }
}
```

A valid group entry could thus look like:

```xml
<group>
  <name>admin</name>
  <user-name>bob</user-name>
  <user-name>joe</user-name>
  <gid xmlns="http://tail-f.com/yang/acm">99</gid>
</group>
```

The above XML data would then mean that users `bob` and `joe` are members of the `admin` group. The users need not necessarily exist as actual users under `/aaa/authentication/users` in order to belong to a group. If for example PAM authentication is used, it does not make sense to have all users listed under `/aaa/authentication/users`.

By default, the user is assigned to groups by using any groups provided by the northbound transport (e.g. via the `ncs_cli` or `netconf-subsys` programs), by consulting data under `/nacm/groups`, by consulting the `/etc/group` file, and by using any additional groups supplied by the authentication method. If `/nacm/enable-external-groups` is set to "false", only the data under `/nacm/groups` is consulted.

The resulting group assignment is the union of these methods, if it is non-empty. Otherwise, the default group is used, if configured ( `/ncs-config/aaa/default-group` in `ncs.conf`).

A user entry has a UNIX uid and UNIX gid assigned to it. Groups may have optional group IDs. When a user is logged in, and NSO tries to execute commands on behalf of that user, the uid/gid for the command execution is taken from the user entry. Furthermore, UNIX supplementary group IDs are assigned according to the `gid`'s in the groups where the user is a member.

## Authorization <a href="#ug.aaa.authorization" id="ug.aaa.authorization"></a>

Once a user is authenticated and group membership is established, when the user starts to perform various actions, each action must be authorized. Normally the authorization is done based on rules configured in the AAA data model as described in this section.

The authorization procedure first checks the value of `/nacm/enable-nacm`. This leaf has a default of `true`, but if it is set to `false`, all access is permitted. Otherwise, the next step is to traverse the `rule-list` list:

```yang
list rule-list {
  key "name";
  ordered-by user;
  description
    "An ordered collection of access control rules.";

  leaf name {
    type string {
      length "1..max";
    }
    description
      "Arbitrary name assigned to the rule-list.";
  }
  leaf-list group {
    type union {
      type matchall-string-type;
      type group-name-type;
    }
    description
      "List of administrative groups that will be
       assigned the associated access rights
       defined by the 'rule' list.

       The string '*' indicates that all groups apply to the
       entry.";
  }

  // ...
}
```

If the `group` leaf-list in a `rule-list` entry matches any of the user's groups, the `cmdrule` list entries are examined for command authorization, while the `rule` entries are examined for RPC, notification, and data authorization.

### Command Authorization <a href="#d5e6440" id="d5e6440"></a>

The `tailf-acm.yang` module augments the `rule-list` entry in `ietf-netconf-acm.yang` with a `cmdrule` list:

```yang
augment /nacm:nacm/nacm:rule-list {

  list cmdrule {
    key "name";
    ordered-by user;
    description
      "One command access control rule. Command rules control access
       to CLI commands and Web UI functions.

       Rules are processed in user-defined order until a match is
       found.  A rule matches if 'context', 'command', and
       'access-operations' match the request.  If a rule
       matches, the 'action' leaf determines if access is granted
       or not.";

    leaf name {
      type string {
        length "1..max";
      }
      description
        "Arbitrary name assigned to the rule.";
    }

    leaf context {
      type union {
        type nacm:matchall-string-type;
        type string;
      }
      default "*";
      description
        "This leaf matches if it has the value '*' or if its value
         identifies the agent that is requesting access, i.e. 'cli'
         for CLI or 'webui' for Web UI.";
    }

    leaf command {
      type string;
      default "*";
      description
        "Space-separated tokens representing the command. Refer
         to the Tail-f AAA documentation for further details.";
    }

    leaf access-operations {
      type union {
        type nacm:matchall-string-type;
        type nacm:access-operations-type;
      }
      default "*";
      description
        "Access operations associated with this rule.

         This leaf matches if it has the value '*' or if the
         bit corresponding to the requested operation is set.";
    }

    leaf action {
      type nacm:action-type;
      mandatory true;
      description
        "The access control action associated with the
         rule.  If a rule is determined to match a
         particular request, then this object is used
         to determine whether to permit or deny the
         request.";
    }

    leaf log-if-permit {
      type empty;
      description
        "If this leaf is present, access granted due to this rule
         is logged in the developer log. Otherwise, only denied
         access is logged. Mainly intended for debugging of rules.";
    }

    leaf comment {
      type string;
      description
        "A textual description of the access rule.";
    }
  }
}
```

Each rule has seven leafs. The first is the `name` list key, the following three leafs are matching leafs. When NSO tries to run a command, it tries to match the command towards the matching leafs and if all of `context`, `command`, and `access-operations` match, the fifth field, i.e. the `action`, is applied.

* `name`: `name` is the name of the rule. The rules are checked in order, with the ordering given by the YANG `ordered-by user` semantics, i.e. independent of the key values.
* `context`: `context` is either of the strings `cli`, `webui`, or `*` for a command rule. This means that we can differentiate authorization rules for which access method is used. Thus if command access is attempted through the CLI, the context will be the string `cli` whereas for operations via the Web UI, the context will be the string `webui`.
* `command`: This is the actual command getting executed. If the rule applies to one or several CLI commands, the string is a space-separated list of CLI command tokens, for example `request system reboot`. If the command applies to Web UI operations, it is a space-separated string similar to a CLI string. A string that consists of just `*` matches any command.\
  \
  In general, we do not recommend using command rules to protect the configuration. Use rules for data access as described in the next section to control access to different parts of the data. Command rules should be used only for CLI commands and Web UI operations that cannot be expressed as data rules.\
  \
  The individual tokens can be POSIX extended regular expressions. Each regular expression is implicitly anchored, i.e. an `^` is prepended and a `$` is appended to the regular expression.
* `access-operations`: `access-operations` is used to match the operation that NSO tries to perform. It must be one or both of the "read" and "exec" values from the `access-operations-type` bits type definition in `ietf-netconf-acm.yang`, or "\*" to match any operation.
* action: If all of the previous fields match, the rule as a whole matches and the value of `action` will be taken. I.e. if a match is found, a decision is made whether to permit or deny the request in its entirety. If `action` is `permit`, the request is permitted, if `action` is `deny`, the request is denied and an entry is written to the developer log.
* `log-if-permit`: If this leaf is present, an entry is written to the developer log for a matching request also when `action` is `permit`. This is very useful when debugging command rules.
* `comment`: An optional textual description of the rule.

For the rule processing to be written to the devel log, the `/ncs-config/logs/developer-log-level` entry in `ncs.conf` must be set to `trace`.

If no matching rule is found in any of the `cmdrule` lists in any `rule-list` entry that matches the user's groups, this augmentation from `tailf-acm.yang` is relevant:

```yang
augment /nacm:nacm {
  leaf cmd-read-default {
    type nacm:action-type;
    default "permit";
    description
      "Controls whether command read access is granted
       if no appropriate cmdrule is found for a
       particular command read request.";
  }

  leaf cmd-exec-default {
    type nacm:action-type;
    default "permit";
    description
      "Controls whether command exec access is granted
       if no appropriate cmdrule is found for a
       particular command exec request.";
  }

  leaf log-if-default-permit {
    type empty;
    description
      "If this leaf is present, access granted due to one of
       /nacm/read-default, /nacm/write-default, or /nacm/exec-default
       /nacm/cmd-read-default, or /nacm/cmd-exec-default
       being set to 'permit' is logged in the developer log.
       Otherwise, only denied access is logged. Mainly intended
       for debugging of rules.";
  }
}
```

* If `read` access is requested, the value of `/nacm/cmd-read-default` determines whether access is permitted or denied.
* If `exec` access is requested, the value of `/nacm/cmd-exec-default` determines whether access is permitted or denied.

If `access` is permitted due to one of these default leafs, the `/nacm/log-if-default-permit`has the same effect as the `log-if-permit` leaf for the `cmdrule` lists.

### RPC, Notification, and Data Authorization <a href="#d5e6526" id="d5e6526"></a>

The rules in the `rule` list are used to control access to rpc operations, notifications, and data nodes defined in YANG models. Access to invocation of actions (`tailf:action`) is controlled with the same method as access to data nodes, with a request for `exec` access. `ietf-netconf-acm.yang` defines a `rule` entry as:

```yang
list rule {
  key "name";
  ordered-by user;
  description
    "One access control rule.

     Rules are processed in user-defined order until a match is
     found.  A rule matches if 'module-name', 'rule-type', and
     'access-operations' match the request.  If a rule
     matches, the 'action' leaf determines if access is granted
     or not.";

  leaf name {
    type string {
      length "1..max";
    }
    description
      "Arbitrary name assigned to the rule.";
  }

  leaf module-name {
    type union {
      type matchall-string-type;
      type string;
    }
    default "*";
    description
      "Name of the module associated with this rule.

       This leaf matches if it has the value '*' or if the
       object being accessed is defined in the module with the
       specified module name.";
  }
  choice rule-type {
    description
      "This choice matches if all leafs present in the rule
       match the request.  If no leafs are present, the
       choice matches all requests.";
    case protocol-operation {
      leaf rpc-name {
        type union {
          type matchall-string-type;
          type string;
        }
        description
          "This leaf matches if it has the value '*' or if
           its value equals the requested protocol operation
           name.";
      }
    }
    case notification {
      leaf notification-name {
        type union {
          type matchall-string-type;
          type string;
        }
        description
          "This leaf matches if it has the value '*' or if its
           value equals the requested notification name.";
      }
    }
    case data-node {
      leaf path {
        type node-instance-identifier;
        mandatory true;
        description
          "Data Node Instance Identifier associated with the
           data node controlled by this rule.

           Configuration data or state data instance
           identifiers start with a top-level data node.  A
           complete instance identifier is required for this
           type of path value.

           The special value '/' refers to all possible
           data-store contents.";
      }
    }
  }

  leaf access-operations {
    type union {
      type matchall-string-type;
      type access-operations-type;
    }
    default "*";
    description
      "Access operations associated with this rule.

       This leaf matches if it has the value '*' or if the
       bit corresponding to the requested operation is set.";
  }

  leaf action {
    type action-type;
    mandatory true;
    description
      "The access control action associated with the
       rule.  If a rule is determined to match a
       particular request, then this object is used
       to determine whether to permit or deny the
       request.";
  }

  leaf comment {
    type string;
    description
      "A textual description of the access rule.";
  }
}
```

`tailf-acm` augments this with two additional leafs:

```yang
augment /nacm:nacm/nacm:rule-list/nacm:rule {

  leaf context {
    type union {
      type nacm:matchall-string-type;
      type string;
    }
    default "*";
    description
      "This leaf matches if it has the value '*' or if its value
       identifies the agent that is requesting access, e.g. 'netconf'
       for NETCONF, 'cli' for CLI, or 'webui' for Web UI.";

  }

  leaf log-if-permit {
    type empty;
    description
      "If this leaf is present, access granted due to this rule
       is logged in the developer log. Otherwise, only denied
       access is logged. Mainly intended for debugging of rules.";
  }
}
```

Similar to the command access check, whenever a user through some agent tries to access an RPC, a notification, a data item, or an action, access is checked. For a rule to match, three or four leafs must match and when a match is found, the corresponding action is taken.

We have the following leafs in the `rule` list entry.

* `name`: The name of the rule. The rules are checked in order, with the ordering given by the YANG `ordered-by user` semantics, i.e., independent of the key values.
* `module-name`: The `module-name` string is the name of the YANG module where the node being accessed is defined. The special value `*` (i.e., the default) matches all modules.\
  **Note**: Since the elements of the path to a given node may be defined in different YANG modules when augmentation is used, rules that have a value other than `*` for the `module-name` leaf may require that additional processing is done before a decision to permit or deny, or the access can be taken. Thus, if an XPath that completely identifies the nodes that the rule should apply to is given for the `path` leaf (see below), it may be best to leave the `module-name` leaf unset.
* `rpc-name / notification-name / path`: This is a choice between three possible leafs that are used for matching, in addition to the `module-name`:
* `rpc-name`: The name of an RPC operation, or `*` to match any RPC.
* `notification-name`: the name of a notification, or `*` to match any notification.
* `path`: A restricted XPath expression leading down into the populated XML tree. A rule with a path specified matches if it is equal to or shorter than the checked path. Several types of paths are allowed.

  1. Tagpaths that do not contain any keys. For example `/ncs/live-device/live-status`.
  2. Instantiated key: as in `/devices/device[name="x1"]/config/interface` matches the interface configuration for managed device "x1" It's possible to have partially instantiated paths only containing some keys instantiated - i.e. combinations of tagpaths and keypaths. Assuming a deeper tree, the path `/devices/device/config/interface[name="eth0"]` matches the `eth0` interface configuration on all managed devices.
  3. The wild card at the end as in: `/services/web-site/*` does not match the website service instances, but rather all children of the website service instances.
  4. The leading/trailing whitespace as in: `" /devices/device/config "` are ignored.

  Thus, the path in a rule is matched against the path in the attempted data access. If the attempted access has a path that is equal to or longer than the rule path - we have a match.\
  \
  If none of the leafs `rpc-name`, `notification-name`, or `path` are set, the rule matches for any RPC, notification, data, or action access.
* `context`: `context` is either of the strings `cli`, `netconf`, `webui`, `snmp`, or `*` for a data rule. Furthermore, when we initiate user sessions from MAAPI, we can choose any string we want. Similarly to command rules, we can differentiate access depending on which agent is used to gain access.
* `access-operations`: `access-operations` is used to match the operation that NSO tries to perform. It must be one or more of the "create", "read", "update", "delete" and "exec" values from the `access-operations-type` bits type definition in `ietf-netconf-acm.yang`, or "\*" to match any operation.
* `action`: This leaf has the same characteristics as the `action` leaf for command access.
* `log-if-permit`: This leaf has the same characteristics as the `log-if-permit` leaf for command access.
* `comment`: An optional textual description of the rule.

If no matching rule is found in any of the `rule` lists in any `rule-list` entry that matches the user's groups, the data model node for which access is requested is examined for the presence of the NACM extensions:

* If the `nacm:default-deny-all` extension is specified for the data model node, the access is denied.
* If the `nacm:default-deny-write` extension is specified for the data model node, and `create`, `update`, or `delete` access is requested, the access is denied.

If examination of the NACM extensions did not result in access being denied, the value (`permit` or `deny`) of the relevant default leaf is examined:

* If `read` access is requested, the value of `/nacm/read-default` determines whether access is permitted or denied.
* If `create`, `update`, or `delete` access is requested, the value of `/nacm/write-default` determines whether access is permitted or denied.
* If `exec` access is requested, the value of `/nacm/exec-default` determines whether access is permitted or denied.

If access is permitted due to one of these default leafs, this augmentation from `tailf-acm.yang` is relevant:

```yang
augment /nacm:nacm {
  ...
  leaf log-if-default-permit {
    type empty;
    description
      "If this leaf is present, access granted due to one of
       /nacm/read-default, /nacm/write-default, /nacm/exec-default
       /nacm/cmd-read-default, or /nacm/cmd-exec-default
       being set to 'permit' is logged in the developer log.
       Otherwise, only denied access is logged. Mainly intended
       for debugging of rules.";
  }
}
```

I.e., it has the same effect as the `log-if-permit` leaf for the `rule` lists, but for the case where the value of one of the default leafs permits access.

When NSO executes a command, the command rules in the authorization database are searched, The rules are tried in order, as described above. When a rule matches the operation (command) that NSO is attempting, the action of the matching rule is applied — whether permit or deny.

When actual data access is attempted, the data rules are searched. E.g., when a user attempts to execute `delete aaa` in the CLI, the user needs delete access to the entire tree `/aaa`.

Another example is if a CLI user writes `show configuration aaa` <kbd>TAB</kbd>, it suffices to have read access to at least one item below `/aaa` for the CLI to perform the <kbd>TAB</kbd> completion. If no rule matches or an explicit deny rule is found, the CLI will not <kbd>TAB</kbd>-complete.

Yet another example is if a user tries to execute `delete aaa authentication users`, we need to perform a check on the paths `/aaa` and `/aaa/authentication` before attempting to delete the sub-tree. Say that we have a rule for path `/aaa/authentication/users` which is a permit rule and we have a subsequent rule for path `/aaa` which is a deny rule. With this rule set the user should indeed be allowed to delete the entire `/aaa/authentication/users` tree but not the `/aaa` tree nor the `/aaa/authentication` tree.

We have two variations on how the rules are processed. The easy case is when we actually try to read or write an item in the configuration database. The execution goes like this:

```
foreach rule {
    if (match(rule, path)) {
       return rule.action;
    }
}
```

The second case is when we execute TAB completion in the CLI. This is more complicated. The execution goes like this:

```
rules = select_rules_that_may_match(rules, path);
if (any_rule_is_permit(rules))
    return permit;
else
    return deny;
```

The idea is that as we traverse (through <kbd>TAB</kbd>) down the XML tree, as long as there is at least one rule that can possibly match later, once we have more data, we must continue. For example, assume we have:

1. `"/system/config/foo" --> permit`
2. `"/system/config" --> deny`

If we in the CLI stand at `"/system/config"` and hit <kbd>TAB</kbd> we want the CLI to show `foo` as a completion, but none of the other nodes that exist under `/system/config`. Whereas if we try to execute `delete /system/config` the request must be rejected.

By default, NACM rules are configured for the entire `tailf:action` or YANG 1.1 `action` statements, but not for `input` statement child leafs. To override this behavior, and enable NACM rules on `input` leafs, set the following parameter to 'true': `/ncs-config/aaa/action-input-rules/enabled`. When enabled all action input leafs given to an action will be validated for NACM rules. If broad 'deny' NACM rules are used, you might need to add 'permit' rules for the affected action input leafs to allow actions to be used with parameters.

### NACM Rules and Services <a href="#d5e6693" id="d5e6693"></a>

By design NACM rules are ignored for changes done by services — FASTMAP, Reactive FASTMAP, or Nano services. The reasoning behind this is that a service package can be seen as a controlled way to provide limited access to devices for a user group that is not allowed to apply arbitrary changes on the devices.

However, there are NSO installations where this behavior is not desired, and NSO administrators want to enforce NACM rules even on changes done by services. For this purpose, the leaf called `/nacm/enforce-nacm-on-services` is provided. By default, it is set to `false`.

Note however that currently, even with this leaf set to true, there are limitations. Namely, the post-actions for nano-services are run in a user session without any access checks. Besides that, NACM rules are not enforced on the read operations performed in the service callbacks.

It might be desirable to deny everything for a user group and only allow access to a specific service. This pattern could be used to allow an operator to provision the service, but deny everything else. While this pattern works for a normal FASTMAP service, there are some caveats for stacked services, Reactive FASTMAP, and Nano services. For these kinds of services, in addition to the service itself, access should be provided to the user group for the following paths:

* In case of stacked services, the user group needs read and write access to the leaf `private/re-deploy-counter` under the bottom service. Otherwise, the user will not be able to redeploy the service.
* In the case of Reactive FASTMAP or Nano services, the user group needs read and write access to the following:
  * `/zombies`
  * `/side-effect-queue`
  * `/kickers`

### Device Group Authorization <a href="#d5e6700" id="d5e6700"></a>

In deployments with many devices, it can become cumbersome to handle data authorization per device. To help with this there is a rule type that works on device group membership (for more on device groups, see [Device Groups](https://nso-docs.cisco.com/guides/administration/management/pages/auKQMOAF2p1jiGYJBweP#user_guide.devicemanager.device_groups)). To do this, devices are added to different device groups, and the rule type `device-group-rule` is used.

The IETF NACM rule type is augmented with a new rule type named `device-group-rule` which contains a leafref to the device groups. See the following example.

{% code title="Device Group Model Augmentation" %}

```yang
augment "/nacm:nacm/nacm:rule-list/nacm:rule/nacm:rule-type" {
  case device-group-rule {
    leaf device-group {
      type leafref {
        path "/ncs:devices/ncs:device-group/ncs:name";
      }
      description
        "Which device group this rule applies to.";
    }
  }
}
```

{% endcode %}

In the example below, we configure two device groups based on different regions and add devices to them.

{% code title="Device Group Configuration" %}

```xml
<devices>
  <device-group>
    <name>us_east</name>
    <device-name>cli0</device-name>
    <device-name>gen0</device-name>
  </device-group>
  <device-group>
    <name>us_west</name>
    <device-name>nc0</device-name>
  </device-group>
</devices>
```

{% endcode %}

In the example below, we configure an operator for the `us_east` region:

{% code title="NACM Group Configuration" %}

```xml
<nacm>
  <groups>
    <group>
      <name>us_east</name>
      <user-name>us_east_oper</user-name>
    </group>
  </groups>
</nacm>
```

{% endcode %}

\
In the example below, we configure the device group rules and refer to the device group and the `us_east` group.

{% code title="Device Group Authorization Rules" %}

```xml
<nacm>
  <rule-list>
    <name>us_east</name>
    <group>us_east</group>
    <rule>
      <name>us_east_read_permit</name>
      <device-group xmlns="http://tail-f.com/yang/ncs-acm/device-group-authorization">us_east</device-group>
      <access-operations>read</access-operations>
      <action>permit</action>
    </rule>
    <rule>
      <name>us_east_create_permit</name>
      <device-group xmlns="http://tail-f.com/yang/ncs-acm/device-group-authorization">us_east</device-group>
      <access-operations>create</access-operations>
      <action>permit</action>
    </rule>
    <rule>
      <name>us_east_update_permit</name>
      <device-group xmlns="http://tail-f.com/yang/ncs-acm/device-group-authorization">us_east</device-group>
      <access-operations>update</access-operations>
      <action>permit</action>
    </rule>
    <rule>
      <name>us_east_delete_permit</name>
      <device-group xmlns="http://tail-f.com/yang/ncs-acm/device-group-authorization">us_east</device-group>
      <access-operations>delete</access-operations>
      <action>permit</action>
    </rule>
  </rule-list>
</nacm>
```

{% endcode %}

In summary device group authorization gives a more compact configuration for deployments where devices can be grouped and authorization can be done on a device group basis.

Modifications on the device-group subtree are recommended to be controlled by a limited set of users.

### Authorization Examples <a href="#d5e6730" id="d5e6730"></a>

Assume that we have two groups, `admin` and `oper`. We want `admin` to be able to see and edit the XML tree rooted at `/aaa`, but we do not want users who are members of the `oper` group to even see the `/aaa` tree. We would have the following rule list and rule entries. Note, here we use the XML data from `tailf-aaa.yang` to exemplify. The examples apply to all data, for all data models loaded into the system.

```xml
<rule-list>
  <name>admin</name>
  <group>admin</group>
  <rule>
    <name>tailf-aaa</name>
    <module-name>tailf-aaa</module-name>
    <path>/</path>
    <access-operations>read create update delete</access-operations>
    <action>permit</action>
  </rule>
</rule-list>
<rule-list>
  <name>oper</name>
  <group>oper</group>
  <rule>
    <name>tailf-aaa</name>
    <module-name>tailf-aaa</module-name>
    <path>/</path>
    <access-operations>read create update delete</access-operations>
    <action>deny</action>
  </rule>
</rule-list>
```

If we do not want the members of `oper` to be able to execute the NETCONF operation `edit-config`, we define the following rule list and rule entries:

```xml
<rule-list>
  <name>oper</name>
  <group>oper</group>
  <rule>
    <name>edit-config</name>
    <rpc-name>edit-config</rpc-name>
    <context xmlns="http://tail-f.com/yang/acm">netconf</context>
    <access-operations>exec</access-operations>
    <action>deny</action>
  </rule>
</rule-list>
```

To spell it out, the above defines four elements to match. If NSO tries to perform a `netconf` operation, which is the operation `edit-config`, and the user who runs the command is a member of the `oper` group, and finally it is an `exec` (execute) operation, we have a match. If so, the action is `deny`.

The `path` leaf can be used to specify explicit paths into the XML tree using XPath syntax. For example the following:

```xml
<rule-list>
  <name>admin</name>
  <group>admin</group>
  <rule>
    <name>bob-password</name>
    <path>/aaa/authentication/users/user[name='bob']/password</path>
    <context xmlns="http://tail-f.com/yang/acm">cli</context>
    <access-operations>read update</access-operations>
    <action>permit</action>
  </rule>
</rule-list>
```

Explicitly allows the `admin` group to change the password for precisely the `bob` user when the user is using the CLI. Had `path` been `/aaa/authentication/users/user/password` the rule would apply to all password elements for all users. Since the `path` leaf completely identifies the nodes that the rule applies to, we do not need to give `tailf-aaa` for the `module-name` leaf.

NSO applies variable substitution, whereby the username of the logged-in user can be used in a `path`. Thus:

```xml
<rule-list>
  <name>admin</name>
  <group>admin</group>
  <rule>
    <name>user-password</name>
    <path>/aaa/authentication/users/user[name='$USER']/password</path>
    <context xmlns="http://tail-f.com/yang/acm">cli</context>
    <access-operations>read update</access-operations>
    <action>permit</action>
  </rule>
</rule-list>
```

The above rule allows all users that are part of the `admin` group to change their own passwords only.

A member of `oper` is able to execute NETCONF operation `action` if that member has `exec` access on NETCONF RPC `action` operation, `read` access on all instances in the hierarchy of data nodes that identifies the specific action in the data store, and `exec` access on the specific action. For example, an action is defined as below.

```yang
container test {
  action double {
    input {
      leaf number {
        type uint32;
      }
    }
    output {
      leaf result {
        type uint32;
      }
    }
  }
}
```

To be able to execute `double` action through NETCONF RPC, the members of `oper` need the following rule list and rule entries.

```xml
<rule-list>
  <name>oper</name>
  <group>oper</group>

  <rule>
    <name>allow-netconf-rpc-action</name>
    <rpc-name>action</rpc-name>
    <context xmlns="http://tail-f.com/yang/acm">netconf</context>
    <access-operations>exec</access-operations>
    <action>permit</action>
  </rule>
  <rule>
    <name>allow-read-test</name>
    <path>/test</path>
    <access-operations>read</access-operations>
    <action>permit</action>
  </rule>
  <rule>
    <name>allow-exec-double</name>
    <path>/test/double</path>
    <access-operations>exec</access-operations>
    <action>permit</action>
  </rule>
</rule-list>
```

Or, a simpler rule set as the following.

```xml
<rule-list>
  <name>oper</name>
  <group>oper</group>

  <rule>
    <name>allow-netconf-rpc-action</name>
    <rpc-name>action</rpc-name>
    <context xmlns="http://tail-f.com/yang/acm">netconf</context>
    <access-operations>exec</access-operations>
    <action>permit</action>
  </rule>
  <rule>
    <name>allow-exec-double</name>
    <path>/test</path>
    <access-operations>read exec</access-operations>
    <action>permit</action>
  </rule>
</rule-list>
```

Finally, if we wish members of the `oper` group to never be able to execute the `request system reboot` command, also available as a `reboot` NETCONF rpc, we have:

```xml
<rule-list>
  <name>oper</name>
  <group>oper</group>

  <cmdrule xmlns="http://tail-f.com/yang/acm">
    <name>request-system-reboot</name>
    <context>cli</context>
    <command>request system reboot</command>
    <access-operations>exec</access-operations>
    <action>deny</action>
  </cmdrule>

  <!-- The following rule is required since the user can -->
  <!-- do "edit system" -->

  <cmdrule xmlns="http://tail-f.com/yang/acm">
    <name>request-reboot</name>
    <context>cli</context>
    <command>request reboot</command>
    <access-operations>exec</access-operations>
    <action>deny</action>
  </cmdrule>

  <rule>
    <name>netconf-reboot</name>
    <rpc-name>reboot</rpc-name>
    <context xmlns="http://tail-f.com/yang/acm">netconf</context>
    <access-operations>exec</access-operations>
    <action>deny</action>
  </rule>

</rule-list>
```

### Troubleshooting NACM Rules

In this section, we list some tips to make it easier to troubleshoot NACM rules.

{% hint style="success" %}
Use `log-if-permit` and `log-if-default-permit` together with the developer log level set to `trace`.
{% endhint %}

Use the `tailf-acm.yang` module augmentation `log-if-permit` leaf for rules with `action` `permit`. When those rules trigger a permit action a trace entry is added to the developer log. To see trace entries make sure the `/ncs-config/logs/developer-log-level` is set to `trace`.

If you have a default rule with `action` `permit` you can use the `log-if-default-permit` leaf instead.

{% hint style="success" %}
NACM rules are read at the start of the session and are used throughout the session.
{% endhint %}

When a user session is created it will gather the authorization rules that are relevant for that user's group(s). The rules are used throughout the user session lifetime. When we update the AAA rules the active sessions are not affected. For example, if an administrator updates the NACM rules in one session the update will not apply to any other currently active sessions. The updates will apply to new sessions created after the update.

{% hint style="success" %}
Explicitly state NACM groups when starting the CLI. For example `ncs_cli -u oper -g oper`.
{% endhint %}

It is the user's group membership that determines what rules apply. Starting the CLI using the `ncs_cli` command without explicitly setting the groups, defaults to the actual UNIX groups the user is a member of. On Darwin, one of the default groups is usually `admin`, which can lead to the wrong group being used.

{% hint style="success" %}
Be careful with namespaces in rulepaths.
{% endhint %}

Unless a rulepath is made explicit by specifying namespace it will apply to that specific path in all namespaces. Below we show parts of an example from [RFC 8341](https://tools.ietf.org/html/rfc8341), where the `path` element has an `xmlns` attribute and the path is namespaced. If these would not have been namespaced, the rules would not behave as expected.

{% code title="Example: Excerpt from RFC 8341 Appendix A.4" %}

```xml
         <rule>
           <name>permit-acme-config</name>
           <path xmlns:acme="http://example.com/ns/netconf">
             /acme:acme-netconf/acme:config-parameters
           </path>
         ...
```

{% endcode %}

\
In the example above (Excerpt from RFC 8341 Appendix A.4), the path is namespaced.

## The AAA Cache <a href="#d5e6799" id="d5e6799"></a>

NSO's AAA subsystem will cache the AAA information in order to speed up the authorization process. This cache must be updated whenever there is a change to the AAA information. The mechanism for this update depends on how the AAA information is stored, as described in the following two sections.

### Populating AAA using CDB <a href="#d5e6802" id="d5e6802"></a>

To start NSO, the data models for AAA must be loaded. The defaults in the case that no actual data is loaded for these models allow all read and exec access, while write access is denied. Access may still be further restricted by the NACM extensions, though — e.g., the `/nacm` container has `nacm:default-deny-all`, meaning that not even read access is allowed if no data is loaded.

The NSO installation ships with an XML initialization file containing AAA configuration. The file is called `aaa_init.xml` and is, by default, copied to the CDB directory by the NSO install scripts.

The local installation variant, targeting development only, defines two users, `admin` and `oper` with passwords set to `admin` and `oper` respectively for authentication. The two users belong to user groups with NACM rules restricting their authorization level. The system installation `aaa_init.xml` variant, targeting production deployment, defines NACM rules only as users are, by default, authenticated using PAM. The NACM rules target two user groups, `ncsadmin` and `ncsoper`. Users belonging to the `ncsoper` group are limited to read-only access.

{% hint style="info" %}
The default `aaa_init.xml` file provided with the NSO system installation must not be used as-is in a deployment without reviewing and verifying that every NACM rule in the file matches
{% endhint %}

Normally the AAA data will be stored as configuration in CDB. This allows for changes to be made through NSO's transaction-based configuration management. In this case, the AAA cache will be updated automatically when changes are made to the AAA data. If changing the AAA data via NSO's configuration management is not possible or desirable, it is alternatively possible to use the CDB operational data store for AAA data. In this case, the AAA cache can be updated either explicitly e.g. by using the `maapi_aaa_reload()` function, see the [confd\_lib\_maapi(3)](/guides/resources/man/confd_lib_maapi.3) in the Manual Pages manual page, or by triggering a subscription notification by using the subscription lock when updating the CDB operational data store, see [Using CDB](/guides/development/core-concepts/using-cdb) in Development.

### Hiding the AAA Tree <a href="#d5e6817" id="d5e6817"></a>

Some applications may not want to expose the AAA data to end users in the CLI or the Web UI. Two reasonable approaches exist here and both rely on the `tailf:export` statement. If a module has `tailf:export none` it will be invisible to all agents. We can then either use a transform whereby we define another AAA model and write a transform program that maps our AAA data to the data that must exist in `tailf-aaa.yang` and `ietf-netconf-acm.yang`. This way we can choose to export and and expose an entirely different AAA model.

Yet another very easy way out, is to define a set of static AAA rules whereby a set of fixed users and fixed groups have fixed access to our configuration data. Possibly the only field we wish to manipulate is the password field.


# NED Administration

Learn about Cisco-provided NEDs and how to manage them.

This section provides necessary information on Network Element Driver (NED) administration with a focus on Cisco-provided NEDs. If you're planning to use NEDs not provided by Cisco, refer to the [NED Development](/guides/development/advanced-development/developing-neds) to build your own NED packages.

## NED Introduction

NED represents a key NSO component that makes it possible for the NSO core system to communicate southbound with network devices in most deployments. NSO has a built-in client that can be used to communicate southbound with NETCONF-enabled devices. Many network devices are, however, not NETCONF-enabled, and there exist a wide variety of methods and protocols for configuring network devices, ranging from simple CLI to HTTP/REST-enabled devices. For such cases, it is necessary to use a NED to allow NSO communicate southbound with the network device.

Even for NETCONF-enabled devices, it is possible that the NSO's built-in NETCONF client cannot be used, for instance, if the devices do not strictly follow the specification for the NETCONF protocol. In such cases, one must also use a NED to seamlessly communicate with the device. See [Managing Cisco-provided third Party YANG NEDs](#sec.managing_thirdparty_neds) for more information on third-party YANG NEDs.

### NED Contents and Capabilities

It's important to understand the functionality of a NED and the capabilities it offers — as well as those it does not. The following summarizes what a NED contains and what it doesn't.

#### **What a NED Provides**

<details>

<summary>YANG Data Model</summary>

The NED provides a YANG data model of the device to NSO and services, enabling standardized configuration management. This applies only to NEDs where Cisco creates and maintains the device data model—commonly referred to as classic NEDs, which includes both the CLI-based and Generic NEDs—and excludes third-party YANG (3PY) NEDs, where the model is provided externally.

Note that for classic NEDs, the device model is typically implemented as a superset, covering multiple versions or variants of a given device type. This approach allows a single NED package to support a broad range of software versions or hardware flavors. The benefit is simplified deployment and upgrade handling across similar devices. However, a side effect is that certain parts of the model may not apply to the specific device instance in use.

</details>

<details>

<summary>Data Translation</summary>

The NED is responsible for transforming outbound data from NSO's internal format into a format understood by the device — whether that format is vendor-specific (e.g., CLI, REST, SOAP) or standards-based (e.g., NETCONF, RESTCONF, gNMI). It also handles the reverse transformation for inbound data from the device back into NSO's format.

</details>

NSO ensures all data modifications occur within a single transaction for consistency and guarantees a transaction is either completely successful or fails, maintaining data integrity.

#### **What a NED Does not Provide**

<details>

<summary>A Data Model of the Entire Set in the Data</summary>

For Classic NEDs, NED development is use-case driven. As a result, a NED, in most cases, does not contain the complete data model of a device. Providing a 100% complete YANG model for a device is not a goal and is not in the scope of NED development. It does not make sense to invest resources into modeling data which is not needed to support the desired use cases. If a NED does not cover a needed use case, please submit an enhancement request via your support channel. For third party NEDs, the models come from third party sources not controlled by Cisco.

</details>

<details>

<summary>An Exact Copy of the Syntax in the Device CLI</summary>

NED development focuses on representing device data for NSO. As a side effect for CLI NEDs, the NSO CLI will get similar behavior as the device CLI, however, in most situations, this will not be perfect and is not the goal of the NED.

</details>

<details>

<summary>Fine-grained Validation of Data (Classic NEDs Only)</summary>

In classic NEDs, adding strict validations in the YANG model (e.g., `mandatory`, `when`, `must`, `range`, `min`, `max`, etc.) can lead to inflexible models. These constraints are interpreted and enforced by NSO at runtime, not the device. Since such validations often need to be updated as devices evolve across versions, NSO's policy is to keep the models relaxed by minimizing the use of these validation constructs. This allows for greater flexibility and forward compatibility.

</details>

<details>

<summary>Convenience Macros in the Device CLI (Only Discrete Data Leaves are Supported)</summary>

Some devices have macro-style functionality in the CLI and users may find it annoying that these are not available in NEDs. The convenience macros have proven very dynamic in the parameters they change, causing frequent out-of-sync situations, but these are generally not available in the NED.

</details>

<details>

<summary>Dynamic Configuration in Devices (Only Data in a Transaction May Change)</summary>

Cisco NEDs do not model device-generated or dynamic configuration, as such behavior varies between device versions and is difficult to standardize. Only configuration explicitly included in a transaction is managed by NSO. If needed, service logic can insert expected dynamic elements during provisioning.

</details>

<details>

<summary>Auto-correction of Parameters with Multiple Syntaxes (i.e., Use Canonical Form)</summary>

The NED does not allow the same value for a parameter to have a different name (e.g., `true` vs. `yes`). The canonical name displayed in `show-running-config` or similar is used.

</details>

<details>

<summary>Handling Out-of-band Changes (Model as Operational Data)</summary>

Leaves that have out-of-band changes will cause NSO and the device to become out- of-sync, and should be made "config false", or not be part of the model at all. Similarly, actions that cause out-of-band changes are not supported.

</details>

<details>

<summary>Splitting a Single Transaction into Several Sub-transactions</summary>

For devices that support the transaction paradigm, the NED will never split an NSO transaction in two or more device transactions. The service must handle this by doing multiple NSO transactions.

</details>

<details>

<summary>Backporting of Fixes to Old NED Releases (i.e., Trunk based Development is Used)</summary>

All NEDs use trunk-based development, i.e., new NED releases are created from the tip of a single branch, develop. New features and fixes are thus delivered to the stakeholders in the latest NED release, not by backporting an old release.

</details>

## Types of NED Packages <a href="#d5e8900" id="d5e8900"></a>

A NED package is a package that NSO uses to manage a particular type of device. A NED is a piece of code that enables communication with a particular type of managed device. You add NEDs to NSO as a special kind of package, called NED packages.

A NED package must provide a device YANG model as well as define means (protocol) to communicate with the device. The latter can either leverage the NSO built-in NETCONF and SNMP support or use a custom implementation. When a package provides custom protocol implementation, typically written in Java, it is called a CLI NED or a Generic NED.

Cisco provides and supports a number of such NEDs. With these Cisco-provided NEDs, a major category are CLI NEDs which communicate with a device through its CLI instead of a dedicated API.

<div data-with-frame="true"><figure><img src="/files/0YTtedJEvrUzWs4XhG1p" alt=""><figcaption><p>NED Package Types</p></figcaption></figure></div>

### NED Types Summary Table

<table><thead><tr><th width="144.1484375" valign="top">NED Category</th><th width="193.3515625" valign="top">Purpose</th><th valign="top">Provider</th><th width="192.66015625" valign="top">YANG Model Provider</th><th width="198.25390625" valign="top">YANG Models Included?</th><th width="173.2890625" valign="top">Device Interface</th><th width="181.953125" valign="top">Protocols Supported</th><th width="275.9375">Key Characteristics</th></tr></thead><tbody><tr><td valign="top"><strong>CLI NED</strong>*</td><td valign="top">Designed for devices with a CLI-based interface. The NED parses CLI commands and translates data to/from YANG.</td><td valign="top">Cisco</td><td valign="top">Cisco NSO NED Team</td><td valign="top">Yes</td><td valign="top">CLI (Command Line Interface)</td><td valign="top">SSH, Telnet</td><td><ul><li>Mimics CLI command hierarchy</li><li>Turbo parser for CLI parsing</li><li>Transform engines for data conversion</li><li>Targets devices using CLI as config interface</li></ul></td></tr><tr><td valign="top"><strong>Generic NED - Cisco YANG Models</strong>*</td><td valign="top">Built for API-based devices (e.g., REST, SOAP, TL1), using custom parsers and data transformation logic maintained by Cisco.</td><td valign="top">Cisco</td><td valign="top">Cisco NSO NED Team</td><td valign="top">Yes</td><td valign="top">Non-CLI (API-based)</td><td valign="top">REST, TL1, CORBA, SOAP, RESTCONF, gNMI, NETCONF</td><td><ul><li>Model-driven devices</li><li>YANG models mimic proprietary protocol messages</li><li>JSON/XML transformers</li><li>Custom protocol implementations</li></ul></td></tr><tr><td valign="top"><strong>Third-party YANG NED</strong></td><td valign="top">Cisco-supplied generic NED packages that do not include any device models.</td><td valign="top">Cisco</td><td valign="top">Third-party Vendors/Organizations (IETF, IEEE, ONF, OpenConfig)</td><td valign="top">No - Must be downloaded separately</td><td valign="top">Model-driven protocols</td><td valign="top">NETCONF, RESTCONF, gNMI</td><td><ul><li>Delivered without YANG models</li><li>Requires download and rebuild process</li><li>Includes recipes for YANG/device fixes</li><li>Legal restrictions prevent Cisco redistribution</li></ul></td></tr></tbody></table>

<sup>\*Also referred to as Classic NED.</sup>

### CLI NED <a href="#d5e8910" id="d5e8910"></a>

This NED category is targeted at devices that use CLI as a configuration interface. Cisco-provided CLI NEDs are available for various network devices from different vendors. Many different CLI syntaxes are supported.

The driver element in a CLI NED implemented by the Cisco NSO NED team typically consists of the following three parts:

* The protocol client, responsible for connecting to and interacting with the device. The protocols supported are SSH and Telnet.
* A fast and versatile CLI parser (+ emitter), usually referred to as the turbo parser.
* Various transform engines capable of converting data between NSO and device formats.

The YANG models in a CLI NED are developed and maintained by the Cisco NSO NED team. Usually, the models for a CLI NED are structured to mimic the CLI command hierarchy on the device.

<div data-with-frame="true"><figure><img src="/files/JjYIgGvB2IvhwOxmEQwI" alt="" width="375"><figcaption><p>CLI NED</p></figcaption></figure></div>

### Generic NED <a href="#d5e8927" id="d5e8927"></a>

A Generic NED is typically used to communicate with non-CLI devices, such as devices using protocols like REST, TL1, Corba, SOAP, RESTCONF, or gNMI as a configuration interface. Even NETCONF-enabled devices in many cases require a generic NED to function properly with NSO.

The driver element in a Generic NED implemented by the Cisco NED team typically consists of the following parts:

* The protocol client, responsible for interacting with the device.
* Various transform engines capable of converting data between NSO and the device formats, usually JSON and/or XML transformers.

There are two types of Generic NEDs maintained by the Cisco NSO NED team:

* NEDs with Cisco-owned YANG models. These NEDs have models developed and maintained by the Cisco NSO NED team.
* NEDs targeted at YANG models from third-party vendors, also known as, third-party YANG NEDs.

### **Generic Cisco-provided NEDs with Cisco-owned YANG Models**

Generic NEDs belonging to the first category typically handle devices that are not model-driven. For instance, devices using proprietary protocols based on REST, SOAP, Corba, etc. The YANG models for such NEDs are usually structured to mimic the messages used by the proprietary protocol of the device.

<div data-with-frame="true"><figure><img src="/files/u1MAoGHDmmrumnnbSwM2" alt="" width="375"><figcaption><p>Generic NED</p></figcaption></figure></div>

### **Third-party YANG NEDs**

As the name implies, this NED category is used for cases where the device YANG models are not implemented, maintained, or owned by the Cisco NSO NED team. Instead, the YANG models are typically provided by the device vendor itself, or by organizations like IETF, IEEE, ONF, or OpenConfig.

This category of NEDs has some special characteristics that set them apart from all other NEDs developed by the Cisco NSO NED team:

* Targeted for devices supporting model-driven protocols like NETCONF, RESTCONF, and gNMI.
* Delivered from the software.cisco.com portal without any device YANG models included. There are several reasons for this, such as legal restrictions that prevent Cisco from re-distributing YANG models from other vendors, or the availability of several different version bundles for open-source YANG, like OpenConfig. The version used by the NED must match the version used by the targeted device.
* The NEDs can be bundled with various fixes to solve shortcomings in the YANG models, the download sources, and/or in the device. These fixes are referred to as recipes.

<div data-with-frame="true"><figure><img src="/files/tsdunTncUVD2NmQFP3mc" alt=""><figcaption><p>Third-Party YANG NEDs</p></figcaption></figure></div>

Since the third-party NEDs are delivered without any device YANG models, there are additional steps required to make this category of NEDs operational:

1. The device models need to be downloaded and copied into the NED package source tree. This can be done by using a special (optional) downloader tool bundled with each third-party YANG NED, or in any custom way.
2. The NED must be rebuilt with the downloaded YANG models.

This procedure is thoroughly described in [Managing Cisco-provided third-Party YANG NEDs](#sec.managing_thirdparty_neds).

#### **Recipes**

A third-party YANG NED can be bundled with up to three types of recipe modules. These recipes are used by the NED to solve various types of issues related to:

* The source of the YANG files.
* The YANG files.
* The device itself.

The recipes represent the characteristics and the real value of a third-party YANG NED. Recipes are typically adapted for a certain bundle of YANG models and/or certain device types. This is why there exist many different third-party YANG NEDs, each one adapted for a specific protocol, a specific model package, and/or a specific device.

{% hint style="info" %}
The NSO NED team does not provide any super third-party YANG NEDs, for instance, a super RESTCONF NED that can be used with any models and any device.
{% endhint %}

**Third-party YANG NED Recipe Types**

<table><thead><tr><th valign="top">Recipe Type</th><th valign="top">Purpose</th><th valign="top">Description</th></tr></thead><tbody><tr><td valign="top"><strong>Download Recipes (DR)</strong></td><td valign="top">YANG Model Sourcing</td><td valign="top"><ul><li>Presets for downloader tool</li><li>Define download sources (device, Git repos, archives)</li><li>Limit scope of YANG files to download</li><li>Multiple profiles per NED</li></ul></td></tr><tr><td valign="top"><strong>YANG Recipes (YR)</strong></td><td valign="top">YANG File Fixes</td><td valign="top"><ul><li>Patch downloaded YANG files before compilation</li><li>Fix compilation errors and YANG construct issues</li><li>Applied automatically during make process</li></ul></td></tr><tr><td valign="top"><strong>Runtime Recipes (RR)</strong></td><td valign="top">Device Behavior Fixes</td><td valign="top"><ul><li>Handle device runtime deviations</li><li>Fix protocol implementation issues</li><li>Clean up "dirty" configuration dumps</li><li>Handle device aliasing issues</li><li>Configurable via runtime profiles</li></ul></td></tr></tbody></table>

**Download Recipes (or Download Profiles)**

When downloading the YANG files, it is first of all important to know which source to use. In some cases, the source is the device itself. For instance, if the device is enabled for NETCONF and sometimes for RESTCONF (in rare cases).

In other cases, the device does not support model download. This applies to all gNMI-enabled devices and most RESTCONF devices too. In this case, the source can be a public Git repository or an archive file provided by the device vendor.

Another important question is what YANG models and what versions to download. To make this task easier, third-party NEDs can be bundled with the download recipes (also known as download profiles). These are presets to be used with the downloader tool bundled with the NED. There can be several profiles, each representing a preset that has been verified to work by the Cisco NSO NED team. A profile can point out a certain source to download from. It can also limit the scope of the download so that only certain YANG files are selected.

**YANG Recipes (YR)**

Third-party YANG files can often contain various types of errors, ranging from real bugs that cause compilation errors to certain YANG constructs that are known to cause runtime issues in NSO. To ensure that the files can be built correctly, the third-party NEDs can be bundled with YANG recipes. These recipes patch the downloaded YANG files before they are built by the NSO compiler. This procedure is performed automatically by the `make` system when the NED is rebuilt after downloading the device YANG files. For more information, refer to the procedure related to rebuilding the NED with a unique NED ID in NED READMEs.

In some cases, YANG recipes are also necessary when a device does not fully conform to the behavior described by its advertised YANG models. This often happens when the device is more permissive than the model suggests—for example, allowing optional parameters that the model marks as mandatory, or omitting data that is expected. Such mismatches can lead to runtime issues in NSO, such as `sync-from` failures or commit errors. YANG recipes allow patching the models to reflect the actual device behavior more accurately.

**Runtime Recipes (RR)**

Many devices enabled for NETCONF, RESTCONF, or gNMI sometimes deviate in their runtime behavior. This can make it impossible to interact properly with NSO. These deviations can be on any level in the runtime behavior, such as:

* The configuration protocol is not properly implemented, i.e., the device lacks support for mandatory parts of, for instance, the RESTCONF RFC.
* The device returns "dirty" configuration dumps, for instance, JSON or XML containing invalid elements.
* Special quirks are required when applying new configuration on a device. May also require additional transforms of the payload before it is relayed by the NED.
* The device has aliasing issues, possibly caused by overlapping YANG models. If leaf X in model A is modified, the device will automatically modify leaf Y in model B as well. While this can be a cause of deviation, note that resolving aliasing issues through runtime recipes is generally avoided by NSO, as it is typically considered a modeling error.

A third-party YANG NED can be bundled with runtime recipes to solve these kinds of issues, if necessary. How this is implemented varies from NED to NED. In some cases, a NED has a fixed set of recipes that are always used. Alternatively, a NED can support several different recipes, which can be configured through a NED setting, referred to as a runtime profile. For example, a multi-vendor third-party YANG NED might have one runtime profile for each device type supported:

```bash
admin@ncs(config)# devices device dev-1 ned-settings
onf-tapi_rc restconf profile vendor-xyz
```

### NED Settings <a href="#d5e9013" id="d5e9013"></a>

NED settings are YANG models augmented as configurations in NSO and control the behavior of the NED. These settings are augmented under:

* `/devices/global-settings/ned-settings`
* `/devices/profiles/ned-settings`
* `/devices/device/ned-settings`

Most NEDs are instrumented with a large number of NED settings that can be used to customize the device instance configured in NSO. The README file in the respective NED contains more information on these.

## Purpose of NED ID <a href="#d5e9027" id="d5e9027"></a>

Each managed device in NSO has a device type that informs NSO how to communicate with the device. When managing NEDs, the device type is either `cli` or `generic`. The other two device types, `netconf` and `snmp`, are used in NETCONF and SNMP packages and are further described in this guide.

In addition, a special identifier, NED ID, is needed. Simply put, this identifier is a handle in NSO pointing to the NED package. NSO uses the identifier when it is about to invoke the driver in a NED package. The identifier ensures that the driver of the correct NED package is called for a given device instance. For more information on how to set up a new device instance, see [Configuring a device with the new Cisco-provided NED](#sec.config_device.with.ciscoid).

Each NED package has a NED ID, which is mandatory. The NED ID is a simple string that can have any format. For NEDs developed by the Cisco NSO NED team, the NED ID is formatted as `<NED NAME>-<gen | cli>-<NED VERSION MAJOR>.<NED VERSION MINOR>`.

**Examples**

* `onf-tapi_rc-gen-2.0`
* `cisco-iosxr-cli-7.43`

The NED ID for a certain NED package stays the same from one version to another, as long as no backward incompatible changes have been introduced to the YANG models. Upgrading a NED from one version to another, where the NED ID is the same, is simple as it only requires replacing the old NED package with the new one in NSO and then reloading all packages. For third-party (3PY) NEDs, such as the `onf-tapi_rc` NED, the situation differs slightly. Since the YANG models originate from external sources, the NED team does not control their evolution or guarantee backward compatibility between revisions. As a result, it is the responsibility of the end user to determine whether changes in the third-party YANG models are backward compatible and to choose an appropriate version and NED ID when rebuilding the NED. Unlike classic NEDs, upgrading a 3PY NED may therefore require more careful validation and potentially a change in NED ID to reflect incompatibilities.

Upgrading a NED package from one version to another, where the NED ID is not the same (typically indicated by a change of major or minor number in the NED version), requires additional steps. The new NED package first needs to be installed side-by-side with the old one. Then, a NED migration needs to be performed. This procedure is thoroughly described in [NED Migration](#sec.ned_migration).

### NED Versioning

As a best practice, a NED is assigned a version number consisting of a sequence of numbers separated by dots. The first two numbers represent the major and minor version, and the third number represents the maintenance version (semantic versioning).

For example, the number 1.2.3 indicates a maintenance release (3) for the minor release 1.2. This scheme updates either the major or minor version number when the YANG model changes significantly or incompatible changes are introduced. Meaning any version within the 1.2.x series is backward compatible with the previous versions.

The Cisco NSO NED team ensures that our CLI NEDs, as well as Generic NEDs with Cisco-owned models, have version numbers and NED ID that indicate any possible backward incompatible YANG model changes. When a NED with such an incompatible change is released, the minor digit in the version is always incremented.

The case is a bit different for our third-party YANG NEDs since it is up to the user to select the NED ID to be used when building the NED. This is further described in [Managing Cisco-provided third-Party YANG NEDs](#sec.managing_thirdparty_neds). In this case, it is possible to build two incompatible NEDs with the same NED ID if not taking sufficient care during the process and should be avoided.

The reason the version numbering is important is to identify a maintenance release. A backward compatible, maintenance release allows NSO to perform a simple data model upgrade to handle stored instance data in the CDB (Configuration Database). This type of upgrade does not pose a risk of data loss.

<div data-with-frame="true"><figure><img src="/files/djpICI7IWHWcWt9gWDQX" alt=""><figcaption><p>Recommended NED Version Scheme</p></figcaption></figure></div>

However, when the new NED is not just a maintenance upgrade (typically identified as a new major/minor release), it becomes a NED migration. These migrations are more complex because the YANG model changes can potentially result in the loss of instance data if not handled correctly. Additionally, services in NSO rely on NEDs to perform network provisioning. These services map service-specific configuration to the device data models, provided by the NEDs. As the NED packages can be upgraded independently, they can introduce changes in the device YANG models that cause issues for the services using them.

## Installing a NED in NSO <a href="#sec.ned_installation_nso" id="sec.ned_installation_nso"></a>

This section describes the NED installation in NSO for Local and System installs.

{% tabs %}
{% tab title="NED Installation on Local Install" %}
{% hint style="info" %}
This procedure below broadly outlines the steps needed to install a NED package on a [Local Install](/guides/administration/installation-and-deployment/local-install). For most up-to-date and specific installation instructions, consult the `README.md` supplied with the NED.
{% endhint %}

General instructions to install a NED package:

1. Download the latest production-grade version of the NED from [software.cisco.com](https://software.cisco.com) using the URLs provided on your NED license certificates. All NED packages are files with the `.signed.bin` extension named using the following rule: `ncs-<NSO VERSION>-<NED NAME>-<NED VERSION>.signed.bin`.
2. Place the NED package in the `/tmp/ned-package-store` directory and configure the environment variable `NSO_RUNDIR` to point to the NSO runtime directory.
3. Unpack the NED package and verify its signature. The result of the unpacking is a `tar.gz` file with the same name as the `.bin` file.
4. Untar the `.tar.gz` file. The result is a subdirectory named like `<NED NAME>-<NED MAJOR VERSION DIGIT>.<NED MINOR VERSION DIGIT>`.
5. Install the NED on NSO, using the `ncs-setup` tool.
6. Finally, open an NSO CLI session and load the new NED package.
   {% endtab %}

{% tab title="NED Installation on System Install" %}
{% hint style="info" %}
This procedure below broadly outlines the steps needed to install a NED package on a [System Install](/guides/administration/installation-and-deployment/system-install). For most up-to-date and specific installation instructions, consult the `README.md` supplied with the NED.
{% endhint %}

General instructions to install a NED package:

1. Download the latest production-grade version of the NED from [software.cisco.com](https://software.cisco.com) using the URLs provided on your NED license certificates. All NED packages are files with the `.signed.bin` extension named using the following rule: `ncs-<NSO VERSION>-<NED NAME>-<NED VERSION>.signed.bin`.
2. Place the NED package in the `/tmp/ned-package-store` directory.
3. Unpack the NED package and verify its signature. The result of the unpacking is a `.tar.gz` file with the same name as the `.bin` file.
4. Perform an NSO backup before installing the new NED package.
5. Start an NSO CLI session.
6. Fetch the NED package.
7. Install the NED package (add the argument `replace-existing` if a previous version has been loaded).
8. Finally, load the NED package.
   {% endtab %}
   {% endtabs %}

## Configuring a Device with an Installed NED <a href="#sec.config_device.with.ciscoid" id="sec.config_device.with.ciscoid"></a>

Once a NED has been installed in NSO, the next step is to create and configure device entries that use this NED. The basic steps for configuring a device instance using a newly installed NED package are described in this section. Only the most basic configuration steps are covered here. Many NEDs also require additional custom configuration to be operational. This applies in particular to Generic NEDs. Information about configuration and such additional configuration can be found in the files `README.md` and `README-ned-settings.md` bundled with the NED package.

The following info is necessary to proceed with the basic setup of a device instance in NSO:

* NED ID of the new NED.
* Connection information for the device to connect to (address and port).
* Authentication information to the device (username and password).

The general steps to configure a device with a NED are:

1. Start an NSO CLI session.
2. Enter the configuration mode.
3. Configure a new authentication group to be used for this device.
4. Configure the new device instance, such as its IP address, port, etc.
5. Check the `README.md` and `README-ned-settings.md` bundled with the NED package for further information on additional settings to make the NED fully operational.
6. Commit the configuration.

## Managing Cisco-provided Third Party YANG NEDs <a href="#sec.managing_thirdparty_neds" id="sec.managing_thirdparty_neds"></a>

The third-party YANG NED type is a special category of the generic NED type targeted for devices supporting protocols like NETCONF, RESTCONF, and gNMI. As the name implies, this NED category is used for cases where the device YANG models are not implemented or maintained by the Cisco NSO NED Team. Instead, the YANG models are typically provided by the device vendor itself or by organizations like IETF, IEEE, ONF, or OpenConfig.

A third-party YANG NED package is delivered from the software.cisco.com portal without any device YANG models included. It is required that the models are first downloaded, followed by a rebuild and reload of the package, before the NED can become fully operational. This task needs to be performed by the NED user.

Detailed NED-specific instructions to manage Cisco-provided third-party YANG NEDs are provided in the respective READMEs.

## NED Migration <a href="#sec.ned_migration" id="sec.ned_migration"></a>

If you upgrade a managed device (such as installing a new firmware), the device data model can change in a significant way. If this is the case, you usually need to use a different or newer NED with an updated YANG model.

When the changes in the NED are not backward compatible, the NED should be assigned a new ned-id to avoid breaking existing code. This allows you to use both versions of the NED at the same time, so some devices can use the new version and some can use the old one. As a result, there is no need to upgrade all devices at the same time. However, NSO doesn't know the two NEDs are related and will not perform any upgrade on its own due to different ned-ids and a NED migration is required.

{% hint style="info" %}
For third-party NEDs, the user is required to configure the ned-id and also be aware of the backward incompatibilities.
{% endhint %}

A potential issue with a new NED is that it can break an existing service (or other packages that rely on it) but NSO provides tools to migrate between backward incompatible NED versions. The tools are designed to give you a structured analysis of which paths will change between two NED versions and visibility into the scope of the potential impact that a change in the NED will drive in the service code. This includes identifying the paths and service instances that may be impacted.

Using the `/ncs:devices/device/migrate` action, you can change the NED of a device. The action migrates all configuration and service meta-data. The example [examples.ncs/device-management/ned-migration](https://github.com/NSO-developer/nso-examples/tree/6.7/device-management/ned-migration) in the NSO examples collection illustrates how to migrate devices between different NED versions using this action. The actual migration procedure consists of:

1. Requesting a maintenance window (if required) and creating an [NSO backup](/guides/administration/management/system-management#backup-and-restore).
2. [Adding](/guides/administration/management/package-mgmt) the new NED package and upgrading existing packages if needed.
3. Updating device templates if used.
4. Running the `migrate` action for affected devices.
5. Redeploying the service instances that reference migrated devices.
6. Optionally removing the old NED.

Some or all of the migration steps could be performed during normal operations and not strictly require a maintenance window, as they only read from and do not write to network devices. However, that depends on your specific operational policies.

### Migration Preparation

For a successful migration, prior preparation is required. In particular, **ensure that all the other packages are compatible with the new NED**. If you are a service developer, leverage the `migrate` action to report what paths have been modified and the services affected by those changes. This information can then be used to prepare the service code to handle the new NED version. Useful `migrate` options for reporting include:

* `dry-run` to report but not migrate yet
* `suppress-modified-paths without-instance-data` to ignore changes in the NED that do not affect your current device configurations
* `report { all }` to produce a list of services that are affected
* `no-networking` to only use the CDB copy of device configurations

For example, when a NED device model has renamed or restructured nodes, such as:

```
   grouping dns {
-    leaf domain {
+    leaf-list search {
       type inet:host;
     }
```

the migrate dry-run can report:

```bash
admin@ncs# devices device ex0 migrate new-ned-id router-nc-1.1 suppress-modified-paths without-instance-data report { all } dry-run
modified-path {
    path /r:sys/dns/domain
    modification {
        info leaf has been removed
        backward-compatible false
    }
    affected-service {
        id /acme-dns
    }
}
```

The service packages should be updated and tested to work with both versions of the NED. This may require updating [service XML templates](/guides/development/core-concepts/templates) or service mapping code to use the new/updated parts of the model.

Likewise, changing a ned-id also affects device templates if you use them. To reuse existing device templates with the new ned-id, you can use the `copy` template action. It will copy the configuration used for one ned-id to another, as long as the schema nodes used haven't changed between the versions. The following example demonstrates the `copy` action usage:

```bash
admin@ncs(config)# devices template acme-ntp ned-id router-nc-1.0
copy ned-id router-nc-1.2
```

However, if the device templates reference e.g. renamed schema nodes, you should prepare new templates beforehand and load them after the new NED package is added to the NSO.

### Migrate Action

Two versions of `migrate` action are available. For individual devices, use the `/devices/device/migrate` action, with the `new-ned-id` parameter. Without additional options, the command will read and update the device configuration in NSO. As part of this process, NSO migrates all the configuration and service meta-data. Use the `dry-run` option to see what the command would do without performing any changes.

You may use the `no-networking` option to prevent NSO from generating any southbound traffic towards the device. In this case, only the device configuration in the CDB is used for the migration but then NSO *cannot* know if the device is in sync. Afterward, you must use the **compare-config** or the **sync-from** action to remedy this.

For migrating multiple devices, use the `/devices/migrate` action, which takes the same options and executes the migration in parallel. However, with this action, you must also specify the `old-ned-id`, which limits the migration to devices using the old NED. You can further restrict the action with the `device` or `device-group` parameter, selecting only specific devices.

{% hint style="info" %}
In case the expected ned-id cannot be selected for `new-ned-id` or `old-ned-id`, use the `show packages` command to verify that both NEDs, the new and the old, are present.
{% endhint %}

It is possible for a NED migration to fail if the device has an active configuration that is incompatible with the new NED version. In such cases, NSO will produce an error with the YANG constraint that is not satisfied. Here, you must first manually adjust the device configuration to make it compatible with the new NED, and then you can perform the migration as usual.

### Redeploy Services Post Migration

After successful device migration, preform a `re-deploy` of all the affected services before removing the old NED package. This step ensures all the old NED references are removed and allows for a smooth future NED upgrade. If you skip re-deploying, the service `get-modifications` output, `deep-check-sync`, and similar operations may no longer work correctly.

It is recommended you start with a service `re-deploy dry-run` to verify the produced configurations are as expected.

The migrate action will output the list of affected services when given `report { all }` option, which you can use, for example, with the [bulk service actions](/guides/operation-and-usage/operations/managing-network-services#bulk-service-actions).

{% hint style="info" %}
If you are using `packages add` command instead of `packages reload` to add the new NED package, and the mapping for service packages had to be updated, you will also have to re-deploy the affected packages for NSO to pick up the new configuration.
{% endhint %}

In case the migrate action has already run, you can still list all the services touching a given device via `/devices/device/services/service`, helping you to identify services that require re-deploy. For example:

```bash
admin@ncs# show devices device ex0 services service
services service /acme-dns
```

## Migrating from Legacy to Third-party NED

{% hint style="info" %}
This section uses `juniper-junos_nc` as an example third-party NED. The process is generally same and applicable to other third-party NEDs.
{% endhint %}

NSO has supported Junos devices from early on. The legacy Junos NED is NETCONF-based, but as Junos devices did not provide YANG modules in the past, complex NSO machinery translated Juniper's XML Schema Description (XSD) files into a single YANG module. This was an attempt to aggregate several Juniper device modules/versions.

Juniper nowadays provides YANG modules for Junos devices. Junos YANG modules can be downloaded from the device and used directly in NSO with the new `juniper-junos_nc` NED.

By downloading the YANG modules using `juniper-junos_nc` NED tools and rebuilding the NED, the NED can provide full coverage immediately when the device is updated instead of waiting for a new legacy NED release.

This guide describes how to replace the legacy `juniper-junos` NED and migrate NSO applications to the `juniper-junos_nc` NED using the NSO MPLS VPN example from the NSO examples collection as a reference.

Prepare the example:

1. Add the `juniper-junos` and `juniper-junos_nc` NED packages to the example.
2. Configure the connection to the Junos device.
3. Add the MPLS VPN service configuration to the simulated network, including the Junos device using the legacy `juniper-junos` NED.

Adapting the service to the `juniper-junos_nc` NED:

1. Un-deploy MPLS VPN service instances with `no-networking`.
2. Delete Junos device config with `no-networking`.
3. Set the Junos device to NETCONF/YANG compliant mode.
4. Download the compliant YANG models, build, and reload the `juniper-junos_nc` NED package.
5. Switch the ned-id for the Junos device to the `juniper-junos_nc` NED package.
6. Sync from the Junos device to get the compliant Junos device config.
7. Update the MPLS VPN service to handle the difference between the non-compliant and compliant configurations belonging to the service.
8. Re-deploy the MPLS VPN service instances with `no-networking` to make the MPLS VPN service instances own the device configuration again.

{% hint style="info" %}
If applying the steps for this example on a production system, you should first take a backup using the `ncs-backup` tool before proceeding.
{% endhint %}

### Prepare the Example <a href="#d5e10954" id="d5e10954"></a>

This guide uses the MPLS VPN example in Python from the NSO example set under [examples.ncs/service-management/mpls-vpn-python](https://github.com/NSO-developer/nso-examples/tree/6.7/service-management/mpls-vpn-python) to demonstrate porting an existing application to use the `juniper-junos_nc` NED. The simulated Junos device is replaced with a Junos vMX 21.1R1.11 container, but other NETCONF/YANG-compliant Junos versions also work.

### **Add the `juniper-junos` and `juniper-junos_nc` NED Packages**

The first step is to add the latest `juniper-junos` and `juniper-junos_nc` NED packages to the example's package directory. The NED tar-balls must be available and downloaded from your <https://software.cisco.com/download/home> account to the `mpls-vpn-python` example directory. Replace the `NSO_VERSION` and `NED_VERSION` variables with the versions you use:

```bash
$ cd $NCS_DIR/examples.ncs/service-management/mpls-vpn-python
$ cp ./ncs-NSO_VERSION-juniper-junos-NED_VERSION.tar.gz packages/
$ cd packages
$ tar xfz ../ncs-NSO_VERSION-juniper-junos_nc-NED_VERSION.tar.gz
$ cd -
```

Build and start the example:

```bash
$ make all start
```

### **Configure the Connection to the Junos Device**

Replace the netsim device connection configuration in NSO with the configuration for connecting to the Junos device. Adjust the `USER_NAME`, `PASSWORD`, and `HOST_NAME/IP_ADDR` variables and the timeouts as required for the Junos device you are using with this example:

```bash
$ ncs_cli -u admin -C
admin@ncs# config
admin@ncs(config)# devices authgroups group juniper umap admin remote-name USER_NAME \
                   remote-password PASSWORD
admin@ncs(config)# devices device pe2 authgroup juniper address HOST_NAME/IP_ADDR port 830
admin@ncs(config)# devices device pe2 connect-timeout 240
admin@ncs(config)# devices device pe2 read-timeout 240
admin@ncs(config)# devices device pe2 write-timeout 240
admin@ncs(config)# commit
admin@ncs(config)# end
admin@ncs# exit
```

Open a CLI terminal or use NETCONF on the Junos device to verify that the `rfc-compliant` and `yang-compliant` modes are not yet enabled. Examples:

```bash
$ ssh USER_NAME@HOST_NAME/IP_ADDR
junos> configure
junos# show system services netconf
ssh;
```

Or:

```bash
$ netconf-console -s plain -u USER_NAME -p PASSWORD --host=HOST_NAME/IP_ADDR \
 --port=830 --get-config
 --subtree-filter=-<<<'<configuration xmlns="http://xml.juniper.net/xnm/1.1/xnm">
                        <system>
                          <services>
                            <netconf/>
                          </services>
                        </system>
                      </configuration>'

<rpc-reply xmlns:junos="http://xml.juniper.net/junos/21.1R0/junos"
           xmlns="urn:ietf:params:xml:ns:netconf:base:1.0" message-id="1">
  <data>
    <configuration xmlns="http://xml.juniper.net/xnm/1.1/xnm">
      <system>
        <services>
          <netconf>
            <ssh>
            </ssh>
          </netconf>
        </services>
      </system>
    </configuration>
  </data>
</rpc-reply>
```

The `rfc-compliant` and `yang-compliant` nodes must not be enabled yet for the legacy Junos NED to work. If enabled, delete in the Junos CLI or using NETCONF. A netconf-console example:

```bash
$ netconf-console -s plain -u USER_NAME -p PASSWORD --host=HOST_NAME/IP_ADDR --port=830
  --db=candidate
  --edit-config=- <<<'<configuration xmlns="http://xml.juniper.net/xnm/1.1/xnm"
                                     xmlns:nc="urn:ietf:params:xml:ns:netconf:base:1.0">
                        <system>
                          <services>
                            <netconf>
                              <rfc-compliant nc:operation="remove"/>
                              <yang-compliant nc:operation="remove"/>
                            </netconf>
                          </services>
                        </system>
                      </configuration>'

$ netconf-console -s plain -u USER_NAME -p PASSWORD --host=HOST_NAME/IP_ADDR \
                  --port=830 --commit
```

Back to the NSO CLI to upgrade the legacy `juniper-junos` NED to the latest version:

```bash
$ ncs_cli -u admin -C
admin@ncs# config
admin@ncs(config)# devices device pe2 ssh fetch-host-keys
admin@ncs(config)# devices device pe2 migrate new-ned-id juniper-junos-nc-NED_VERSION
admin@ncs(config)# devices sync-from
admin@ncs(config)# end
```

### **Add the MPLS VPN Service Configuration to the Simulated Network**

Turn off `autowizard` and `complete-on-space` to make it possible to paste configs:

```cli
admin@ncs# autowizard false
admin@ncs# complete-on-space false
```

The example service config for two MPLS VPNs where the endpoints have been selected to pass through the `PE` node `PE2`, which is a Junos device:

```
vpn l3vpn ikea
as-number 65101
endpoint branch-office1
  ce-device    ce1
  ce-interface GigabitEthernet0/11
  ip-network   10.7.7.0/24
  bandwidth    6000000
!
endpoint branch-office2
  ce-device    ce4
  ce-interface GigabitEthernet0/18
  ip-network   10.8.8.0/24
  bandwidth    300000
!
endpoint main-office
  ce-device    ce0
  ce-interface GigabitEthernet0/11
  ip-network   10.10.1.0/24
  bandwidth    12000000
!
qos qos-policy GOLD
!
vpn l3vpn spotify
as-number 65202
endpoint branch-office1
  ce-device    ce5
  ce-interface GigabitEthernet0/1
  ip-network   10.2.3.0/24
  bandwidth    10000000
!
endpoint branch-office2
  ce-device    ce3
  ce-interface GigabitEthernet0/4
  ip-network   10.4.5.0/24
  bandwidth    20000000
!
endpoint main-office
  ce-device    ce2
  ce-interface GigabitEthernet0/8
  ip-network   10.0.1.0/24
  bandwidth    40000000
!
qos qos-policy GOLD
!
```

To verify that the traffic passes through `PE2`:

```cli
admin@ncs(config)# commit dry-run outformat native
```

Toward the end of this lengthy output, observe that some config changes are going to the `PE2` device using the `http://xml.juniper.net/xnm/1.1/xnm` legacy namespace:

```
device {
    name pe2
    data <rpc xmlns="urn:ietf:params:xml:ns:netconf:base:1.0" message-id="1">
          <edit-config xmlns:nc="urn:ietf:params:xml:ns:netconf:base:1.0">
            <target>
              <candidate/>
            </target>
            <test-option>test-then-set</test-option>
            <error-option>rollback-on-error</error-option>
            <with-inactive xmlns="http://tail-f.com/ns/netconf/inactive/1.0"/>
            <config>
              <configuration xmlns="http://xml.juniper.net/xnm/1.1/xnm">
                <interfaces>
                  <interface>
                    <name>xe-0/0/2</name>
                    <unit>
                      <name>102</name>
                      <description>Link to CE / ce5 - GigabitEthernet0/1</description>
                      <family>
                        <inet>
                          <address>
                            <name>192.168.1.22/30</name>
                          </address>
                        </inet>
                      </family>
                      <vlan-id>102</vlan-id>
                    </unit>
                  </interface>
                </interfaces>
      ...
```

Looks good. Commit to the network:

```cli
admin@ncs(config)# commit
```

### Adapting the Service to the `juniper-junos_nc` NED <a href="#d5e11047" id="d5e11047"></a>

Now that the service's configuration is in place using the legacy `juniper-junos` NED to configure the `PE2` Junos device, proceed and switch to using the `juniper-junos_nc` NED with `PE2` instead. The service template and Python code will need a few adaptations.

### **Un-deploy MPLS VPN Services Instances with `no-networking`**

To keep the NSO service meta-data information intact when bringing up the service with the new `juniper-junos_nc` NED, first `un-deploy` the service instances in NSO, only keeping the configuration on the devices:

```cli
admin@ncs(config)# vpn l3vpn * un-deploy no-networking
```

### **Delete Junos Device Config with `no-networking`**

First, save the legacy Junos non-compliant mode device configuration to later diff against the compliant mode config:

```cli
admin@ncs(config)# show full-configuration devices device pe2 config \
                                   configuration | display xml | save legacy.xml
```

Delete the `PE2` configuration in NSO to prepare for retrieving it from the device in a NETCONF/YANG compliant format using the new NED:

```cli
admin@ncs(config)# no devices device pe2 config
admin@ncs(config)# commit no-networking
admin@ncs(config)# end
admin@ncs# exit
```

### **Set the Junos Device to NETCONF/YANG Compliant Mode**

Using the Junos CLI:

```bash
$ ssh USER_NAME@HOST_NAME/IP_ADDR
junos> configure
junos# set system services netconf rfc-compliant
junos# set system services netconf yang-compliant
junos# show system services netconf
ssh;
rfc-compliant;
ÿang-compliant;
junos# commit
```

Or, using the NSO `netconf-console` tool:

```bash
$ netconf-console -s plain -u USER_NAME -p PASSWORD --host=HOST_NAME/IP_ADDR --port=830 \
  --db=candidate
  --edit-config=- <<<'<configuration xmlns="http://xml.juniper.net/xnm/1.1/xnm">
                        <system>
                          <services>
                            <netconf>
                              <rfc-compliant/>
                              <yang-compliant/>
                            </netconf>
                          </services>
                        </system>
                      </configuration>'

$ netconf-console -s plain -u USER_NAME -p PASSWORD --host=HOST_NAME/IP_ADDR --port=830 \
                  --commit
```

### **Switch the NED ID for the Junos Device to the `juniper-junos_nc` NED Package**

```bash
$ ncs_cli -u admin -C
admin@ncs# config
admin@ncs(config)# devices device pe2 device-type generic ned-id juniper-junos_nc-gen-1.0
admin@ncs(config)# commit
admin@ncs(config)# end
```

### **Download the Compliant YANG models, Build, and Load the `juniper-junos_nc` NED Package**

The `juniper-junos_nc` NED is delivered without YANG modules, enabling populating it with device-specific YANG modules. The YANG modules are retrieved directly from the Junos device:

```bash
$ ncs_cli -u admin -C
admin@ncs# devices device pe2 connect
admin@ncs# devices device pe2 rpc rpc-get-modules get-modules
admin@ncs# exit
```

See the `juniper-junos_nc` `README` for more options and details.

Build the YANG modules retrieved from the Junos device with the `juniper-junos_nc` NED:

```bash
$ make -C packages/juniper-junos_nc-gen-1.0/src
```

Reload the packages to load the `juniper-junos_nc` NED with the added YANG modules:

```bash
$ ncs_cli -u admin -C
admin@ncs# packages reload
```

### **Sync From the Junos Device to Get the Device Configuration in NETCONF/YANG Compliant Format**

```cli
admin@ncs# devices device pe2 sync-from
```

### **Update the MPLS VPN Service**

The service must be updated to handle the difference between the Junos device's non-compliant and compliant configuration. The NSO service uses Python code to configure the Junos device using a service template. One way to find the required updates to the template and code is to check the difference between the non-compliant and compliant configurations for the parts covered by the template.

<div data-with-frame="true"><figure><img src="/files/v8Ot33aRF04TjlBFGgnB" alt=""><figcaption><p>Side by Side, Running Config on the Left, Template on the Right.</p></figcaption></figure></div>

Checking the `packages/l3vpn/templates/l3vpn-pe.xml` service template Junos device part under the legacy `http://xml.juniper.net/xnm/1.1/xnm` namespace, you can observe that it configures `interfaces`, `routing-instances`, `policy-options`, and `class-of-service`.

You can save the NETCONF/YANG compliant Junos device configuration and diff it against the non-compliant configuration from the previously stored `legacy.xml` file:

```cli
admin@ncs# show running-config devices device pe2 config configuration \
                          | display xml | save new.xml
```

Examining the difference between the configuration in the `legacy.xml` and `new.xml` files for the parts covered by the service template:

1. There is no longer a single namespace covering all configurations. The configuration is now divided into multiple YANG modules with a namespace for each.
2. The `/configuration/policy-options/policy-statement/then/community` node choice identity is no longer provided with a leaf named `key1`. Instead, the leaf name is `choice-ident`, and a `choice-value` leaf is set.
3. The `/configuration/class-of-service/interfaces/interface/unit/shaping-rate/rate` leaf format has changed from using an `int32` value to a string with either no suffix or a "k", "m" or "g" suffix. This differs from the other devices controlled by the template, so a new template `BW_SUFFIX` variable set from the Python code is needed.

To enable the template to handle a Junos device in NETCONF/YANG compliant mode, add the following to the `packages/l3vpn/templates/l3vpn-pe.xml` service template:

```xml
            </interfaces>
          </class-of-service>
        </configuration>
+
+        <configuration xmlns="http://yang.juniper.net/junos/conf/root" tags="merge">
+          <interfaces xmlns="http://yang.juniper.net/junos/conf/interfaces">
+            <interface>
+              <name>{$PE_INT_NAME}</name>
+              <no-traps/>
+              <vlan-tagging/>
+              <per-unit-scheduler/>
+              <unit>
+                <name>{$VLAN_ID}</name>
+                <description>Link to CE / {$CE} - {$CE_INT_NAME}</description>
+                <vlan-id>{$VLAN_ID}</vlan-id>
+                <family>
+                  <inet>
+                    <address>
+                      <name>{$LINK_PE_ADR}/{$LINK_PREFIX}</name>
+                    </address>
+                  </inet>
+                </family>
+              </unit>
+            </interface>
+          </interfaces>
+          <routing-instances xmlns="http://yang.juniper.net/junos/conf/routing-instances">
+            <instance>
+              <name>{/name}</name>
+              <instance-type>vrf</instance-type>
+              <interface>
+                <name>{$PE_INT_NAME}.{$VLAN_ID}</name>
+              </interface>
+              <route-distinguisher>
+                <rd-type>{/as-number}:1</rd-type>
+              </route-distinguisher>
+              <vrf-import>{/name}-IMP</vrf-import>
+              <vrf-export>{/name}-EXP</vrf-export>
+              <vrf-table-label>
+              </vrf-table-label>
+              <protocols>
+                <bgp>
+                  <group>
+                    <name>{/name}</name>
+                    <local-address>{$LINK_PE_ADR}</local-address>
+                    <peer-as>{/as-number}</peer-as>
+                    <local-as>
+                      <as-number>100</as-number>
+                    </local-as>
+                    <neighbor>
+                      <name>{$LINK_CE_ADR}</name>
+                    </neighbor>
+                  </group>
+                </bgp>
+              </protocols>
+            </instance>
+          </routing-instances>
+          <policy-options xmlns="http://yang.juniper.net/junos/conf/policy-options">
+            <policy-statement>
+              <name>{/name}-EXP</name>
+              <from>
+                <protocol>bgp</protocol>
+              </from>
+              <then>
+                <community>
+                  <choice-ident>add</choice-ident>
+                  <choice-value/>
+                  <community-name>{/name}-comm-exp</community-name>
+                </community>
+                <accept/>
+              </then>
+            </policy-statement>
+            <policy-statement>
+              <name>{/name}-IMP</name>
+              <from>
+                <protocol>bgp</protocol>
+                <community>{/name}-comm-imp</community>
+              </from>
+              <then>
+                <accept/>
+              </then>
+            </policy-statement>
+            <community>
+              <name>{/name}-comm-imp</name>
+              <members>target:{/as-number}:1</members>
+            </community>
+            <community>
+              <name>{/name}-comm-exp</name>
+              <members>target:{/as-number}:1</members>
+            </community>
+          </policy-options>
+          <class-of-service xmlns="http://yang.juniper.net/junos/conf/class-of-service">
+            <interfaces>
+              <interface>
+                <name>{$PE_INT_NAME}</name>
+                <unit>
+                  <name>{$VLAN_ID}</name>
+                  <shaping-rate>
+                    <rate>{$BW_SUFFIX}</rate>
+                  </shaping-rate>
+                </unit>
+              </interface>
+            </interfaces>
+          </class-of-service>
+        </configuration>
      </config>
    </device>
  </devices>
```

The Python file changes to handle the new `BW_SUFFIX` variable to generate a string with a suffix instead of an `int32`:

```bash
# of the service. These functions can be useful e.g. for
# allocations that should be stored and existing also when the
# service instance is removed.
+
+    @staticmethod
+    def int32_to_numeric_suffix_str(val):
+        for suffix in ["", "k", "m", "g", ""]:
+            suffix_val = int(val / 1000)
+            if suffix_val * 1000 != val:
+                return str(val) + suffix
+            val = suffix_val
+
@ncs.application.Service.create
def cb_create(self, tctx, root, service, proplist):
    # The create() callback is invoked inside NCS FASTMAP and must
```

Code that uses the function and set the string to the service template:

```
            tv.add('LOCAL_CE_NET', getIpAddress(endpoint.ip_network))
            tv.add('CE_MASK', getNetMask(endpoint.ip_network))
+            tv.add('BW_SUFFIX', self.int32_to_numeric_suffix_str(endpoint.bandwidth))
            tv.add('BW', endpoint.bandwidth)
            tmpl = ncs.template.Template(service)
            tmpl.apply('l3vpn-pe', tv)
```

After making the changes to the service template and Python code, reload the updated package(s):

```bash
$ ncs_cli -u admin -C
admin@ncs# packages reload
```

### **Re-deploy the MPLS VPN Service Instances**

The service instances need to be re-deployed to own the device configuration again:

```cli
admin@ncs# vpn l3vpn * re-deploy no-networking
```

The service is now in sync with the device configuration stored in NSO CDB:

```cli
admin@ncs# vpn l3vpn * check-sync
vpn l3vpn ikea check-sync
in-sync true
vpn l3vpn spotify check-sync
in-sync true
```

When re-deploying the service instances, any issues with the added service template section for the compliant Junos device configuration, such as the added namespaces and nodes, are discovered.

As there is no validation for the rate leaf string with a suffix in the Junos device model, no errors are discovered if it is provided in the wrong format until updating the Junos device. Comparing the device configuration in NSO with the configuration on the device shows such inconsistencies without having to test the configuration with the device:

```cli
admin@ncs# devices device pe2 compare-config
```

If there are issues, correct them and redo the `re-deploy no-networking` for the service instances.

When all issues have been resolved, the service configuration is in sync with the device configuration, and the NSO CDB device configuration matches to the configuration on the Junos device:

```bash
$ ncs_cli -u admin -C
admin@ncs# vpn l3vpn * re-deploy
```

The NSO service instances are now in sync with the configuration on the Junos device using the `juniper-junos_nc` NED.

## Revision Merge Functionality <a href="#d5e9642" id="d5e9642"></a>

The YANG modeling language supports the notion of a module `revision`. It allows users to distinguish between different versions of a module, so the module can evolve over time. If you wish to use a new revision of a module for a managed device, for example, to access new features, you generally need to create a new NED.

When a model evolves quickly and you have many devices that require the use of a lot of different revisions, you will need to maintain a high number of NEDs, which are mostly the same. This can become especially burdensome during NSO version upgrades, when all NEDs may need to be recompiled.

When a YANG module is only updated in a backward-compatible way (following the upgrade rules in RFC6020 or RFC7950), the NSO compiler, `ncsc`, allows you to pack multiple module revisions into the same package. This way, a single NED with multiple device model revisions can be used, instead of multiple NEDs. Based on the capabilities exchange, NSO will then use the correct revision for communication with each device.

However, there is a major downside to this approach. While the exact revision is known for each communication session with the managed device, the device model in NSO does not have that information. For that reason, the device model always uses the latest revision. When pushing configuration to a device that only supports an older revision, NSO silently drops the unsupported parts. This may have surprising results, as the NSO copy can contain configuration that is not really supported on the device. Use the `no-revision-drop` commit parameter when you want to make sure you are not committing config that is not supported by a device. See [Commit Parameters](/guides/operation-and-usage/operations/lifecycle-operations#d5e5048) for the shared model and interface mappings.

If you still wish to use this functionality, you can create a NED package with the `ncs-make-package --netconf-ned` command as you would otherwise. However, the supplied source YANG directory should contain YANG modules with different revisions. The files should follow the *`module-or-submodule-name`*`@`*`revision-date`*`.yang` naming convention, as specified in the RFC6020. Some versions of the compiler require you to use the `--no-fail-on-warnings` option with the `ncs-make-package` command or the build process may fail.

The [examples.ncs/device-management/ned-yang-revision](https://github.com/NSO-developer/nso-examples/tree/6.7/device-management/ned-yang-revision) example shows how you can perform a YANG model upgrade. The original, 1.0 version of the router NED uses the `router@2020-02-27.yang` YANG model. First, it is updated to the version 1.0.1 `router@2020-09-18.yang` using a revision merge approach. This is possible because the changes are backward-compatible.

In the second part of the example, the updates in `router@2022-01-25.yang` introduce breaking changes, therefore the version is increased to 1.1 and a different NED-ID is assigned to the NED. In this case, you can't use revision merge and the usual NED migration procedure is required.


# Advanced Topics

Deep-dive into advanced NSO concepts.


# Locks

Learn about different transaction locks in NSO and their interactions.

This section explains the different locks that exist in NSO and how they interact. It is important to understand the architecture of NSO with its management backplane and the transaction state machine as described in [Package Development](/guides/development/advanced-development/developing-packages) to be able to understand how the different locks fit into the picture.

## Global Locks

The NSO management backplane keeps a lock on the datastore running. This lock is usually referred to as the global lock, and it provides a mechanism to grant exclusive access to the datastore.

The global is the only lock that can explicitly be taken through a northbound agent, for example, by the NETCONF `<lock>` operation, or by calling `Maapi.lock()`.

A global lock can be taken for the whole datastore, or it can be a partial lock (for a subset of the data model). Partial locks are exposed through NETCONF and MAAPI and are only supported for operations toward the running datastore.

An agent can request a global lock to ensure that it has exclusive write access. When a global lock is held by an agent, it is not possible for anyone else to write to the datastore that the lock guards—this is enforced by the transaction engine. A global lock on running is granted to an agent if there are no other holders of it (including partial locks) and if all data providers approve the lock request. Each data provider (CDB and/or external data providers) will have its `lock()` callback invoked to get a chance to refuse or accept the lock. The output of `ncs --status` includes the locking status. For each user session, locks (if any) per datastore, is listed.

## Transaction Locks <a href="#d5e4207" id="d5e4207"></a>

A northbound agent starts a user session towards NSO's management backplane. Each user session can then start multiple transactions. A transaction is either read/write or read-only.

The transaction engine has its internal locks towards the running datastore. These transaction locks exist to serialize configuration updates towards the datastore and are separate from the global locks.

As a northbound agent wants to update the running datastore with a new configuration, it will implicitly grab and release the transactional lock. The transaction engine takes care of managing the locks as it moves through the transaction state machine, and there is no API that exposes the transactional locks to the northbound agents.

When the transaction engine wants to take a lock for a transaction (for example, when entering the validate state), it first checks that no other transaction has the lock. Then it checks that no user session has a global lock on that datastore. Finally, each data provider is invoked by its `transLock()` callback.

## Northbound Agents and Global Locks <a href="#d5e4214" id="d5e4214"></a>

In contrast to the implicit transactional locks, some northbound agents expose explicit access to the global locks. This is done a bit differently by each agent.

The management API exposes the global locks by providing `Maapi.lock()` and `Maapi.unlock()` methods (and the corresponding `Maapi.lockPartial()` `Maapi.unlockPartial()` for partial locking). Once a user session is established (or attached to), these functions can be called.

In the CLI, the global locks are taken when entering different configure modes as follows:

* `config exclusive`: The running datastore global lock will be taken.
* `config terminal`: Does not grab any locks.

The global lock is then kept by the CLI until the configure mode is exited.

The Web UI behaves in the same way as the CLI (it presents three edit tabs called **Edit private**, **Edit exclusive**, and which correspond to the CLI modes described above).

The NETCONF agent translates the `<lock>` operation into a request for the global lock for the requested datastore. Partial locks are also exposed through the partial-lock RPC.

## External Data Providers <a href="#d5e4238" id="d5e4238"></a>

Implementing the `lock()` and `unlock()` callbacks is not required of an external data provider. NSO will never try to initiate the `transLock()` state transition (see the transaction state diagram in [Package Development](/guides/development/advanced-development/developing-packages)) towards a data provider while a global lock is taken—so the reason for a data provider to implement the locking callbacks is if someone else can write (or lock, for example, to take a backup) to the data provider's database.

## CDB and Locks <a href="#d5e4245" id="d5e4245"></a>

CDB ignores the `lock()` and `unlock()` callbacks (since the data-provider interface is the only write interface towards it).

CDB has its own internal locks on the database. The running datastore has a single write and multiple read locks. It is not possible to grab the write lock on a datastore while there are active read locks on it. The locks in CDB exist to make sure that a reader always gets a consistent view of the data (in particular, it becomes very confusing if another user is able to delete configuration nodes in between calls to `getNext()` on YANG list entries).

During a transaction, `transLock()` takes a CDB read lock towards the transaction's datastore, and `writeStart()` tries to release the read lock and grab the write lock instead.

A CDB external reader client implicitly takes a CDB read lock between `Cdb.startSession()` and `Cdb.endSession()` This means that while a CDB client is reading, a transaction can not pass through `writeStart()` (and conversely, a CDB reader can not start while a transaction is in between `writeStart()` and `commit()` or `abort()`).

The operational store in CDB does not have any locks. NSO's transaction engine can only read from it, and the CDB client writes are atomic per write operation.

## Lock Impact on User Sessions <a href="#d5e4261" id="d5e4261"></a>

When a session tries to modify a data store that is locked in some way, it will fail. For example, the CLI might print:

```bash
admin@ncs(config)# commit
Aborted: the configuration database is locked
```

Since some of the locks are short-lived (such as a CDB read-lock), NSO is by default configured to retry the failing operation for a short period of time. If the data store still is locked after this time, the operation fails.

To configure this, set `/ncs-config/commit-retry-timeout` in `ncs.conf`.


# CDB Persistence

Select the optimal CDB persistence mode for your use case.

The Configuration Database (CDB) is a built-in datastore for NSO, specifically designed for network automation use cases and backed by the YANG schema. Since NSO 6.4, the CDB can be configured to operate in one of the two distinct modes: `in-memory-v1` and `on-demand-v1`. From NSO 6.7, the default mode for CDB persistence is `on-demand-v1` and the previous default mode `in-memory-v1` has been deprecated. When upgrading to NSO 6.7, if the `/ncs-config/cdb/persistence/format` leaf is not specified in the `ncs.conf` file, an automatic conversion of the CDB persistence layer occurs to transition to the `on-demand-v1` mode.

The `in-memory-v1` mode keeps all the configuration data in RAM for the fastest access time. New data is persisted to disk in the form of journal (WAL) files, which the system uses on every restart to reconstruct the RAM database. But the amount of RAM needed is proportional to the number of managed devices and services. When NSO is used to manage a large network, the amount of needed RAM can be quite large. This is the only CDB persistence mode available before NSO 6.4.

The `on-demand-v1` mode loads data on demand from the disk into the RAM and supports offloading the least-used data to free up memory. Loading only the compiled YANG schema initially (in the form of .fxs files) results in faster system startup times. This mode was first introduced in NSO 6.4.

{% hint style="warning" %}
For reliable storage of the configuration on disk, regardless of the persistence mode, the CDB requires that the file system correctly implements the standard primitives for file synchronization and truncation. For this reason (as well as for performance), NFS or other network file systems are unsuitable for use with the CDB - they may be acceptable for development, but using them in production is unsupported and strongly discouraged.
{% endhint %}

Compared to `in-memory-v1`, `on-demand-v1` mode has a number of benefits:

* **Faster startup time**: Data is not loaded into memory at startup; only the schema is.
* **Lower memory requirements**: Data is loaded into memory only when needed and offloaded when not.
* **Faster sync of high-availability nodes**: Only subscribed data on the followers is loaded at once.
* **Background compaction**: The compaction process no longer locks the CDB, allowing writes to proceed uninterrupted.

While the `on-demand-v1` mode is as fast for reads of "hot" data (already in memory) as the `in-memory-v1` mode, reads are slower for "cold" data (not loaded in memory), since the data first has to be read from disk. In turn, this results in a bigger variance in the time that a read takes in the `on-demand-v1` mode, based on whether the data is already available in RAM or not. The variance could express in different ways, for example, by taking a longer time to produce the service mapping or creating a rollback for the first request. To lessen the effect, we highly recommend fast storage, such as NVMe flash drives.

Furthermore, the two modes differ in the way they internally organize and store data, resulting in different performance characteristics. If sufficient RAM is available, in some cases, `in-memory-v1` performs better, while in others, `on-demand-v1` performs better. One known case where the performance of `on-demand-v1` does not reach that of `in-memory-v1` is deleting large trees of data. But in general, only extensive testing of the specific use case can tell which mode performs better.

As a rule of thumb, we recommend the `on-demand-v1` mode, as it has typical performance comparable to `in-memory-v1` but has better maintainability properties. However, if performance requirements and testing favor the `in-memory-v1` mode, that may be a viable choice. Discounting the migration time, you can easily switch between the two modes with automatic migration at system startup.

## Configuring Persistence Mode

The CDB persistence is configured under `/ncs-config/cdb/persistence` in the `ncs.conf` file. The `format` leaf selects the desired persistence mode, either `on-demand-v1` or `in-memory-v1` (default `on-demand-v1` from NSO 6.7), and the system automatically migrates the data on the next start, if needed. Note that the system will not be available for the migration duration.

{% hint style="info" %}
Before switching persistence mode, ensure CDB is compacted first. You can start the compaction process manually by stopping the ncs daemon and running the `ncs --cdb-compact` command.
{% endhint %}

With the `on-demand-v1` mode, additional offloading configuration under `offload` container becomes relevant (`in-memory-v1` keeps all data in RAM and does not perform any offloading). The `offload/interval` specifies how often the system checks its memory consumption and starts the offload process if required.

During the offloading process, data is evicted from memory:

1. If the piece of data was last accessed more than `offload/threshold/max-age` ago (the default value of infinity disables this check).
2. The least-recently-used items are evicted until their usage drops below the allowed amount.

The allowed amount is defined either by the absolute value `offload/threshold/megabytes` or by `offload/threshold/system-memory-percentage`, where the value is calculated dynamically based on the available system RAM. We recommend using the latter unless testing has shown specific requirements.

The actual value should be adjusted according to the use case and system requirements; there is no single optimal setting for all cases. We recommend you start with defaults and then adjust according to observations. You can enable the new `/ncs-config/cdb/persistence/db-statistics` property to aid you in this task (producing `LOG` files inside the CDB directory), as well as the counters and gauges that are available under `/ncs:metric/sysadmin/*/cdb`.

## Compaction

For durability, improved performance, and snapshot isolation, CDB writes in NSO use data structures, such as a write-ahead log (WAL), that require periodic compaction.

For example, the `in-memory-v1` persistence mode appends a new log entry for each CDB transaction to the target datastore WAL file (`A.cdb` for configuration, `O.cdb` for operational, and `S.cdb` for snapshot datastore). Depending on the size and number of transactions towards the system, these files will grow in size leading to increased disk utilization, longer boot times, and longer initial data synchronization time when setting up a high-availability cluster using this persistence mode.

Compaction is a mechanism used to reduce the size of the write-ahead logs to a minimum. In `on-demand-v1` mode, it is automatic, non-configurable, and runs in the background without affecting the ongoing transactions.

But in `in-memory-v1` mode, it works by replacing an existing write-ahead log, which is composed of a number of consecutive transaction logs created in run-time, with a single transaction log representing the full current state of the datastore. From this perspective, a compaction acts similarly to a write transaction towards a datastore. To ensure data integrity, 'write' transactions towards the datastore are not permitted during the time compaction takes place. For this reason, NSO exposes a number of settings to control the compaction process in `in-memory-v1` mode (these have no effect for `on-demand-v1`).

### Compacting In-Memory CDB

By default, compaction is handled automatically by the CDB. After each transaction, CDB evaluates whether compaction is required for the affected datastore.

This is done by examining the number of added nodes as well as the file size changes since the last performed compaction. The thresholds used can be modified in the `ncs.conf` file by configuring the `/ncs-config/compaction/file-size-relative`, `/ncs-config/compaction/file-size-absolute`, and `/ncs-config/compaction/num-node-relative` settings.

It is also possible to automatically trigger compaction after a set number of transactions by setting the `/ncs-config/compaction/num-transaction` property.

In the configuration datastore, compaction is by default delayed by 5 seconds when the threshold is reached to prevent any upcoming write transaction from being blocked. If the system is idle during these 5 seconds, meaning that there is no new transaction, the compaction will initiate. Otherwise, compaction is delayed by another 5 seconds. The delay time can be configured in `ncs.conf` by setting the `/ncs-config/compaction/delayed-compaction-timeout` property.

As compaction may require a significant amount of time, it may be preferable to disable automatic compaction by CDB and instead trigger compaction manually according to specific needs. If doing so, it is highly recommended to have another automated system in place. Automation of compaction can be done by using a scheduling mechanism such as CRON or by using the NCS scheduler. See [Scheduler](/guides/development/connected-topics/scheduler) for more information.

By default, CDB may perform compaction during its boot process. This may be disabled, if required, by starting NSO with the flag `--disable-compaction-on-start`.

Additionally, CDB CAPI provides a set of functions that may be used to create an external mechanism for compaction. See `cdb_initiate_journal_compaction()`, `cdb_initiate_journal_dbfile_compaction()`, and `cdb_get_compaction_info()` in [confd\_lib\_cdb(3)](/guides/resources/man/confd_lib_cdb.3) in Manual Pages.


# IPC Connection

Connect client libraries to NSO with IPC.

Client libraries connect to NSO for inter-process communication (IPC) using Unix domain sockets (Local IPC) or TCP. By default, NSO uses Local IPC over a Unix domain socket at the path `/tmp/nso/nso-ipc`, controlled by the `/ncs-config/ncs-local-ipc/path` element in `ncs.conf`. If you change the socket path, you can tell clients to use the new path through the `NCS_IPC_PATH` environment variable. Clients must also have filesystem permission to access the IPC socket path, or they will not be able to communicate with the NSO daemon process. Local IPC requires clients to run on the same host as NSO; for clients running on a remote host, use TCP IPC instead.

Alternatively, NSO can be configured to use TCP sockets for IPC by setting `/ncs-config/ncs-local-ipc/enabled` to `false` and configuring the address and port through the `/ncs-config/ncs-ipc-address/ip` (default value 127.0.0.1) and `/ncs-config/ncs-ipc-address/port` elements in `ncs.conf`. If you change these values, you will likely need to configure the clients accordingly. Note that these values have security implications; see [Security Issues](/guides/administration/installation-and-deployment/development-to-production-deployment/secure-deployment#securing-ipc-access). In particular, changing the address away from 127.0.0.1 may allow unauthenticated remote connections.

Many of the clients read the environment variable `NCS_IPC_PATH` to determine the Local IPC socket path, or `NCS_IPC_ADDR` and `NCS_IPC_PORT` for TCP IPC, but others might need source code changes. When both `NCS_IPC_PATH` and `NCS_IPC_PORT` are set, `NCS_IPC_PATH` takes precedence and TCP is not used. This is a list of clients that communicate with NSO and what needs to be done when the IPC configuration is changed.

<table><thead><tr><th width="218">Client</th><th>Changes required</th></tr></thead><tbody><tr><td>Remote commands via the <code>ncs</code> command</td><td>Remote commands, such as <code>ncs --reload</code>, check the environment variables <code>NCS_IPC_PATH</code> and <code>NCS_IPC_ADDR</code>/<code>NCS_IPC_PORT</code>.</td></tr><tr><td>CLI tools</td><td>The Command Line Interface (CLI) client <strong>ncs_cli</strong> and similar commands, such as <code>ncs_cmd</code> and <code>ncs_load</code>, check the environment variables <code>NCS_IPC_PATH</code> and <code>NCS_IPC_ADDR</code>/<code>NCS_IPC_PORT</code>. Alternatively, many of them also support command-line options (e.g. <code>-S</code> for socket path).</td></tr><tr><td>CDB and MAAPI clients</td><td>The address or path supplied to <code>Cdb.connect()</code> and <code>Maapi.connect()</code> must be changed.</td></tr><tr><td>Data provider API clients</td><td>The address or path supplied to <code>Dp</code> constructor socket must be changed.</td></tr><tr><td>Notification API clients</td><td>The new address or path must be supplied to the socket for the <code>Notif</code> constructor.</td></tr></tbody></table>

To run more than one instance of NSO on the same host (which can be useful in development scenarios), each instance needs its own IPC socket. Set `/ncs-config/ncs-local-ipc/path` in `ncs.conf` to different values for each instance. If, instead, you are using TCP for IPC, set `/ncs-config/ncs-ipc-address/port` in `ncs.conf` to different values. In either case, you may also need to change the NETCONF and CLI over SSH ports under `/ncs-config/netconf/transport` and `/ncs-config/cli/ssh` by either disabling them or changing their values.

## Restricting Access to the IPC Socket

By default, NSO uses Local IPC (Unix domain sockets) and relies on Unix filesystem permissions on the socket path to prevent unauthorized access. In case this is not sufficient, such as when using TCP IPC or untrusted users have shell access on the system where NSO runs, it is possible to further restrict the access to the IPC socket.

For Local IPC, you can leverage Unix filesystem permissions for the socket path to limit which OS users and groups can initiate connections to the socket. NSO may also perform additional authentication of the connecting users based on their UID; see [Authenticating IPC Access](/guides/administration/management/aaa-infrastructure#authenticating-ipc-access).

For TCP sockets, you can enable an access check by setting the `ncs.conf` element `/ncs-config/ncs-ipc-access-check/enabled` to `true`, and specifying a filename for `/ncs-config/ncs-ipc-access-check/filename`. The file should contain a shared secret, i.e., a random (printable ASCII) character string. Clients connecting to the IPC socket will then be required to prove that they have knowledge of the secret through a challenge handshake before they are allowed access to the NSO functions provided via the IPC socket.

{% hint style="info" %}
The access permissions on this file must be restricted via OS file permissions, such that it can only be read by the NSO daemon and client processes that are allowed to connect to the IPC socket. E.g. if both the daemon and the clients run as root, the file can be owned by root and have only "read by owner" permission (i.e. mode 0400). Another possibility is to have a group that only the daemon and the clients belong to, set the group ID of the file to that group, and have only "read by group" permission (i.e. mode 040).
{% endhint %}

To provide the secret to the client libraries and inform them that they need to use the access check handshake, you have to set the environment variable `NCS_IPC_ACCESS_FILE` to the full pathname of the file containing the secret. This is sufficient for all the clients mentioned above, i.e., there is no need to change the application code to support or enable this check.

{% hint style="info" %}
The access check must be either enabled or disabled for both the daemon and the clients. E.g., if `/ncs-config/ncs-ipc-access-check/enabled` in `ncs.conf` is not set to `true` but clients are started with the environment variable `NCS_IPC_ACCESS_FILE` pointing to a file with a secret, the client connections will fail.
{% endhint %}


# Cryptographic Keys

Store strings in NSO that are encrypted and decrypted using cryptographic keys.

By using the NSO built-in encrypted YANG extension types `tailf:aes-cfb-128-encrypted-string` or `tailf:aes-256-cfb-128-encrypted-string`, it is possible to store encrypted string values in NSO. See the [tailf\_yang\_extensions(5)](/guides/resources/man/tailf_yang_extensions.5#yang-types-2) man page for more details on the encrypted string YANG extension types.

## Providing Keys

NSO supports defining one or more sets of cryptographic keys directly in `ncs.conf` or using an external command. Three methods can be used to configure the keys in `ncs.conf`:

* External command providing keys under `/ncs-config/encrypted-strings/external-keys`.
* Key rotation under `/ncs-config/encrypted-strings/key-rotation`.
* Legacy (single generation) format: `/ncs-config/encrypted-strings/AESCFB128` and `/ncs-config/encrypted-strings/AES256CFB128` .

### NSO Installer-Provided Cryptography Keys

* **Local installation**: Dummy keys are provided in legacy format in `ncs.conf` for development purposes. For deployment, the keys must be changed to random values. Example local installation `ncs.conf` (do not reuse):

  ```xml
  <ncs-config xmlns="http://tail-f.com/yang/tailf-ncs-config">
    <encrypted-strings>
      <AESCFB128>
        <key>0123456789abcdef0123456789abcdeg</key>
      </AESCFB128>
      <AES256CFB128>
        <key>0123456789abcdef0123456789abcdef0123456789abcdef0123456789abcdeg</key>
      </AES256CFB128>
    </encrypted-strings>
  </ncs-config>
  ```
* **System installation**: Random keys are generated in the legacy format stored in `${NCS_CONFIG_DIR}/ncs.crypto_keys`, and read using the `${NCS_DIR}/bin/ncs_crypto_keys` external command as configured in `${NCS_CONFIG_DIR}/ncs.conf`. Example system installation `ncs.conf:`

  ```xml
  <ncs-config xmlns="http://tail-f.com/yang/tailf-ncs-config">
    <encrypted-strings>
      <external-keys>
        <command>${NCS_DIR}/bin/ncs_crypto_keys</command>
        <command-argument>${NCS_CONFIG_DIR}/ncs.crypto_keys</command-argument>
      </external-keys>
    </encrypted-strings>
  </ncs-config>
  ```

  Example system installation`ncs.crypto_keys` file (do not reuse):

  ```
  AESCFB128_KEY=40f7c3b5222c1458be3411cdc0899fg
  AES256CFB128_KEY=5a08b6d78b1ce768c67e13e76f88d8af7f3d925ce5bfedf7e3169de6270bb6eg
  ```

  For details on using a custom external command to read the encryption keys, see [Encrypted Strings](/guides/development/connected-topics/encryption-keys).

You can generate a new set of keys, e.g. for use within the `ncs.crypto_keys` file, with the following command (requires `openssl` to be present):

```sh
#!/bin/sh
cat <<EOF
AESCFB128_KEY=$(openssl rand -hex 16)
AES256CFB128_KEY=$(openssl rand -hex 32)
EOF
```

### Providing Keys for Key Rotation

To provide keys that can be rotated in `ncs.conf`, each generation of cryptographic keys must be encapsulated in a `/ncs-config/encrypted-strings/key-rotation` list, and a `/ncs-config/encrypted-strings/key-rotation/generation` list key starting from `0` must be included and incremented for each set of cryptographic keys. Example (do not reuse):

```xml
<ncs-config xmlns="http://tail-f.com/yang/tailf-ncs-config">
  <encrypted-strings>
    <key-rotation>
      <generation>0</generation>
      <AESCFB128>
        <key>0123456789abcdef0123456789abcdeg</key>
      </AESCFB128>
      <AES256CFB128>
        <key>3c687d564e250ad987198d179537af563341357493ed2242ef3b16a881dd608g</key>
      </AES256CFB128>
    </key-rotation>
    <key-rotation>
      <generation>1</generation>
      <AESCFB128>
        <key>0123456789abcdef0123456789abcdeh</key>
      </AESCFB128>
      <AES256CFB128>
        <key>3c687d564e250ad987198d179537af563341357493ed2242ef3b16a881dd608h</key>
      </AES256CFB128>
    </key-rotation>
  </encrypted-strings>
</ncs-config>
```

External keys that can be rotated must be provided with the initial line `EXTERNAL_KEY_FORMAT=2` and the `generation` within square brackets. Example (do not reuse):

```
EXTERNAL_KEY_FORMAT=2
AESCFB128_KEY[0]=0123456789abcdef0123456789abcdeg
AES256CFB128_KEY[0]=3c687d564e250ad987198d179537af563341357493ed2242ef3b16a881dd608g
AESCFB128_KEY[1]=0123456789abcdef0123456789abcdeh
AES256CFB128_KEY[1]=3c687d564e250ad987198d179537af563341357493ed2242ef3b16a881dd608h
```

There is always an active generation:

* Active generation is the generation in the set of keys currently used to encrypt and decrypt all leafs with an encrypted string type.
* The active generation is persisted.
* If using the legacy method of providing keys in `ncs.conf` or when providing keys using the `/ncs-config/encrypted-strings/key-rotation` method without providing the initial line `EXTERNAL_KEY_FORMAT=2` in the application, the active generation will be `-1`.
* If starting NSO without any previous keys using the `/ncs-config/encrypted-strings/key-rotation` method or the `external-keys` method with the initial line `EXTERNAL_KEY_FORMAT=2`, the highest provided generation will be selected as the active generation.

For `ncs.conf` details, see the [ncs.conf(5) man page](/guides/resources/man/ncs.conf.5) under `/ncs-config/encrypted-strings`.

## Key Rotation

Rotating cryptographic keys means replacing an old cryptographic key with a new one while maintaining the functionality of the encryption and decryption of encrypted string values in NSO. It is a standard practice in cryptography and key management to enhance security and mitigate risks associated with key exposure or compromise.\
Key rotation helps ensure that sensitive data remains secure over time. It reduces the impact of potential key compromise and adheres to best practices for cryptographic hygiene. Key benefits:

* If a cryptographic key is compromised, rotating it reduces the amount of data exposed to the attacker since previously encrypted values can be re-encrypted with a new key.
* Regular rotation minimizes the time a single key is in use, thereby reducing the potential damage an attacker could do if they gain access to it.
* Reusing the same key for a prolonged period increases the risk of data correlation attacks (e.g., frequency analysis). Rotation ensures unique keys are used for encrypting strings, reducing this risk.
* Regularly rotating keys helps organizations maintain and test their key management processes. This ensures the system is prepared to handle key management tasks effectively in an emergency.

To rotate to a new generation of keys and re-encrypt the data:

1. Always [take a backup](/guides/administration/management/system-management#backup-and-restore) using [ncs-backup](/guides/resources/man/ncs-backup.1).
2. Check the currently active generation using the `/key-rotation/get-active-generation` action.
3. Re-encrypt all encrypted values with a new set of keys using the `/key-rotation/apply-new-key` action with the `new-key-generation` to rotate to as input.\
   The commit queue must be empty before running the action, or the action will fail, as the snapshot database is re-initialized. To wait for the commit queue to become empty, use the `wait-commit-queue` argument with the number of seconds to wait before failing.

CLI example:

```
$ ${NCS_DIR}/bin/ncs-backup
$ ncs_cli -Cu admin
# key-rotation get-active-generation 
active-generation -1
# key-rotation apply-new-keys new-key-generation 0 wait-commit-queue 10
result true
new-active-key-generation 0
```

The data in CDB that is subject to re-encryption when executing the `/key-rotation/apply-new-key` action:

* Encrypted types.
* Unions of encrypted types.
* Service metadata (original attribute, reverse and forward diff set).
* NED secrets.
* Rollback files.
* History log.

Under the hood, the`/key-rotation/apply-new-keys` action, when executed, performs the following steps:

1. Starts an upgrade transaction that will be used when re-encrypting the datastore.
2. Load the new active cryptographic keys into CDB and persist them.
3. Sync HA.
4. Re-encrypt data.
5. Drops the CDB snapshot database.
6. Commits data.
7. Restart NSO VMs.
8. End upgrade.

## Reloading After Changes to the Cryptographic Keys

1. Before changing the cryptographic keys, always [take a backup](/guides/administration/management/system-management#backup-and-restore) using [ncs-backup](/guides/resources/man/ncs-backup.1). Also, back up the external key file, default `${NCS_CONFIG_DIR}/ncs.crypto_keys`, or the `${NCS_CONFIG_DIR}/ncs.conf` file, depending on where the keys are stored.
2. Suppose you have previously provided keys in the legacy format and wish to switch to `/ncs-config/encrypted-strings/key-rotation` or `external-keys` with the initial line `EXTERNAL_KEY_FORMAT=2`. In that case, you must provide the currently used keys as generation `-1`. The new keys can have any non-negative generation number.
3. Replace the external key file or `ncs.conf` file depending on where the keys are stored.
4. Issue `ncs --reload` to reload the cryptographic keys.
5. Ensure commit queues are empty or wait for them to become empty.
6. Execute`/key-rotation/apply-new-keys` action to change the active generation, for example, from `-1` to `new-key-generation 0` as shown in the CLI example above.

{% hint style="info" %}
In a high-availability setting, keys must be identical on all nodes before attempting key rotation. Otherwise, the action will abort. The node executing the action will initiate the key reload for all nodes.
{% endhint %}

## Migrating 3DES Encrypted Values

NSO 6.5 removed support for 3DES encryption since the algorithm is no longer deemed sufficiently secure. If you are migrating from an older version and you have data using the `tailf:des3-cbc-encrypted-string` YANG type, NSO will no longer be able to read this data. In fact, compiling a YANG module using this type will produce an error.

To avoid losing data when upgrading to NSO 6.5 or later, you must first update all the YANG data models and change the `tailf:des3-cbc-encrypted-string` type to either `tailf:aes-cfb-128-encrypted-string` or `tailf:aes-256-cfb-128-encrypted-string`. Compile the updated models and then perform a package upgrade for the affected packages.

While upgrading the packages, the automatic CDB schema upgrade will re-encrypt the data in the new (AES) format. At this point you are ready to upgrade to the new NSO version that no longer supports 3DES.


# Service Manager Restart

Restart strategy for the service manager.

The service manager executes in a Java VM outside of NSO. The `NcsMux` initializes a number of sockets to NSO at startup. These are Maapi sockets and data provider sockets. NSO can choose to close any of these sockets whenever NSO requests the service manager to perform a task, and that task is not finished within the stipulated timeout. If that happens, the service manager must be restarted. The timeout(s) are controlled by several `ncs.conf` parameters found under `/ncs-config/japi`.


# IPv6 on Northbound Interfaces

Learn about using IPv6 on NSO's northbound interfaces.

NSO supports access to all northbound interfaces via IPv6, and in the most simple case, i.e., IPv6-only access, this is just a matter of configuring an IPv6 address (typically the wildcard address `::`) instead of IPv4 for the respective agents and transports in `ncs.conf`, e.g., `/ncs-config/cli/ssh/ip` for SSH connections to the CLI or `/ncs-config/netconf-north-bound/transport/ssh/ip` for SSH to the NETCONF agent. The SNMP agent configuration is configured via one of the other northbound interfaces rather than via `ncs.conf`, see [NSO SNMP Agent](/guides/development/core-concepts/northbound-apis#the-nso-snmp-agent) in Northbound APIs. For example, via the CLI, we would set `snmp agent ip` to the desired address. All these addresses default to the IPv4 wildcard address `0.0.0.0`.

In most IPv6 deployments, it will, however, be necessary to support IPv6 and IPv4 access simultaneously. This requires that both IPv4 and IPv6 addresses are configured, typically `0.0.0.0` plus `::`. To support this, there is in addition to the `ip` and `port` leafs also a list `extra-listen` for each agent and transport, where additional IP addresses and port pairs can be configured. Thus, to configure the CLI to accept SSH connections to port 2024 on any local IPv6 address, in addition to the default (port 2024 on any local IPv4 address), we can add an `<extra-listen>` section under `/ncs-config/cli/ssh` in `ncs.conf`:

```xml
  <cli>
    <enabled>true</enabled>

    <!-- Use the built-in SSH server -->
    <ssh>
      <enabled>true</enabled>
      <ip>0.0.0.0</ip>
      <port>2024</port>

      <extra-listen>
        <ip>::</ip>
        <port>2024</port>
      </extra-listen>

    </ssh>

    ...
  </cli>
```

To configure the SNMP agent to accept requests to port 161 on any local IPv6 address, we could similarly use the CLI and give the command:

```bash
admin@ncs(config)# snmp agent extra-listen :: 161
```

The `extra-listen` list can take any number of address/port pairs; thus, this method can also be used when we want to accept connections/requests on several specified (IPv4 and/or IPv6) addresses instead of the wildcard address or when we want to use multiple ports.


# Layered Service Architecture

Design large and scalable NSO applications using LSA.

Layered Service Architecture (LSA) is a design approach for massively large and scalable NSO applications. Large service providers and enterprises can use it to manage services for millions of users, ranging over several hundred thousand managed devices. Such scale requires special consideration since a single NSO instance no longer suffices and LSA helps you address this challenge.

## Going Big <a href="#d5e35" id="d5e35"></a>

At some point, scaling up hits the law of diminishing returns. Effectively, adding more resources to the NSO server becomes prohibitively expensive. To further increase the throughput of the whole system, you can share the load across multiple instances, in a scale-out fashion.

You achieve this by splitting a service into a main, upper-layer part, and one or more lower-layer parts. The upper part controls and dispatches work to the lower parts. This is the same approach as using a customer-facing service (CFS) and a resource-facing service (RFS). However, here the CFS code (the upper-layer part) runs in a different NSO node than the RFS code (the lower-layer parts). What is more, the lower-layer parts can be spread across multiple NSO nodes.

Each RFS node is responsible for its own set of managed devices, mounted under its `/devices` tree, and the upper-layer, CFS node only concerns itself with the RFS nodes. So, the CFS node only mounts the RFS nodes under its `/devices` tree, not managed devices directly. The main advantage of this architecture is that you can add many device RFS nodes that collectively manage a huge number of actual devices—much more than a single node could.

<div data-with-frame="true"><figure><img src="/files/pz9ZjwM5v0asWJ7NwkQI" alt="" width="563"><figcaption><p>Layered CFS/RFS architecture</p></figcaption></figure></div>

## Is LSA for Me?

While it is tempting to design the system in the most scalable way from the start, it comes with a cost. Compared to a single, non-LSA setup, the automation system now becomes distributed across multiple nodes, with all the complexity that entails. For example, in a non-distributed system, the communication between different parts has mostly negligible latency and hardly ever fails. That is certainly not true anymore for distributed systems as we know them today, including LSA.

More practically, taking a service in NSO and deploying a single instance on an LSA system is likely to take longer and have a higher chance of failure compared to a non-LSA system, because additional network communication is involved.

Moreover, multiple NSO nodes present a higher operational complexity and administrative burden. There is no longer a “single pane of glass” view of all the individual devices. That's why you must weigh the benefits of the LSA approach against the scale at which you operate. When LSA starts making sense will depend on the type of devices you manage, the services you have, the geographical distribution of resources, and so on.

A distributed system can push the overall throughput way beyond what a single instance can do. But you will achieve a much better outcome by first focusing on eliminating the bottlenecks in the provisioning code, as discussed in [Scaling and Performance Optimization](/guides/development/advanced-development/scaling-and-performance-optimization). Only when that proves insufficient, consider deploying LSA.

LSA also addresses the memory limitations of NSO when device configurations become very large (individually or all together). If the NSO server is memory-constrained and more memory cannot be added, the LSA approach can be a solution.

Another challenge that LSA may help you overcome is scaling organizationally. When many teams share the same NSO instance, it can get hard to separate the different concerns and responsibilities. Teams may also have different cadences or preferences for upgrades, resulting in friction. With LSA, it becomes possible to create a clearer separation. The CFS node and the RFS nodes can have different release cycles (as long as the YANG upgrade rules are followed) and each can be upgraded independently. If a bug is found or a feature is missing in the RFS nodes, it can be fixed without affecting the CFS node, and vice versa.

To summarize, the major advantage of this architecture is scalability. The solution scales horizontally, both at the upper and the lower layer, thus catering for truly massive deployments, but at the expense of the increased complexity.

## Layered Service Design <a href="#d5e58" id="d5e58"></a>

To take advantage of the scalability potential of LSA, your services must be designed in a layered fashion. Once the automation logic in NSO reaches a certain level of complexity, a stacked service design tends to emerge naturally. Often, you can extend it to LSA with relatively little change. The same is true for brand-new, green field designs.

In other situations, you might need to invest some additional effort to split and orchestrate the work across multiple groups of devices. Examples are existing monolithic services or stacked service designs that require all RFSs to access all devices.

### New, Greenfield Design <a href="#d5e62" id="d5e62"></a>

If you are designing the service from scratch, you have the most freedom in choosing the partitioning of logic between CFS and RFS. The CFS must contain the YANG definition for the service and its configurable options that are available to the customer, perhaps through an order capture system north of the NSO. On the other hand, the RFS YANG models are internal to the service, that is, they are not used directly by the customer. So, you are free to design them in a way that makes the provisioning code as simple as possible.

As an example, you might have a VLAN provisioning service where the CFS lets users select if the hosts on the VLAN can access the internet. Then you can divide provisioning into, let's say, an RFS service that configures the VLAN and the appropriate IP subnet across the data center switches, and another RFS service that configures the firewall to allow the traffic from the subnet to reach the internet. This design clearly separates the provisioned devices into two groups: firewalls and data center switches. Each group can be managed by a separate lower-layer NSO.

### Existing Monolithic Application with Stacked Services <a href="#d5e66" id="d5e66"></a>

Similar to a brand new design, an existing monolithic application that uses stacked services has already laid the groundwork for LSA-compatible design because of the existing division into two layers (upper and lower).

A possible complication, in this case, is when each existing RFS touches all of the affected devices, and that makes it hard to partition devices across multiple lower-layer NSO nodes. For example, if one RFS manages the VLAN interface (the VLAN ID and layer 2 settings) and another RFS manages the IP configuration for this interface, that configuration very likely happens on the same devices. The solution in this situation could be to partition RFS services based on the data center that they operate in, such as one lower-layer NSO node for one data center, another lower-layer NSO for another data center, and so on. If that is not possible, an alternative is to redesign each RFS and split their responsibilities differently.

#### Existing Monolithic Application <a href="#d5e70" id="d5e70"></a>

The most complex, yet common case is when a single node NSO installation grows over time and you are faced with performance problems due to the new size. To leverage the LSA functionality, you must first split the service into upper- and lower-layer parts, which require a certain amount of effort. That is why the decision to use LSA should always be accompanied by a thorough analysis to determine what makes the system too slow. Sometimes, it is a result of a bad "must" expression in the service YANG code or similar. Fixing that is much easier than re-architecting the application.

### Orchestrating the Work <a href="#d5e73" id="d5e73"></a>

Regardless of whether you start with a green field design or extend an existing application, you must tackle the problem of dispatching the RFS instantiation to the correct lower-layer NSO node.

Imagine a VPN application that uses a managed device on each site to securely connect to the private network. In a service provider network, this is usually done by the CPE. When a customer orders connectivity to an additional site (another leg of the VPN), the service needs to configure the site-local device (the CPE). As there will be potentially many such devices, each will be managed by one of the RFS nodes. However, the VPN service is managed centrally, through the CFS, which must:

* Figure out which RFS node is responsible for the device for the new site (CPE).
* Dispatch the RFS instantiation to that particular RFS node, making sure the device is properly configured.

NSO provides a mechanism to facilitate the second part, the actual dispatch, but the service logic must somehow select the correct RFS node. If the RFS nodes are geographically separated across different countries or different data centers, the CFS could simply infer or calculate the right RFS node based on service instance parameters, such as the physical location of the new site.

A more flexible alternative is to use dynamic mapping. It can be as simple as a list of 2-tuples that map a device name to an RFS node, stored in the CDB. The trade-off is that the list must be maintained. It is straightforward to automate the maintenance of the list though, for example through NETCONF notifications whenever `/devices/device` on the RFS nodes is manipulated or by explicitly asking the CFS node to query the RFS nodes for their list of devices.

Ultimately, the right approach to dispatch will depend on the complexity of your service and operational procedures.

### Provisioning of an LSA Service Request <a href="#d5e86" id="d5e86"></a>

Having designed a layered service with the CFS and RFS parts, the CFS must now communicate with the RFS that resides on a different node. You achieve that by adding the lower-layer (RFS) node as a managed device to the upper-layer (CFS) node. The CFS node must access the RFS data model on the lower-layer node, just like it accesses any other configuration on any managed device. But don't you need a NED to do this? Indeed, you do. That's why the RFS model needs to be specially compiled for the upper-layer node to use as part of NED and not a standalone service. A model compiled in this way is called a 'device compiled'.

Let's then see how the LSA setup affects the whole service provisioning process. Suppose a new request arrives at the CFS node, such as a new service instance being created through RESTCONF by a customer order portal. The CFS runs the service mapping logic as usual; however, instead of configuring the network devices directly, the CFS configures the appropriate RFS nodes with the generated RFS service instance data. This is the dispatch logic in action.

<div data-with-frame="true"><figure><img src="/files/rBCHgDwS9LI6JQPyh2ZZ" alt="" width="375"><figcaption><p>LSA Request Flow</p></figcaption></figure></div>

As the configuration for the lower-layer nodes happens under the `/devices/device` tree, it is picked up and pushed to the relevant NSO instances by the NED. The NED sends the appropriate NETCONF edit-config RPCs, which trigger the RFS FASTMAP code at the RFS nodes. The RFS mapping logic constructs the necessary network configuration for each RFS instance and the RFS nodes update the actual network devices.

In case the commit queue feature is not being used, this entire sequence is serialized through the system as a whole. It means that if another northbound request arrives at the CFS node while the first request is being processed, the second request is synchronously queued at the CFS node, waiting for the currently running transaction to either succeed or fail.

If the code on the RFS nodes is reactive, it will likely return without much waiting, since the RFM applications are usually very fast during their first round of execution. But that will still have a lower performance than using the commit queue since the execution is serialized eventually when modifying devices. To maximize throughput, you also need to enable the commit queue functionality throughout the system.

### Implementation Considerations <a href="#d5e100" id="d5e100"></a>

The main benefit of LSA is that it scales horizontally at the RFS node layer. If one RFS node starts to become overloaded, it's easy to bring up an additional one, to share the load. Thus LSA caters to scalability at the level of the number of managed devices. However, each RFS node needs to host all the RFSs that touch the devices it manages under its `/devices/device` tree. There is still one, and only one, NSO node that directly manages a single device.

Dividing a provisioning application into upper and lower-layer services also increases the complexity of the application itself. For example, to follow the execution of a reactive or nano RFS, typically an additional NETCONF notification code must be written. The notifications have to be sent from the RFS nodes and received and processed by the CFS code. This way, if something goes wrong at the device layer, the information is relayed all the way to the top level of the system.

Furthermore, it is highly recommended that LSA applications enable the commit queue on all NSO nodes. If the commit queue is not enabled, the slowest device on the network will limit the overall throughput, significantly reducing the benefits of LSA.

Finally, if the two-layer approach proves to be insufficient due to requirements at the CFS node, you can extend it to three layers, with an additional layer of NSO nodes between the CFS and RFS layers.

## LSA Examples

### Greenfield LSA Application

This section describes a small LSA application, which exists as a running example in the [examples.ncs/layered-services-architecture/lsa-single-version-deployment](https://github.com/NSO-developer/nso-examples/tree/6.7/layered-services-architecture/lsa-single-version-deployment) directory.

The application is a slight variation on the [examples.ncs/service-management/rfs-service](https://github.com/NSO-developer/nso-examples/tree/6.7/service-management/rfs-service) example where the YANG code has been split up into an upper-layer and a lower-layer implementation. The example topology (based on netsim for the managed devices, and NSO for the upper/lower layer NSO instances) looks like the following:

<div data-with-frame="true"><figure><img src="/files/iwFFjo0SM2ZvosjQZpB7" alt="" width="563"><figcaption><p>Example LSA architecture</p></figcaption></figure></div>

The upper layer of the YANG service data for this example looks like the following:

```yang
module cfs-vlan {
  ...
  list cfs-vlan {
    key name;
    leaf name {
      type string;
    }

    uses ncs:service-data;
    ncs:servicepoint cfs-vlan;

    leaf a-router {
      type leafref {
        path "/dispatch-map/router";
      }
      mandatory true;
    }
    leaf z-router {
      type leafref {
        path "/dispatch-map/router";
      }
      mandatory true;
    }
    leaf iface {
      type string;
      mandatory true;
    }
    leaf unit {
      type int32;
      mandatory true;
    }
    leaf vid {
      type uint16;
      mandatory true;
    }
  }
}
```

Instantiating one CFS we have:

```
admin@upper-nso% show cfs-vlan
cfs-vlan v1 {
    a-router ex0;
    z-router ex5;
    iface    eth3;
    unit     3;
    vid      77;
}
```

The provisioning code for this CFS has to make a decision on where to instantiate what. In this example the "what" is trivial, it's the accompanying RFS, whereas the "where" is more involved. The two underlying RFS nodes, each manage 3 netsim routers, thus given the input, the CFS code must be able to determine which RFS node to choose. In this example, we have chosen to have an explicit map, thus on the `upper-nso` we also have:

```
admin@upper-nso% show dispatch-map
dispatch-map ex0 {
    rfs-node lower-nso-1;
}
dispatch-map ex1 {
    rfs-node lower-nso-1;
}
dispatch-map ex2 {
    rfs-node lower-nso-1;
}
dispatch-map ex3 {
    rfs-node lower-nso-2;
}
dispatch-map ex4 {
    rfs-node lower-nso-2;
}
dispatch-map ex5 {
    rfs-node lower-nso-2;
}
```

So, we have a template CFS code that does the dispatching to the right RFS node.

```xml
<config-template xmlns="http://tail-f.com/ns/config/1.0"
                 servicepoint="cfs-vlan">
  <devices xmlns="http://tail-f.com/ns/ncs">
    <!-- Do this for the two leafs a-router and z-router -->
    <?foreach {a-router|z-router}?>
    <device>
      <!--
      Pick up the name of the rfs-node from the dispatch-map
      and do not change the current context thus the string()
      -->
      <name>{string(deref(current())/../rfs-node)}</name>
      <config>
        <vlan xmlns="http://com/example/rfsvlan">
          <!-- We do not want to change the current context here either -->
          <name>{string(/name)}</name>
          <!-- current() is still a-router or z-router -->
          <router>{current()}</router>
          <iface>{/iface}</iface>
          <unit>{/unit}</unit>
          <vid>{/vid}</vid>
          <description>Interface owned by CFS: {/name}</description>
        </vlan>
      </config>
    </device>
    <?end?>
  </devices>
</config-template>
```

This technique for dispatching is simple and easy to understand. The dispatching might be more complex, it might even be determined at execution time dependent on CPU load. It might be (as in this example) inferred from input parameters or it might be computed.

The result of the template-based service is to instantiate the RFS, at the RFS nodes.

First, let's have a look at what happened in the upper-nso. Look at the modifications but ignore the fact that this is an LSA service:

```
admin@upper-nso% request cfs-vlan v1 get-modifications no-lsa
cli {
    local-node {
        data  devices {
                   device lower-nso-1 {
                       config {
              +            rfs-vlan:vlan v1 {
              +                router ex0;
              +                iface eth3;
              +                unit 3;
              +                vid 77;
              +                description "Interface owned by CFS: v1";
              +            }
                       }
                   }
                   device lower-nso-2 {
                       config {
              +            rfs-vlan:vlan v1 {
              +                router ex5;
              +                iface eth3;
              +                unit 3;
              +                vid 77;
              +                description "Interface owned by CFS: v1";
              +            }
                       }
                   }
               }
    }
}
```

Just the dispatched data is shown. As `ex0` and `ex5` reside on different nodes, the service instance data has to be sent to both `lower-nso-1` and `lower-nso-2`.

Now let's see what happened in the `lower-nso`. Look at the modifications and take into account that these are LSA nodes (this is the default):

```
admin@upper-nso% request cfs-vlan v1 get-modifications
cli {
  local-node {
    .....
  }
  lsa-service {
    service-id /devices/device[name='lower-nso-1']/config/rfs-vlan:vlan[name='v1']
    data devices {
      device ex0 {
        config {
          r:sys {
            interfaces {
   +          interface eth3 {
   +            enabled;
   +            unit 3 {
   +              enabled;
   +              description "Interface owned by CFS: v1";
   +              vlan-id 77;
   +            }
   +          }
            }
          }
        }
      }
    }
  }
  lsa-service {
    service-id /devices/device[name='lower-nso-2']/config/rfs-vlan:vlan[name='v1']
    data devices {
      device ex5 {
        config {
          r:sys {
            interfaces {
   +          interface eth3 {
   +            enabled;
   +            unit 3 {
   +              enabled;
   +              description "Interface owned by CFS: v1";
   +              vlan-id 77;
   +            }
   +          }
            }
          }
        }
      }
    }
  }
```

Both the dispatched data and the modification of the remote service are shown. As `ex0` and `ex5` reside on different nodes, the service modifications of the service `rfs-vlan` on both `lower-nso-1` and `lower-nso-2` are shown.

The communication between the NSO nodes is of course NETCONF.

```
admin@upper-nso% set cfs-vlan v1 a-router ex0 z-router ex5 iface eth3 unit 3 vid 78
[ok][2016-10-20 16:52:45]

[edit]
admin@upper-nso% commit dry-run outformat native
native {
    device {
        name lower-nso-1
        data <rpc xmlns="urn:ietf:params:xml:ns:netconf:base:1.0"
                  message-id="1">
               <edit-config xmlns:nc="urn:ietf:params:xml:ns:netconf:base:1.0">
                 <target>
                   <running/>
                 </target>
                 <test-option>test-then-set</test-option>
                 <error-option>rollback-on-error</error-option>
                 <with-inactive xmlns="http://tail-f.com/ns/netconf/inactive/1.0"/>
                 <config>
                   <vlan xmlns="http://com/example/rfsvlan">
                     <name>v1</name>
                     <vid>78</vid>
                     <private>
                       <re-deploy-counter>-1</re-deploy-counter>
                     </private>
                   </vlan>
                 </config>
               </edit-config>
             </rpc>
    }
               ...........
               ....
```

The YANG model at the lower layer, also known as the RFS layer, is similar to the CFS, but slightly different:

```yang
module rfs-vlan {

  ...

  list vlan {
    key name;
    leaf name {
      tailf:cli-allow-range;
      type string;
    }

    uses ncs:service-data;
    ncs:servicepoint "rfs-vlan";

    leaf router {
      type string;
    }
    leaf iface {
      type string;
      mandatory true;
    }
    leaf unit {
      type int32;
      mandatory true;
    }
    leaf vid {
      type uint16;
      mandatory true;
    }
    leaf description {
      type string;
      mandatory true;
    }
  }
}
```

The task for the RFS provisioning code here is to actually provision the designated router. If we log into one of the lower layer NSO nodes, we can check the following.

```
admin@lower-nso-1> show configuration vlan
vlan v1 {
    router      ex0;
    iface       eth3;
    unit        3;
    vid         77;
    description "Interface owned by CFS: v1";
}
[ok][2016-10-20 17:01:08]
admin@lower-nso-1> request vlan v1 get-modifications
cli {
  local-node {
    data  devices {
             device ex0 {
               config {
                 r:sys {
                   interfaces {
    +                interface eth3 {
    +                  enabled;
    +                  unit 3 {
    +                    enabled;
    +                    description "Interface owned by CFS: v1";
    +                    vlan-id 77;
    +                  }
    +                }
                   }
                 }
               }
             }
    }
  }
}
```

To conclude this section, the final remark here is that to design a good LSA application, the trick is to identify a good layering for the service data models. The upper layer, the CFS layer is what is exposed northbound, and thus requires a model that is as forward-looking as possible since that model is what a system north of NSO integrates to, whereas the lower layer models, the RFS models can be viewed as "internal system models" and they can be more easily changed.

### Greenfield LSA Application Designed for Easy Scaling <a href="#d5e148" id="d5e148"></a>

In this section, we'll describe a lightly modified version of the example in the previous section. The application we describe here exists as a running example under [examples.ncs/layered-services-architecture/lsa-scaling](https://github.com/NSO-developer/nso-examples/tree/6.7/layered-services-architecture/lsa-scaling).

Sometimes it is desirable to be able to easily move devices from one lower LSA node to another. This makes it possible to easily expand or shrink the number of lower LSA nodes. Additionally, it is sometimes desirable to avoid HA pairs for replication but instead use a common store for all lower LSA devices, such as a distributed database, or a common file system.

The above is possible provided that the LSA application is structured in certain ways.

* The lower LSA nodes only expose services that manipulate the configuration of a single device. We call these devices RFSs, or dRFS for short.
* All services are located in a way that makes it easy to extract them, for example in /drfs:dRFS/device

  ```yang
  container dRFS {
    list device {
      key name;
      leaf name {
        type string;
      }
    }
  }
  ```
* No RFS takes place on the lower LSA nodes. This avoids the complication with locking and distributed event handling.
* The LSA nodes need to be set up with the proper NEDs and with auth groups such that a device can be moved without having to install new NEDs or update auth groups.

Provided that the above requirements are met, it is possible to move a device from one lower LSA node by extracting the configuration from the source node and installing it on the target node. This, of course, requires that the source node is still alive, which is normally the case when HA-pairs are used.

An alternative to using HA-pairs for the lower LSA nodes is to extract the device configuration after each modification to the device and store it in some central storage. This would not be recommended when high throughput is required but may make sense in certain cases.

In the example application, there are two packages on the lower LSA nodes that provide this functionality. The package `inventory-updater` installs a database subscriber that is invoked every time any device configuration is modified, both in the preparation phase and in the commit phase of any such transaction. It extracts the device and dRFS configuration, including service metadata, during the preparation phase. If the transaction proceeds to a full commit, the package is again invoked and the extracted configuration is stored in a file in the directory `db_store`.

The other package is called `device-actions`. It provides three actions: `extract-device`, `install-device`, and `delete-device`. They are intended to be used by the upper LSA node when moving a device either from a lower LSA node or from `db_store`.

In the upper LSA node, there is one package for coordinating the movement, called `move-device`. It provides an action for moving a device from one lower LSA node to another. For example when invoked to move device `ex0` from `lower-1` to `lower-2` using the action

```cli
request move-device move src-nso lower-1 dest-nso lower-2 device-name ex0
```

it goes through the following steps:

* A partial lock is acquired on the upper-nso for the path `/devices/device[name=lower-1]/config/dRFS/device[name=ex0]` to avoid any changes to the device while the device is in the process of being moved.
* The device and dRFS configuration are extracted in one of two ways:

  * Read the configuration from `lower-1` using the action

    ```cli
    request device-action extract-device name ex0
    ```
  * Read the configuration from some central store, in our case the file system in the directory. `db_store`.

  The configuration will look something like this

  ```
  devices {
      device ex0 {
          address   127.0.0.1;
          port      12022;
          ssh {
          ...
             /* Refcount: 1 */
              /* Backpointer: [ /drfs:dRFS/drfs:device[drfs:name='ex0']/rfs-vlan:vlan[rfs-vlan:name='v1'] ] */
              interface eth3 {
              ...
              }
          ...
      }
  }
  dRFS {
      device ex0 {
          vlan v1 {
              private {
              ...
              }
          }
      }
  }
  ```
* Install the configuration on the `lower-2` node. This can be done by running the action:

  ```cli
  request device-action install-device name ex0 config <cfg>
  ```

  This will load the configuration and commit using the flags `no-deploy` and `no-networking`.
* Delete the device from `lower-1` by running the action

  ```cli
  request device-action delete-device name ex0
  ```
* Update mapping table

  ```
  dispatch-map ex0 {
      rfs-node lower-nso-2;
  }
  ```
* Release the partial lock for `/devices/device[name=lower-1]/config/dRFS/device[name=ex0]`.
* Re-deploy all services that have touched the device. The services all have backpointers from `/devices/device{lower-1}/config/dRFS/device{ex0}`. They are `re-deployed` using the flags `no-lsa` and `no-networking`.
* Finally, the action runs `compare-config` on `lower-1` and `lower-2`.

With this infrastructure in place, it is fairly straightforward to implement actions for re-balancing devices among lower LSA nodes, as well as evacuating all devices from a given lower LSA node. The example contains implementations of those actions as well.

### Re-architecting an Existing VPN Application for LSA <a href="#d5e230" id="d5e230"></a>

If we do not have the luxury of designing our NSO service application from scratch, but rather are faced with extending/changing an existing, already deployed application into the LSA architecture we can use the techniques described in this section.

Usually, the reasons for re-architecting an existing application are performance-related.

In the NSO example collection, two popular examples are the [examples.ncs/service-management/mpls-vpn-java](https://github.com/NSO-developer/nso-examples/tree/6.7/service-management/mpls-vpn-java) and [examples.ncs/service-management/mpls-vpn-python](https://github.com/NSO-developer/nso-examples/tree/6.7/service-management/mpls-vpn-python) examples. Those example contains an almost "real" VPN provisioning example whereby VPNs are provisioned in a network of CPEs, PEs, and P routers according to this picture:

<figure><img src="/files/PddnV8PVgxk8jLdRXaVv" alt=""><figcaption><p>VPN network</p></figcaption></figure>

The service model in this example roughly looks like this:

```yang
   list l3vpn {
      description "Layer3 VPN";

      key name;
      leaf name {
        type string;
      }

      leaf route-distinguisher {
        description "Route distinguisher/target identifier unique for the VPN";
        mandatory true;
        type uint32;
      }

      list endpoint {
        key "id";
        leaf id {
          type string;
        }
        leaf ce-device {
          mandatory true;
          type leafref {
            path "/ncs:devices/ncs:device/ncs:name";
          }
        }

        leaf ce-interface {
          mandatory true;
          type string;
        }

        ....

        leaf as-number {
          tailf:info "CE Router as-number";
          type uint32;
        }
      }
      container qos {
        leaf qos-policy {
           ......
```

There are several interesting observations on this model code related to the Layered Service Architecture.

* Each instantiated service has a list of endpoints and CPE routers. These are modeled as a leafref into the /devices tree. This has to be changed if we wish to change this application into an LSA application since the /devices tree at the upper layer doesn't contain the actual managed routers. Instead, the /devices tree contains the lower layer RFS nodes.
* There is no connectivity/topology information in the service model. Instead, the `mpls-vpn` example has topology information on the side, and that data is used by the provisioning code. That topology information for example contains data on which CE routers are directly connected to which PE router.

  Remember from the previous section, that one of the additional complications of an LSA application is the dispatching part. The dispatching problem fits well into the pattern where we have topology information stored on the side and let the provisioning FASTMAP code use that data to guide the provisioning. One straightforward way would be to augment the topology information with additional data, indicating which RFS node is used to manage a specific managed device.

By far the easiest way to change an existing monolithic NSO application into the LSA architecture is to keep the service model at the upper layer and lower layer almost identical, only changing things like leafrefs directly into the /devices tree which obviously breaks.

In this example, the topology information is stored in a separate container `share-data` and propagated to the LSA nodes by means of service code.

The example [examples.ncs/layered-services-architecture/mpls-vpn-lsa](https://github.com/NSO-developer/nso-examples/tree/6.7/layered-services-architecture/mpls-vpn-lsa) example does exactly this, the upper layer data model in `upper-nso/packages/l3vpn/src/yang/l3vpn.yang` now looks as:

```yang
   list l3vpn {
      description "Layer3 VPN";

      key name;
      leaf name {
        type string;
      }

      leaf route-distinguisher {
        description "Route distinguisher/target identifier unique for the VPN";
        mandatory true;
        type uint32;
      }

      list endpoint {
        key "id";
        leaf id {
          type string;
        }
        leaf ce-device {
          mandatory true;
          type string;
        }
        .......
```

The `ce-device` leaf is now just a regular string, not a leafref.

So, instead of an NSO topology that looks like:

<div data-with-frame="true"><figure><img src="/files/CJjQC4XKMzTXsGl5CMSI" alt="" width="563"><figcaption><p>NSO topology</p></figcaption></figure></div>

\
We want an NSO architecture that looks like this:

<div data-with-frame="true"><figure><img src="/files/B0uP2un0dhVXbuxiUW68" alt="" width="563"><figcaption><p>NSO LSA topology</p></figcaption></figure></div>

The task for the upper layer FastMap code is then to instantiate a copy of itself on the right lower layer NSO nodes. The upper layer FastMap code must:

* Determine which routers, (CE, PE, or P) will be touched by its execution.
* Look in its dispatch table, which lower-layer NSO nodes are used to host these routers.
* Instantiate a copy of itself on those lower layer NSO nodes. One extremely efficient way to do that is to use the `Maapi.copyTree()` method. The code in the example contains code that looks like this:

  ```java
          public Properties create(
              ....
              NavuContainer lowerLayerNSO = ....

              Maapi maapi = service.context().getMaapi();
              int tHandle = service.context().getMaapiHandle();
              NavuNode dstVpn = lowerLayerNSO.container("config").
                      container("l3vpn", "vpn").
                      list("l3vpn").
                      sharedCreate(serviceName);
              ConfPath dst = dstVpn.getConfPath();
              ConfPath src = service.getConfPath();

              maapi.copyTree(tHandle, true, src, dst);
  ```

Finally, we must make a minor modification to the lower layer (RFS) provisioning code too. Originally, the FastMap code wrote all config for all routers participating in the VPN, now with the LSA partitioning, each lower layer NSO node is only responsible for the portion of the VPN that involves devices that reside in its /devices tree, thus the provisioning code must be changed to ignore devices that do not reside in the /devices tree.

### Re-architecting Details <a href="#d5e283" id="d5e283"></a>

In addition to conceptual changes of splitting into upper- and lower-layer parts, migrating an existing monolithic application to LSA may also impact the models used. In the new design, the upper-layer node contains the (more or less original) CFS model as well as the device-compiled RFS model, which it requires for communication with the RFS nodes. In a typical scenario, these are two separate models. So, for example, they must each use a unique namespace.

To illustrate the different YANG files and namespaces used, the following text describes the process of splitting up an example monolithic service. Let's assume that the original service resides in a file, `myserv.yang`, and looks like the following:

```yang
module myserv {

  namespace "http://example.com/myserv";
  prefix ms;

  .....

  list srv {
    key name;
    leaf name {
      type string;
    }

    uses ncs:service-data;
    ncs:servicepoint vlanspnt;

    leaf router {
       type leafref {
         path "/ncs:devices/ncs:device/ncs:name";
    .....
    }
}
```

In an LSA setting, we want to keep this module as close to the original as possible. We clearly want to keep the namespace, the prefix, and the structure of the YANG identical to the original. This is to not disturb any provisioning systems north of the original NSO. Thus with only minor modifications, we want to run this module at the CFS node, but with non-applicable leafrefs removed, thus at the CFS node we would get:

```yang
module myserv {

  namespace "http://example.com/myserv";
  prefix ms;

  .....

  list srv {
    key name;
    leaf name {
      type string;
    }

    uses ncs:service-data;
    ncs:servicepoint vlanspnt;

    leaf router {
       type string;
    .....
    }
}
```

Now, we want to run almost the same YANG module at the RFS node, however, the namespace must be changed. For the sake of the CFS node, we're going to NED compile the RFS and NSO doesn't like the same namespace to occur twice, thus for the RFS node, we would get a YANG module `myserv-rfs.yang` that looks like the following:

```yang
module myserv-rfs {

  namespace "http://example.com/myserv-rfs";
  prefix ms-rfs;

  .....

  list srv {
    key name;
    leaf name {
      type string;
    }

    uses ncs:service-data;
    ncs:servicepoint vlanspnt;

    leaf router {
       type leafref {
         path "/ncs:devices/ncs:device/ncs:name";
    .....
    }
}
```

This file can, and should, keep the leafref as is.

The final and last file we get is the compiled NED, which should be loaded in the CFS node. The NED is directly compiled from the RFS model, as an LSA NED.

```bash
$ ncs-make-package --lsa-netconf-ned /path/to-rfs-yang  myserv-rfs-ned
```

Thus, we end up with three distinct packages from the original one:

1. The original, slated for the CFS node, with leafrefs removed.
2. The modified original, slated for the RFS node, with the namespace and the prefix changed.
3. The NED, compiled from the RFS node code, slated for the CFS node.

## Deploying LSA

The purpose of the upper CFS node is to manage all CFS services and to push the resulting service mappings to the RFS services. The lower RFS nodes are configured as devices in the device tree of the upper CFS node and the RFS services are created under the `/devices/device/config` accordingly. This is almost identical to the relation between a normal NSO node and the normal devices. However, there are differences when it comes to commit parameters and the commit queue, as well as some other LSA-specific features. For the shared commit-parameter model used by northbound interfaces and SDK APIs, see [Commit Parameters](/guides/operation-and-usage/operations/lifecycle-operations#d5e5048).

Such a design allows you to decide whether you will run the same version of NSO on all nodes or not. Since some differences arise between the two options, this document distinguishes a single-version deployment from a multi-version one.

Deployment of an LSA cluster where all the nodes have the same major version of NSO running is called a single version deployment. If the versions are different, then it is a multi-version deployment, since the packages on the CFS node must be managed differently.

The choice between the two deployment options depends on your functional needs. The single version is easier to maintain and is a good starting point but is less flexible. While it is possible to migrate from one to the other, the migration from a single version to a multi-version is typically easier than the other way around. Still, every migration requires some effort, so it is best to pick one approach and stick to it.

You can find working examples of both deployment types in the [examples.ncs/layered-services-architecture/lsa-single-version-deployment](https://github.com/NSO-developer/nso-examples/tree/6.7/layered-services-architecture/lsa-single-version-deployment) and [examples.ncs/layered-services-architecture/lsa-multi-version-deployment](https://github.com/NSO-developer/nso-examples/tree/6.7/layered-services-architecture/lsa-multi-version-deployment) folders, respectively.

### RFS Nodes Setup <a href="#d5e320" id="d5e320"></a>

The type of deployment does not affect the RFS nodes. In general, the RFS nodes act very much like ordinary standalone NSO instances but only support the RFS services.

Configure and set up the lower RFS nodes as you would a standalone node, by making sure the necessary NED and RFS packages are loaded and the managed network devices added. This requires you to have already decided on the distribution of devices to lower RFS nodes. The RFS packages are ordinary service packages.

The only LSA-specific requirement is that these nodes enable NETCONF communication northbound, as this is how the upper CFS node will interact with them. To enable NETCONF northbound, ensure that a configuration similar to the following is present in the `ncs.conf` of every RFS node:

```xml
  <netconf-north-bound>
    <enabled>true</enabled>
    <transport>
      <ssh>
        <enabled>true</enabled>
        <ip>0.0.0.0</ip>
        <port>2022</port>
      </ssh>
    </transport>
  </netconf-north-bound>
```

One thing to note is that you do not need to explicitly enable the commit queue on the RFS nodes, even if you intend to use LSA with the commit queue feature. The upper CFS node is aware of the LSA setup and will propagate the relevant shared commit parameters to the lower RFS nodes automatically.

This also applies to developer-defined commit parameters added by augmenting the shared `tailf-ncs-commit-params` structure. Such augmented parameters can be forwarded to lower RFS nodes together with the built-in parameters.

If you wish to enable the commit queue by default, that is, even for transactions originating on the RFS node (non-LSA), you are strongly encouraged to enable it globally, through the `/devices/global-settings/commit-queue/enabled-by-default` setting on all the RFS nodes and, importantly, the upper CFS node too. Otherwise, you may end up in a situation where only a part of the transaction runs through the commit queue. In that case, the `rollback-on-error` commit queue error option will not work correctly, as it can't roll back the full original transaction but just the part that went through the commit queue. This can result in an inconsistent network state.

### CFS Node Setup <a href="#d5e332" id="d5e332"></a>

Regardless of single or multi-version deployment, the upper CFS node has the lower RFS nodes configured as devices under the `/devices/device` tree. The CFS node communicates with these devices through NETCONF and must have the correct `ned-id` configured for each lower RFS node. The `ned-id` is set under `/devices/device/device-type/netconf/ned-id`, as for any NETCONF device.

The part that is specific to LSA is the actual `ned-id` used. This has to be `ned:lsa-netconf` or a `ned-id` derived from it. What is more, the `ned-id` depends on the deployment type. For a single-version deployment, you can use the `lsa-netconf` value directly. This `ned-id` is built-in (defined in `tailf-ncs-ned.yang`) and available in NSO without any additional packages.

So the configuration for the RFS device in the CFS node would look similar to:

```cli
admin@upper-nso% show devices device | display-level 4
device lower-nso-1 {
    lsa-remote-node lower-nso-1;
    authgroup       default;
    device-type {
        netconf {
            ned-id lsa-netconf;
        }
    }
    state {
        admin-state unlocked;
    }
}
```

Notice the use of the `lsa-remote-node` instead of the `address` (and `port`) as is usually done. This setting identifies the device as a lower-layer LSA node and instructs NSO to use connection information provided under `cluster` configuration.

The value of `lsa-remote-node` references a `cluster remote-node`, such as the following:

```cli
admin@upper-nso% show cluster remote-node
remote-node lower-nso-1 {
    address   127.0.2.1;
    authgroup default;
}
```

In addition to `devices device`, the `authgroup` value is again required here and refers to `cluster authgroup`, not the device one. Both authgroups must be configured correctly for LSA to function.

Having added device and cluster configuration for all RFS nodes, you should update the SSH host keys for both, the `/devices/device` and `/cluster/remote-node` paths. For example:

```cli
admin@upper-nso% request devices device lower-nso-* ssh fetch-host-keys
admin@upper-nso% request cluster remote-node lower-nso-* ssh fetch-host-keys
```

Moreover, the RFS NSO nodes have an extra configuration that may not be visible to the CFS node, resulting in out-of-sync behavior. You are strongly encouraged to set the `out-of-sync-commit-behaviour` value to `accept`, with a command such as:

```cli
admin@upper-nso% set devices device lower-nso-* out-of-sync-commit-behaviour accept
```

At the same time you should also enable the `/cluster/device-notifications`, which will allow the CFS node to receive the forwarded device notifications from the RFS nodes, and `/cluster/commit-queue`, to enable the commit queue support for LSA. Without the latter, you will not be able to use the `commit commit-queue async` command, for example.

If you wish to enable the commit queue by default, you should do so by setting the `/devices/global-settings/commit-queue/enabled-by-default` on the CFS node. Do not use per device or per device group configuration, for the same reason you should avoid it on the RFS nodes.

#### Multi-Version Deployment <a href="#ncs_lsa.lsa_setup.multi_version" id="ncs_lsa.lsa_setup.multi_version"></a>

If you plan a single-version deployment, the preceding steps are sufficient. For a multi-version deployment, on the other hand, there are two additional tasks to perform.

First, you will need to install the correct Cisco-NSO LSA NED package (or packages if you need to support more versions). Each NSO release includes these packages that are specifically tailored for LSA. They are used by the upper CFS node if the lower RFS nodes are running a different version than the CFS node itself. The packages are named `cisco-nso-nc-X.Y` where X.Y are the two most significant numbers of the NSO release (the major version) that the package supports. So, if your RFS nodes are running NSO 5.7.2, for example, you should use `cisco-nso-nc-5.7`.

These packages are found in the `$NCS_DIR/packages/lsa` directory. Each package contains the complete model of the `ncs` namespace for the corresponding NSO version, compiled as an LSA NED. Please always use the `cisco-nso` package included with the NSO version of the upper CFS node and not some older variant (such as the one from the lower RFS node) as it may not work correctly.

Second, installing the cisco-nso LSA NED package will make the corresponding `ned-id` available, such as `cisco-nso-nc-5.7` (`ned-id` matches the package name). Use this `ned-id` for the RFS nodes instead of `lsa-netconf`. For example:

```cli
admin@upper-nso% show devices device | display-level 4
device lower-nso-1 {
    lsa-remote-node lower-nso-1;
    authgroup       default;
    device-type {
        netconf {
            ned-id cisco-nso-nc-5.7;
        }
    }
    state {
        admin-state unlocked;
    }
}
```

This configuration allows the CFS node to communicate with a different NSO version but there are still some limitations. The upper CFS node must have the same or newer version than the managed RFS nodes. For all the currently supported versions of the lower node, the packages can be found in the `$NCS_DIR/packages/lsa` directory, but you may also be able to build an older one yourself.

In case you already have a single-version deployment using the `lsa-netconf` `ned-id'`s, you can use the NED migrate procedure to switch to the new `ned-id` and multi-version deployment.

### Device Compiled RFS Services <a href="#d5e392" id="d5e392"></a>

Besides adding managed lower-layer nodes, the upper-layer node also requires packages for the services. Obviously, you must add the CFS package, which is an ordinary service package, to the CFS node. But you must also provide the device compiled RFS YANG models to allow provisioning of RFSs on the remote RFS nodes.

The process resembles the way you create and compile device YANG models in normal NED packages. The `ncs-make-package` tool provides the `--lsa-netconf-ned` option, where you specify the location of the RFS YANG model and the tool creates a NED package for you. This is a new package that is separate from the RFS package used in the RFS nodes, so you might want to name it differently to avoid confusion. The following text uses the `-ned` suffix.

Usually, you would also provide the `--no-netsim`, `--no-java`, and `--no-python` switches to the invocation, as the package is used with the NETCONF protocol and doesn't need any additional code. The `--no-netsim` option is required because netsim is not supported for these types of packages. For example:

```bash
ncs-make-package --no-netsim --no-java --no-python    \
    --lsa-netconf-ned ./path/to/rfs/src/yang          \
    myrfs-service-ned
```

In this case, there is no explicit `--lsa-lower-nso` option specified and `ncs-make-package` will by default be set up to compile the package for the single version deployment, tied to the `lsa-netconf` `ned-id`. That means the models in the NED can be used with devices that have a `lsa-netconf` `ned-id` configured.

To compile it for the multi-version deployment, which uses a different `ned-id`, you must select the target NSO version with the `--lsa-lower-nso cisco-nso-nc-X.Y` option, for example:

```bash
ncs-make-package --no-netsim --no-java --no-python    \
    --lsa-netconf-ned ./path/to/rfs/src/yang          \
    --lsa-lower-nso cisco-nso-nc-5.7
    myrfs-service-ned
```

Depending on the RFS model, the package may fail to compile, even though the model compiles fine as a service. A typical error would indicate some node from a module, such as `tailf-ncs`, is not found. The reason is that the original RFS service YANG model has dependencies on other YANG models that are not included in the compilation process.

One solution to this problem is to remove the dependencies in the YANG model before compilation. Normally this can be solved by changing the datatype in the NED compiled copy of the YANG model, for example from `leafref` or `instance-identifier` to string. This is only needed for the NED compiled copy, the lower RFS node YANG model can remain the same. There will then be an implicit conversion between types, at runtime, in the communication between the upper CFS node and the lower RFS node.

An alternate solution, if you are doing a single version deployment and there are dependencies on the `tailf-ncs` namespace, is to switch to a multi-version deployment because the `cisco-nso` package includes this namespace (device compiled). Here, the NSO versions match but you are still using the `cisco-nso-nc-X.Y` `ned-id` and have to follow the instructions for the multi-version deployment.

Once you have both, the CFS and device-compiled RFS service packages are ready; add them to the CFS node, then invoke a `sync-from` action to complete the setup process.

### Example Walkthrough <a href="#d5e425" id="d5e425"></a>

You can see all the required setup steps for a single version deployment performed in the example [examples.ncs/layered-services-architecture/lsa-single-version-deployment](https://github.com/NSO-developer/nso-examples/tree/6.7/layered-services-architecture/lsa-single-version-deployment) and the [examples.ncs/layered-services-architecture/lsa-multi-version-deployment](https://github.com/NSO-developer/nso-examples/tree/6.7/layered-services-architecture/lsa-multi-version-deployment) has the steps for the multi-version one. The two are quite similar but the multi-version deployment has additional steps, so it is the one described here.

First, build the example for manual setup.

```bash
$ make clean manual
$ make start-manual
$ make cli-upper-nso
```

Then configure the nodes in the cluster. This is needed so that the upper CFS node can receive notifications from the lower RFS node and prepare the upper CFS node to be used with the commit queue.

```cli
> configure

% set cluster device-notifications enabled
% set cluster remote-node lower-nso-1 authgroup default username admin
% set cluster remote-node lower-nso-1 address 127.0.0.1 port 2023
% set cluster remote-node lower-nso-2 authgroup default username admin
% set cluster remote-node lower-nso-2 address 127.0.0.1 port 2024
% set cluster commit-queue enabled
% commit
% request cluster remote-node lower-nso-* ssh fetch-host-keys
```

To be able to handle the lower NSO node as an LSA node, the correct version of the `cisco-nso-nc` package needs to be installed. In this example, 5.4 is used.

Create a link to the `cisco-nso` package in the packages directory of the upper CFS node:

```bash
$ ln -sf ${NCS_DIR}/packages/lsa/cisco-nso-nc-5.4 upper-nso/packages
```

Reload the packages:

```cli
% exit
> request packages reload

e>>> System upgrade is starting.
>>> Sessions in configure mode must exit to operational mode.
>>> No configuration changes can be performed until upgrade has completed.
>>> System upgrade has completed successfully.
reload-result {
    package cisco-nso-nc-5.4
    result true
}
```

Now when the `cisco-nso-nc` package is in place, configure the two lower NSO nodes and `sync-from` them:

```cli
> configure
Entering configuration mode private

% set devices device lower-nso-1 device-type netconf ned-id cisco-nso-nc-5.4
% set devices device lower-nso-1 authgroup default
% set devices device lower-nso-1 lsa-remote-node lower-nso-1
% set devices device lower-nso-1 state admin-state unlocked
% set devices device lower-nso-2 device-type netconf ned-id cisco-nso-nc-5.4
% set devices device lower-nso-2 authgroup default
% set devices device lower-nso-2 lsa-remote-node lower-nso-2
% set devices device lower-nso-2 state admin-state unlocked

% commit
Commit complete.

% request devices fetch-ssh-host-keys
fetch-result {
    device lower-nso-1
    result updated
    fingerprint {
        algorithm ssh-ed25519
        value 4a:c6:5d:91:6d:4a:69:7a:4e:0d:dc:4e:51:51:ee:e2
    }
}
fetch-result {
    device lower-nso-2
    result updated
    fingerprint {
        algorithm ssh-ed25519
        value 4a:c6:5d:91:6d:4a:69:7a:4e:0d:dc:4e:51:51:ee:e2
    }
}

% request devices sync-from
sync-result {
    device lower-nso-1
    result true
}
sync-result {
    device lower-nso-2
    result true
}
```

Now, for example, the configured devices of the lower nodes can be viewed:

```cli
% show devices device config devices device | display xpath | display-level 5

/devices/device[name='lower-nso-1']/config/ncs:devices/device[name='ex0']
/devices/device[name='lower-nso-1']/config/ncs:devices/device[name='ex1']
/devices/device[name='lower-nso-1']/config/ncs:devices/device[name='ex2']
/devices/device[name='lower-nso-2']/config/ncs:devices/device[name='ex3']
/devices/device[name='lower-nso-2']/config/ncs:devices/device[name='ex4']
/devices/device[name='lower-nso-2']/config/ncs:devices/device[name='ex5']
```

Or, alarms inspected:

```cli
% run show devices device lower-nso-1 live-status alarms summary

live-status alarms summary indeterminates 0
live-status alarms summary criticals 0
live-status alarms summary majors 0
live-status alarms summary minors 0
live-status alarms summary warnings 0
```

Now, create a netconf package on the upper CFS node which can be used towards the `rfs-vlan` service on the lower RFS node, in the shell terminal window, do the following:

```bash
$ ncs-make-package --no-netsim --no-java --no-python                \
    --lsa-netconf-ned package-store/rfs-vlan/src/yang               \
    --lsa-lower-nso cisco-nso-nc-5.4                                \
    --package-version 5.4 --dest upper-nso/packages/rfs-vlan-nc-5.4 \
    --build rfs-vlan-nc-5.4
```

The created NED is an `lsa-netconf-ned` based on the YANG files of the `rfs-vlan` service:

```
--lsa-netconf-ned package-store/rfs-vlan/src/yang
```

The version of the NED reflects the version of the nso on the lower node:

```
--package-version 5.4
```

The package will be generated in the packages directory of the upper NSO CFS node:

```
--dest upper-nso/packages/rfs-vlan-nc-5.4
```

And, the name of the package will be:

```
rfs-vlan-nc-5.4
```

Install the `cfs-vlan` service on the upper CFS node. In the shell terminal window, do the following:

```bash
$ ln -sf ../../package-store/cfs-vlan     upper-nso/packages
```

Reload the packages once more to get the `cfs-vlan` package. In the CLI terminal window, do the following:

```cli
% exit

> request packages reload

>>> System upgrade is starting.
>>> Sessions in configure mode must exit to operational mode.
>>> No configuration changes can be performed until upgrade has completed.
>>> System upgrade has completed successfully.
reload-result {
    package cfs-vlan
    result true
}
reload-result {
    package cisco-nso-nc-5.4
    result true
}
reload-result {
    package rfs-vlan-nc-5.4
    result true
}

> configure
Entering configuration mode private
```

Now, when all packages are in place a `cfs-vlan` service can be configured. The `cfs-vlan` service will dispatch service data to the right lower RFS node depending on the device names used in the service.

In the CLI terminal window, verify the service:

```cli
% set cfs-vlan v1 a-router ex0 z-router ex5 iface eth3 unit 3 vid 77

% commit dry-run
.....
    local-node {
        data  devices {
                  device lower-nso-1 {
                      config {
                          services {
             +                vlan v1 {
             +                    router ex0;
             +                    iface eth3;
             +                    unit 3;
             +                    vid 77;
             +                    description "Interface owned by CFS: v1";
             +                }
                          }
                      }
                  }
                  device lower-nso-2 {
                      config {
                          services {
             +                vlan v1 {
             +                    router ex5;
             +                    iface eth3;
             +                    unit 3;
             +                    vid 77;
             +                    description "Interface owned by CFS: v1";
             +                }
                          }
                      }
                  }
              }
.....
```

As `ex0` resides on `lower-nso-1` that part of the configuration goes there and the `ex5` part goes to `lower-nso-2`.

### Migration and Upgrades <a href="#d5e478" id="d5e478"></a>

Since an LSA deployment consists of multiple NSO nodes (or HA pairs of nodes), each can be upgraded to a newer NSO version separately. While that offers a lot of flexibility, it also makes upgrades more complex in many cases. For example, performing a major version upgrade on the upper CFS node only will make the deployment Multi-Version even if it was Single-Version before the upgrade, requiring additional action on your part.

In general, staying with the Single-Version Deployment is the simplest option and does not require any further LSA-specific upgrade action (except perhaps recompiling the packages). However, the main downside is that, at least for a major upgrade, you must upgrade all the nodes at the same time (otherwise, you no longer have a Single-Version Deployment).

If that is not feasible, the solution is to run a Multi-Version Deployment. Along with all of the requirements, the section [Multi-Version Deployment](#ncs_lsa.lsa_setup.multi_version) describes a major difference from the Single Version variant: the upper CFS node uses a version-specific `cisco-nso-nc-X.Y` NED ID to refer to lower RFS nodes. That means, if you switch to a Multi-Version Deployment, or perform a major upgrade of the lower-layer RFS node, the `ned-id` should change accordingly. However, do not change it directly but follow the correct NED upgrade procedure described in the section called [NED Migration](https://nso-docs.cisco.com/guides/administration/advanced-topics/pages/vGfK2qplvC1wu0670OgS#sec.ned_migration). Briefly, the procedure consists of these steps:

1. Keep the currently configured ned-id for an RFS device and the corresponding packages. If upgrading the CFS node, you will need to recompile the packages for the new NSO version.
2. Compile and load the packages that are device-compiled with the new `ned-id`, alongside the old packages.
3. Use the `migrate` action on a device to switch over to the new `ned-id`.

The procedure requires you to have two versions of the device-compiled RFS service packages loaded in the upper CFS node when calling the `migrate` action: one version compiled by referencing the old (current) NED ID and the other one by referencing the new (target) NED ID.

To illustrate, suppose you currently have an upper-layer and a lower-layer node both running NSO 5.4. The nodes were set up as described in the Single-Version Deployment option, with the upper CFS node using the `tailf-ncs-ned:lsa-netconf` NED ID for the lower-layer RFS node. The CFS node also uses the `rfs-vlan-ned` NED package for the `rfs-vlan` service.

Now you wish to upgrade the CFS node to NSO 5.7 but keep the RFS node on the existing version 5.4. Before upgrading the CFS node, you create a backup and recompile the `rfs-vlan-ned` package for NSO 5.7. Note that the package references the `lsa-netconf` `ned-id`, which is the `ned-id` configured for the RFS device in the CFS node's CDB. Then, you perform the CFS node upgrade as usual.

At this point the CFS node is running the new, 5.7 version and the RFS node is running 5.4. Since you now have a Multi-Version Deployment, you should migrate to the correct `ned-id` as well. Therefore, you prepare the `rfs-vlan-nc-5.4` package, as described in the Multi-Version Deployment option, compile the package, and load it into the CFS node. Thanks to the NSO CDM feature, both packages, `rfs-vlan-nc-5.4` and `rfs-vlan-ned`, can be used at the same time.

With the packages ready, you execute the `devices device lower-nso-1 migrate new-ned-id cisco-nso-nc-5.4` command on the CFS node. The command configures the RFS device entry on CFS to use the new `cisco-nso-nc-5.4 ned-id`, as well as migrates the device configuration and service meta-data to the new model. Having completed the upgrade, you can now remove the `rfs-vlan-ned` if you wish.

Later on, you may decide to upgrade the RFS node to NSO 5.6. Again, you prepare the new `rfs-vlan-nc-5.6` package for the CFS node in a similar way as before, now using the `cisco-nso-nc-5.6` ned-id instead of `cisco-nso-nc-5.4`. Next, you perform the RFS node upgrade to 5.6 and finally migrate the RFS device on the CFS node to the `cisco-nso-nc-5.6 ned-id`, with the `migrate` action.

Likewise, you can return to the Single-Version Deployment, by upgrading the RFS node to the NSO 5.7, reusing the old, or preparing anew, the `rfs-vlan-ned` package and migrating to the `lsa-netconf ned-id`.

All these `ned-id` changes stem from the fact that the upper-layer CFS node treats the lower-layer RFS node as a managed device, requiring the correct model, just like it does for any other device type. For the same reason, maintenance (bug fix or patch) NSO upgrades do not result in a changed `ned-id`, so for those, no migration is necessary.

The [NSO example set](https://github.com/NSO-developer/nso-examples/tree/6.7/layered-services-architecture) illustrates different aspects of LSA deployment including working with single- and multi-version deployments.

### User Authorization Passthrough

In LSA, northbound users are authenticated on the CFS, and the request is re-authenticated on the RFS using either a system user or user/pass passthrough.

For token-based authentication using external auth/package auth, this becomes a problem as the user and password are not expected to be locally provisioned and hence cannot be used for authentication towards the RFS, which leaves the option of a system user.

Using a system user has two major limitations:

* Auditing on the RFS becomes hard, as system sessions are not logged in the `audit.log`.
* Device-level RBAC becomes challenging as the devices reside in the RFS and the user information is lost.

To handle this scenario, one can enable the passthrough of the user name and its groups to lower layer nodes to allow the session on the RFS to assume the same user as used on the CFS (similar to use of "sudo"). This will allow for the use of a system user between the CFS and RFS while allowing for auditing and RBAC on the RFS using the locally authenticated user on the CFS.

On the CFS node, create an authgroup under `/devices/authgroups/group` with the `/devices/authgroups/group/{umap,default-map}/passthrough` empty leaf set, then select this authgroup on the configured RFS nodes by setting the `/devices/device/authgroup` leaf. When the passthrough leaf is set and a user (e.g., alice) on the CFS node connects to an RFS node, she will authenticate using the credentials specified in the `/devices/device/authgroup` authgroup (e.g., `lsa_passthrough_user` : `ahVaesai8Ahn0AiW`). Once the authentication completes successfully, the user `lsa_passthrough_user` changes into alice on the RFS node.

{% code overflow="wrap" %}

```bash
admin@cfs% set devices authgroups group rfs-east default-map remote-name lsa_passthrough_user remote-password ahVaesai8Ahn0AiW passthrough
admin@cfs% set devices device rfs1 authgroup rfs-east
admin@cfs% set devices device rfs2 authgroup rfs-east
admin@cfs% commit
```

{% endcode %}

On the RFS node, configure the mapping of permitted users in the `/cluster/global-settings/passthrough/permit` list. The key of the permit list specifies what user may change into a different user. The different possible users to change into are specified by the `as-user` leaf-list, and the `as-group` leaf-list specifies valid groups. The user will end up with the intersection of groups in the user session on the CFS and the groups specified by the `as-group` leaf-list. Only users in the permit list will be allowed to change into the users set in the permit list elements `as-user` list.

{% code overflow="wrap" %}

```bash
admin@rfs1% set cluster global-settings passthrough permit lsa_passthrough_user as-user [ alice bob carol ] as-group [ oper dev ]
admin@rfs1% commit
```

{% endcode %}

To allow the passthrough user to change into any user, set the `as-any-user` leaf, or for any group, set the `as-any-group` leaf. Use this with care as setting these leafs will allow the `lsa_passthrough_user` to elevate privileges by changing to `user admin` / `group admin`.

{% code overflow="wrap" %}

```bash
admin@rfs1% set cluster global-settings passthrough permit lsa_passthrough_user as-any-user as-any-group
admin@rfs1% commit
```

{% endcode %}


# Get Started

Operate and use NSO.

## CLI

<table data-view="cards"><thead><tr><th></th><th></th><th data-hidden data-card-target data-type="content-ref"></th></tr></thead><tbody><tr><td><strong>Introduction to NSO CLI</strong></td><td>Familiarize yourself with the NSO CLI.</td><td><a href="/pages/oYfaaJQb8GzGog8TcxF7">/pages/oYfaaJQb8GzGog8TcxF7</a></td></tr><tr><td><strong>CLI Commands</strong></td><td>List of available CLI commands.</td><td><a href="/pages/cuf7W6RI3em3mzfy0w8b">/pages/cuf7W6RI3em3mzfy0w8b</a></td></tr></tbody></table>

## Web UI

<table data-view="cards"><thead><tr><th></th><th></th><th data-hidden data-card-target data-type="content-ref"></th></tr></thead><tbody><tr><td><strong>Home</strong></td><td>Intro to Web UI home page and extension packages.</td><td><a href="/pages/rR0EFcUovKimyyUwmbNb">/pages/rR0EFcUovKimyyUwmbNb</a></td></tr><tr><td><strong>Devices</strong></td><td>Manage devices and device groups in the Web UI.</td><td><a href="/pages/4AoIvh1N3BIPiFvOxyrX">/pages/4AoIvh1N3BIPiFvOxyrX</a></td></tr><tr><td><strong>Services</strong></td><td>Manage NSO services using the Web UI.</td><td><a href="/pages/gIg4G9X6fNBPI6oD2wE0">/pages/gIg4G9X6fNBPI6oD2wE0</a></td></tr><tr><td><strong>Config Editor</strong></td><td>Traverse and configure NSO using the YANG model.</td><td><a href="/pages/sb4VMzAGSNNHxwrW73FA">/pages/sb4VMzAGSNNHxwrW73FA</a></td></tr><tr><td><strong>Tools</strong></td><td>Tools to perform specialized tasks on NSO.</td><td><a href="/pages/YbD21HK8ZwE5HK5p6bay">/pages/YbD21HK8ZwE5HK5p6bay</a></td></tr></tbody></table>

## Operations

<table data-view="cards" data-full-width="false"><thead><tr><th></th><th></th><th data-hidden data-card-target data-type="content-ref"></th></tr></thead><tbody><tr><td><strong>Basic Operations</strong></td><td>Learn NSO's basic command line operations.</td><td><a href="/pages/Vd92i4ejI48e7azpmNEQ">/pages/Vd92i4ejI48e7azpmNEQ</a></td></tr><tr><td><strong>NEDs and Adding Devices</strong></td><td>Learn about NEDs and how to add devices in NSO.</td><td><a href="/pages/L71OxJvXZB7sa8kUVNPH">/pages/L71OxJvXZB7sa8kUVNPH</a></td></tr><tr><td><strong>Manage Network Services</strong></td><td>Manage network services and configure life cycle ops.</td><td><a href="/pages/YcLFVRIXkCdYCOpZhoCw">/pages/YcLFVRIXkCdYCOpZhoCw</a></td></tr><tr><td><strong>Device Manager</strong></td><td>Explore device management, related ops.</td><td><a href="/pages/auKQMOAF2p1jiGYJBweP">/pages/auKQMOAF2p1jiGYJBweP</a></td></tr><tr><td><strong>Out-of-band Interoperation</strong></td><td>Manage out-of-band changes.</td><td><a href="/pages/d9u5OLpEXHQxHLv1k8XA">/pages/d9u5OLpEXHQxHLv1k8XA</a></td></tr><tr><td><strong>SSH Key Management</strong></td><td>Use NSO as an SSH server or a client.</td><td><a href="/pages/DU9Fz2hQ6utFrKf0dNdW">/pages/DU9Fz2hQ6utFrKf0dNdW</a></td></tr><tr><td><strong>Alarm Manager</strong></td><td>Explore NSO alarm management &#x26; related ops.</td><td><a href="/pages/unFvfPfmgshmNTEr5OYr">/pages/unFvfPfmgshmNTEr5OYr</a></td></tr><tr><td><strong>Plug-and-Play Scripting</strong></td><td>Use scripting to add new functionality to NSO.</td><td><a href="/pages/6IC1MdDxBMr17RnXFMMS">/pages/6IC1MdDxBMr17RnXFMMS</a></td></tr><tr><td><strong>Compliance Reporting</strong></td><td>Implement network compliance in NSO.</td><td><a href="/pages/HgnWDvRBZucvvFH137Dx">/pages/HgnWDvRBZucvvFH137Dx</a></td></tr><tr><td><strong>Listing Packages</strong></td><td>View and list NSO packages.</td><td><a href="/pages/nCFFm27jQrfiC0pBcJ4d">/pages/nCFFm27jQrfiC0pBcJ4d</a></td></tr><tr><td><strong>Lifecycle Operations</strong></td><td>Manipulate existing services and devices.</td><td><a href="/pages/l27TLK7q4SdzGuanVM1Q">/pages/l27TLK7q4SdzGuanVM1Q</a></td></tr><tr><td><strong>Network Simulator</strong></td><td>Simulate a network to be managed by NSO.</td><td><a href="/pages/fyUNRY3AxRen4qJ6kzXK">/pages/fyUNRY3AxRen4qJ6kzXK</a></td></tr></tbody></table>


# CLI

Get started with NSO CLI.


# Introduction to NSO CLI

Get started with the NSO CLI.

The NSO CLI (command line interface) provides a unified CLI towards the complete network. The NSO CLI is a northbound interface to the NSO representation of the network devices and network services. Do not confuse this with a cut-through CLI that reaches the devices directly. Although the network might be a mix of vendors and device interfaces with different CLI flavors, NSO provides one northbound CLI.

Starting the CLI:

```bash
$> ncs_cli -C -u admin
```

{% hint style="info" %}
Note the use of the `-u` parameter which tells NSO which user to authenticate towards NSO. It is a common mistake to forget this. This user must be configured in NSO AAA (Authentication, Authorization, and Accounting).
{% endhint %}

Like many CLI's there is an operational mode and a configuration mode. Show commands display different data in those modes. A show in configuration mode displays network configuration data from the NSO configuration database, the CDB. Show in operational mode shows live values from the devices and any operational data stored in the CDB. The CLI starts in operational mode. Note that different prompts are used for the modes (these can be changed in `ncs.conf` configuration file).

NSO organizes all managed devices as a list of devices. The path to a specific device is `devices device DEVICE-NAME`. The CLI sequence below does the following:

1. Show operational data for all devices: fetches operational data from the network devices like interface statistics, and also operational data that is maintained by NSO like alarm counters.
2. Move to configuration mode. Show configuration data for all devices: In this example, this is done before the configuration from the real devices has been loaded in the network to NSO. At this point, only the NSO-configured data like IP Address, port, etc. are shown.

Show device operational data and configuration data:

```bash
admin@ncs# show devices device
devices device ce0
 ...
 alarm-summary indeterminates 0
 alarm-summary criticals 0
 alarm-summary majors 0
 alarm-summary minors 0
 alarm-summary warnings 0
devices device ce1
 ...
admin@ncs# config
Entering configuration mode terminal
admin@ncs(config)# show full-configuration devices device
devices device ce0
 address   127.0.0.1
 port      10022
 ssh host-key ssh-dss
 ...
!
devices device ce1
 ...
!
...
```

It can be annoying to move between modes to display configuration data and operational data. The CLI has ways around this.

Show config data in operational mode and vice versa:

```bash
admin@ncs# show running-config devices device
admin@ncs(config)# do show running-config devices device
```

Look at the device configuration above, no configuration relates to the actual configuration on the devices. To boot-strap NSO and discover the device configuration, it is possible to perform an action to synchronize NSO from the devices, `devices sync-from`. This reads the configuration over available device interfaces and populates the NSO data store with the corresponding configuration. The device-specific configuration is populated below the device's entry in the configuration tree and can be listed specifically.

Perform the action to synchronize from devices:

```bash
admin@ncs(config)# devices sync-from
sync-result {
    device ce0
    result true
}
sync-result {
    device ce1
    result true
}
...
```

Display the device configuration after the synchronization:

```bash
admin@ncs(config)# show full-configuration devices device ce0 config
devices device ce0
 config
  no ios:service pad
  no ios:ip domain-lookup
  no ios:ip http secure-server
  ios:ip source-route
  ios:interface GigabitEthernet0/1
  exit
  ios:interface GigabitEthernet0/10
  exit
  ios:interface GigabitEthernet0/11
  exit
  ios:interface GigabitEthernet0/12
  exit
  ios:interface GigabitEthernet0/13
  exit
  ...
 !
!
...
```

NSO provides a network CLI in two different styles (selectable by the user): J-style and C-style. The CLI is automatically rendered using the data models described by the YANG files. There are three distinctly different types of YANG files, the built-in NSO models describing the device manager and the service manager, models imported from the managed devices, and finally service models. Regardless of model type, the NSO CLI seamlessly handles all models as a whole.

This creates an auto-generated CLI, without any extra effort, except the design of our YANG files. The auto-generated CLI supports the following features:

* Unified CLI across the complete network, devices, and network services.
* Command line history and command line editor.
* Tab completion for the content of the configuration database.
* Monitoring and inspecting log files.
* Inspecting the system configuration and system state.
* Copying and comparing different configurations, for example, between two interfaces or two devices.
* Configuring common settings across a range of devices.

The CLI contains commands for manipulating the network configuration.

An alias provides a shortcut for a complex command.

Alias expansion is performed when a command line is entered. Aliases are part of the configuration and are manipulated accordingly. This is done by manipulating the nodes in the alias configuration tree.

Actions in the YANG files are mapped into actual commands. In J-style CLI actions are mapped to the `request` commands.

Even though the auto-generated CLI is fully functional it can be customized and extended in numerous ways:

* Built-in commands can be moved, hidden, deleted, reordered, and extended.
* Confirmation prompts can be added to built-in commands.
* New commands can be implemented using the Java API, ordinary executables, and shell scripts.
* New commands can be mounted freely in the existing command hierarchy.
* The built-in tab completion mechanism can be overridden using user-defined callbacks.
* New command hierarchies can be created.
* A command timeout can be added, both a global timeout for all commands and command-specific timeouts.
* Actions and parts of the configuration tree can be hidden and can later be made visible when the user enters a password.

How to customize and extend the auto-generated CLI is described in [Plug-and-play Scripting](/guides/operation-and-usage/operations/plug-and-play-scripting).

## CLI Modes <a href="#d5e1216" id="d5e1216"></a>

The CLI is entirely data model-driven. The YANG model(s) defines a hierarchy of configuration elements. The CLI follows this tree. The NSO CLI provides various commands for configuring and monitoring software, hardware, and network connectivity of managed devices.

The CLI supports two modes:

* **Operational** **mode**: For monitoring the state of the NSO node.
* **Configure** **mode**: For changing the state of the network.

The prompt indicates which mode the CLI is in. When moving from operational mode to configure mode using the `configure` command, the prompt is changed from `host#` to `host(config)#`. The prompts can be configured using the `c-prompt1` and `c-prompt2` settings in the `ncs.conf` file.

For example:

```bash
admin@ncs# configure
Entering configuration mode terminal
admin@ncs(config)#
```

{% tabs %}
{% tab title="Operational Mode" %}
The operational mode is the initial mode after successful login to the CLI. It is primarily used for viewing the system status, controlling the CLI environment, monitoring and troubleshooting network connectivity, and initiating the configure mode.

A list of base commands available in the operational mode is listed below in the [Operational Mode Commands](/guides/operation-and-usage/cli/cli-commands#d5e1943) section. Additional commands are rendered from the loaded YANG files.
{% endtab %}

{% tab title="Configure Mode" %}
The configure mode can be initiated by entering the `configure` command in operational mode. All changes to the network configuration are done to a copy of the active configuration. These changes do not take effect until a successful `commit` or `commit confirm` command is entered.

A list of base commands available in `configure` mode is listed below in the [Configure Mode Commands](/guides/operation-and-usage/cli/cli-commands#d5e2199) section. Additional commands are rendered from the loaded YANG files.

{% hint style="info" %}
When using the `config` mode to enter/set passwords, you may face issues if you are using special characters in your password (e.g., `!`, `""`, `\`, etc.). Some characters are automatically escaped by the CLI, while others require manual escaping. Therefore, the recommendation is to always enclose your password in double quotes `" "` and avoid using quotes `"` and backslash `\` characters in your password. If you prefer including quotes and backslash in your password, remember to manually escape them, as shown in the example below:

```cli
admin@ncs(config)# devices authgroups group default umap
admin remote-name admin remote-password "admin\"admin"
```

{% endhint %}
{% endtab %}
{% endtabs %}

## Starting the CLI <a href="#d5e1253" id="d5e1253"></a>

The CLI is started using the `ncs_cli` program. It can be used as a login program (replacing the shell for a user), started manually once the user has logged in, or used in scripts for performing CLI operations.

In some NSO installations, ordinary users would have the `ncs_cli` program as a login shell, and the root user would have to log in and then start the CLI using `ncs_cli`, whereas in others, the `ncs_cli` can be invoked freely as a normal shell command.

The `ncs_cli` program supports a range of options, primarily intended for debugging and development purposes (see description below).

The `ncs_cli` program can also be used for batch processing of CLI commands, either by storing the commands in a file and running `ncs_cli` on the file, or by having the following line at the top of the file (with the location of the program modified appropriately):

```
#!/bin/ncs_cli
```

When the CLI is run non-interactively it will terminate at the first error and will only show the output of the commands executed. It will not output the prompt or echo the commands. This is the same behavior as for shell scripts.

To run a script non-interactively, such as a script or through a pipe, and still produce prompts and echo commands, use the `--interactive` option.

### Command Line Options

```bash
ncs_cli --help
Usage: ncs_cli [options] [file]
Options:
--help, -h            display this help
--host, -H <host>     current host name (used in prompt)
--address, -A <addr>  cli address to connect to
--port, -P <port>     cli port to connect to
  < ... output omitted ... >
```

<table data-full-width="false"><thead><tr><th>Command</th><th>Description</th></tr></thead><tbody><tr><td><code>-h</code>, <code>--help</code></td><td>Display help text.</td></tr><tr><td><code>-H</code>, <code>--host</code> <em><code>HostName</code></em></td><td>Gives the name of the current host. The <code>ncs_cli</code> program will use the value of the system call <code>gethostbyname()</code> by default. The hostname is used in the CLI prompt.</td></tr><tr><td><code>-A</code>, <code>--address</code> <em><code>Address</code></em></td><td>TCP address to connect to. When TCP IPC is used, the default is 127.0.0.1. This can be controlled by either this flag or the UNIX environment variable <code>NCS_IPC_ADDR</code>. The <code>-A</code> flag takes precedence.</td></tr><tr><td><code>-P</code>, <code>--port</code> <em><code>PortNumber</code></em></td><td>CLI port to connect to when using TCP IPC. This can be controlled by either this flag, or the UNIX environment variable <code>NCS_IPC_PORT</code>. The <code>-P</code> flag takes precedence.</td></tr><tr><td><code>-S</code></td><td>Path of the UNIX domain socket to connect to for Local IPC, used in place of the TCP address and port. This can be controlled by either this flag or the UNIX environment variable <code>NCS_IPC_PATH</code>. The <code>-S</code> flag takes precedence.</td></tr><tr><td><code>-c</code>, <code>--cwd</code> <em><code>Directory</code></em></td><td>The current working directory (CWD) for the user once in the CLI. All file references from the CLI will be relative to the CWD. By default, the value will be the actual CWD where <code>ncs_cli</code> is invoked.</td></tr><tr><td><code>-p</code>, <code>--proto</code> <code>ssh</code> | <code>tcp</code> | <code>console</code></td><td>The protocol the user is using to connect. This value is used in the audit logs. Defaults to <code>ssh</code> if <code>SSH_CONNECTION</code> environment variable is set; <code>console</code> otherwise.</td></tr><tr><td><code>-i</code>, <code>--ip</code> <em><code>IpAddress</code></em> | <em><code>IpAddress/Port</code></em></td><td>The IP (or IP address and port) which NSO reports that the user is connecting from. This value is used in the audit logs. Defaults to the information in the <code>SSH_CONNECTION</code> environment variable if set, 127.0.0.1 otherwise.</td></tr><tr><td><code>-v</code>, <code>--verbose</code></td><td>Produce additional output about the execution of the command, in particular during the initial handshake phase.</td></tr><tr><td><code>-n</code>, <code>--interactive</code></td><td>Force the CLI to echo prompts and commands. Useful when <code>ncs_cli</code> auto-detects it is not running in a terminal, e.g. when executing as a script, reading input from a file, or through a pipe.</td></tr><tr><td><code>-N</code>, <code>--noninteractive</code></td><td>Force the CLI to only show the output of the commands executed. Do not output the prompt or echo the commands, much like a shell does for a shell script.</td></tr><tr><td><code>-s</code>, <code>--stop-on-error</code></td><td>Force the CLI to terminate at the first error and use a non-zero exit code.</td></tr><tr><td><code>-E</code>, <code>--escape-char</code> <em><code>C</code></em></td><td>A special character that forcefully terminates the CLI when repeated three times in a row. Defaults to control underscore (Ctrl-_).</td></tr><tr><td><code>-J</code>, <code>-C</code></td><td>This flag sets the mode of the CLI. <code>-J</code> is Juniper style CLI, <code>-C</code> is Cisco XR style CLI.</td></tr><tr><td><code>-u</code>, <code>--user</code> <em><code>User</code></em></td><td>The username of the connecting user. Used for access control and group assignment in NSO (if the group mapping is kept in NSO). The default is to use the login name of the user.</td></tr><tr><td><code>-g</code>, <code>--groups</code> <em><code>GroupList</code></em></td><td>A comma-separated list of groups the connecting user is a member of. Used for access control by the AAA system in NSO to authorize data and command access. Defaults to the UNIX groups that the user belongs to, i.e., the same as the <code>groups</code> shell command returns.</td></tr><tr><td><code>-U</code>, <code>--uid</code> <em><code>Uid</code></em></td><td>The numeric user ID the user shall have. Used for executing OS commands on behalf of the user, when checking file access permissions, and when creating files. Defaults to the effective user ID (euid) in use for running the command. Note that NSO needs to run as root for this to work properly.</td></tr><tr><td><code>-G</code>, <code>--gid</code> <em><code>Gid</code></em></td><td>The numeric group ID of the user shall have. Used for executing OS commands on behalf of the user, when checking file access permissions, and when creating files. Defaults to the effective group ID (egid) in use for running the command. Note that NSO needs to run as root for this to work properly.</td></tr><tr><td><code>-D</code>, <code>--gids</code> <em><code>GidList</code></em></td><td>A comma-separated list of supplementary numeric group IDs the user shall have. Used for executing OS commands on behalf of the user and when checking file access permissions. Defaults to the supplementary UNIX group IDs in use for running the command. Note that NSO needs to run as root for this to work properly.</td></tr><tr><td><code>-a</code>, <code>--noaaa</code></td><td>Completely disables all AAA checks for this CLI. This can be used as a disaster recovery mechanism if the AAA rules in NSO have somehow become corrupted.</td></tr><tr><td><code>-O</code>, <code>--opaque</code> <em><code>Opaque</code></em></td><td>Pass an opaque string to NSO. The string is not interpreted by NSO, only made available to application code. See built-in variables in <a href="/pages/nC4ZtV8G20z08QfSWaCP">clispec(5)</a> and <code>maapi_get_user_session_opaque()</code> in <a href="/pages/J4Y2hFhkdXndfOu1Jz29">confd_lib_maapi(3)</a>. The string can be given either via this flag or via the UNIX environment variable <code>NCS_CLI_OPAQUE</code>. The <code>-O</code> flag takes precedence.</td></tr></tbody></table>

For `clispec(5)` and `confd_lib_maapi(3)` refer to [Manual Pages](/guides/resources/man).

### CLI Styles <a href="#d5e1366" id="d5e1366"></a>

The CLI comes in two flavors: C-Style (Cisco XR style) and the J-style. It is possible to choose one specifically or switch between them.

{% tabs %}
{% tab title="C-Style" %}
Starting the CLI (C-style, Cisco XR style):

```bash
$> ncs_cli -C -u admin
```

{% endtab %}

{% tab title="J-Style" %}
Starting the CLI (J-style):

```bash
$> ncs_cli -J -u admin
```

{% endtab %}
{% endtabs %}

It is possible to interactively switch between these styles while inside the CLI using the builtin `switch` command:

```bash
admin@ncs# switch cli
```

C-style is mainly used throughout the documentation for examples etc., except when otherwise stated.

### **Starting the CLI in an Overloaded System** <a href="#d5e1383" id="d5e1383"></a>

If the number of ongoing sessions has reached the configured system limit, no more CLI sessions will be allowed until one of the existing sessions has been terminated.

This makes it impossible to get into the system — a situation that may not be acceptable. The CLI therefore has a mechanism for handling this problem. When the CLI detects that the session limit has been reached, it will check if the new user has the privileges to execute the `logout` command. If the user does, it will display a list of the current user sessions in NSO and ask the user if one of the sessions should be terminated to make room for the new session.

## Modifying the Configuration <a href="#d5e1387" id="d5e1387"></a>

Once NSO is synchronized with the devices' configuration, done by using the `devices sync-from` command, it is possible to modify the devices. The CLI is used to modify the NSO representation of the device configuration and then committed as a transaction to the network.

As an example, to change the speed setting on the interface GigabitEthernet0/1 across several devices:

```bash
admin@ncs(config)# devices device ce0..1 config ios:interface GigabitEthernet0/1 speed auto
admin@ncs(config-if)# top
admin@ncs(config)# show configuration
devices device ce0
 config
  ios:interface GigabitEthernet0/1
   speed auto
  exit
 !
!
devices device ce1
 config
  ios:interface GigabitEthernet0/1
   speed auto
  exit
 !
!
admin@ncs(config)# commit ?
Possible completions:
  and-quit               Exit configuration mode
  check                  Validate configuration
  comment                Add a commit comment
  commit-queue           Commit through commit queue
  label                  Add a commit label
  no-confirm             No confirm
  no-networking          Send nothing to the devices
  no-out-of-sync-check   Commit even if out of sync
  no-overwrite           Do not overwrite modified data on the device
  no-revision-drop       Fail if device has too old data model
  save-running           Save running to file
  ---
  dry-run                Show the diff but do not perform commit
  [<cr>
admin@ncs(config)# commit
Commit complete.
```

Note the availability of commit flags. These CLI flags correspond to the shared commit-parameter model described in [Commit Parameters](/guides/operation-and-usage/operations/lifecycle-operations#d5e5048). See also the [examples.ncs/northbound-interfaces/commit-parameters](https://github.com/NSO-developer/nso-examples/tree/6.7/northbound-interfaces/commit-parameters) example for an end-to-end CLI and RESTCONF walkthrough.

Any failure on any device will make the whole transaction fail. It is also possible to perform a manual rollback, a rollback is the undoing of a commit.

This is operational data and the CLI is in configuration mode so the way of showing operational data in config mode is used.

The command `show configuration rollback changes` can be used to view rollback changes in more detail. It will show what will be done when the rollback file is loaded, similar to loading the rollback and using `show configuration`:

```bash
admin@ncs(config)# show configuration rollback changes 10019
devices device ce0
 config
  ios:interface GigabitEthernet0/1
   no speed auto
  exit
 !
!
devices device ce1
 config
  ios:interface GigabitEthernet0/1
   no speed auto
  exit
 !
!
```

The command `show configuration commit changes` can be used to see which changes were done in a given commit, i.e. the roll-forward commands performed in that commit:

```bash
admin@ncs(config)# show configuration commit changes 10019
!
! Created by: admin
! Date: 2015-02-03 12:29:08
! Client: cli
!
devices device ce0
 config
  ios:interface GigabitEthernet0/1
   speed auto
  exit
 !
!
devices device ce1
 config
  ios:interface GigabitEthernet0/1
   speed auto
  exit
 !
!
```

The command `rollback-files apply-rollback-file` can be used to perform the rollback:

```bash
admin@ncs(config)# rollback-files apply-rollback-file fixed-number 10019
admin@ncs(config)# show configuration
devices device ce0
 config
  ios:interface GigabitEthernet0/1
   no speed auto
  exit
 !
!
devices device ce1
 config
  ios:interface GigabitEthernet0/1
   no speed auto
  exit
 !
!
```

And now the `commit` the rollback:

```bash
admin@ncs(config)# commit
Commit complete.
```

When the command `rollback-files apply-rollback-file fixed-number 10019` is run the changes recorded in rollback 10019-N (where N is the highest, thus the most recent rollback number) will all be undone. In other words, the configuration will be rolled back to the state it was in before the commit associated with rollback 10019 was performed.

It is also possible to undo individual changes by running the command `rollback-files apply-rollback-file selective`. E.g., to undo the changes recorded in rollback 10019, but not the changes in 10020-N run the command `rollback-files apply-rollback-file selective fixed-number 10019`.

This operation may fail if the commits following rollback 10019 depend on the changes made in rollback 10019.

## Command Output Processing <a href="#d5e1430" id="d5e1430"></a>

It is possible to process the output from a command using an output redirect. This is done using the | character (a pipe character):

```bash
admin@ncs# show running-config | ?
Possible completions:
  annotation      Show only statements whose annotation matches a pattern
  append          Append output text to a file
  begin           Begin with the line that matches
  best-effort     Display data even if data provider is unavailable or
                  continue loading from file in presence of failures
  context-match   Context match
  count           Count the number of lines in the output
  csv             Show table output in CSV format
  de-select       De-select columns
  details         Display show/commit details
  display         Display options
  exclude         Exclude lines that match
  extended        Display referring entries
  hide            Hide display options
  include         Include lines that match
  linnum          Enumerate lines in the output
  match-all       All selected filters must match
  match-any       At least one filter must match
  more            Paginate output
  nomore          Suppress pagination
  save            Save output text to a file
  select          Select additional columns
  sort-by         Select sorting indices
  tab             Enforce table output
  tags            Show only statements whose tags matches a pattern
  until           End with the line that matches
```

The precise list of pipe commands depends on the command executed. Some pipe commands, like `select` and `de-select`, are only available for the `show` command, whereas others are universally available.

{% hint style="info" %}
Note that the `tab` pipe target is used to enforce table output which is only suitable for the list element. Naturally, the table format is not suitable for displaying arbitrary data output since it needs to map the data to columns and rows.

For example, the following are clearly not suitable because the data has a nested structure. It could take an incredibly long time to display it if you use the `tab` pipe target on a huge amount of data which is not a list element.

```bash
show running-config | tab
show running-config | include aaa | tab
```

{% endhint %}

### Count the Number of Lines in the Output <a href="#d5e1443" id="d5e1443"></a>

This redirect target counts the number of lines in the output. For example:

```bash
admin@ncs# show running-config | count
Count: 1783 lines
admin@ncs# show running-config aaa | count
Count: 28 lines
```

### Search for a String in the Output <a href="#d5e1449" id="d5e1449"></a>

The `include` targets is used to only include lines matching a regular expression:

```bash
admin@ncs# show running-config aaa | include aaa
aaa authentication users user admin
aaa authentication users user oper
aaa authentication users user private
aaa authentication users user public
```

In the example above only lines containing aaa are shown. Similarly lines not containing a regular expression can be included. This is done using the `exclude` target:

```bash
admin@ncs# show running-config aaa authentication | exclude password
aaa authentication users user admin
 uid        1000
 gid        1000
 ssh_keydir /var/ncs/homes/admin/.ssh
 homedir    /var/ncs/homes/admin
!
aaa authentication users user oper
 uid        1000
 gid        1000
 ssh_keydir /var/ncs/homes/oper/.ssh
 homedir    /var/ncs/homes/oper
!
aaa authentication users user private
 uid        1000
 gid        1000
 ssh_keydir /var/ncs/homes/private/.ssh
 homedir    /var/ncs/homes/private
!
aaa authentication users user public
 uid        1000
 gid        1000
 ssh_keydir /var/ncs/homes/public/.ssh
 homedir    /var/ncs/homes/public
 !
```

It is possible to display the context for a match using the pipe command `include -c`. Matching lines will be prefixed by `<line no>`: and context lines with `<line no>-`. For example:

```bash
admin@ncs# show running-config aaa authentication | include -c 3 homes/admin
 2- uid        1000
 3- gid        1000
 4- password   $1$brH6BYLy$iWQA2T1I3PMonDTJOd0Y/1
 5: ssh_keydir /var/ncs/homes/admin/.ssh
 6: homedir    /var/ncs/homes/admin
 7-!
 8-aaa authentication users user oper
 9- uid        1000
```

It is possible to display the context for a match using the pipe command `context-match`:

```bash
admin@ncs# show running-config aaa authentication | context-match homes/admin
aaa authentication users user admin
 ssh_keydir /var/ncs/homes/admin/.ssh
aaa authentication users user admin
 homedir    /var/ncs/homes/admin
```

It is possible to display the output starting at the first match of a regular expression. This is done using the `begin` pipe command:

```bash
admin@ncs# show running-config aaa authentication users | begin public
aaa authentication users user public
 uid        1000
 gid        1000
 password   $1$DzGnyJGx$BjxoqYEj0QKxwVX5fbfDx/
 ssh_keydir /var/ncs/homes/public/.ssh
 homedir    /var/ncs/homes/public
!
```

### Saving the Output to a File <a href="#d5e1478" id="d5e1478"></a>

The output can also be saved to a file using the `save` or `append` redirect target:

```bash
admin@ncs# show running-config aaa | save /tmp/saved
```

Or to save the configuration, except all passwords:

```bash
admin@ncs# show running-config aaa | exclude password | save /tmp/saved
```

### Regular Expressions <a href="#ug.ncs.cli.regexp" id="ug.ncs.cli.regexp"></a>

The regular expressions are a subset of the regular expressions found in egrep and in the AWK programming language. Some common operators are:

<table><thead><tr><th width="198">Operator</th><th>Description</th></tr></thead><tbody><tr><td><code>.</code></td><td>Matches any character.</td></tr><tr><td><code>^</code></td><td>Matches the beginning of a string.</td></tr><tr><td><code>$</code></td><td>Matches the end of a string.</td></tr><tr><td><code>[abc...]</code></td><td>Character class, which matches any of the characters abc... Character ranges are specified by a pair of characters separated by a <code>-</code>.</td></tr><tr><td><code>[^abc...]</code></td><td>Negated character class, which matches any character except abc... .</td></tr><tr><td><code>r1 | r2</code></td><td>Alternation. It matches either <code>r1</code> or <code>r2</code>.</td></tr><tr><td><code>r1r2</code></td><td>Concatenation. It matches <code>r1</code> and then <code>r2</code>.</td></tr><tr><td><code>r+</code></td><td>Matches one or more <code>rs</code>.</td></tr><tr><td><code>r*</code></td><td>Matches zero or more <code>rs</code>.</td></tr><tr><td><code>r?</code></td><td>Matches zero or one <code>rs</code>.</td></tr><tr><td><code>(r)</code></td><td>Grouping. It matches <code>r</code>.</td></tr></tbody></table>

For example, to only display `uid` and `gid` do the following:

```bash
admin@ncs# show running-config aaa | include "(uid)|(gid)"
 uid        1000
 gid        1000
 uid        1000
 gid        1000
 uid        1000
 gid        1000
 uid        1000
 gid        1000
```

## Displaying the Configuration <a href="#d5e1541" id="d5e1541"></a>

There are several options for displaying the configuration and stats data in NSO. The most basic command consists of displaying a leaf or a subtree of the configuration by giving the path to the element.

To display the configuration of a device do:

```bash
admin@ncs# show running-config devices device ce0 config
devices device ce0
 config
  no ios:service pad
  no ios:ip domain-lookup
  no ios:ip http secure-server
  ios:ip source-route
  ios:interface GigabitEthernet0/1
  exit
  ios:interface GigabitEthernet0/10
  exit
  ...
 !
 !
```

This can also be done for a group of devices by substituting the instance name (`ce0` in this case) with [Regular Expressions](#ug.ncs.cli.regexp).

To display the config of all devices:

```bash
admin@ncs# show running-config devices device * config
devices device ce0
 config
  no ios:service pad
  no ios:ip domain-lookup
  no ios:ip http secure-server
  ios:ip source-route
  ios:interface GigabitEthernet0/1
  exit
  ios:interface GigabitEthernet0/10
  exit
  ...
 !
!
devices device ce1
 config
  ...
 !
!
...
```

It is possible to limit the output even further. View only the HTTP settings on each device:

```bash
admin@ncs# show running-config devices device * config ios:ip http
devices device ce0
 config
  no ios:ip http secure-server
 !
!
devices device ce1
 config
  no ios:ip http secure-server
 !
!
...
```

There is an alternative syntax for this using the `select` pipe command:

```bash
admin@ncs# show running-config devices device * | \
    select config ios:ip http
devices device ce0
 config
  no ios:ip http secure-server
 !
!
devices device ce1
 config
  no ios:ip http secure-server
 !
!
...
```

The `select` pipe command can be used multiple times for adding additional content:

```bash
admin@ncs# show running-config devices device * | \
    select config ios:ip http | \
    select config ios:ip domain-lookup
devices device ce0
 config
  no ios:ip domain-lookup
  no ios:ip http secure-server
 !
!
devices device ce1
 config
  no ios:ip domain-lookup
  no ios:ip http secure-server
 !
!
...
```

There is also a `de-select` pipe command that can be used to instruct the CLI to not display certain parts of the config. The above printout could also be achieved by first selecting the `ip` container, and then de-selecting the `source-route` leaf:

```bash
admin@ncs# show running-config devices device * | \
    select config ios:ip | \
    de-select config ios:ip source-route
devices device ce0
 config
  no ios:ip domain-lookup
  no ios:ip http secure-server
 !
!
devices device ce1
 config
  no ios:ip domain-lookup
  no ios:ip http secure-server
 !
!
...
```

A use-case for the `de-select` pipe command is to de-select the `config` container to only display the device settings without actually displaying their config:

```bash
admin@ncs# show running-config devices device * | de-select config
devices device ce0
 address   127.0.0.1
 port      10022
 ssh host-key ssh-dss
  ...
 !
 authgroup default
 device-type cli ned-id cisco-ios
 state admin-state unlocked
!
devices device ce1
 ...
!
...
```

The above statements also work for the `save` command. To save the devices managed by NSO, but not the contents of their `config` container:

```bash
admin@ncs# show running-config devices device * | \
    de-select config | save /tmp/devices
```

It is possible to use the `select` command to select which list instances to display. To display all devices that have the interface `GigabitEthernet 0/0/0/4`:

```bash
admin@ncs# show running-config devices device * | \
    select config cisco-ios-xr:interface GigabitEthernet 0/0/0/4
devices device p0
 config
  cisco-ios-xr:interface GigabitEthernet 0/0/0/4
   shutdown
  exit
 !
!
devices device p1
 config
  cisco-ios-xr:interface GigabitEthernet 0/0/0/4
   shutdown
  exit
 !
!
...
```

This means to display all device instances that have the interface GigabitEthernet 0/0/0/4. Only the subtree defined by the select path will be displayed. It is also possible to display the entire content of the `config` container for each instance by using an additional select statement:

```bash
admin@ncs# show running-config devices device * | \
    select config cisco-ios-xr:interface GigabitEthernet 0/0/0/4 | \
    select config | match-all
devices device p0
 config
  cisco-ios-xr:hostname PE1
  cisco-ios-xr:interface MgmtEth 0/0/CPU0/0
  exit
  ...
  cisco-ios-xr:interface GigabitEthernet 0/0/0/4
   shutdown
  exit
 !
!
devices device p1
 config
  ...
  cisco-ios-xr:interface GigabitEthernet 0/0/0/4
   shutdown
  exit
 !
!
...
```

The `match-all` pipe command is used for telling the CLI to only display instances that match all select commands. The default behavior is `match-any` which means to display instances that match any of the given `select` commands.

The `display` command is used to format configuration and statistics data. There are several output formats available, and some of these are unique to specific modes, such as configuration or operational mode. The output formats `json`, `keypath`, `xml`, and `xpath` are available in most modes and CLI styles (J, I, and C). The output formats `netconf` and `maagic` are only available if `devtools` has been set to `true` in the CLI session settings.

For instance, assuming we have a data model featuring a set of hosts, each containing a set of servers, we can display the configuration data as JSON. This is depicted in the example below.

```bash
admin@ncs# show running-config hosts | display json
{
  "data": {
    "pipetargets_model:hosts": {
      "host": [
        {
          "name": "host1",
          "enabled": true,
          "numberOfServers": 2,
          "servers": {
            "server": [
              {
                "name": "serv1",
                "ip": "192.168.0.1",
                "port": 5001
              },
              {
                "name": "serv2",
                "ip": "192.168.0.1",
                "port": 5000
              }
            ]
          }
        },
        {
          "name": "host2",
          "enabled": false,
          "numberOfServers": 0
...
```

Still working with the same data model as used in the example above, we might want to see the current configuration in keypath format.

The following example shows how to do that and shows the resulting output:

```bash
admin@ncs# show running-config hosts | display keypath
/hosts/host{host1} enabled
/hosts/host{host1}/numberOfServers 2
/hosts/host{host1}/servers/server{serv1}/ip 192.168.0.1
/hosts/host{host1}/servers/server{serv1}/port 5001
/hosts/host{host1}/servers/server{serv2}/ip 192.168.0.1
/hosts/host{host1}/servers/server{serv2}/port 5000
/hosts/host{host2} disabled
/hosts/host{host2}/numberOfServers 0
```

The `ignore-display-when` pipe command is intended to expose data nodes that have been hidden by the `tailf:display-when` expression. This pipe command can also be used in combination with various display targets.

The following example is for a leaf `hidden-empty-leaf` that has a `tailf:display-when` expression that evaluates to `false` but can still be forced to get exposed in the CLI by use of the `ignore-display-when` pipe command:

```bash
admin@ncs(config)# show configuration
devices device foo0
 config
  bar
 !
!

admin@ncs(config)# show configuration | ignore-display-when
devices device foo0
 config
  bar hidden-empty-leaf
 !
!

admin@ncs(config)# show configuration | ignore-display-when | display curly-braces
 devices {
     device foo0 {
         config {
+            bar {
+                hidden-empty-leaf;
+            }
         }
     }
 }
```

## Range Expressions <a href="#d5e1612" id="d5e1612"></a>

To modify a range of instances, at the same time, use range expressions or display a specific range of instances.

Basic range expressions are written with a combination of x..y (meaning from x to y), x,y (meaning x and y) and \* (meaning any value), example:

```
1..4,8,10..18
```

It is possible to use range expressions for all key elements of integer type, both for setting values, executing actions, and displaying status and config.

Range expressions are also supported for key elements of non-integer types as long as they are restricted to the pattern \[a-zA-Z-]\*\[0-9]+/\[0-9]+/\[0-9]+/.../\[0-9]+ and the annotation `tailf:cli-allow-range` is used on the key leaf. This is the case for the device list.

The following can be done in the CLI to display a subset of the devices (`ce0`, `ce1`, `ce3`):

```bash
admin@ncs# show running-config devices device ce0..1,3
```

If the devices have names with slashes, for example, Firewall/1/1, Firewall/1/2, Firewall/1/3, Firewall/2/1, Firewall/2/2, and Firewall/2/3, expressions like this are possible:

```bash
admin@ncs# show running-config devices device Firewall/1-2/*
admin@ncs# show running-config devices device Firewall/1-2/1,3
```

In configure mode, it is possible to edit a range of instances in one command:

```bash
admin@ncs(config)# devices device ce0..2 config ios:ethernet cfm ieee
```

Or, like this:

```bash
admin@ncs(config)# devices device ce0..2 config
admin@ncs(config-config)# ios:ethernet cfm ieee
admin@ncs(config-config)# show config
devices device ce0
 config
  ios:ethernet cfm ieee
 !
!
devices device ce1
 config
  ios:ethernet cfm ieee
 !
!
devices device ce2
 config
  ios:ethernet cfm ieee
 !
!
```

## Command History <a href="#d5e1639" id="d5e1639"></a>

Command history is maintained separately for each mode. When entering configure mode from operational for the first time, an empty history is used. It is not possible to access the command history from operational mode when in configure mode and vice versa. When exiting back into operational mode access to the command history from the preceding operational mode session will be used. Likewise, the old command history from the old configure mode session will be used when re-entering configure mode.

## Command Line Editing <a href="#d5e1642" id="d5e1642"></a>

The default keystrokes for editing the command line and moving around the command history are as follows.

### Moving the Cursor <a href="#d5e1645" id="d5e1645"></a>

* Move the cursor back by one character: Ctrl-b or Left Arrow.
* Move the cursor back by one word: Esc-b or Alt-b.
* Move the cursor forward one character: Ctrl-f or Right Arrow.
* Move the cursor forward one word: Esc-f or Alt-f.
* Move the cursor to the beginning of the command line: Ctrl-a or Home.
* Move the cursor to the end of the command line: Ctrl-e or End.

### Delete Characters <a href="#d5e1672" id="d5e1672"></a>

* Delete the character before the cursor: Ctrl-h, Delete, or Backspace.
* Delete the character following the cursor: Ctrl-d.
* Delete all characters from the cursor to the end of the line: Ctrl-k.
* Delete the whole line: Ctrl-u or Ctrl-x.
* Delete the word before the cursor: Ctrl-w, Esc-Backspace, or Alt-Backspace.
* Delete the word after the cursor: Esc-d or Alt-d.

### Insert Recently Deleted Text <a href="#d5e1699" id="d5e1699"></a>

* Insert the most recently deleted text at the cursor: Ctrl-y.

### Display Previous Command Lines <a href="#d5e1706" id="d5e1706"></a>

* Scroll backward through the command history: Ctrl-p or Up Arrow.
* Scroll forward through the command history: Ctrl-n or Down Arrow.
* Search the command history in reverse order: Ctrl-r.
* Show a list of previous commands: run the `show cli history` command.

### Capitalization <a href="#d5e1725" id="d5e1725"></a>

* Capitalize the word at the cursor, i.e. make the first character uppercase and the rest of the word lowercase: Esc-c.
* Change the word at the cursor to lowercase: Esc-l.
* Change the word at the cursor to uppercase: Esc-u.

### Special <a href="#d5e1740" id="d5e1740"></a>

* Abort a command/Clear line: Ctrl-c.
* Quote insert character, i.e. do not treat the next keystroke as an edit command: Ctrl-v/ESC-q.
* Redraw the screen: Ctrl-l.
* Transpose characters: Ctrl-t.
* Enter multi-line mode. Enables entering multi-line values when prompted for a value in the CLI: ESC-m.
* Exit configuration mode: Ctrl-z.

## CLI Completion <a href="#d5e1767" id="d5e1767"></a>

It is not necessary to type the full command or option name for the CLI to recognize it. To display possible completions, type the partial command followed immediately by `<tab>` or `<space>`.

If the partially typed command uniquely identifies a command, the full command name will appear. Otherwise, a list of possible completions is displayed.

Long lines can be broken into multiple lines using the backslash (`\`) character at the end of the line. This is primarily useful inside scripts.

Completion is disabled inside quotes. To type an argument containing spaces either quote them with a \ (e.g. `file show foo\ bar`) or with a " (e.g. `file show "foo bar"`). Space completion is disabled when entering a filename.

Command completion also applies to filenames and directories:

```bash
admin@ncs# <space>
Possible completions:
  alarms                 Alarm management
  autowizard             Automatically query for mandatory elements
  cd                     Change working directory
  clear                  Clear parameter
  cluster                Cluster configuration
  compare                Compare running configuration to another
                         configuration or a file
  complete-on-space      Enable/disable completion on space
  compliance             Compliance reporting
  config                 Manipulate software configuration information
  describe               Display transparent command information
  devices                The managed devices and device communication settings
  display-level          Configure show command display level
  exit                   Exit the management session
  file                   Perform file operations
  help                   Provide help information
  ...
admin@ncs# dev<space>ices <space>
Possible completions:
  check-sync            Check if the NCS config is in sync with the device
  check-yang-modules    Check if NCS and the devices have compatible YANG
                        modules
  clear-trace           Clear all trace files
  commit-queue          List of queued commits
  ...
admin@ncs# devices check-s<space>ync
```

## Comments, Annotations, and Tags <a href="#d5e1780" id="d5e1780"></a>

All characters following a **`!`**, up to the next new line, are ignored. This makes it possible to have comments in a file containing CLI commands, and still be able to paste the file into the command-line interface. For example:

```bash
! Command file created by Joe Smith
! First show the configuration before we change it
show running-config
! Enter configuration mode and configure an ethernet setting on the ce0 device
config
devices device ce0 config ios:ethernet cfm global
commit
top
exit
exit
! Done
```

To enter the comment character as an argument, it has to be prefixed with a backslash (\\) or used inside quotes (").

The `/* ... */` comment style is also supported.

When using large configurations it may make sense to be able to associate comments (annotations) and tags with the different parts. Then filter the configuration with respect to the annotations or tags. For example, tagging parts of the configuration that relate to a certain department or customer.

NSO has support for both tags and annotations. There is a specific set of commands available in the CLI for annotating and tagging parts of the configuration. There is also a set of pipe commands for controlling whether the tags and annotations should be displayed and for filtering depending on annotation and tag content.

The commands are:

* `annotate <statement> <text>`
* `tag add <statement> <tag>`
* `tag clear <statement> <tag>`
* `tag del <statement> <tag>`

Example:

```bash
admin@ncs(config)# annotate aaa authentication users user admin \
"Only allow the XX department access to this user."
admin@ncs(config)# tag add aaa authentication users user oper oper_tag
admin@ncs(config)# commit
Commit complete.
```

To view the placement of tags and annotations in the configuration it is recommended to use the pipe command `display curly-braces`. The annotations and tags will be displayed as comments where the tags are prefixed by `Tags:`. For example:

```bash
admin@ncs(config)# do show running-config aaa authentication users user | \
    tags oper_tag | display curly-braces
/* Tags: oper_tag */
user oper {
    uid        1000;
    gid        1000;
    password   $1$9qV138GJ$.olmolTfRbFGQhWJMZ9kA0;
    ssh_keydir /var/ncs/homes/oper/.ssh;
    homedir    /var/ncs/homes/oper;
}
admin@ncs(config)# do show running-config aaa authentication users user | \
    annotation XX | display curly-braces
/* Only allow the XX department access to this user. */
user admin {
    uid        1000;
    gid        1000;
    password   $1$EcQwYvnP$Rvq3MPTMSz29UaVOHA/511;
    ssh_keydir /var/ncs/homes/admin/.ssh;
    homedir    /var/ncs/homes/admin;
}
```

It is possible to hide the tags and annotations when viewing the configuration or to explicitly include them in the listing. This is done using the `display annotations/tags` and `hide annotations/tags` pipe commands. To hide all attributes (annotations, tags, and FASTMAP attributes) use the `hide attributes` pipe command.

Annotations and tags are part of the configuration. When adding, removing, or modifying an annotation or a tag, the configuration needs to be committed similar to any other change to the configuration.

## CLI Messages <a href="#d5e1824" id="d5e1824"></a>

Messages appear when entering and exiting configure mode, when committing a configuration, and when typing a command or value that is not valid:

```bash
admin@ncs# show c
-----------------^
syntax error:
Possible alternatives starting with c:
  cli           - Display cli settings
  configuration - Commit configuration changes
admin@ncs# show configuration
------------------------------^
syntax error: expecting
  commit - Commit configuration changes
```

When committing a configuration, the CLI first validates the configuration, and if there is a problem it will indicate what the problem is.

If a missing identifier or a value is out of range a message will indicate where the errors are:

```bash
admin@ncs# config
Entering configuration mode terminal
admin@ncs(config)# nacm rule-list any-group rule allowrule
admin@ncs(config-rule-allowrule)# commit
Aborted: 'nacm rule-list any-group rule allowrule action' is not configured
```

## `ncs.conf` Settings <a href="#d5e1837" id="d5e1837"></a>

Parts of the CLI behavior can be controlled from the `ncs.conf` file. See the [ncs.conf(5)](/guides/resources/man/ncs.conf.5) in Manual Pages manual page for a comprehensive description of all the options.

## CLI Environment <a href="#d5e1842" id="d5e1842"></a>

There are a number of session variables in the CLI. They are only used during the session and are not persistent. Their values are inspected using `show cli` in operational mode, and set using **set** in operational mode. Their initial values are in order derived from the content of the `ncs.conf` file, and the global defaults as configured at `/aaa:session` and user-specific settings configured at `/aaa:user{<user>}/setting`.

```bash
admin@ncs# show cli
autowizard                   false
commit-prompt                false
complete-on-space            true
devtools                     false
display-level                99999999
dry-run-drift-detection      false;
dry-run-drift-detection-mode warn;
dry-run-duration             0
dry-run-outformat            cli
history                      100
idle-timeout                 1800
ignore-leading-space         false
output-file                  terminal
paginate                     true
prompt1                      \h\M#
prompt2                      \h(\m)#
screen-length                71
screen-width                 80
service prompt config        true
show-defaults                false
terminal                     xterm-256color
...
```

The different values control different parts of the CLI behavior:

<details>

<summary><code>autowizard (true | false)</code></summary>

When enabled, the CLI will prompt the user for the required settings when a new identifier is created.\
\
For example:

```bash
admin@ncs(config)# aaa authentication users user John
Value for 'uid' (<int>): 1006
Value for 'gid' (<int>): 1006
Value for 'password' (<hash digest string>): ******
Value for 'ssh_keydir' (<string>): /var/ncs/homes/john/.ssh
Value for 'homedir' (<string>): /var/ncs/homes/john
```

This helps the user set all mandatory settings.\
\
It is recommended to disable the autowizard before pasting in a list of commands in order to avoid prompting. A good practice is to start all such scripts with a line that disables the `autowizard`:

```
autowizard false
...
autowizard true
```

</details>

<details>

<summary><code>commit-prompt (true | false)</code></summary>

When enabled, the CLI will display dry-run output of the configuration changes and prompt the user to confirm before the commit operation or actions using the ncs-commit-params grouping. This setting is effective on the following actions.

* Service actions
  * `re-deploy`
  * `un-deploy`
* Device actions
  * `sync-to`
  * `partial-sync-to`
  * `migrate`
  * `rollback`

For example with commit:

```bash
admin@ncs(config)# devices global-settings commit-retries attempts 3
admin@ncs(config)# commit
cli {
    local-node {
        data  devices {
                  global-settings {
                      commit-retries {
             +            attempts 3;
                      }
                  }
              }
    }
}
Warning: Please review the changes before commit.
Proceed? [yes,no]
```

For example with action:

```bash
admin@ncs(config)# dns-config test re-deploy
cli {
    local-node {
        data  devices {
                   device ex1 {
                       config {
                           sys {
                               dns {
              +                    # after server 10.2.3.4
              +                    server 192.0.2.1;
                               }
                           }
                       }
                   }
               }
              
    }
}
Warning: Please review the changes before 're-deploy'.
Proceed? [yes,no]
```

{% hint style="info" %}
Note that dry-run output could be very long if the configuration changes are large.
{% endhint %}

</details>

<details>

<summary><code>complete-on-space (true | false)</code></summary>

Controls if command completion should be attempted when `<space>` is entered. Entering `<tab>` always results in command completion.

</details>

<details>

<summary><code>devtools (true | false)</code></summary>

Controls if certain commands that are useful for developers should be enabled. The command `xpath` and `timecmd` are examples of such a command.

</details>

<details>

<summary><code>dry-run-drift-detection (true | false)</code></summary>

When enabled, the user is notified during commit if the transaction’s changeset has changed since their last dry-run, helping prevent unintended changes from being committed. This setting is available in `warn` and `strict` modes; see below.

</details>

<details>

<summary><code>dry-run-drift-detection-mode (warn | strict)</code></summary>

This is the mode for the `dry-run-drift-detection` setting. In `warn` mode, the CLI warns if the changeset differs between dry-run and commit, and prompts the user whether to proceed. In `strict` mode, the commit is aborted, and the user must perform a new dry-run before committing.

For example ("warn" mode):

```bash
admin@ncs(config)# commit
The following warnings were generated:
  Commit changeset does not match dry-run changeset
Proceed? [yes,no]
```

"strict" mode will instead give the following error:

```bash
admin@ncs(config)# commit
Aborted: Commit changeset does not match dry-run changeset
```

</details>

<details>

<summary><code>dry-run-duration (&#x3C;seconds>)</code></summary>

Valid period of dry-run output before prompting the user to confirm commit or action if `commit-prompt` is set to `true`.

Setting this to 0 (zero) means the same dry-run output will be displayed instantly each time before prompting the user to proceed.

If it is not set to 0 (zero), the CLI will not display dry-run output for the same configuration changes repeatedly within this time period. After it expires, dry-run output will be displayed again.

For example with `dry-run-duration` set to 0 (zero):

<pre class="language-bash"><code class="lang-bash"><strong>admin@ncs(config)# devices global-settings commit-retries attempts 3
</strong>admin@ncs(config)# commit
cli {
    local-node {
        data  devices {
                  global-settings {
                      commit-retries {
             +            attempts 3;
                      }
                  }
              }
    }
}
Warning: Please review the changes before commit.
Proceed? [yes,no] no
Aborted: by user
admin@ncs(config)# commit
cli {
    local-node {
        data  devices {
                  global-settings {
                      commit-retries {
             +            attempts 3;
                      }
                  }
              }
    }
}
Warning: Please review the changes before commit.
Proceed? [yes,no]
</code></pre>

For example with `dry-run-duration` set to 5 (seconds):

```bash
admin@ncs# dry-run-duration 5
admin@ncs# config
Entering configuration mode terminal
admin@ncs(config)# devices global-settings commit-retries attempts 3
admin@ncs(config)# commit
cli {
    local-node {
        data  devices {
                  global-settings {
                      commit-retries {
             +            attempts 3;
                      }
                  }
              }
    }
}
Warning: Please review the changes before commit.
Proceed? [yes,no] no
Aborted: by user
admin@ncs(config)# commit
Proceed? [yes,no] no
Aborted: by user
... <User waits for 6 seconds> ...
admin@ncs(config)# commit
cli {
    local-node {
        data  devices {
                  global-settings {
                      commit-retries {
             +            attempts 3;
                      }
                  }
              }
    }
}
Warning: Please review the changes before commit.
Proceed? [yes,no]
```

The same situation applies to the case if `dry-run` flag is used with commit or `dry-run` option is used on action.

For example with `dry-run-duration` set to 5 (seconds) and run `commit dry-run` first:

```bash
admin@ncs# dry-run-duration 5
admin@ncs# config
Entering configuration mode terminal
admin@ncs(config)# devices global-settings commit-retries attempts 3
admin@ncs(config)# commit dry-run
cli {
    local-node {
        data  devices {
                  global-settings {
                      commit-retries {
             +            attempts 3;
                      }
                  }
              }
    }
}
admin@ncs(config)# commit
Proceed? [yes,no] no
Aborted: by user
... <User waits for 6 seconds> ...
admin@ncs(config)# commit
cli {
    local-node {
        data  devices {
                  global-settings {
                      commit-retries {
             +            attempts 3;
                      }
                  }
              }
    }
}
Warning: Please review the changes before commit.
Proceed? [yes,no]
```

For example with `dry-run-duration` set to 5 (seconds) and run `re-deploy dry-run` first:

```bash
admin@ncs# dry-run-duration 5
admin@ncs# config
Entering configuration mode terminal
admin@ncs(config)# dns-config test re-deploy dry-run
cli {
    local-node {
        data  devices {
                   device ex1 {
                       config {
                           sys {
                               dns {
              +                    # after server 10.2.3.4
              +                    server 192.0.2.1;
                               }
                           }
                       }
                   }
               }
              
    }
}
admin@ncs(config)# dns-config test re-deploy
Proceed? [yes,no] no
Aborted: by user
... <User waits for 6 seconds> ...
admin@ncs(config)# dns-config test re-deploy
cli {
    local-node {
        data  devices {
                   device ex1 {
                       config {
                           sys {
                               dns {
              +                    # after server 10.2.3.4
              +                    server 192.0.2.1;
                               }
                           }
                       }
                   }
               }
              
    }
}
Warning: Please review the changes before 're-deploy'.
Proceed? [yes,no]
```

</details>

<details>

<summary><code>dry-run-outformat (&#x3C;string>)</code></summary>

Format of dry-run output for the configuration changes before prompting the user to confirm commit or action if `commit-prompt` is set to `true`.

The supported formats are: `cli`, `xml`, `native` and `cli-c`.

For example with `dry-run-outformat` set to `xml` first and then set to `cli-c`:

```bash
admin@ncs# dry-run-outformat xml
admin@ncs# config
Entering configuration mode terminal
admin@ncs(config)# devices global-settings commit-retries attempts 3
admin@ncs(config)# commit
result-xml {
    local-node {
        data <devices xmlns="http://tail-f.com/ns/ncs">
               <global-settings>
                 <commit-retries>
                   <attempts>3</attempts>
                 </commit-retries>
               </global-settings>
             </devices>
    }
}
Warning: Please review the changes before commit.
Proceed? [yes,no] no
Aborted: by user
admin@ncs(config)# do dry-run-outformat cli-c
admin@ncs(config)# commit
cli-c {
    local-node {
        data devices global-settings commit-retries attempts 3
    }
}
Warning: Please review the changes before commit.
Proceed? [yes,no]
```

</details>

<details>

<summary><code>history (&#x3C;integer>)</code></summary>

Size of CLI command history.

</details>

<details>

<summary><code>idle-timeout (&#x3C;seconds>)</code></summary>

Maximum idle time before being logged out. Use 0 (zero) for infinity.

</details>

<details>

<summary><code>ignore-leading-space (true | false)</code></summary>

Controls if leading spaces should be ignored or not. This is useful to turn off when pasting commands into the CLI.

</details>

<details>

<summary><code>paginate (true | false)</code></summary>

Some commands paginate (or MORE process) the output, for example, `show running-config`. This can be disabled or enabled. It is enabled by default. Setting the screen length to 0 has the same effect as turning off pagination.

</details>

<details>

<summary><code>screen length (&#x3C;integer>)</code></summary>

The current length of the terminal. This is used when paginating output to get the proper line count. Setting this to 0 (zero) means it becomes the maximum length and turns off pagination.

</details>

<details>

<summary><code>screen width (&#x3C;integer>)</code></summary>

The current width of the terminal. This is used when paginating output to get the proper line count. Setting this to 0 (zero) means it becomes the maximum width.

</details>

<details>

<summary><code>service prompt config</code></summary>

Controls whether a prompt should be displayed in configure mode. If set to false, then no prompt will be displayed. The setting is changed using the commands `no service prompt config` and `service prompt config` in configure mode.

</details>

<details>

<summary><code>terminal (string)</code></summary>

Terminal type. This setting is used for controlling how line editing is performed. Supported terminals are: `dumb`, `vt100`, `xterm`, `linux` and `ansi`. Other terminals may also work but have no explicit support.

</details>

## Customizing the CLI

### Adding New Commands <a href="#d5e2650" id="d5e2650"></a>

New commands can be added by placing a script in the `scripts/command` directory. See [Plug-and-play Scripting](/guides/operation-and-usage/operations/plug-and-play-scripting).

### File Access <a href="#d5e2655" id="d5e2655"></a>

The default behavior is to enforce Unix-style access restrictions. That is, the user's `uid`, `gid`, and `gids` are used to control what the user has read and write access to.

However, it is also possible to jail a CLI user to their home directory (or the directory where `ncs_cli` is started). This is controlled using the `ncs.conf` parameter `restricted-file-access`. If this is set to `true`, then the user only has access to the home directory.

### Help Texts <a href="#d5e2664" id="d5e2664"></a>

Help and information texts are specified in several places. In the Yang files, the `tailf:info` element is used to specify a descriptive text that is shown when the user enters `?` in the CLI. The first sentence of the `info` text is used when showing one-line descriptions in the CLI.

## Quoting and Escaping Scheme <a href="#d5e2669" id="d5e2669"></a>

### **Canonical Quoting Scheme** <a href="#d5e2671" id="d5e2671"></a>

NCS understands multiple quoting schemes on input and de-quotes a value when parsing the command. Still, it uses what it considers a canonical quoting scheme when printing out this value, e.g., when pushing a configuration change to the device. However, different devices may have different quoting schemes, possibly not compatible with the NCS canonical quoting scheme. For example, the following value cannot be printed out by NCS as two backslashes `\\` match `\` in the quoting scheme used by NCS when encoding values.

```
"foo\\/bar\\?baz"
```

General rules for NCS to represent backslash are as follows, and so on. It can only get an odd number of backslashes output from NCS.

* `\` and `\\` are represented as `\`.
* `\\\` and `\\\\` are represented as `\\\`.
* `\\\\\` and `\\\\\\` are represented as `\\\\\`.

A backslash `\` is represented as a backslash `\` when it is followed by a character that does not need to be escaped but is represented as double backslashes `\\` if the next character could be escaped. With remote passwords, if you are using special characters, be sure to follow recommended guidelines, see [Configure Mode](#d5e1216) for more information.

### **Escape Backslash Handling** <a href="#d5e2688" id="d5e2688"></a>

To let NCS pass through a quoted string verbatim, one can do as stated below:

* Enable the NCS configuration parameter `escapeBackslash` in the `ncs.conf` file. This is a global setting on NCS which affects all the NEDs.
* Alternatively, a certain NED may be updated on request to be able to transform the value printed by NCS to what the device expects if one only wants to affect a certain device instead of all the connected ones.

### **Octal Numbers Handling** <a href="#d5e2697" id="d5e2697"></a>

If there are numeric triplets following a backslash `\`, NCS will treat them as octal numbers and convert them to one character based on ASCII code. For example:

* `\123` is converted to `S`.
* `\067` is converted to `7`.


# CLI Commands

CLI command reference.

## Commands

To get a full XML listing of the commands available in a running NSO instance, use the `ncs` option `--cli-c-dump <file>`. The generated file is only intended for documentation purposes and cannot be used as input to the `ncsc` compiler. The command `show parser dump` can be used to get a command listing.

### Operational Mode Commands <a href="#d5e1943" id="d5e1943"></a>

#### Invoke an Action

<details>

<summary><code>&#x3C;path> &#x3C;parameters></code></summary>

Invokes the action found at `<path>` using the supplied parameters.

This command is auto-generated from the YANG file.

For example, given the following action specification in a YANG file:

```yang
tailf:action shutdown {
  tailf:actionpoint actions;
  input {
    tailf:constant-leaf flags {
      type uint64 {
        range "1 .. max";
      }
      tailf:constant-value 42;
    }
    leaf timeout {
      type xs:duration;
      default PT60S;
    }
    leaf message {
      type string;
    }
    container options {
      leaf rebootAfterShutdown {
        type boolean;
        default false;
      }
      leaf forceFsckAfterReboot {
        type boolean;
        default false;
      }
      leaf powerOffAfterShutdown {
        type boolean;
        default true;
      }
    }
  }
}
```

The action can be invoked in the following way:

```bash
admin@ncs> shutdown timeout 10s message reboot options { \
    forceFsckAfterReboot true }
```

</details>

#### Builtin Commands

<details>

<summary><code>commit (abort | confirm)</code></summary>

Abort or confirm a pending confirming commit. A pending confirming commit will also be aborted if the CLI session is terminated without doing `commit confirm`. The default is confirm.

For the general CLI commit parameters available with `commit ?`, see [Commit Parameters](/guides/operation-and-usage/operations/lifecycle-operations#d5e5048) and the [examples.ncs/northbound-interfaces/commit-parameters](https://github.com/NSO-developer/nso-examples/tree/6.7/northbound-interfaces/commit-parameters) example.

Example:

```bash
admin@ncs# commit abort
```

</details>

<details>

<summary><code>config (exclusive | terminal) [no-confirm]</code></summary>

Enter configure mode. The default is `terminal`.

</details>

<details>

<summary><code>terminal</code></summary>

Edit a private copy of the running configuration; no lock is taken.

</details>

<details>

<summary><code>no-confirm</code></summary>

Enter configure mode ignoring any confirm dialog

Example:

```bash
admin@ncs# config terminal
Entering configuration mode terminal
```

</details>

<details>

<summary><code>file list &#x3C;directory></code></summary>

List files in `<directory>.`

Example:

```bash
admin@ncs# file list /config
rollback10001
rollback10002
rollback10003
rollback10004
rollback10005
```

</details>

<details>

<summary><code>file show &#x3C;file></code></summary>

Display contents of a `<file>`.

Example:

```bash
admin@ncs# file show /etc/skel/.bash_profile
# /etc/skel/.bash_profile

# This file is sourced by bash for login shells.  The following line
# runs our .bashrc and is recommended by the bash info pages.
[[ -f ~/.bashrc ]] && . ~/.bashrc
```

</details>

<details>

<summary><code>help &#x3C;command></code></summary>

Display help text related to `<command>.`

Example:

```bash
admin@ncs# help job
Help for command: job
    Job operations
```

</details>

<details>

<summary><code>job stop &#x3C;job id></code></summary>

Stop a specific background job. In the default CLI, the only command that creates background jobs is `monitor start`.

Example:

```bash
admin@ncs# monitor start /var/log/messages
[ok][...]
admin@ncs# show jobs
JOB COMMAND
3   monitor start /var/log/messages
admin@ncs# job stop 3
admin@ncs# show jobs
JOB COMMAND
```

</details>

<details>

<summary><code>logout session &#x3C;session></code></summary>

Log out a specific user session from NSO. If the user holds the `configure exclusive` lock, it will be released.

`<sessionid>`

Log out a specific user session.

Example:

```bash
admin@ncs# who
Session User  Context From         Proto Date     Mode
 25     oper  cli     192.168.1.72 ssh   12:10:40 operational
*24     admin cli     192.168.1.72 ssh   12:05:50 operational
admin@ncs# logout session 25
admin@ncs# who
Session User  Context From         Proto Date     Mode
*24     admin cli     192.168.1.72 ssh   12:05:50 operational
```

</details>

<details>

<summary><code>logout user &#x3C;username></code></summary>

Log out a specific user from NSO. If the user holds the `configure exclusive` lock, it will be released.

`<username>`

Log out a specific user.

Example:

```bash
admin@ncs# who
Session User  Context From         Proto Date     Mode
 25     oper  cli     192.168.1.72 ssh   12:10:40 operational
*24     admin cli     192.168.1.72 ssh   12:05:50 operational
admin@ncs# logout user oper
admin@ncs# who
Session User  Context From         Proto Date     Mode
*24     admin cli     192.168.1.72 ssh   12:05:50 operational
```

</details>

<details>

<summary><code>script reload</code></summary>

Reload scripts found in the `scripts/command`directory. New scripts will be added, and if a script file has been removed, the corresponding CLI command will be purged. See [Plug-and-Play Scripting](/guides/operation-and-usage/operations/plug-and-play-scripting).

</details>

<details>

<summary><code>send (all | &#x3C;user>) &#x3C;message></code></summary>

Display a message on the screens of all users who are logged in to the device or on a specific screen.

`all`

Display the message to all currently logged-in users.

`<user>`

Display the message to a specific user.

Example:

<pre class="language-bash"><code class="lang-bash"><strong>admin@ncs# send oper "I will reboot system in 5 minutes."
</strong></code></pre>

In the oper's session:

```bash
oper@ncs# Message from admin@ncs at 13:16:41...
I will reboot system in 5 minutes.
EOF
```

</details>

<details>

<summary><code>show cli</code></summary>

Display CLI properties.

Example:

```bash
admin@ncs# show cli
autowizard                   false
complete-on-space            true
display-level                99999999
dry-run-drift-detection      false;
dry-run-drift-detection-mode warn;
history                      100
idle-timeout                 1800
ignore-leading-space         false
output-file                  terminal
paginate                     true
prompt1                      \h\M#
prompt2                      \h(\m)#
screen-length                71
screen-width                 80
service prompt config        true
show-defaults                false
terminal                     xterm-256color
timestamp                    disable
```

</details>

<details>

<summary><code>show history [ &#x3C;limit> ]</code></summary>

Display CLI command history. By default, the last 100 commands are listed. The size of the history list is configured using the history CLI setting. If a history limit has been specified, only the last number of commands up to that limit will be shown.

Example:

```bash
admin@ncs# show history
06-19 14:34:02 -- ping router
06-20 14:42:35 -- show running-config
06-20 14:42:37 -- who
06-20 14:42:40 -- show history
admin@ncs# show history 3
14:42:37 -- who
14:42:40 -- show history
14:42:46 -- show history 3
```

</details>

<details>

<summary><code>show jobs</code></summary>

Display currently running background jobs.

Example:

```bash
admin@ncs# show jobs
JOB COMMAND
3   monitor start /var/log/messages
```

</details>

<details>

<summary><code>show parser dump &#x3C;command prefix></code></summary>

Shows all possible commands starting with the `<command prefix>`.

</details>

<details>

<summary><code>show running-config [ &#x3C;pathfilter> [ sort-by &#x3C;idx> ] ]</code></summary>

Display the current configuration. By default, the whole configuration is displayed. It is possible to limit what is shown by supplying a pathfilter.

The `<pathfilter>` maybe either a path pointing to a specific instance or, if an instance ID is omitted, the part following the omitted instance is treated as a filter.

The `sort-by` argument can be given when the `<pathfilter>` points to a list element with secondary indexes. `<idx>` is the name of a secondary index. When given, the table will be sorted in the order defined by the secondary index. This makes it possible for the CLI user to control in which order instances should be displayed.

To show the `aaa` settings for the `admin` user:

```bash
admin@ncs# show running-config aaa authentication users user admin
aaa authentication users user admin
 uid        1000
 gid        1000
 password   $1$JA.1O3Tx$Zt1ycpnMlg1bVMqM/zSZ7/
 ssh_keydir /var/ncs/homes/admin/.ssh
 homedir    /var/ncs/homes/admin
!
```

To show all users that have group ID 1000, omit the user ID and instead specify `gid` `1000`:

<pre class="language-bash"><code class="lang-bash"><strong>admin@ncs# show running-config aaa authentication users user * gid 1000
</strong>...
</code></pre>

</details>

<details>

<summary><code>show &#x3C;path> [ sort-by &#x3C;idx> ]</code></summary>

This command shows the configuration as a table provided that `<path>` leads to a list element, and the data can be rendered as a table (i.e., the table fits on the screen). It is also possible to force table formatting of a list by using the `| tab` pipe command.

The `sort-by` argument can be given when the *path* points to a list element with secondary indexes. `<idx>` is the name of a secondary index. When given, the table will be sorted in the order defined by the secondary index. This makes it possible for the CLI user to control in which order instances should be displayed.

Example:

```bash
admin@ncs# show devices device ce0 module
NAME                       REVISION    FEATURE  DEVIATION
-----------------------------------------------------------
tailf-ned-cisco-ios        2015-03-16  -        -
tailf-ned-cisco-ios-stats  2015-03-16  -        -
```

</details>

<details>

<summary><code>source &#x3C;file></code></summary>

Execute commands from \<file> as if they had been entered by the user. The `autowizard` is disabled when executing commands from the file; also, any commands that require input from the user (commands added by clispec, for example) will receive an interrupt signal upon an attempt to read from stdin.

</details>

<details>

<summary><code>timecmd &#x3C;command></code></summary>

Time command. It measures and displays the execution time of `<command>`.

Note that this command will only be available if `devtools` has been set to `true` in the CLI session settings.

Example:

```bash
admin@ncs# timecmd id
user = admin(501), gid=20, groups=admin, gids=12,20,33,61,79,80,81,98,100
Command executed in 0.00 sec
admin@ncs#
```

</details>

<details>

<summary><code>who</code></summary>

Display currently logged-on users. The current session, i.e., the session running the show status command, is marked with an asterisk.

Example:

```bash
admin@ncs# who
Session User  Context From         Proto Date     Mode
 25     oper  cli     192.168.1.72 ssh   12:10:40 operational
*24     admin cli     192.168.1.72 ssh   12:05:50 operational
admin@ncs#
```

</details>

### Configure Mode Commands <a href="#d5e2199" id="d5e2199"></a>

#### **Configure a Value**

<details>

<summary><code>&#x3C;path> [&#x3C;value>]</code></summary>

Set a parameter. If a new identifier is created and `autowizard` is enabled, then the CLI will prompt the user for all mandatory sub-elements of that identifier.

This command is auto-generated from the YANG file.

If no `<value>` is provided, then the CLI will prompt the user for the value. No echo of the entered value will occur if `<path>` is an encrypted value, i.e. of the type `ianach:crypt-hash` or one of `md5-digest-string`, `aes-cfb-128-encrypted-string`, or `aes-256-cfb-128-encrypted-string` as documented in the `tailf-common.yang` data model.

</details>

#### **Builtin Commands**

<details>

<summary><code>annotate &#x3C;statement> &#x3C;text></code></summary>

Associate an annotation with a given configuration. To remove an annotation, leave the text empty.

Only available when the system has been configured with attributes enabled.

</details>

<details>

<summary><code>commit (check | and-quit | confirmed | to-startup)</code><br><code>[comment &#x3C;text>] [label &#x3C;text>]</code></summary>

Commit the current configuration to "running".

* `check`: Validate current configuration.
* `and-quit`: Commit to running and quit configure mode.
* `comment <text>`: Associate a comment with the commit. The comment can later be seen when examining rollback files.
* `label <text>`: Associate a label with the commit. The label can later be seen when examining rollback files.

</details>

<details>

<summary><code>copy &#x3C;instance path> &#x3C;new id></code></summary>

Make a copy of an instance.

Copying between different `ned-id` versions works as long as the schema nodes being copied have not changed between the versions.

</details>

<details>

<summary><code>copy cfg [ merge | overwrite] &#x3C;src path> to &#x3C;dest path></code></summary>

Copy data from one configuration tree to another. Only data that makes sense at the destination will be copied. No error message will be generated for data that cannot be copied, and the operation can fail completely without any error messages being generated.

For example, to create a template from a part of a device config. First, configure the device, then copy the config into the template configuration tree.

```bash
admin@ncs(config)# devices template host_temp
admin@ncs(config-template-host_temp)# exit
admin@ncs(config)# copy cfg merge devices device ce0 config \
    ios:ethernet to devices template host_temp config ios:ethernet
admin@ncs(config)# show configuration diff
+devices template host_temp
+ config
+  ios:ethernet cfm global
+ !
+!
```

</details>

<details>

<summary><code>copy compare &#x3C;src path> to &#x3C;dest path></code></summary>

Compare two arbitrary configuration trees. Items that only appear in the `src` tree are ignored.

</details>

<details>

<summary><code>delete &#x3C;path></code></summary>

Delete a data element.

</details>

<details>

<summary><code>do &#x3C;command></code></summary>

Run the command in operational mode.

</details>

<details>

<summary><code>edit &#x3C;path></code></summary>

Edit a sub-element. Missing elements in the `<path>` will be created.

</details>

<details>

<summary><code>exit (level | configuration-mode)level</code></summary>

* `level`\
  Exit from this level. If performed on the top level, it will exit configure mode. This is the default if no option is given.
* `configuration-mode`\
  Exit from configuration mode regardless of which edit level.

</details>

<details>

<summary><code>help &#x3C;command></code></summary>

Shows help text for `<command>`.

</details>

<details>

<summary><code>hide &#x3C;hide-group></code></summary>

Re-hides the elements and actions belonging to the hide groups. No password is required for hiding. This command is hidden and not shown during command completion.

</details>

<details>

<summary><code>insert &#x3C;path></code></summary>

Inserts a new element. If the element already exists and has the `indexedView` option set in the data model, then the old element will be renamed to element+1, and the new element will be inserted in its place.

</details>

<details>

<summary><code>insert &#x3C;path>[ first| last| before key| after key]</code></summary>

Inject a new element into an ordered list. The element can be added first, last (default), before, or after another element.

</details>

<details>

<summary><code>load (merge | override | replace) (terminal | &#x3C;file>)</code></summary>

Load configuration from file or terminal.

* `merge`\
  Merge the content of the file/terminal with the current configuration.
* `override`\
  Configuration from file/terminal overwrites the current configuration.
* `replace`\
  Configuration from file/terminal replaces the current configuration.

If this is the current configuration:

```
devices device p1
 config
  cisco-ios-xr:interface GigabitEthernet 0/0/0/0
   shutdown
  exit
  cisco-ios-xr:interface GigabitEthernet 0/0/0/1
   shutdown
 !
!
```

The `shutdown` value for the entry `GigabitEthernet 0/0/0/0` should be deleted. As the configuration file is basically just a sequence of commands with comments in between, the configuration file should look like this:

```
devices device p1
 config
  cisco-ios-xr:interface GigabitEthernet 0/0/0/0
   no shutdown
  exit
 !
!
```

The file can then be used with the command ` load merge`` `` `*`FILENAME`* to achieve the desired results.

</details>

<details>

<summary><code>move &#x3C;path>[ first | last| before key | after key]</code></summary>

Move an existing element to a new position in an ordered list. The element can be moved first, last (default), before, or after another element.

</details>

<details>

<summary><code>rename &#x3C;instance path> &#x3C;new id></code></summary>

Rename an instance.

</details>

<details>

<summary><code>revert [no-confirm] [&#x3C;path>]</code></summary>

Copy the running configuration into the current configuration. Without a path, this removes all uncommitted changes in the current transaction.

If a path is provided, only changes under that path are reverted. The path is resolved relative to the current CLI sub-mode. For example, from the top level, `revert foo bar` reverts changes below `/foo/bar`, while from inside a `foo a` sub-mode, `revert no-confirm baz` reverts only the changes below that `baz` child in the current transaction.

The `no-confirm` option suppresses the confirmation prompt and can be used both with and without a path, for example `revert no-confirm foo bar`.

Contextual subtree revert is supported in private configuration mode. In shared or exclusive mode, use `revert` without a path to discard the whole transaction.

</details>

<details>

<summary><code>rload (merge | override | replace) (terminal | &#x3C;file>)</code></summary>

Load the file relative to the current sub-mode. For example, given a file with a device config, it is possible to enter one device and issue the `rload merge/override/replace <file>` command to load the config for that device, then enter another device and load the same config file using `rload`. See also the `load` command.

* `merge`\
  Merge the content of the file/terminal with the current configuration.
* `override`\
  Configuration from file/terminal overwrites the current configuration.
* `replace`\
  Configuration from file/terminal replaces the current configuration.

</details>

<details>

<summary><code>rollback-files apply-rollback-file (id | fixed-number)</code><br><code>&#x3C;number> [path &#x3C;path>] [selective]</code></summary>

Return the configuration to a previously committed configuration. The system stores a limited number of old configurations. The number of old configurations to store is configured in the `ncs.conf` file. If more than the configured number of configurations is stored, then the oldest configuration is removed before creating a new one.

The configuration changes are stored in rollback files where the most recent changes are stored in the file rollbackN with the highest number N.

Only the deltas are stored in the rollback files. When rolling back the configuration to rollback N, all changes stored in rollback10001-rollbackN are applied.

There are two ways to address which rollback file to use, either `fixed-number <number>` to address an absolute rollback number or `id <number>` to address a relative number. For example, the latest commit has a relative rollback ID of 0, the second-latest has ID 1, and so on.

The optional path argument allows subtrees to be rolled back while the rest of the configuration tree remains unchanged.

Instead of undoing all changes from rollback10001 to rollbackN it is possible to undo only the changes stored in a specific rollback file. This may or may not work depending on which changes have been made to the configuration after the rollback was created. In some cases applying the rollback file may fail, or the configuration may require additional changes in order to be valid. E.g., to undo the changes recorded in rollback 10019, but not the changes in 10020-N run the command `rollback-files apply-rollback-file selective fixed-number 10019`.

Example:

```bash
admin@ncs(config)# rollback-files apply-rollback-file fixed-number 10005
```

This command is only available if rollback has been enabled in `ncs.conf`.

</details>

<details>

<summary><code>show full-configuration [&#x3C;pathfilter> [sort-by &#x3C;idx>]]</code></summary>

Show the current configuration, taking local changes into account. The `show` command can be limited to a part of the configuration by providing a `<pathfilter>`.

The `sort-by` argument can be given when the `<pathfilter>` points to a list element with secondary indexes. `<idx>` is the name of a secondary index. When given, the table will be sorted in the order defined by the secondary index. This makes it possible for the CLI user to control in which order instances should be displayed.

</details>

<details>

<summary><code>show configuration [&#x3C;pathfilter>]</code></summary>

Show current edits to the configuration.

</details>

<details>

<summary><code>show configuration merge [&#x3C;pathfilter> [sort-by &#x3C;idx>]]</code></summary>

Show the current configuration, taking local changes into account. The `show` command can be limited to a part of the configuration by providing a `<pathfilter>`.

The `sort-by` argument can be given when the `<pathfilter>` points to a list element with secondary indexes. `<idx>` is the name of a secondary index. When given, the table will be sorted in the order defined by the secondary index. This makes it possible for the CLI user to control in which order instances should be displayed.

</details>

<details>

<summary><code>show configuration commit changes [&#x3C;number> [&#x3C;path>]]</code></summary>

Display edits associated with a commit, identified by the rollback number created for the commit. The changes are displayed as forward changes, as opposed to `show configuration rollback changes`, which displays the commands for undoing the changes.

The optional path argument allows only edits related to a given subtree to be listed.

</details>

<details>

<summary><code>show configuration commit list [&#x3C;path>]</code></summary>

List rollback files

The optional path argument allows only rollback files related to a given subtree to be listed.

</details>

<details>

<summary><code>show configuration rollback listed [&#x3C;number>]</code></summary>

Display the operations needed to undo the changes performed in a commit associated with a rollback file. These are the changes that will be applied if the configuration is rolled back to that rollback number.

</details>

<details>

<summary><code>show configuration running [&#x3C;pathfilter>]</code></summary>

Display the "running" configuration without taking uncommitted changes into account. An optional `<pathfilter>` can be provided to limit what is displayed.

</details>

<details>

<summary><code>show configuration diff [&#x3C;pathfilter>]</code></summary>

Display uncommitted changes to the running config in diff-style, i.e., with + and - in front of added and deleted configuration lines.

</details>

<details>

<summary><code>show parser dump &#x3C;command prefix></code></summary>

Shows all possible commands starting with `<command>` prefix.

</details>

<details>

<summary><code>tag add &#x3C;statement> &#x3C;tag></code></summary>

Add a tag to a configuration statement.

Only available when the system has been configured with attributes enabled.

</details>

<details>

<summary><code>tag del &#x3C;statement> &#x3C;tag></code></summary>

Remove a tag from a configuration statement.

Only available when the system has been configured with attributes enabled.

</details>

<details>

<summary><code>tag clear &#x3C;statement></code></summary>

Remove all tags from a configuration statement.

Only available when the system has been configured with attributes enabled.

</details>

<details>

<summary><code>timecmd &#x3C;command></code></summary>

Time command. It measures and displays the execution time of `<command>`.

Note that this command will only be available if `devtools` has been set to `true` in the CLI session settings.

Example:

```bash
admin@ncs# timecmd id
user = admin(501), gid=20, groups=admin, gids=12,20,33,61,79,80,81,98,100
Command executed in 0.00 sec
admin@ncs#
```

</details>

<details>

<summary><code>top [command]</code></summary>

Exit to the top level of the configuration, or execute a command at the top level of the configuration.

</details>

<details>

<summary><code>unhide &#x3C;hide-group></code></summary>

Unhides all elements and actions belonging to the `<hide-group>`. It may be required to enter a password. This command is hidden and not shown during command completion

</details>

<details>

<summary><code>validate</code></summary>

Validates current configuration. This is the same operation as `commit check`.

</details>

<details>

<summary><code>xpath [ctx &#x3C;path>] (eval | must | when) &#x3C;expression></code></summary>

Evaluate an XPath expression. A context path may be given to be used as the current context for the evaluation of the expression. If no context path is given, the current sub-mode will be used as the context path. The pipe command `trace` may be used to display debug/trace information during the execution of the command.

Note that this command will only be available if `devtools` has been set to `true` in the CLI session settings.

* `eval`\
  Evaluate an XPath expression.
* `must`\
  Evaluate the expression as a YANG must expression.
* `when`\
  Evaluate the expression as a YANG when expression.

</details>

<details>

<summary><code>reapply-commands [best-effort | list]</code></summary>

Reapply entered config commands since the latest commit. The command will stop on the first error by default.

Commands that may have unknown side effects, will be skipped and thus not reapplied, such as actions, custom commands, etc. To display all commands, including those that will be skipped, the pipe command `details` can be used.

Note that this command will only be available if there is a conflict.

`best-effort`

Do not stop on the first error but continue to process the rest of the commands.

`list`

Display the current set of commands.

</details>


# Web UI

Operate NSO using the Web UI.

The NSO Web UI provides an intuitive northbound interface to your NSO deployment. The UI consists of individual views, each with a different purpose to perform operations such as device management, service management, commit handling, etc.

The main components of the Web UI are shown in the figure below.

<div data-with-frame="true"><figure><img src="/files/ICYxtqtDeVfbCKXmY5sU" alt=""><figcaption><p>NSO Web UI Overview</p></figcaption></figure></div>

The UI works by auto-rendering the underlying device and service models. This gives the benefit that the Web UI is immediately updated when new devices or services are added to the system. For example, say you have added support for a new device vendor. Then, without any programming requirements, the NSO Web UI provides the capability to configure those devices.

{% hint style="info" %}
It's important to understand that the bulk of concepts and configuration options in Web UI are shared with the NSO CLI. The rest of the documentation covers these in detail. You need to be familiar with the fundamental concepts to work with the Web UI.
{% endhint %}

## Browser Requirements <a href="#d5e5676" id="d5e5676"></a>

All modern web browsers are supported, and no plug-ins are needed. The interface itself is a JavaScript client.

## Accessing the Web UI <a href="#d5e5679" id="d5e5679"></a>

By default, the Web UI is accessible on port 8080 of the NSO server for an NSO Local Install and port 8888 for a System Install. The port can be changed in the `ncs.conf` file. Users are required to authenticate before accessing the Web UI.

## Basic Operations <a href="#d5e5683" id="d5e5683"></a>

### **Log In**

Log in to the NSO Web UI by using the username and password provided by your administrator. SSO SAML or OIDC login is available if set up by your administrator. If applicable, use the SSO option to log in.

### **Log Out**

Log out by clicking your username on the top-right corner and choosing **Logout**.

### Theme

Apply a theme for the user interface by clicking your username and selecting from **Light**, **Dark**, or **System** **default**.

### **Help Options**

Access the help options by clicking the help options icon in the UI banner. The following options are available:

* **Online documentation**: Access the Web UI's online help.
* **Manage hidden groups**: Administer hidden groups, e.g., for debugging. Read more about hide groups in [CLI Commands](/guides/operation-and-usage/cli/cli-commands).
* **NSO version**: Information about the version of NSO you are running.

In the Web UI, supplementary help text, whenever applicable, is available on the configuration fields and can be accessed by clicking the info icons.

## Dirty State

Anytime a configuration is changed in the Web UI (such as a device or service configuration change), the UI reflects the change with a so-called color-coded "dirty state" with the following meanings:

* <mark style="color:blue;">Blue</mark> color: An addition or a modification to an already-committed list element was made.
* <mark style="color:red;">Red</mark> color: A deletion was made.

## Commit Management <a href="#d5e5718" id="d5e5718"></a>

Commit options are accessible at all times from the UI header. A number, corresponding to the number of changes in a transaction, is displayed next to the <img src="/files/8ovrmr5SxeUoWMlfMzwS" alt="" data-size="line"> icon when changes are available for review. These changes can be reviewed in the **Transactions** view. For certain actions, it is possible to skip the commit review and apply the changes directly.

{% hint style="warning" %}
**Transactions and Commits**

Take special note of commit management. Whenever a transaction has started, the active configuration data changes can be inspected and evaluated before they are committed and pushed to the network. The data is saved to the NSO datastore and pushed to the network when a user presses **Commit**.

Any network-wide configuration change can be picked up as a rollback file. The rollback can then be applied to undo whatever happened to the network.
{% endhint %}

### **Review a Configuration Change**

To review available configuration changes:

1. Access commit management by clicking its icon <img src="/files/Jd939sKUXxRwrQAfvpro" alt="" data-size="line"> in the banner.
2. Review the available changes by clicking the **Changes** or **Transactions** option. This action redirects you to the **Transactions** view.
3. Press **Validate** to check for errors. All changes must be validated before they can be committed.
4. Click **Revert** to undo or **Commit** to confirm the changes in the transaction.
5. If you are committing a change, **Commit Settings** are shown after pressing **Commit**. Examples of commit settings include: **No revision drop**, **No deploy**, **No networking**, etc. Commit options are described in detail in the JSON-RPC API documentation under [Methods - transaction - commit changes](/guides/development/advanced-development/web-ui-development/json-rpc-api#methods-transaction-commit-changes).


# Home

Home page of NSO Web UI.

The **Home** view is the default view after logging in. It provides shortcuts to **Devices**, **Services**, **Config editor**, and **Tools**.

<div data-with-frame="true"><figure><img src="/files/GpchtHVK3i68WGyYtFYa" alt=""><figcaption><p>Home View</p></figcaption></figure></div>

## Web UI Extension Packages

Currently loaded Web UI extension packages are shown in this view under **Packages**. Web UI packages are used to extend the functionality of your Web UI, for example, to create additional views and functionalities. Examples are creating a view to visualize your MPLS network, etc.


# Devices

Manage devices, device groups, and authgroups in your NSO deployment.

The **Devices** section provides options to manage devices, device groups, and authgroups in the NSO network.

## Device Management <a href="#d5e5752" id="d5e5752"></a>

The **Device management** view lists the devices in the network and provides options to manage them. Expand a device in the list to view more details about it.

<div data-with-frame="true"><figure><img src="/files/4YD7Hf7Ls4s6CuF5jnfp" alt=""><figcaption><p>Device Management View</p></figcaption></figure></div>

### **Search and Filter Devices**

You can search for a device by its name, IP address, or other parameters. You can also narrow down the results by using the **Select device group** filter.

### **Add a Device**

To add a new device to NSO:

1. Click the **Add device** button. You will be redirected to the **Configuration editor**.
2. Click the **Add list item** button.
3. Enter the name of the device.
4. Click the device name in the list to configure the device further.
5. Review and commit the changes in the **Transactions** view.

### **Apply an Action on a Device**

Actions can be applied on a device from the **Device management** view or the **Configuration editor**.

{% tabs %}
{% tab title="From the Device Management View" %}
An action can be applied to a single or multiple devices at once.

1. Select the device(s) from the list using the checkbox.
2. Using the **Choose actions** button, select the desired action. The result of the action is returned momentarily.

{% hint style="info" %}
In the **Device management** view, you can also apply actions on a device using the more options <img src="/files/sstclOvFs1yjOiJgltpV" alt="" data-size="line"> button.
{% endhint %}

**Actions Possible in the Device Management View**

Available actions include **Connect**, **Ping**, **Sync from**, **Sync to**, **Check sync**, **Compare config**, **Fetch ssh host keys**, and **Apply template**, and. See [Lifecycle Operations](/guides/operation-and-usage/operations/lifecycle-operations) for the details of these actions.

{% hint style="info" %}
The **Modify in Config Editor** and **Delete** are GUI-specific operations accessible by clicking the more options <img src="/files/sstclOvFs1yjOiJgltpV" alt="" data-size="line"> button on the device row.
{% endhint %}
{% endtab %}

{% tab title="From the Configuration Editor" %}
Additional actions are applied to an individual device. Use this option if you want to run an action with additional parameters.

1. Click the device name in the list. You will be redirected to the **Configuration editor** view.
2. Click the **Actions** button.
3. Click the desired action in the list.
4. At this point, you can configure different parameters.
5. Click **Run** to initiate the action.

**Actions Possible in the Configuration Editor -> Actions Tab**

If you access the device in the **Configuration editor**, the following additional actions are available:

**migrate**, **instantiate-from-other-device**, **check-yang-modules**, **scp-to**, **copy-capabilities**, **compare-config**, **connect**, **scp-from**, **find-capabilities**, **sync-from**, **disconnect**, **rename**, **add-capability**, **sync-to**, **ping**, **load-native-config**, **apply-template**, **check-sync**, **delete-config**, **clear-trace**, and **fetch-host-keys**,

See [Lifecycle Operations](/guides/operation-and-usage/operations/lifecycle-operations) for the details of these actions.
{% endtab %}
{% endtabs %}

### **Edit Device Configuration**

To edit the configuration of an existing device:

1. In the **Devices** view, locate the desired device.
2. Open the device in the **Configuration editor** by either clicking the device name in the list and then enabling **Edit mode**, or by clicking the more options button <img src="/files/sstclOvFs1yjOiJgltpV" alt="" data-size="line"> on the device row and selecting **Modify in Config Editor**, which opens the device with edit mode already enabled.
3. Make the desired changes. Depending on the field type, changes are applied automatically when you update the value or when you leave the field.
4. Review and commit the change in the **Transactions** view.

## Device Groups <a href="#d5e5978" id="d5e5978"></a>

The **Device groups** view lists all the available groups and devices belonging to them. You can add new device groups in this view as well as carry out actions on devices belonging to a group.

<div data-with-frame="true"><figure><img src="/files/BqZUoJyHPuoaKPmzh0A9" alt=""><figcaption><p>Device Groups View</p></figcaption></figure></div>

### **Create a Device Group**

Device groups allow for the grouping and collective management of devices.

1. Click **Add device group**.
2. In the **Create device group** pop-up, specify the group name.
   * If you want to place the new device group under a parent group, select the **Place under parent device group** option and specify the parent group.
3. Click **Create**. You will be redirected to the group's details page. Here, the following panes are available:
   * **Details**: Displays basic details of the group, i.e., its name and parent/subgroup information. To link a sub-group, use the **Connect sub device group** option.
   * **Devices in this group**: Displays currently added devices in the group and provides the option to remove them from the group.
   * **Add devices**: Displays all available NSO devices and provides the option to add them to the group.
4. In the **Add devices** pane, select the device(s) that you want to add to the new group and click **Add to device group**. The added devices become visible under the **Devices in this group** pane.
5. Finally, click **Create device group**.

### **Remove Device(s) from a Device Group**

1. Click the desired device group to access the group's detail page.
2. In the **Devices in this group** pane, select the device(s) to be removed from the group.
3. Click **Remove from device group**. The devices are removed immediately (without a Commit Manager review).
4. Click **Save device group**.

### **Apply an Action on a Device Group**

Device group actions let you perform an action on all the devices belonging to a group.

1. Select the desired device group from the list. It is possible to select multiple groups at once.
2. Choose the desired action from the **Choose actions** button.

{% hint style="info" %}
In the **Device groups** view, you can also apply actions on a device group using the more options <img src="/files/sstclOvFs1yjOiJgltpV" alt="" data-size="line"> button.
{% endhint %}

**Actions Possible in the Device Groups View**

The available group actions are the same as in the section called [Apply an Action on a Device](#apply-an-action-on-a-device) (e.g., **Connect**, **Sync from**, **Sync to**, etc.) and are described in [Lifecycle Operations](/guides/operation-and-usage/operations/lifecycle-operations).

{% hint style="info" %}
The **Modify in Config editor** option is accessible by clicking the more options <img src="/files/sstclOvFs1yjOiJgltpV" alt="" data-size="line"> button on a device group.
{% endhint %}

## Authgroups

The **Authgroups** view display device authentication groups and provides ways to manage them. Concepts and settings involved in the authentication groups setup are discussed in [NSO Device Management](https://nso-docs.cisco.com/guides/operation-and-usage/webui/pages/auKQMOAF2p1jiGYJBweP#user_guide.devicemanager.authgroups).

The **Authgroups** views include:

* The **Authgroups** - **group** view
* The **Authgroups** - **SNMP group** view

### Authgroups - group

The **Authgroups** - **group** view is used to view, search, and manage device authentication groups for CLI and NETCONF-managed devices.

<div data-with-frame="true"><figure><img src="/files/qZeyKnCR26TBYlo7C3rE" alt=""><figcaption><p>Authgroups View (Group)</p></figcaption></figure></div>

#### Create an Authgroup

To create a new group:

1. Click the **Add authgroup** button.
2. Enter the **Authgroup name** and click **Continue**.
3. In the group details page, add users to the newly created group. If a default map is desired for unknown/unmapped users, use the **Set default-map** option.
   1. Click the **Add user** button to bring up the **Add user** overlay window. Here, you have the option to add the user with the authentication type set to 'remote mapping' or 'callback':
      * Remote mapping: If remote mapping is desired, specify the **local-user** that is to be mapped to remote authentication credentials and configure the following settings:
        * **remote-user**: Choose between **same-user** or **remote-name** options.
        * **remote-auth**: Choose between **same-pass**, **remote-password**, or **public-key** options.
        * **remote-secondary-auth** (optional): Choose between **same-secondary-password** or **remote-secondary-password** options.
      * Callback: If a callback-type authentication is desired to retrieve login credentials, specify the **local-user**, set the **Use callback** flag, and configure the following settings:
        * **callback-node**
        * **action-name**
   2. Click **Add**. This adds the newly created user to the group and displays it in the list.
4. Click **Create** **authgroup** to save and finish creating the group.

#### View/Edit Authgroup Details

To view/edit details of a group:

1. Click the group name to access the group details page.
2. Make the desired changes, such as adding/removing a user from the group, editing existing user settings, or configuring general group settings.
3. Click the **Save authgroup** button to save and apply the changes.

#### Delete an Authgroup

To delete a group:

{% hint style="warning" %}
Proceed with caution as the changes are applied immediately.
{% endhint %}

1. Select the desired group using the checkbox.
2. Click **Delete**.
3. Confirm the intent by pressing **Delete** in the pop-up.

### Authgroups - SNMP group

The **Authgroups** **-** **SNMP group** view is used to view, search, and manage device authentication groups for SNMP-managed devices.

<div data-with-frame="true"><figure><img src="/files/Cm8tNtxvMSAG6ZEJl9VK" alt=""><figcaption><p>Authgroups View (SNMP Group)</p></figcaption></figure></div>

#### Create an SNMP Group

To add a new group:

1. Click the **Add SNMP group** button.
2. Enter the **SNMP group name** and click **Continue**.
3. In the group details page, add users to the newly created group. If a default map is desired for unknown/unmapped users, use the **Set default-map** option.
   1. Click the **Add user** button to bring up the **Add user** overlay window.
   2. Specify the **local-user** and configure the following settings:
      * **remote-user**: Choose between **same-user** or **remote-name** options.
      * **security-level**: Use **auth-priv**. Then configure SHA as the authentication protocol and AES as the privacy protocol, together with the required remote passwords.
   3. Click **Add**. This adds the newly created user to the group and displays it in the list.
4. Click **Create** **SNMP** **group** to save and finish creating the group.

#### View/Edit SNMP Group Details

To view/edit details of a group:

1. Click the group name to access the group details page.
2. Make the desired changes, such as adding/removing a user from the group, editing existing user settings, or configuring general group settings.
3. Click the **Save SNMP group** button to save and apply the changes.

#### Delete an SNMP Group

To delete a group:

{% hint style="warning" %}
Proceed with caution as the changes are applied immediately.
{% endhint %}

1. Select the desired group using the checkbox.
2. Click **Delete**.
3. Confirm the intent by pressing **Delete** in the pop-up.


# Services

Create and manage service deployment.

The **Services** view is used to view, create, and manage services in your NSO deployment. The default **Services** view displays the existing services.

<div data-with-frame="true"><figure><img src="/files/vncz6ZKtTxVNtPKTOGJC" alt=""><figcaption><p>Services View</p></figcaption></figure></div>

## Search <a href="#d5e6128" id="d5e6128"></a>

If you have multiple services configured, you can use the **Search** to filter down results to the service of your choice. The search filter matches the entered characters to the service name and shows the results accordingly. Results are shown only for the service point that you have selected.

To filter the service list:

1. In the **Select service type** drop-down, select the service point to populate all the services under it.
2. Enter a partial or full name of the service you are searching for.
3. Press **Enter**.

## Create a Service <a href="#d5e6142" id="d5e6142"></a>

To create and deploy a service:

1. In the **Select service type** drop-down, select the service point.
2. Click the **Add service** button. You will be redirected to the **Configuration editor** view.
3. Click the <img src="/files/o97oqkgV6qWz83p1qTYI" alt="" data-size="line"> button.
4. In the pop-up, enter the name of the service that will identify it.
5. Click **Create**.
6. You can configure additional service data in the **Configuration editor**.
7. Review and commit the service to NSO in the **Transactions** view. Committing the service deploys it to NSO and displays it in the **Services** view.

## Edit Service Configuration <a href="#d5e6291" id="d5e6291"></a>

Service configuration is viewed and carried out in the Configuration Editor. In the **Services** view, you can use the **Modify in Config Editor** option on the desired service to access its config in the Configuration Editor. You need to have **Edit mode** enabled to perform edits.

{% hint style="warning" %}
The **Configuration editor** view shows a host of options when configuring a service. You are expected to be well-versed with these options (and service concepts in general) before you delve into service configuration. Refer to the [Services](/guides/development/core-concepts/services) and [Developing Services](/guides/development/advanced-development/developing-services) documentation for more information.
{% endhint %}

## Apply an Action on a Service <a href="#d5e6164" id="d5e6164"></a>

You can apply actions on a service from the **Services** view or the **Configuration editor**.

Start by selecting the service point to populate all services under it and then follow the instructions below:

{% tabs %}
{% tab title="From the Services View" %}
To apply an action on a service:

1. On the desired service in the list, click the more options <img src="/files/sstclOvFs1yjOiJgltpV" alt="" data-size="line"> button.
2. Choose the preferred action from the list, i.e., **Re-deploy**, **Un-deploy**, **Check sync**, **Deep check sync**, or **get modifications**.

{% hint style="info" %}
The **Check sync** action can be run on multiple services at once by selecting them using the checkbox and then running the action using the **Choose actions** button.
{% endhint %}

**Actions Possible in the Services View**

Available actions include **Re-deploy**, **Un-deploy**, **Check sync**, **Deep check sync**, and **get modifications**. See [Lifecycle Operations](/guides/operation-and-usage/operations/lifecycle-operations) for the details of these actions.

{% hint style="info" %}
The **Modify in Config Editor** and **Delete** are GUI-specific operations accessible on the service row.
{% endhint %}
{% endtab %}

{% tab title="From the Configuration Editor " %}
Additional actions are applied to an individual service. Use this option if you want to run an action with additional parameters.

1. Access the service in the Configuration Editor by selecting the **Modify in Config Editor** option on a service.
2. Click the **Actions** button and select the desired action in the list.
3. Configure parameters to run.
4. Click **Run** to initiate the action.

**Actions Possible in the Configuration Editor -> Actions Tab**

Access the service in the **Configuration editor** to run the following actions: **check-sync**, **reactive-re-deploy**, **un-deploy**, **deep-check-sync**, **touch**, **set-rank**, **re-deploy**, **get-modifications**, and **purge.** See [Lifecycle Operations](/guides/operation-and-usage/operations/lifecycle-operations) for the details of these actions.
{% endtab %}
{% endtabs %}

## View Service Details

To view details of a service:

1. In the **Select service type** drop-down, select the service point.
2. Click the desired service. This opens up the service details view.
3. Browse service details using the following tabs:
   * **Details**
   * **Plan** (hidden if no plan configured)
   * **Log**

## Delete a Service <a href="#d5e6324" id="d5e6324"></a>

To delete a service instance:

1. In the **Select service type** drop-down list, select the service point.
2. Select, using the checkbox, the service to be deleted. You can select multiple services at once.
3. Click **Delete**.
4. Confirm the intent in the pop-up.
5. Review and commit the change in the **Transactions** view.

{% hint style="info" %}
To skip the commit management (Transactions) review, use the **Commit changes directly** option in the **Delete service instance** pop-up.
{% endhint %}


# Config Editor

Traverse and edit NSO configuration using the YANG model.

The **Configuration editor** view is the main interface for browsing and managing NSO configuration using the underlying YANG model. It displays the loaded YANG modules and configuration data in a hierarchical tree and provides a form-based view of the selected node.

In this view, you can browse configuration, inspect operational data, edit configurable nodes, invoke actions, and review metadata for the selected node. Depending on how you navigate in the Web UI, you may also be directed to the **Configuration editor** to continue viewing or editing a specific device, service, package, or other NSO object.

The **Configuration editor** consists of the following main parts:

* A navigation tree for browsing the YANG hierarchy.
* A form-based content area that renders the selected node.
* A metadata panel that shows details about the selected node.
* Context-sensitive actions for nodes that support actions or presence operations.

<div data-with-frame="true"><figure><img src="/files/WYe9AJ44rlbvtAk68kIG" alt=""><figcaption><p>Configuration Editor Main View</p></figcaption></figure></div>

## View Options

The following options are available in the **Configuration editor** view:

* **Include oper data**: Displays operational data together with the configuration tree. Use this option when you want to inspect runtime or status information in addition to configured data.
* **Edit mode**: Enables editing in the **Configuration editor**. When edit mode is enabled, configurable nodes can be modified directly in the rendered form or list view. When it is disabled, the view is read-only.

## Configuration Navigation

The Configuration Editor displays configuration defined by the YANG model in a hierarchical tree structure that you can browse and navigate. Expand nodes to reveal their child nodes, and use the filter field to narrow the visible nodes in the current tree.

Selecting a node in the tree updates the content area to show the selected node. The rendered view depends on the type of YANG node selected. The tree itself shows the hierarchy, while the content area shows the data and controls available for the selected node.

For example, to access a specific device, enter **devices** in the filter field, expand **ncs:devices**, select **device** to load the device list, and then select the **ce0** entry to view or configure it.

<div data-with-frame="true"><figure><img src="/files/eZKvEF4OnMiSGcFfrEJV" alt=""><figcaption><p>Configuration Navigation Example</p></figcaption></figure></div>

### Breadcrumb

The breadcrumb above the tree shows the current navigation path, for example **/ > ncs:devices > device > {ce1}**. This helps identify the selected node and the current root of the rendered view as you move deeper into the YANG hierarchy.

### Tree Icons

The tree shows the following icons depending on the selection of view options:

| Icon                                                                              | Meaning                                                  |
| --------------------------------------------------------------------------------- | -------------------------------------------------------- |
| Key icon <img src="/files/VQMuWhzilSehcTVVQQ4m" alt="" data-size="line">          | A key leaf, meaning a value that identifies a list entry |
| Leaf icon <img src="/files/GUC0w69l8rb0RoWNvoIC" alt="" data-size="line">         | A regular leaf node                                      |
| List icon <img src="/files/sLWlWIogruiQojSikJWr" alt="" data-size="line">         | A list and leaf-list node                                |
| Split/arrows icon <img src="/files/qlPGK0yw8sTR0Dlq6P24" alt="" data-size="line"> | A choice node                                            |
| Lightning icon <img src="/files/sQlZa0NU0GJqTaMOVgyy" alt="" data-size="line">    | An action node                                           |
| Database icon <img src="/files/6D7wkU3aHcNNefvf5i1a" alt="" data-size="line">     | A node with operational data                             |

{% hint style="info" %}
Nodes that contain uncommitted changes are marked with a visual indicator (blue dot) in the tree. This help users identify pending edits before committing without needing to open each node individually.
{% endhint %}

### Rendered Node View

When a node is selected, the content area on the right renders that node as a form or list-based view. Leaf values are shown directly as fields, while nested containers, lists, and leaf-lists must be selected from the tree to be rendered separately.

#### Input Fields and Behavior

Input fields in the rendered view depend on the YANG type of the selected node. For example, fields may be shown as text inputs, selection controls, or other widgets depending on the allowed values and schema definition.

Field descriptions and default values are shown in the rendered view where applicable. Additional node details, such as type and access information, are available in the **Metadata** panel.

Depending on the field type, changes are applied automatically when you update the value or when you leave the field. If a value is invalid or cannot be applied, the UI indicates the error on the relevant field so that you can correct it before committing the change.

#### Choice

YANG choices are rendered as grouped choice widgets in the content area. In the tree, a choice or case branch can be expanded for navigation, but selecting the parent context renders the choice as a whole widget, while selecting an individual child node renders only that specific node.

#### Metadata

The **Metadata** panel displays schema and node information for the currently selected tree node. Expand this panel to inspect additional details about the selected node.

Depending on the node, the metadata can include information such as node type, kind, access, description, default value, etc. The metadata panel is the primary way to inspect node details that are not shown directly in the rendered field view.

#### Actions

For selected container nodes that define actions, the **Actions** button is shown above the metadata panel and remains available even when **Edit mode** is disabled. In the tree, however, actions are shown only when **Edit mode** is enabled. If no actions are available for the selected container, the **Actions** button is hidden.

#### Presence Containers

If the selected node is a presence container, the rendered view provides controls to create or delete that container.


# Transactions

Review uncommitted changes, commit queue entries, and rollback files in the Web UI.

The **Transactions** view lets you view and manage current NSO transactions. It provides a centralized way to inspect and manage configuration changes in your NSO deployment. You can review uncommitted changes, monitor commit queue activity, validate or revert active changes, and work with rollback files.

<figure><img src="/files/nRSViWEohaeeEuG5IErc" alt=""><figcaption></figcaption></figure>

The **Transactions** view is further divided into the following tabs:

* **Uncommitted**: Shows the changes currently present in the active transaction and provides options to validate, export, load, or revert them.
* **Commit queue**: Displays queued commit operations and their status.
* **Rollback files**: Lists rollback files created for committed changes so that earlier transactions can be inspected or undone.

## Uncommitted Changes

The **Uncommitted** tab shows the configuration changes currently made in NSO but not yet committed. The changes are presented with their path, operation, old value, and new value.

From this view, you can review the active transaction before committing it.

### Validate Changes

Use **Validate** to verify the current transaction before committing it. Validation helps detect issues in the pending configuration so they can be corrected before the changes are applied.

### Revert Changes

Use **Revert changes** to discard the current uncommitted transaction. This removes the pending changes from the active transaction.

### Export Transaction

Use **Export transaction** to save the current uncommitted transaction for later use or review.

### Load Transaction

Use **Load transaction** to load configuration data into the current transaction. This can be used to continue working with previously saved changes or to import changes from another source. Options include:

* **Use** **File**: Use a file from your local disk.
* **Use** **Data**: Paste in the configuration date.

## Commit Queue

The **Commit queue** tab displays configuration changes that have been committed to NSO and queued for delivery to devices. Use this tab to monitor queued operations and their progress.

This tab is divided into the following subtabs:

* **Queue**: Displays commit queue items that are waiting to be processed or are currently being processed.
* **Completed**: Displays commit queue items that have finished processing and are kept as completed results.

For more information about commit queue behavior, see [Commit Queue](https://nso-docs.cisco.com/guides/operation-and-usage/webui/pages/auKQMOAF2p1jiGYJBweP#user_guide.devicemanager.commit-queue).

### Queue

The **Queue** subtab displays commit queue entries that are pending or in progress. Use this subtab to monitor queued operations and their progress before they complete.

### Completed

The **Completed** subtab displays commit queue entries that have finished processing. This view can be used to review previously-processed queue items and their associated result information.

The completed results list includes the following information:

* **Status**: The final status of the completed queue item.
* **ID**: The identifier assigned to the queue item.
* **Label**: The label associated with the queue item, if available.
* **Devices**: The devices affected by the queue item.
* **Date**: The date and time when the result was recorded.

You can search the completed results list by using the **Search** field.

#### Purge Completed Results

Use **Purge** in the **Completed** subtab to remove completed commit results from the list. This action is useful when you want to clean up historical commit queue results that are no longer needed for review.

When you click **Purge**, the **Purge completed queue items** dialog is displayed. In this dialog, you can specify which queue items to remove. Options include:

* **status**: Select the result status of the queue items to purge. Available values are **completed**, **deleted**, and **failed**.
* **older-than**: Purge items older than specified seconds, minutes, hours, days, or weeks. The time fields can be used to define the age of queue items that should be removed.

To purge completed queue items:

1. Open the **Commit queue** tab and select the **Completed** subtab.
2. Click **Purge**.
3. In the **Purge completed queue items** dialog, specify the desired **status** and age criteria.
4. Click **Run**.

Use **Close** to exit the dialog without removing any completed queue items.

## Rollback Files

The **Rollback files** tab lists rollback files generated by NSO for committed transactions. These files can be used to inspect earlier changes and, when needed, restore configuration to a previous state.

Rollback files provide a convenient way to undo network-wide configuration changes after they have been committed.


# Tools

Tools to view NSO status and perform specialized tasks.

The **Tools** view includes utilities that you can use to run specific tasks on your deployment.

<div data-with-frame="true"><figure><img src="/files/pwQU299PSmH9OwsE5yUB" alt=""><figcaption><p>Tools View</p></figcaption></figure></div>

The following tools are available:

* [**Insights**](#d5e6470): Gathers and displays useful statistics of your deployment.
* [**Packages**](#d5e6487): Used to perform upgrades to the packages running in NSO.
* [**High availability**](#d5e6538): Used to manage a High Availability (HA) setup in your deployment.
* [**Alarms**](#d5e6565): Shows current alarms/events in your deployment and provides options to manage them.
* [**Compliance reporting**](#sec.webui_compliance): Used to run compliance checks on your NSO network.

## Insights <a href="#d5e6470" id="d5e6470"></a>

The **Insights** view collects and displays the following types of operational information using the `/ncs:metrics` data model to present useful statistics:

* Real-time data about transactions, commit queues, and northbound sessions.
* Sessions created and closed towards northbound interfaces since the last restart (CLI, JSON-RPC, NETCONF, RESTCONF, SNMP).
* Transactions since the last restart (committed, aborted, and conflicting). You can select between the running and operational data stores.
* Devices and their sync statuses.
* CDB info about its size, compaction, etc.

## Packages <a href="#d5e6487" id="d5e6487"></a>

In the **Packages** view, you can upload, install, and view the operational state of custom packages in NSO.

<div data-with-frame="true"><figure><img src="/files/eCKu2zE92QCf3wCLEuky" alt=""><figcaption><p>Packages View</p></figcaption></figure></div>

### Add a Package

Adding a new package via the Web UI entails uploading the package and then installing it. You can add multiple packages at once.

A package can be in one of the following states:

* **Up**: Package is installed and operational.
* **Not installed**: The package is uploaded but not installed.
* **Error**: Information that an error has occurred.

To add a new package:

1. Click the **Add package** button.
2. In the **Add package** dialog, browse the package using the **Add** button. The file format must be `.tar`, `.tar.gz`, or `.tgz`. You can add multiple packages at once.
3. Click **Upload**. A result is shown whether the operation was successful or not.
4. Once the upload has finished successfully, select the packages to install. If you want to replace an existing package with a new one, use the **Replace package if already exists** option, and to bypass or ignore version mismatches, use the **Allow NSO mismatch** option.
5. Click **Install**. A result is shown whether the operation was successful or not. For more details and troubleshooting the errors, see the trace output.
6. Perform a reload of packages if required. This can, for example, be if you uploaded and installed a new package version (e.g., Version 2.0) that subsequently requires a package reload to become operational. After running the package reload, the state of the package changes to **Up**.

### View Package Details

To view package details:

* Click the package name. This reveals information about the package, such as its status, version, location, etc. You can also uninstall a package in this view.

### Reload Packages

The reload action is the equivalent of the `packages reload` command in CLI and is used to load new/updated packages. If NSO is used in an HA or Raft setup, the `packages ha sync` action is invoked instead of the usual `packages reload` action, i.e., the packages will be synced in the cluster. Read more about the `reload` action in [NSO Packages](/guides/operation-and-usage/operations/listing-packages) and for HA in [High Availability](/guides/administration/management/high-availability#packages-upgrades-in-raft-cluster). General package concepts are covered in [Package Management](/guides/administration/management/package-mgmt).

To reload the packages:

1. Click the **Reload all packages** button.
2. In the dialog, set the **Max wait time (sec)** for the commit queue to empty before proceeding with the reload. The default is 10 seconds if you leave the field unset.
3. Set the **Timeout action** behavior to define what happens after the maximum wait time is over, i.e., kill the open transactions and continue, or cancel (fail) the package reload operation altogether. The default for this setting is **fail**.
4. Apply additional action parameters from the following (optional): **Force** (to force package reload overriding issues or warnings), **Dry run** (to simulate the package reload process without making any actual changes), and **Wait commit queue empty** (to wait until the commit queue is empty before proceeding).
5. Click **Reload**. A live trace of the reload operation is displayed while the packages are being reloaded.
6. Click **Done** when the operation has finished.

### Deinstall a Package

To deinstall a package:

* Go to the package details view and click the **Deinstall** button, or use the more options <img src="/files/sstclOvFs1yjOiJgltpV" alt="" data-size="line"> button in the packages list.

## High Availability <a href="#d5e6538" id="d5e6538"></a>

The **High Availability** view is used to visualize your HA setup ([Rule-based](https://nso-docs.cisco.com/guides/operation-and-usage/webui/pages/qG3CMifhI63daJ1BZfmB#ug.ha.builtin) or [Raft](https://nso-docs.cisco.com/guides/operation-and-usage/webui/pages/qG3CMifhI63daJ1BZfmB#ug.ha.raft)). Depending on the type of HA configured (shown under the **High availability** title), the view displays available management options, current operational status, and actions for your cluster.

### Rule-based HA

The Rule-based HA view displays the general information and operational status of your cluster. Actions on the Rule-based cluster can be performed using the **Configuration editor** -> **Actions** tab.

Available Rule-based HA actions are described further under [Actions](/guides/administration/management/high-availability#d5e5031). Specific parameters and field definitions shown in the view are covered in detail in the rest of the [HA documentation](/guides/administration/management/high-availability).

An example cluster of a Rule-based HA setup is shown below.

<div data-with-frame="true"><figure><img src="/files/vAbhIy66OsMUupb6O8K1" alt=""><figcaption><p>High Availability View (Rule-based)</p></figcaption></figure></div>

### Raft HA

The Raft HA view displays overview of your cluster and provides options to manage them.

Available Raft HA actions are described further under [Actions](https://nso-docs.cisco.com/guides/operation-and-usage/webui/pages/qG3CMifhI63daJ1BZfmB#ch_ha.raft_actions), and can be run directly in the Web UI. Specific parameters and field definitions shown in the view are covered in detail in the rest of the [HA documentation](/guides/administration/management/high-availability).

<div data-with-frame="true"><figure><img src="/files/j3TbNFnoy4Wmi08JKNz5" alt=""><figcaption><p>High Availability View (Raft)</p></figcaption></figure></div>

#### Handover Cluster Leadership <a href="#d5e6565" id="d5e6565"></a>

The **Handover** option allows you to hand over the leadership of your Raft cluster to another node.

Perform the handover as follows:

1. Click the **Handover** button.
2. Select the new leader from the list.
3. Click **Save**. A message is shown if the handover was successful or not.

#### Actions on a Node

Actions on a node, such as **Add node**, **Remove node**, **Disconnect**, etc., are available by accessing the more options <img src="/files/sstclOvFs1yjOiJgltpV" alt="" data-size="line"> button on a node. Most of the actions in Raft HA can only be executed from the leader node.

#### Logs and Certificates

The **Logs** and **Certificates** tabs provide detailed insights into the state and configuration of the Raft cluster.

* **Logs** tab: Provides detailed logs pertaining to Raft operation and displays information about the internal RAFT replication process and its operational status. This includes the synchronization state of configuration data across HA nodes.
  * **Log status** – Summarizes the current state of the RAFT log on this node:

    * **Current index**: The index of the latest log entry stored on the node.
      * **Applied index**: The index of the latest log entry that has been committed and applied to the system’s state machine.
      * **Num entries**: The total number of log entries currently held by the node.

    When all three values are equal (for example, 27), it indicates that the node is fully synchronized and up to date with the cluster leader.
* **Certificates** tab: Lists the SSL/TLS certificates used for secure communication between nodes in the HA cluster. This ensures encrypted and authenticated synchronization of data across the RAFT network.

## Alarms <a href="#d5e6565" id="d5e6565"></a>

The **Alarms** view displays alerts in the system for your NSO-managed objects and provides options to manage them.

<div data-with-frame="true"><figure><img src="/files/0a27GGL68I8FXBCTu9N0" alt=""><figcaption><p>Alarms View</p></figcaption></figure></div>

An alarm is raised when an NSO object undergoes a state change that requires attention. The alarms, depending on their severity, are categorized as **Critical**, **Major**, **Minor**, **Warning**, and **Indeterminate**. Detailed alarm management concepts are covered in [Alarm Manager](/guides/operation-and-usage/operations/alarm-manager) and different alarm types are described in [Alarm Types](/guides/administration/management/system-management/alarms).

### Viewing Options

You can search and sort the alarm list to display alarm results according to your need.

* To search for an alarm against an object, search for the object name (e.g., device name).
* To sort the alarms list, use one of the specified criteria from **Alarm type** , **Severity**, **Is cleared**, or **Handling state**.

**Alarm Details**

Individual alarm details are accessible by clicking the severity level icon on an alarm. This brings up the alarm's details, its status (severity) changes, and historical handling information.

### Compress and Purge Alarms

The Web UI provides additional options to compress and purge alarms.

* The **Compress alarms** action streamlines the alarm entries by deleting their historical state changes that occured before the last one (i.e., the only the latest state change is kept), while keeping the alarm entries intact.
* The **Purge alarms** action completely removes the alarm entries according to the specified criteria.

To utilize these features, click the respective button and follow the on-screen instructions.

### Alarm Handling

Alarm handling refers to attending to an alarm. This usually entails reviewing the alarm and setting a state on it, for example, **Acknowledged**. Historical handling state changes are accessible in alarm details.

To set an alarm handling state:

1. In the **Alarms** main view, click the more options <img src="/files/sstclOvFs1yjOiJgltpV" alt="" data-size="line"> button on the desired alarm and click **Set alarm handling state**.
2. Set the alarm state to one of the following: **None**, **Acknowledged**, **Investigation**, **Observation**, and **Closed**.
3. Enter a description (optional).
4. Click **Set state**. This sets the alarm handling state as well as records the state change under the **Alarm handling** tab in alarm details.
5. Access the Commit Manager by clicking its icon <img src="/files/Jd939sKUXxRwrQAfvpro" alt="" data-size="line"> in the banner.
6. Review the available changes appearing as **Current transaction**. If there are errors in the change, the Commit Manager alerts you and suggests possible corrections. You can then fix them and press **Re-validate** to clear the errors.
7. Click **Revert** to undo or **Commit** to confirm the changes in the transaction.
   * **Commit Options**: When committing a transaction, you have the possibility to choose **Commit options** and perform a commit with the specified commit option(s). Examples of commit options are: **No revision drop**, **No deploy**, **No networking**, etc. Commit options are described centrally in [Commit Parameters](/guides/operation-and-usage/operations/lifecycle-operations#d5e5048).
8. Access the Commit Manager by clicking its icon <img src="/files/Jd939sKUXxRwrQAfvpro" alt="" data-size="line"> in the banner.
9. Review the available changes appearing as **Current transaction**. If there are errors in the change, the Commit Manager alerts you and suggests possible corrections. You can then fix them and press **Re-validate** to clear the errors.
10. Click **Revert** to undo or **Commit** to confirm the changes in the transaction.
    * **Commit Options**: When committing a transaction, you have the possibility to choose **Commit options** and perform a commit with the specified commit option(s). Examples of commit options are: **No revision drop**, **No deploy**, **No networking**, etc. Commit options are described in detail in the JSON-RPC API documentation under [Methods - transaction - commit changes](/guides/development/advanced-development/web-ui-development/json-rpc-api#methods-transaction-commit-changes).

## Compliance Reporting <a href="#sec.webui_compliance" id="sec.webui_compliance"></a>

The **Compliance reporting** view is used to create and run compliance reports to check the current situation, check historical events, or both. The conceptual aspects of the compliance reporting feature are discussed in greater depth in the [Compliance Reports](/guides/operation-and-usage/operations/compliance-reporting) section.

{% hint style="success" %}
Web UI is the recommended way of running the compliance reports.
{% endhint %}

The following tabs are available in this view:

* **Compliance reports**
* **Report results**
* **Compliance templates**

### Compliance Reports

The **Compliance reports** tab is used to view, create, run, and manage the existing compliance reports.

<div data-with-frame="true"><figure><img src="/files/tkc83elsQLcCPYeGbhoj" alt=""><figcaption><p>Compliance Reports View</p></figcaption></figure></div>

#### **Create a Compliance Report**

To create a new compliance report:

1. In the **Compliance reporting** view -> **Compliance reports** tab, click **New report**.
2. In the **Create new report** pop-up, enter the report name and click **Create**.
3. Next, set up the compliance report using the following tabs. For a more detailed description of Compliance Reporting concepts and related configuration options, see [Compliance Reporting](/guides/operation-and-usage/operations/compliance-reporting).
   * **General** tab: To configure the report name. Configuration options include:
     * **Report name**: Displays the report name and allows editing of the report name.
   * **Devices** tab: To configure device compliance checks. Configuration options include:
     * **Device choice**: Include **All devices** or only **Some devices** to include in compliance checks. If **Some devices** is selected, specify the devices using a device group, an XPath expression, or individual devices.
     * **Device checks**:
       * **Current out of sync**: Check the device's current status and report if the device is in sync or out of sync. Possible values are **true** (yes, request a check-sync) and **false** (no, do not request a check-sync).
       * **Historic changes**: Include or exclude previous changes to devices using the commit log. Possible values are **true** (yes, include) and **false** (no, exclude).
     * **Compliance templates**: If a compliance template should be used to check for compliance (see [Device Configuration Checks](/guides/operation-and-usage/operations/compliance-reporting#device-configuration-checks)). You have the option to add a compliance template using the **Add template** button or create a new compliance template (see [Compliance Templates](#compliance-templates)). To enforce devices to comply exactly with the template's configuration, use **Strict** mode; see [Additional Configuration Checks](/guides/operation-and-usage/operations/compliance-reporting#additional-configuration-checks) for more information.
   * **Services** tab: To configure service compliance checks. Configuration options include:
     * **Service choice**: Include **All services** or only **Some services**. If **Some services** is selected, specify the services using service type, an XPath expression, or individual service instances.
     * **Service checks**:
       * **Current out of sync**: Check the service's current status and report if the service is in sync or out of sync. Possible values are **true** (yes, request a check-sync) and **false** (no, do not request a check-sync).
       * **Historic changes**: Include or exclude previous changes to services using the commit log. Possible values are **true** (yes, include) and **false** (no, exclude).
4. Click **Create report** when the report setup is complete. The changes are saved and applied immediately.

{% hint style="info" %}
In the **Compliance reports** tab, you can apply the following actions on the report by selecting it using the checkbox and using the more options <img src="/files/sstclOvFs1yjOiJgltpV" alt="" data-size="line"> button.

* **Copy as new report**: Copy an existing report as a new report.
* **Run**: Run the report.
* **Delete**: Delete the report.
* **Edit name**: Edit the report name.
  {% endhint %}

#### **Run a Compliance Report**

To run a compliance report:

1. In the **Compliance reports** tab, click the desired report and then click **Run report**.
2. Specify the following in the **Run report** pop-up:
   * **Report title**: A title for this specific report run.
   * **Historical time interval**. Select the time range. The report runs with the maximum possible interval if you do not specify an interval.
3. Click **Run report**.

### Report Results

The **Reports results** tab is used to view the status and results of the compliance reports that have been run.

<div data-with-frame="true"><figure><img src="/files/u0mDQdd3AX2zks8EivDU" alt=""><figcaption><p>Reports Results View</p></figcaption></figure></div>

#### View Compliance Report Results

The report's results show if the devices/services included in the report are compliant/in-sync or have violations. A summary of the report status is readily available in the **Report results** tab. To fetch detailed information on the report, click the report name. The following information panes are then available:

* **Details**: Includes specifics about the report that was run, such as report name, date/time it was run, time range, and contents analyzed (i.e., services, devices, and rollback files).
* **Results overview**: Shows a summary of results with visuals on the number of devices and services that are presently compliant/in-sync.
* **Historic compliance**: Shows a history of compliance (in percentages) for the devices and services that were included in the report run. The graph is presented based on the previous report runs and you can narrow down the graph to show data from specific periods (e.g., last 10 runs only). Predefined time ranges include last 30 days, last month, last 6 months, and last year, whereas custom time ranges allow users to define their own time ranges. The default preset is set to last 30 days.
* **Devices**/**Services**/**Errors**: Displays individual compliance and error information for analyzed devices and services. In case of non-compliance, a 'diff view' is available.

{% hint style="info" %}
Use the **Export to file** button to export the report results to a downloadable file (PDF).
{% endhint %}

### Compliance Templates

The **Compliance Templates** tab is used to create new compliance templates and manage existing ones.

<div data-with-frame="true"><figure><img src="/files/4pblGBhDTSX7Az5xgUl6" alt=""><figcaption><p>Compliance Templates View</p></figcaption></figure></div>

There are two ways to create a compliance template:

* **From device template**: Build a new template from an existing predefined device template.
* **From config**: Build a new template directly from an existing device configuration.

{% hint style="info" %}
**Template Creation using Config Editor**

A third way to create a compliance template from scratch is by using the Config Editor. With this option, you will need to manually type in your desired configuration model to create a compliance template.
{% endhint %}

{% tabs %}
{% tab title="From device template" %}
Use this option to base your new template on an existing [device template](/guides/operation-and-usage/operations/basic-operations#d5e228).

To create a compliance template from a device template:

1. In the **Compliance templates** tab, click **Create template**.
2. In the **Create template** window -> **Source** category, continue with the default option, **From device template**.
3. Next choose a device template using the **Select device template** drop-down list. A device template should exist prior to this selection.
4. Name your compliance template in the **New compliance template name** field (optional). Leaving this blank retains the device template title for the compliance template.
5. Click **Create**.
   {% endtab %}

{% tab title="From config" %}
Use this option to build a new template from configuration.

To create a compliance template from config:

1. In the **Compliance templates** tab, click **Create template**.
2. In the **Create template** window -> **Source** category, select **From config**.
3. Provide a **Template name** that will be used to reference the template in NSO.
4. In the **Path** field, enter the an XPath to target for extracting config data.
5. Click **Add to list** to add the path.
6. **Match rate**: Enter a value between 0 - 100 to determine how often a configuration recurrence must appear in device configurations to be included in the template.
   * A value of **100** means that configuration must be identical across all devices.
   * A lower value allows partial commonality.
7. **Exclude service config**: Enable this option to ensure that NSO will exclude configurations already managed by services from the template. This option ensures that the service-managed configurations are not duplicated.
8. **Collapse list keys**: Enable this option to determine how lists in the XPath configuration are collapsed into single entries when keys do not match. The options in this category are:
   * **Automatic**: Automatically find non-matching lists to collapse. Lists on the same path in `/devices/device/config` that do not compare equal will be collapsed.
   * **All**: All list keys are collapsed into a single entry, regardless of matching rules.
   * **List path**: Use a user-provided list of paths to collapse. Allows manual control.
   * **Disabled**: Disable list collapsing entirely, displaying all list entries and differences in full detail.
9. Click **Create**.
   {% endtab %}
   {% endtabs %}


# Operations

Manage the network with NSO.


# Basic Operations

Learn basic operational scenarios and common CLI commands.

This section helps you to get started with NSO, learn basic operational scenarios, and get acquainted with the most common CLI commands.

## Setup <a href="#d5e47" id="d5e47"></a>

Make sure that you have installed NSO and that you have sourced the `ncsrc` file in `$NCS_DIR`. This sets up the paths and environment variables to run NSO. As this must be done every time before running NSO, it is recommended to add it to your profile.

We will use the NSO network simulator to simulate three Cisco IOS routers. NSO will talk Cisco CLI to those devices. You will use the NSO CLI and Web UI to perform the tasks. Sometimes you will use the native Cisco device CLI to inspect configuration or do out-of-band changes.

<div data-with-frame="true"><figure><img src="/files/4Drqogf8R0TCdyPfg0oH" alt="" width="375"><figcaption><p>The First Example</p></figcaption></figure></div>

\
Note that both the NSO software (NCS) and the simulated network devices run on your local machine.

## Starting the Simulator <a href="#d5e59" id="d5e59"></a>

To start the simulator:

1. Go to [examples.ncs/device-management/simulated-devices](https://github.com/NSO-developer/nso-examples/tree/6.7/device-management/simulated-devices). First of all, we will generate a network simulator with three Cisco devices. They will be called `c0`, `c1`, and `c2`.

{% hint style="info" %}
Most of this section follows the procedure in the `README` file, so it is useful to have it opened as well.
{% endhint %}

Perform the following command:

```bash
$ ncs-netsim create-network $NCS_DIR/packages/neds/cisco-ios 3 c
```

This creates three simulated devices all running Cisco IOS and they will be named `c0`, `c1`, `c2`.

2. Start the simulator.

```bash
$ ncs-netsim start
DEVICE c0 OK STARTED
DEVICE c1 OK STARTED
DEVICE c2 OK STARTED
```

3. Run the CLI toward one of the simulated devices.

```bash
$ ncs-netsim cli-i c1
admin connected from 127.0.0.1 using console *

c1> enable
c1# show running-config
class-map m
match mpls experimental topmost 1
match packet length max 255
match packet length min 2
match qos-group 1
!
...
c1# exit
```

This shows that the device has some initial configurations.

## Starting NSO and Reading Device Configuration <a href="#d5e80" id="d5e80"></a>

The previous step started the simulated Cisco devices. It is now time to start NSO.

1. The first action is to prepare directories needed for NSO to run and populate NSO with information on the simulated devices. This is all done with the `ncs-setup` command. Make sure that you are in the [examples.ncs/device-management/simulated-devices](https://github.com/NSO-developer/nso-examples/tree/6.7/device-management/simulated-devices) directory. (Again, ignore the details for the time being).

```bash
$ ncs-setup --netsim-dir ./netsim --dest .
```

{% hint style="info" %}
Note the `.` at the end of the command referring to the current directory. What the command does is to create directories needed for NSO in the current directory and populate NSO with devices that are running in netsim. We call this the "run-time" directory.
{% endhint %}

2. Start NSO.

```bash
$ ncs
```

3. Start the NSO CLI as the user `admin` with a Cisco XR-style CLI.

```bash
$ ncs_cli -C -u admin
```

NSO also supports a J-style CLI, that is started by using a -J modification to the command like this.

```bash
$ ncs_cli -J -u admin
```

Throughout this user guide, we will show the commands in Cisco XR style.

4. At this point, NSO only knows the address, port, and authentication information of the devices. This management information was loaded to NSO by the setup utility. It also tells NSO how to communicate with the devices by using NETCONF, SNMP, Cisco IOS CLI, etc. However, at this point, the actual configuration of the individual devices is unknown.

```bash
admin@ncs# show running-config devices device
devices device c0
 address   127.0.0.1
 port      10022
...
 authgroup default
 device-type cli ned-id cisco-ios
 state admin-state unlocked
 config
  no ios:service pad
  no ios:ip domain-lookup
  no ios:ip http secure-server
  ios:ip source-route
 !
!  ...
```

Let us analyze the above CLI command. First of all, when you start the NSO CLI it starts in operational mode, so to show configuration data, you have to explicitly run `show running-config`.

NSO manages a list of devices, each device is reached by the path `devices device "name"` . You can use standard tab completion in the CLI to learn this.

The `address` and `port` fields tells NSO where to connect to the device. For now, they all live in local host with different ports. The `device-type` structure tells NSO it is a CLI device and the specific CLI is supported by the Network Element Driver (NED) `cisco-ios`. A more detailed explanation of how to configure the device-type structure and how to choose NEDs will be addressed later in this guide.

So now NSO can try to connect to the devices:

```bash
admin@ncs# devices connect
connect-result {
    device c0
    result true
    info (admin) Connected to c0 - 127.0.0.1:10022
}
connect-result {
    device c1
    result true
    info (admin) Connected to c1 - 127.0.0.1:10023
}
connect-result {
    device c2
    result true
    info (admin) Connected to c2 - 127.0.0.1:10024
}....
```

NSO does not need to have the connections active continuously, instead, NSO will establish a connection when needed and connections are pooled to conserve resources. At this time, NSO can read the configurations from the devices and populate the configuration database, CDB.

The following command will synchronize the configurations of the devices with the CDB and respond with `true` if successful:

```bash
admin@ncs# devices sync-from
sync-result {
    device c0
    result true
}....
```

The NSO data store, CDB, will store the configuration for every device at the path `devices device "name" config` . Everything after this path is the configuration in the device. Normally, NSO keeps this synchronized with the device. The synchronization is managed with the following principles:

1. At initialization, NSO can discover the configuration as shown above.
2. In day-to-day operations on the network, the network engineer uses NSO (CLI, WebUI, REST,...) to modify the representation of device configuration in the NSO CDB. The changes are committed to the network as a transaction that includes the actual devices. Only if all changes happen on the actual devices, they are committed to the NSO data store. The transaction also covers the devices, so if any device participating in the transaction fails, NSO will roll back the configuration changes on all modified devices. This works even in the case of devices that do not natively support roll-back, such as Cisco IOS CLI.
3. NSO can detect out-of-band changes and reconcile them by either updating the CDB or modifying the configuration on the devices to reflect the currently stored configuration.

NSO only needs to be synchronized with the devices in the event of a change being made outside of NSO. Changes made using NSO are reflected in both the CDB and the devices. The following actions do not need to be taken:

1. Perform configuration change via NSO.
2. Perform sync-from action.

The above incorrect (or not necessary) sequence stems from the assumption that the NSO CLI talks directly to the devices. This is not the case; the northbound interfaces in NSO modify the configuration in the NSO data store, NSO calculates a minimum difference between the current configuration and the new configuration, giving only the changes to the configuration to the NEDs that runs the commands to the devices. All this is done as one single change-set.

The one exception to the above are devices that change their own configuration. For example, you only configure A but also B appears in the device configuration. These are so-called "auto-configs". In this case, the NED needs to implement special code to handle each such scenario individually. If the NED does not fully cover all of these device quirks, the device may get out of sync when you make configuration changes through NSO.

<div data-with-frame="true"><figure><img src="/files/wNh5HgsgwAFnMzbMZX2a" alt="" width="375"><figcaption><p>Device Transaction</p></figcaption></figure></div>

View the configuration of the `c0` device using the command:

```bash
admin@ncs# show running-config devices device c0 config
devices device c0
 config
  no ios:service pad
  ios:ip vrf my-forward
   bgp next-hop Loopback 1
  !
...
```

Or, show a particular piece of configuration from several devices:

```bash
admin@ncs# show running-config devices device c0..2 config ios:router
devices device c0
 config
  ios:router bgp 64512
   aggregate-address 10.10.10.1 255.255.255.251
   neighbor 1.2.3.4 remote-as 1
   neighbor 1.2.3.4 ebgp-multihop 3
   neighbor 2.3.4.5 remote-as 1
   neighbor 2.3.4.5 activate
   neighbor 2.3.4.5 capability orf prefix-list both
   neighbor 2.3.4.5 weight 300
  !
 !
!
devices device c1
 config
  ios:router bgp 64512
...
```

Or, show a particular piece of configuration from all devices:

```bash
admin@ncs# show running-config devices device config ios:router
```

The CLI can pipe commands, try <kbd>TAB</kbd> after `|` to see various pipe targets:

```bash
admin@ncs# show running-config devices device config ios:router \
                     | display xml | save router.xml
```

The above command shows the router config of all devices as XML and then saves it to a file `router.xml`.

## Writing Device Configuration <a href="#d5e156" id="d5e156"></a>

1. To change the configuration, enter configure mode.

```bash
admin@ncs# config
Entering configuration mode terminal
admin@ncs(config)#
```

2. Change or add some configuration across the devices, for example:

```bash
 admin@ncs(config)# devices device c0..2 config ios:router bgp 64512
                       neighbor 10.10.10.0 remote-as 64502
admin@ncs(config-router)#
```

### Transaction Commit

It is important to understand how NSO applies configuration changes to the network. At this point, the changes are local to NSO, no configurations have been sent to the devices yet. Since the NSO Configuration Database, CDB is in sync with the network, NSO can calculate the minimum diff to apply the changes to the network.

The command below compares the ongoing changes with the running database:

```bash
admin@ncs(config-router)# top
admin@ncs(config)# show configuration
devices device c0
 config
  ios:router bgp 64512
   neighbor 10.10.10.0 remote-as 64502
...
```

It is possible to dry-run the changes to see the native Cisco CLI output (in this case almost the same as above):

```bash
admin@ncs(config)# commit dry-run outformat native
native {
    device {
        name c0
        data router bgp 64512
              neighbor 10.10.10.0 remote-as 64502
             !
...
```

The changes can be committed to the devices and the NSO CDB simultaneously with a single commit. In the commit command below, we pipe to details to understand the actions being taken.

```bash
admin@ncs% commit | details
```

### Transaction Rollback

Changes are committed to the devices and the NSO database as one transaction. If any of the device configurations fail, all changes will be rolled back and the devices will be left in the state that they were in before the commit and the NSO CDB will not be updated.

There are numerous options to the commit command which will affect the behavior of the atomic transactions:

```bash
admin@ncs(config)# commit TAB
Possible completions:
  and-quit               Exit configuration mode
  check                  Validate configuration
  comment                Add a commit comment
  commit-queue           Commit through commit queue
  label                  Add a commit label
  no-confirm             No confirm
  no-networking          Send nothing to the devices
  no-out-of-sync-check   Commit even if out of sync
  no-overwrite           Do not overwrite modified data on the device
  no-revision-drop       Fail if device has too old data model
  save-running           Save running to file
  ---
  dry-run                Show the diff but do not perform commit
```

As seen by the details output, NSO stores a roll-back file for every commit so that the whole transaction can be rolled back manually. The following is an example of a rollback file:

```bash
admin@ncs(config)# do file show logs/rollback1000
Possible completions:
     rollback10001  rollback10002  rollback10003  \
                           rollback10004  rollback10005
admin@ncs(config)# do file show logs/rollback10005
# Created by: admin
# Date: 2014-09-03 14:35:10
# Via: cli
# Type: delta
# Label:
# Comment:
# No: 10005

ncs:devices {
    ncs:device c0 {
        ncs:config {
            ios:router {
                ios:bgp 64512 {
                    delete:
                    ios:neighbor 10.10.10.0;
                }
            }
        }
    }
```

(Viewing files as an operational command, prefixing a command in configuration mode with `do` executes in operational mode.) To perform a manual rollback, first load the rollback file:

```bash
admin@ncs(config)# rollback-files apply-rollback-file fixed-number 10005
```

`apply-rollback-file` by default restores to that saved configuration, adding `selective` as a parameter allows you to just roll back the delta in that specific rollback file. Show the differences:

```bash
admin@ncs(config)# show configuration
devices device c0
 config
  ios:router bgp 64512
   no neighbor 10.10.10.0 remote-as 64502
  !
 !
!
devices device c1
 config
  ios:router bgp 64512
   no neighbor 10.10.10.0 remote-as 64502
  !
 !
!
devices device c2
 config
  ios:router bgp 64512
   no neighbor 10.10.10.0 remote-as 64502
  !
 !
!
```

Commit the rollback:

```bash
admin@ncs(config)# commit
Commit complete.
```

### Trace Log

A trace log can be created to see what is going on between NSO and the device CLI enable trace. Use the following command to enable trace:

```bash
admin@ncs(config)# devices global-settings trace raw trace-dir logs
admin@ncs(config)# commit
Commit complete.
admin@ncs(config)# devices disconnect
```

Note that the trace settings only take effect for new connections, so is important to disconnect the current connections. Make a change to for example `c0`:

```bash
admin@ncs(config)# devices device c0 config ios:interface FastEthernet
                                1/2 ip address  192.168.1.1 255.255.255.0
admin@ncs(config-if)# commit dry-run outformat native
admin@ncs(config-if)# commit
```

Note the use of the command `commit dry-run outformat native`. This will display the net result device commands that will be generated over the native interface without actually committing them to the CDB or the devices. In addition, there is the possibility to append the `reverse` flag that will display the device commands for getting back to the current running state in the network if the commit is successfully executed.

Exit from the NSO CLI and return to the Unix Shell. Inspect the CLI trace:

```bash
 less logs/ned-cisco-ios-c0.trace
```

## More on Device Management <a href="#d5e207" id="d5e207"></a>

### Device Groups <a href="#d5e209" id="d5e209"></a>

As seen above, ranges can be used to send configuration commands to several devices. Device groups can be created to allow for group actions that do not require naming conventions. A group can reference any number of devices. A device can be part of any number of groups, and groups can be hierarchical.

The command sequence below creates a group of core devices and a group with all devices. Note that you can use tab completion when adding the device names to the group. Also, note that it requires configuration mode. (If you are still in the Unix Shell from the steps above, do `$ncs_cli -C -u admin`).

```bash
admin@ncs(config)# devices device-group core device-name [ c0 c1 ]
admin@ncs(config-device-group-core)# commit

admin@ncs(config)# devices device-group all device-name c2 device-group core
admin@ncs(config-device-group-all)# commit

admin@ncs(config)# show full-configuration devices device-group
devices device-group all
 device-name  [ c2 ]
 device-group [ core ]
!
devices device-group core
 device-name [ c0 c1 ]
!

admin@ncs(config)# do show devices device-group
NAME  MEMBER        INDETERMINATES  CRITICALS  MAJORS  MINORS  WARNINGS
-------------------------------------------------------------------------
all   [ c0 c1 c2 ]  0               0          0       0       0
core  [ c0 c1 ]     0               0          0       0       0
```

Note well the `do show` which shows the operational data for the groups. Device groups have a member attribute that shows all member devices, flattening any group members.

Device groups can contain different devices as well as devices from different vendors. Configuration changes will be committed to each device in its native language without needing to be adjusted in NSO.

You can, for example, at this point use the group to check if all `core` are in sync:

```bash
admin@ncs# devices device-group core check-sync
sync-result {
    device c0
    result in-sync
}
sync-result {
    device c1
    result in-sync
}
```

### Device Templates <a href="#d5e228" id="d5e228"></a>

Assume that we would like to manage permit lists across devices. This can be achieved by defining templates and applying them to device groups. The following CLI sequence defines a tiny template, called `community-list` :

```bash
admin@ncs(config)# devices template community-list
                                ned-id cisco-ios-cli-3.0
                                config ios:ip
                                community-list standard test1
                                permit permit-list 64000:40

admin@ncs(config-permit-list-64000:40)# commit
Commit complete.
admin@ncs(config-permit-list-64000:40)# top

admin@ncs(config)# show full-configuration devices template
devices template community-list
 config
  ios:ip community-list standard test1
   permit permit-list 64000:40
   !
  !
 !
!
[ok][2013-08-09 11:27:28]
```

This can now be applied to a device group:

```bash
admin@ncs(config)# devices device-group core apply-template \
                                 template-name community-list
admin@ncs(config)# show configuration
devices device c0
 config
  ios:ip community-list standard test1 permit 64000:40
 !
!
devices device c1
 config
  ios:ip community-list standard test1 permit 64000:40
 !
!
admin@ncs(config)# commit dry-run outformat native
native {
    device {
        name c0
        data ip community-list standard test1 permit 64000:40
    }
    device {
        name c1
        data ip community-list standard test1 permit 64000:40
    }
}
admin@ncs(config)# commit
Commit complete.
```

What if the device group `core` contained different vendors? Since the configuration is written in IOS the above template would not work on Juniper devices. Templates can be used on different device types (read NEDs) by using a prefix for the device model. The template would then look like:

```
template community-list {
  config {
    junos:configuration {
    ...
    }
    ios:ip {
    ...
    }
```

The above indicates how NSO manages different models for different device types. When NSO connects to the devices, the NED checks the device type and revision and returns that to NSO. This can be inspected (note, in operational mode):

```bash
admin@ncs# show devices device module
NAME  NAME                       REVISION    FEATURES  DEVIATIONS
-------------------------------------------------------------------
c0    tailf-ned-cisco-ios        2014-02-12  -         -
      tailf-ned-cisco-ios-stats  2014-02-12  -         -
c1    tailf-ned-cisco-ios        2014-02-12  -         -
      tailf-ned-cisco-ios-stats  2014-02-12  -         -
c2    tailf-ned-cisco-ios        2014-02-12  -         -
      tailf-ned-cisco-ios-stats  2014-02-12  -         -
```

So here we see that `c0` uses a `tailf-ned-cisco-ios` module which tells NSO which data model to use for the device. Every NED package comes with a YANG data model for the device (except for third-party YANG NED for which the YANG device model must be downloaded and fixed before it can be used). This renders the NSO data store (CDB) schema, the NSO CLI, WebUI, and southbound commands.

The model introduces namespace prefixes for every configuration item. This also resolves issues around different vendors using the same configuration command for different configuration elements. Note that every item is prefixed with `ios`:

```bash
admin@ncs# show running-config devices device c0 config ios:ip community-list
devices device c0
 config
  ios:ip community-list 1 permit
  ios:ip community-list 2 deny
  ios:ip community-list standard s permit
  ios:ip community-list standard test1 permit 64000:40
 !
!
```

Another important question is how to control if the template merges the list or replaces the list. This is managed via tags. The default behavior of templates is to merge the configuration. Tags can be inserted at any point in the template. Tag values are `merge`, `replace`, `delete`, `create` and `nocreate`.

Assume that `c0` has the following configuration:

```bash
admin@ncs# show running-config devices device c0 config ios:ip community-list
devices device c0
 config
  ios:ip community-list 1 permit
  ios:ip community-list 2 deny
  ios:ip community-list standard s permit}
```

If we apply the template the default result would be:

```bash
admin@ncs# show running-config devices device c0 config ios:ip community-list
devices device c0
 config
  ios:ip community-list 1 permit
  ios:ip community-list 2 deny
  ios:ip community-list standard s permit
  ios:ip community-list standard test1 permit 64000:40
 !
!
```

We could change the template in the following way to get a result where the permit list would be replaced rather than merged. When working with tags in templates, it is often helpful to view the template as a tree rather than a command view. The CLI has a display option for showing a curly-braces tree view that corresponds to the data-model structure rather than the command set. This makes it easier to see where to add tags.

```bash
admin@ncs(config)# show full-configuration devices template
devices template community-list
 config
  ios:ip community-list standard test1
   permit permit-list 64000:40
   !
  !
 !
!
admin@ncs(config)# show full-configuration devices \
                                 template | display curly-braces
template community-list {
    config {
        ios:ip {
            community-list {
                standard test1 {
                    permit {
                        permit-list 64000:40;
                    }
                }
            }
        }
    }
}


admin@ncs(config)# tag add devices template community-list
                                ned-id cisco-ios-cli-3.0
                                config ip community-list replace
admin@ncs(config)# commit
Commit complete.
admin@ncs(config)# show full-configuration devices
                                 template | display curly-braces
template community-list {
    config {
        ios:ip {
            /* Tags: replace */
            community-list {
                standard test1 {
                    permit {
                        permit-list 64000:40;
                    }
                }
            }
        }
    }
}
```

Different tags can be added across the template tree. If we now apply the template to the device `c0` which already have community lists, the following happens:

```bash
admin@ncs(config)# show full-configuration devices device c0 \
                                 config ios:ip community-list
devices device c0
 config
  ios:ip community-list 1 permit
  ios:ip community-list 2 deny
  ios:ip community-list standard s permit
  ios:ip community-list standard test1 permit 64000:40
 !
!
admin@ncs(config)# devices device c0 apply-template \
                                 template-name community-list
admin@ncs(config)# show configuration
devices device c0
 config
  no ios:ip community-list 1 permit
  no ios:ip community-list 2 deny
  no ios:ip community-list standard s permit
 !
!
```

Any existing values in the list are replaced in this case. The following tags are available:

* `merge` (default): the template changes will be merged with the existing template.
* `replace`: the template configuration will be replaced by the new configuration.
* `create`: the template will create those nodes that do not exist. If a node already exists this will result in an error.
* `nocreate`: the merge will only affect configuration items that already exist in the template. It will never create the configuration with this tag, or any associated commands inside it. It will only modify existing configuration structures.
* `delete`: delete anything from this point.

Note that a template can have different tags along the tree nodes.

A problem with the above template is that every value is hard-coded. What if you wanted a template where the `community-list` name and `permit-list` value are variables passed to the template when applied? Any part of a template can be a variable, (or actually an XPATH expression). We can modify the template to use variables in the following way:

```bash
admin@ncs(config)# no devices template community-list config ios:ip \
                                community-list standard test1
admin@ncs(config)# devices template community-list config ios:ip \
                                community-list standard \
                                {$LIST-NAME} permit permit-list {$AS}

admin@ncs(config-permit-list-{$AS})# commit
Commit complete.

admin@ncs(config-permit-list-{$AS})# top
admin@ncs(config)# show full-configuration devices template
devices template community-list
 config
  ios:ip community-list standard {$LIST-NAME}
   permit permit-list {$AS}
   !
  !
 !
!
```

The template now requires two parameters when applied (<kbd>tab</kbd> completion will prompt for the variable):

```bash
admin@ncs(config)# devices device-group all apply-template
template-name community-list variable { name LIST-NAME value 'test2' }
variable { name AS value '60000:30' }

admin@ncs(config)# commit
```

Note, that the `replace` tag was still part of the template and it would delete any existing community lists, which is probably not the desired outcome in the general case.

The template mechanism described so far is "fire-and-forget". The templates do not have any memory of what happened to the network, or which devices they touched. A user can modify the templates without anything happening to the network until an explicit `apply-template` action is performed. (Templates are of course, as all configuration changes, applied as a transaction). NSO also supports service templates that are more advanced in many ways, more information on this will be presented later in this guide.

Also, note that device templates have some additional restrictions on the values that can be supplied when applying the template. In particular, a value must either be a number or a single-quoted string. It is currently not possible to specify a value that contains a single quote (`'`).

### Policies <a href="#d5e319" id="d5e319"></a>

To make sure that configuration is applied according to site or corporate rules, you can use policies. Policies are validated at every commit, they can be of type `error` that implies that the change cannot go through or a `warning` which means that you have to confirm a configuration that gives a warning.

A policy is composed of:

1. Policy name.
2. Iterator: loop over a path in the model, for example, all devices, all services of a specific type.
3. Expression: a boolean expression that must be true for every node returned from the iterator, for example, SNMP must be turned on.
4. Warning or error: a message displayed to the user. If it is of the type warning, the user can still commit the change, if of type error the change cannot be made.

An example is shown below:

```bash
admin@ncs(config)# policy rule class-map
Possible completions:
  error-message     Error message to print on expression failure
  expr              XPath 1.0 expression that returns a boolean
  foreach           XPath 1.0 expression that returns a node set
  warning-message   Warning message to print on expression failure

admin@ncs(config)# policy rule class-map foreach /devices/device \
       expr config/ios:class-map[name='a'] \
       warning-message "Device {name} must have a class-map a"

admin@ncs(config-rule-class-map)# top

admin@ncs(config)# commit
Commit complete.

admin@ncs(config)# show full-configuration policy
policy rule class-map
 foreach         /devices/device
 expr            config/ios:class-map[ios:name='a']
 warning-message "Device {name} must have a class-map a"
!
```

Now, if we try to delete a `class-map` `a`, we will get a policy violation:

```bash
admin@ncs(config)# no devices device c2 config ios:class-map match-all a
admin@ncs(config)# validate
Validation completed with warnings:
  Device c2 must have a class-map a

admin@ncs(config)# commit
The following warnings were generated:
  Device c2 must have a class-map a
Proceed? [yes,no] yes
Commit complete.

admin@ncs(config)# validate
Validation completed with warnings:
  Device c2 must have a class-map a
```

The `{name}` variable refers to the node set from the iterator. This node-set will be the list of devices in NSO and the devices have an attribute called 'name'.

To understand the syntax for the expressions a pipe target in the CLI can be used:

```bash
admin@ncs(config)# show full-configuration devices device c2 config \
                                 ios:class-map | display xpath
/ncs:devices/ncs:device[ncs:name='c2']/ncs:config/ \
ios:class-map[ios:name='cmap1']/ios:prematch match-all
...
```

To debug policies look at the end of `logs/xpath.trace`. This file will show all validated XPATH expressions and any errors.

```log
4-Sep-2014::11:05:30.103 Evaluating XPath for policy: class-map:
  /devices/device
get_next(/ncs:devices/device) = {c0}
XPath policy match: /ncs:devices/device{c0}
get_next(/ncs:devices/device{c0}) = {c1}
XPath policy match: /ncs:devices/device{c1}
get_next(/ncs:devices/device{c1}) = {c2}
XPath policy match: /ncs:devices/device{c2}
get_next(/ncs:devices/device{c2}) = false
exists("/ncs:devices/device{c2}/config/class-map{a}") = true
exists("/ncs:devices/device{c1}/config/class-map{a}") = true
exists("/ncs:devices/device{c0}/config/class-map{a}") = true
```

Validation scripts can also be defined in Python, see more about that in [Plug-and-Play Scripting](/guides/operation-and-usage/operations/plug-and-play-scripting).

### Out-of-band Changes, Transactions, and Pre-Provisioning <a href="#d5e363" id="d5e363"></a>

In reality, network engineers might still modify configurations using other tools like out-of-band CLI or other management interfaces. It is important to understand how NSO manages this.

The NSO network simulator supports CLI towards the devices. For example, we can use the IOS CLI on say `c0` and delete a `permit-list`.

From the UNIX shell, start a CLI session towards `c0`.

```bash
$ ncs-netsim cli-i c0

c0> enable
c0# configure
Enter configuration commands, one per line. End with CNTL/Z.

c0(config)# show full-configuration ip community-list
ip community-list standard test1 permit
ip community-list standard test2 permit 60000:30
c0(config)# no ip community-list standard test2
c0(config)#
c0# exit
$
```

Start the NSO CLI again:

```bash
$ ncs_cli -C -u admin
```

NSO detects if its configuration copy in CDB differs from the configuration in the device. Various strategies are used depending on device support: transaction IDs, time stamps, and configuration hash-sums. For example, an NSO user can request a `check-sync` operation:

```bash
admin@ncs# devices check-sync
sync-result {
    device c0
    result out-of-sync
    info got: e54d27fe58fda990797d8061aa4d5325 expected: 36308bf08207e994a8a83af710effbf0

}
sync-result {
    device c1
    result in-sync
}
sync-result {
    device c2
    result in-sync
}

admin@ncs# devices device-group core check-sync
sync-result {
    device c0
    result out-of-sync
    info got: e54d27fe58fda990797d8061aa4d5325 expected: 36308bf08207e994a8a83af710effbf0

}
sync-result {
    device c1
    result in-sync
}
```

NSO can also compare the configurations with the CDB and show the difference:

```bash
admin@ncs# devices device c0 compare-config
diff
 devices {
     device c0 {
         config {
             ios:ip {
                 community-list {
+                    standard test1 {
+                        permit {
+                        }
+                    }
-                    standard test2 {
-                        permit {
-                            permit-list 60000:30;
-                        }
-                    }
                 }
             }
         }
     }
 }
```

At this point, we can choose if we want to use the configuration stored in the CDB as the valid configuration or the configuration on the device:

```bash
admin@ncs# devices sync-
Possible completions:
  sync-from   Synchronize the config by pulling from the devices
  sync-to     Synchronize the config by pushing to the devices

admin@ncs# devices sync-to
```

In the above example, we chose to overwrite the device configuration from NSO.

NSO will also detect out-of-sync when committing changes. In the following scenario, a local `c0` CLI user adds an interface. Later the NSO user tries to add an interface:

```bash
$ ncs-netsim cli-i c0

c0> enable
c0# configure
Enter configuration commands, one per line. End with CNTL/Z.
c0(config)#  interface FastEthernet 1/0 ip address 192.168.1.1 255.255.255.0
c0(config-if)#
c0# exit

$ ncs_cli -C -u admin

admin@ncs# config
Entering configuration mode terminal

admin@ncs(config)# devices device c0 config ios:interface \
       FastEthernet1/1 ip address 192.168.1.1 255.255.255.0

admin@ncs(config-if)# commit
Aborted: Network Element Driver: device c0: out of sync
```

At this point, we have two diffs:

1. The device and NSO CDB (`devices device compare-config`).
2. The ongoing transaction and CDB (`show configuration`).

```bash
admin@ncs(config)# devices device c0 compare-config
diff
 devices {
     device c0 {
         config {
             ios:interface {
                 FastEthernet 1/0 {
                     ip {
                         address {
                             primary {
+                                mask 255.255.255.0;
+                                address 192.168.1.1;
                             }
                         }
                     }
                 }
             }
         }
     }
 }

admin@ncs(config)# show configuration
devices device c0
 config
  ios:interface FastEthernet1/1
   ip address 192.168.1.1 255.255.255.0
  exit
 !
!
```

To resolve this, you can choose to synchronize the configuration between the devices and the CDB before committing. In setups where it is normal for engineers or other systems to make out-of-band changes, you may want to configure NSO to automatically bring in these changes, so you can avoid performing `sync-to` or `sync-from` explicitly. See [Out-of-band Interoperation](/guides/operation-and-usage/operations/out-of-band-interoperation) section for details.

There is also an option to override the out-of-sync check but beware that this could result in NSO inadvertently overwriting some device configuration:

```bash
admin@ncs(config)# commit no-out-of-sync-check
```

Or:

```bash
admin@ncs(config)# devices global-settings out-of-sync-commit-behaviour
Possible completions:
  accept  reject
```

As noted before, all changes are applied as complete transactions of all configurations on all of the devices. Either all configuration changes are completed successfully or all changes are removed entirely. Consider a simple case where one of the devices is not responding. For the transaction manager, an error response from a device or a non-responding device, are both errors and the transaction should automatically rollback to the state before the commit command was issued.

Stop `c0`:

```bash
$ ncs-netsim stop c0
DEVICE c0 STOPPED
```

Go back to the NSO CLI and perform a configuration change over `c0` and `c1`:

```bash
admin@ncs(config)# devices device c0 config ios:ip community-list \
                                 standard test3 permit 50000:30
admin@ncs(config-config)# devices device c1 config ios:ip \
                                community-list standard test3 permit 50000:30

admin@ncs(config-config)# top
admin@ncs(config)# show configuration
devices device c0
 config
  ios:ip community-list standard test3 permit 50000:30
 !
!
devices device c1
 config
  ios:ip community-list standard test3 permit 50000:30
 !
!

admin@ncs(config)# commit
Aborted: Failed to connect to device c0: connection refused: Connection refused
admin@ncs(config)# *** ALARM connection-failure: Failed to connect to
device c0: connection refused: Connection refused
```

NSO sends commands to all devices in parallel, not sequentially. If any of the devices fail to accept the changes or report an error, NSO will issue a rollback to the other devices. Note, that this works also for non-transactional devices like IOS CLI and SNMP. This works even for non-symmetrical cases where the rollback command sequence is not just the reverse of the commands. NSO does this by treating the rollback as it would any other configuration change. NSO can use the current configuration and previous configuration and generate the commands needed to roll back from the configuration changes.

The diff configuration is still in the private CLI session, it can be restored, modified (if the error was due to something in the config), or in some cases, fix the device.

NSO is not a best-effort configuration management system. The error reporting coupled with the ability to completely rollback failed changes to the devices, ensures that the configurations stored in the CDB and the configurations on the devices are always consistent and that no failed or orphan configurations are left on the devices.

First of all, if the above was not a multi-device transaction, meaning that the change should be applied independently device per device, then it is just a matter of performing the commit between the devices.

Second, NSO has a commit flag `commit-queue async` or `commit-queue sync`. The commit queue should primarily be used for throughput reasons when doing configuration changes in large networks. Atomic transactions come with a cost, the critical section of the database is locked when committing the transaction on the network. So, in cases where there are northbound systems of NSO that generate many simultaneous large configuration changes these might get queued. The commit queue will send the device commands after the lock has been released, so the database lock is much shorter. If any device fails, an alarm will be raised.

```bash
admin@ncs(config)# commit commit-queue async
commit-queue-id 2236633674
Commit complete.

admin@ncs(config)# do show devices commit-queue | notab
devices commit-queue queue-item 2236633674
 age       11
 status    executing
 devices   [ c0 c1 c2 ]
 transient c0
  reason "Failed to connect to device c0: connection refused"
 is-atomic true
```

Go to the UNIX shell, start the device, and monitor the commit queue:

```bash
$ncs-netsim start c0
DEVICE c0 OK STARTED

$ncs_cli -C -u admin

admin@ncs# show devices commit-queue
devices commit-queue queue-item 2236633674
 age       11
 status    executing
 devices   [ c0 c1 c2 ]
 transient c0
  reason "Failed to connect to device c0: connection refused"
 is-atomic true

admin@ncs# show devices commit-queue
devices commit-queue queue-item 2236633674
 age       11
 status    executing
 devices   [ c0 c1 c2 ]
 is-atomic true

admin@ncs# show devices commit-queue
% No entries found.
```

Devices can also be pre-provisioned, this means that the configuration can be prepared in NSO and pushed to the device when it is available. To illustrate this, we can start by adding a new device to NSO that is not available in the network simulator:

```bash
admin@ncs(config)# devices device c3 address 127.0.0.1 port 10030 \
                                authgroup default device-type cli
                                ned-id cisco-ios
admin@ncs(config-device-c3)# state admin-state southbound-locked
admin@ncs(config-device-c3)# commit
```

Above, we added a new device to NSO with an IP address local host, and port 10030. This device does not exist in the network simulator. We can tell NSO not to send any commands southbound by setting the `admin-state` to `southbound-locked` (actually the default). This means that all configuration changes will succeed, and the result will be stored in CDB. At any point in time when the device is available in the network, the state can be changed and the complete configuration pushed to the new device. The CLI sequence below also illustrates a powerful copy configuration command that can copy any configuration from one device to another. The from and to paths are separated by the keyword `to`.

```bash
admin@ncs(config)# copy cfg merge devices device c0 config \
                                ios:ip community-list to \
                                devices device c3 config ios:ip community-list
admin@ncs(config)# show configuration
devices device c3
 config
  ios:ip community-list standard test2 permit 60000:30
  ios:ip community-list standard test3 permit 50000:30
 !
!


admin@ncs(config)# commit

admin@ncs(config)# devices check-sync
...

sync-result {
    device c3
    result locked
}
```

As shown above, `check-sync` operations will tell the user that the device is southbound locked. When the device is available in the network, the device can be synchronized with the current configuration in the CDB using the `sync-to` action.

### About Conflicts <a href="#d5e462" id="d5e462"></a>

Different users or management tools can of course run parallel sessions to NSO. All ongoing sessions have a logical copy of CDB. An important case needs to be understood if there is a conflict when multiple users attempt to modify the same device configuration at the same time with different changes. First, let's look at the CLI sequence below, user `admin` to the left, user `joe` to the right.

```bash
admin@ncs(config)# devices device c0 config ios:snmp-server community fozbar

      joe@ncs(config)# devices device c0 config ios:snmp-server community fezbar

admin@ncs(config-config)# commit

      System message at 2014-09-04 13:15:19...
      Commit performed by admin via console using cli.
      joe@ncs(config-config)# commit
      joe@ncs(config)# show full-configuration devices device c0 config ios:snmp-server
      devices device c0
        config
          ios:snmp-server community fezbar
          ios:snmp-server community fozbar
        !
      !
```

There is no conflict in the above sequence, `community` is a list so both `joe` and `admin` can add items to the list. Note that user `joe` gets information about the user `admin` committing.

On the other hand, if two users modify an ordered-by user list in such a way that one user rearranges the list, along with other non-conflicting modifications, and one user deletes the entire list, the following happens:

```bash
admin@ncs(config)# no devices device c0 config access-list 10

      joe@ncs(config)# move devices device c0 config access-list 10 permit 168.215.202.0 0.0.0.255 first
      joe@ncs(config)# devices device c0 config logging history informational
      joe@ncs(config)# devices device c0 config logging source-interface Vlan512
      joe@ncs(config)# devices device c0 config logging 10.1.22.122
      joe@ncs(config)# devices device c0 config logging 66.162.108.21
      joe@ncs(config)# devices device c0 config logging 50.58.29.21

admin@ncs% commit

      System message at 2022-09-01 14:17:59...
      Commit performed by admin via console using cli.
      joe@ncs(config-config)# commit
      Aborted: Transaction 542 conflicts with transaction 562 started by user admin: 'devices device c0 config access-list 10' read-op on-descendant write-op delete in work phase(s)
      --------------------------------------------------------------------------
      This transaction is in a non-resolvable state.
      To attempt to reapply the configuration changes made in the CLI,
      in a new transaction, revert the current transaction by running
      the command 'revert' followed by the command 'reapply-commands'.
      --------------------------------------------------------------------------
```

In this case, `joe` commits a change to `access-list` after `admin` and a conflict message is displayed. Since the conflict is non-resolvable, the transaction has to be reverted. To reapply the changes made by `joe` to `logging` in a new transaction, the following commands are entered:

```bash
      joe@ncs(config)# revert no-confirm
      joe@ncs(config)# reapply-commands best-effort
      move devices device c0 config access-list 10 permit 168.215.202.0 0.0.0.255 first
      Error: on line 1: move devices device c0 config access-list 10 permit 168.215.202.0 0.0.0.255 first
      devices device c0 config
      logging history informational
      logging facility local0
      logging source-interface Vlan512
      logging 10.1.22.122
      logging 66.162.108.21
      logging 50.58.29.21
      joe@ncs(config-config)# show config
      logging facility local0
      logging history informational
      logging 10.1.22.122
      logging 50.58.29.21
      logging 66.162.108.21
      logging source-interface Vlan512
      joe@ncs(config-config)# commit
      Commit complete.
```

In this case, `joe` tries to reapply the changes made in the previous transaction and since `access-list 10` has been removed, the move command will fail when applied by the `reapply-commands` command. Since the mode is `best-effort`, the next command will be processed. The changes to `logging` will succeed and `joe` then commits the transaction.


# NEDs and Adding Devices

Learn about NEDs, their types, and how to work with them.

Network Element Drivers, NEDs, provides the connectivity between NSO and the devices. NEDs are installed as NSO packages. For information on how to add a package for a new device type, see NSO [Package Management](/guides/administration/management/package-mgmt).

To see the list of installed packages (you will not see the F5 BigIP):

```cli
admin@ncs# show packages
packages package cisco-ios
 package-version 3.0
 description     "NED package for Cisco IOS"
 ncs-min-version [ 3.0.2 ]
 directory       ./state/packages-in-use/1/cisco-ios
 component upgrade-ned-id
  upgrade java-class-name com.tailf.packages.ned.ios.UpgradeNedId
 component cisco-ios
  ned cli ned-id  cisco-ios
  ned cli java-class-name com.tailf.packages.ned.ios.IOSNedCli
  ned device vendor Cisco
NAME      VALUE
---------------------
show-tag  interface

 oper-status up
packages package f5-bigip
 package-version 1.3
 description     "NED package for the F5 BigIp FW/LB"
 ncs-min-version [ 3.0.1 ]
 directory       ./state/packages-in-use/1/bigip
 component f5-bigip
  ned generic java-class-name com.tailf.packages.ned.bigip.BigIpNedGeneric
  ned device vendor F5
 oper-status up
!
```

The core parts of a NED are:

* **A Driver Element**: Running in a Java VM.
* **Data Model:** Independent of the underlying device interface technology, NEDs come with a data model in YANG that specifies configuration data and operational data that is supported for the device.

  * For native NETCONF devices, the YANG comes from the device.
  * For JunOS, NSO generates the model from the JunOS XML schema.
  * For SNMP devices, NSO generates the model from the MIBs.
  * For CLI devices, the NED designer writes the YANG to map the CLI.

  NSO only cares about the data that is in the model for the NED. The rest is ignored. See the [NED documentation](/guides/development/advanced-development/developing-neds) to learn more about what is covered by the NED.
* **Code:** For NETCONF and SNMP devices, there is no code. For CLI devices there is a minimum of code managing connecting over SSH/Telnet and looking for version strings. The rest is auto-rendered from the data model.

There are four categories of NEDs depending on the device interface:

1. **NETCONF NED**: The device supports NETCONF, for example, Juniper.
2. **CLI NED**: Any device with a CLI that resembles a Cisco CLI.
3. **Generic NED**: Proprietary protocols like REST, and non-Cisco CLIs.
4. **SNMP NED**: An SNMP device.

## Device Authentication <a href="#d5e524" id="d5e524"></a>

Every device needs an auth group that tells NSO how to authenticate to the device:

```cli
admin@ncs(config)# show full-configuration devices authgroups
devices authgroups group default
 umap admin
  remote-name     admin
  remote-password $4$wIo7Yd068FRwhYYI0d4IDw==
 !
 umap oper
  remote-name     oper
  remote-password $4$zp4zerM68FRwhYYI0d4IDw==
 !
!
devices authgroups snmp-group default
 umap admin
  usm remote-name admin
  usm security-level auth-priv
  usm auth sha remote-password $4$wIo7Yd068FRwhYYI0d4IDw==
  usm priv aes remote-password $4$wIo7Yd068FRwhYYI0d4IDw==
 !
!
```

The CLI snippet above shows that there is a mapping from the NSO users `admin` and `oper` to the remote user and password to be used on the devices. There are two options, either a mapping from the local user to the remote user or to pass the credentials. Below is a CLI example to create a new `authgroup foobar` and map NSO user `jim`:

```cli
admin@ncs(config)# devices authgroups group foobar umap joe same-pass same-user
admin@ncs(config-umap-joe)# commit
```

This auth group will pass on `joe`'s credentials to the device.

There is a similar structure for SNMP `devices authgroups snmp-group` that, in the examples, is used with SNMPv3 `auth-priv`, SHA authentication, and AES privacy.

## Connecting Devices for Different NED Types <a href="#d5e537" id="d5e537"></a>

Make sure you know the authentication information and created authgroups as above. Also, try all information like port numbers and authentication information, and that you can read and set the configuration over for example CLI if it is a CLI NED. So if it is a CLI device try to ssh (or telnet) to the device and do show and set configuration first of all.

All devices have a `admin-state` with default value `southbound-locked`. This means that if you do not set this value to unlocked no commands will be sent to the device.

### CLI NEDs <a href="#d5e543" id="d5e543"></a>

(See also [examples.ncs/device-management/real-device-cisco-ios](https://github.com/NSO-developer/nso-examples/tree/6.7/device-management/real-device-cisco-ios)). Straightforward, adding a new device on a specific address, standard SSH port:

```cli
admin@ncs(config)# devices device c7 address 1.2.3.4 port 22 \
                                device-type cli ned-id cisco-ios-cli-3.0
admin@ncs(config-device-c7)# authgroup
Possible completions:
  default  foobar
admin@ncs(config-device-c7)# authgroup default
admin@ncs(config-device-c7)# state admin-state unlocked
admin@ncs(config-device-c7)# commit
```

### NETCONF NEDs, JunOS <a href="#d5e554" id="d5e554"></a>

See also [examples.ncs/device-management/real-device-juniper](https://github.com/NSO-developer/nso-examples/tree/6.7/device-management/real-device-juniper). Make sure that NETCONF over SSH is enabled on the JunOS device:

```
junos1% show system services
ftp;
ssh;
telnet;
netconf {
    ssh {
        port 22;
    }
}
```

Then you can create a NSO netconf device as:

```cli
admin@ncs(config)# devices device junos1 address junos1.lab port 22 \
                                 authgroup foobar device-type netconf
admin@ncs(config-device-junos1)# state admin-state unlocked
admin@ncs(config-device-junos1)# commit
```

### SNMP NEDs <a href="#d5e566" id="d5e566"></a>

(See also [examples.ncs/device-management/snmp-ned](https://github.com/NSO-developer/nso-examples/tree/6.7/device-management/snmp-ned).) First of all, let's explain SNMP NEDs a bit. By default all read-only objects are mapped to operational data in NSO and read-write objects are mapped to configuration data. This means that a sync-from operation will load read-write objects into NSO. How can you reach read-only objects? Note the following is true for all NED types that have modeled operational data. The device configuration exists at `devices device config` and has a copy in CDB. NSO can speak live to the device to fetch for example counters by using the path `devices device live-status`:

```cli
admin@ncs# show devices device r1 live-status SNMPv2-MIB
live-status SNMPv2-MIB system sysDescr "Tail-f ConfD agent - r1"
live-status SNMPv2-MIB system sysObjectID 1.3.6.1.4.1.24961
live-status SNMPv2-MIB system sysUpTime 4253
live-status SNMPv2-MIB system sysContact ""
live-status SNMPv2-MIB system sysName ""
live-status SNMPv2-MIB system sysLocation ""
live-status SNMPv2-MIB system sysServices 72
live-status SNMPv2-MIB system sysORLastChange 0
live-status SNMPv2-MIB snmp snmpInPkts 3
live-status SNMPv2-MIB snmp snmpInBadVersions 0
live-status SNMPv2-MIB snmp snmpInBadCommunityNames 0
live-status SNMPv2-MIB snmp snmpInBadCommunityUses 0
live-status SNMPv2-MIB snmp snmpInASNParseErrs 0
live-status SNMPv2-MIB snmp snmpEnableAuthenTraps disabled
live-status SNMPv2-MIB snmp snmpSilentDrops 0
live-status SNMPv2-MIB snmp snmpProxyDrops 0
live-status SNMPv2-MIB snmpSet snmpSetSerialNo 2161860
```

In many cases, SNMP NEDs are used for reading operational data in parallel with a CLI NED for writing and reading configuration data. More on that later.

Before trying NSO use net-snmp command line tools or your favorite SNMP Browser to try that all settings are ok.

Adding an SNMP device assuming that NED is in place:

```cli
admin@ncs(config)# show full-configuration devices device r1
devices device r1
 address 127.0.0.1
 port    11023
 device-type snmp version v3
 device-type snmp snmp-authgroup default
 state admin-state unlocked
!
admin@ncs(config)# show full-configuration devices device r2
devices device r2
 address 127.0.0.1
 port    11024
 device-type snmp version v3
 device-type snmp snmp-authgroup default
 device-type snmp mib-group [ basic snmp ]
 state admin-state unlocked
!
```

MIB Groups are important. A MIB group is just a named collection of SNMP MIB Modules. If you do not specify any MIB group for a device, NSO will try with all known MIBs. It is possible to create MIB groups with wild cards such as `CISCO*`.

```cli
admin@ncs(config)# show full-configuration devices mib-group
devices mib-group basic
 mib-module [ BASIC-CONFIG-MIB ]
!
devices mib-group snmp
 mib-module [ SNMP* ]
!
```

### Generic NEDs <a href="#d5e585" id="d5e585"></a>

Generic devices are typically configured like a CLI device. Make sure you set the right address, port, protocol, and authentication information.

Below is an example of setting up NSO with F5 BigIP:

```cli
admin@ncs(config)# devices device bigip01 address 192.168.1.162 \
                                 port 22 device-type generic ned-id f5-bigip
admin@ncs(config-device-bigip01)# state admin-state southbound-locked
admin@ncs(config-device-bigip01)# authgroup
Possible completions:
  default  foobar
admin@ncs(config-device-bigip01)# authgroup default
admin@ncs(config-device-bigip01)# commit
```

### Live Status Protocol <a href="#d5e596" id="d5e596"></a>

Assume that you have a Cisco device that you would like NSO to configure over CLI but read statistics over SNMP. This can be achieved by adding settings for `live-device-protocol`:

```cli
admin@ncs(config)# devices device c0 live-status-protocol snmp \
                                device-type snmp version v3 \
                                snmp-authgroup default mib-group [ snmp ]
admin@ncs(config-live-status-protocol-snmp)# commit


admin@ncs(config)# show full-configuration devices device c0
devices device c0
 address   127.0.0.1
 port      10022
 !
 authgroup default
 device-type cli ned-id cisco-ios
 live-status-protocol snmp
  device-type snmp version v3
  device-type snmp snmp-authgroup default
  device-type snmp mib-group [ snmp ]
 !
```

Device `c0` has a config tree from the CLI NED and a live-status tree (read-only) from the SNMP NED using all MIBs in the group `snmp`.

#### Multi-NEDs for Statistics

Sometimes we wish to use a different protocol to collect statistics from the live tree than the protocol that is used to configure a managed device. There are many interesting use cases where this pattern applies. For example, if we wish to access SNMP data as statistics in the live tree on a Juniper router, or alternatively, if we have a CLI NED to a Cisco-type device, and wish to access statistics in the live tree over SNMP.

The solution is to configure additional protocols for the live tree. We can have an arbitrary number of NEDs associated to statistics data for an individual managed device.

The additional NEDs are configured under `/devices/device/live-status-protocol`.

In the configuration snippet below, we have configured two additional NEDs for statistics data.

```
devices {
    authgroups {
        snmp-group g1 {
            umap admin {
                usm {
                    remote-name    admin;
                    security-level auth-priv;
                    auth {
                        sha {
                            remote-password $4$wIo7Yd068FRwhYYI0d4IDw==;
                        }
                    }
                    priv {
                        aes {
                            remote-password $4$wIo7Yd068FRwhYYI0d4IDw==;
                        }
                    }
                }
            }
        }
    }
    mib-group m1 {
        mib-module [ SIMPLE-MIB ];
    }
    device device0 {
        live-status-protocol x1 {
            port 4001;
            device-type {
                snmp {
                    version        v3;
                    snmp-authgroup g1;
                    mib-group      [ m1 ];
                }
            }
        }
        live-status-protocol x2 {
            authgroup default;
            device-type {
                cli {
                    ned-id xstats;
                }
            }
        }
     }
```

## Administrative State for Devices <a href="#d5e605" id="d5e605"></a>

Devices have an `admin-state` with following values:

* **unlocked**: the device can be modified and changes will be propagated to the real device.
* **southbound-locked**: the device can be modified but changes will not be propagated to the real device. Can be used to prepare configurations before the device is available in the network.
* **locked**: the device can only be read.

The admin-state value southbound-locked is the default. This means if you create a new device without explicitly setting this value configuration changes will not propagate to the network. To see default values, use the pipe target `details`

```cli
admin@ncs(config)# show full-configuration devices device c0 | details
```

## Troubleshooting NEDs <a href="#d5e626" id="d5e626"></a>

To analyze NED problems, turn on the tracing for a device and look at the trace file contents.

```cli
admin@ncs(config)# show full-configuration devices global-settings
devices global-settings trace-dir ./logs

admin@ncs(config)# devices device c0 trace raw
admin@ncs(config-device-c0)# commit

admin@ncs(config)# devices device c0 disconnect
admin@ncs(config)# devices device c0 connect
```

NSO pools SSH connections and trace settings are only affecting new connections so therefore any open connection must be closed before the trace setting will take effect. Now you can inspect the raw communication between NSO and the device:

```bash
$ less logs/ned-c0.trace

admin connected from 127.0.0.1 using ssh on HOST-17
c0>
  *** output 8-Sep-2014::10:05:39.673 ***
enable

  *** input 8-Sep-2014::10:05:39.674 ***
 enable
c0#
  *** output 8-Sep-2014::10:05:39.713 ***
terminal length 0

  *** input 8-Sep-2014::10:05:39.714 ***
 terminal length 0
c0#
  *** output 8-Sep-2014::10:05:39.782 ***
terminal width 0

  *** input 8-Sep-2014::10:05:39.783 ***
 terminal width 0
0^M
c0#
  *** output 8-Sep-2014::10:05:39.839 ***
-- Requesting version string --
show version

  *** input 8-Sep-2014::10:05:39.839 ***
 show version
Cisco IOS Software, 7200 Software (C7200-JK9O3S-M), Version 12.4(7h), RELEASE SOFTWARE (fc1)^M
Technical Support: http://www.cisco.com/techsupport^M
Copyright (c) 1986-2007 by Cisco Systems, Inc.^M
...
```

### Device Communication Failure

If NSO fails to talk to the device, the typical root causes are:

<details>

<summary>Timeout Problems</summary>

Some devices are slow to respond, latency on connections, etc. Fine-tune the connect, read, and write timeouts for the device:

```cli
admin@ncs(config)# devices device c0
Possible completions:
  ...
  connect-timeout           - Timeout in seconds for new connections
  ...
  read-timeout              - Timeout in seconds used when reading data
  ...
  write-timeout             - Timeout in seconds used when writing data
```

\
These settings can be set in profiles shared by devices.

```cli
admin@ncs(config)# devices profiles profile good-profile
Possible completions:
  connect-timeout   Timeout in seconds for new connections
  ned-settings      Control which device capabilities NCS uses
  read-timeout      Timeout in seconds used when reading data
  trace             Trace the southbound communication to devices
  write-timeout     Timeout in seconds used when writing data
```

</details>

<details>

<summary>Device Management Interface Problems</summary>

Examples, not enabling the NETCONF SSH subsystem on Juniper, not enabling the SNMP agent, using the wrong port numbers, etc. Use standalone tools to make sure that you can connect, read configuration, and write configuration over the device interface that NSO is using

</details>

<details>

<summary>Access Rights</summary>

The NSO-mapped user does not have access rights to do the operation on the device. Make sure the `authgroups` settings are OK and test them manually to read and write configuration with those credentials.

</details>

<details>

<summary>NED Data Model and Device Version Problems</summary>

If the device is upgraded and existing commands actually change in an incompatible way, the NED has to be updated. This can be done by editing the YANG data model for the device or by using Cisco support.

</details>


# Manage Network Services

Manage the life-cycle of network services.

NSO can also manage the life-cycle for services like VPNs, BGP peers, and ACLs. It is important to understand what is meant by service in this context:

* NSO abstracts the device-specific details. The user only needs to enter attributes relevant to the service.
* The service instance has configuration data itself that can be represented and manipulated.
* A service instance configuration change is applied to all affected devices.

## Service Configuration Features

The following are the features that NSO uses to support service configuration:

* **Service Modeling**: Network engineers can model the service attributes and the mapping to device configurations. For example, this means that a network engineer can specify at data-model for VPNs with router interfaces, VLAN ID, VRF, and route distinguisher.
* **Service Life-cycle**: While less sophisticated configuration management systems can only create an initial service instance in the network they do not support changing or deleting a service instance. With NSO you can at any point in time modify service elements like the VLAN id of a VPN and NSO can generate the corresponding changes to the network devices.
* **Service Instance**: The NSO service instance has configuration data that can be represented and manipulated. The service model on run-time updates all NSO northbound interfaces so that a network engineer can view and manipulate the service instance over CLI, WebUI, REST, etc.
* **References between Service Instances and Device Configuration**: NSO maintains references between service instances and device configuration. This means that a VPN instance knows exactly which device configurations it created or modified. Every configuration stored in the CDB is mapped to the service instance that created it.

## Service Example <a href="#d5e684" id="d5e684"></a>

An example is the best method to illustrate how services are created and used in NSO. As described in the sections about devices and NEDs, it was said that NEDs come in packages. The same is true for services, either if you design the services yourself or use ready-made service applications, it ends up in a package that is loaded into NSO.

{% hint style="success" %}
Watch a video presentation of this demo on [YouTube](https://www.youtube.com/watch?v=sYuETSuTsrM).
{% endhint %}

The example [examples.ncs/service-management/mpls-vpn-java](https://github.com/NSO-developer/nso-examples/tree/6.7/service-management/mpls-vpn-java) will be used to explain NSO Service Management features. This example illustrates Layer-3 VPNs in a service provider MPLS network. The example network consists of Cisco ASR 9k and Juniper core routers (P and PE) and Cisco IOS-based CE routers. The Layer-3 VPN service configures the CE/PE routers for all endpoints in the VPN with BGP as the CE/PE routing protocol. The layer-2 connectivity between CE and PE routers is expected to be done through a Layer-2 ethernet access network, which is out of scope for this example. The Layer-3 VPN service includes VPN connectivity as well as bandwidth and QOS parameters.

<div data-with-frame="true"><figure><img src="/files/PddnV8PVgxk8jLdRXaVv" alt="" width="563"><figcaption><p>A L3 VPN Example</p></figcaption></figure></div>

The service configuration only has references to CE devices for the end-points in the VPN. The service mapping logic reads from a simple topology model that is configuration data in NSO, outside the actual service model and derives what other network devices to configure.

The topology information has two parts:

* The first part lists connections in the network and is used by the service mapping logic to find out which PE router to configure for an endpoint. The snippets below show the configuration output in the Cisco-style NSO CLI.

  ```
   topology connection c0
   endpoint-1 device ce0 interface GigabitEthernet0/8 ip-address 192.168.1.1/30
   endpoint-2 device pe0 interface GigabitEthernet0/0/0/3 ip-address 192.168.1.2/30
   link-vlan 88
  !
  topology connection c1
   endpoint-1 device ce1 interface GigabitEthernet0/1 ip-address 192.168.1.5/30
   endpoint-2 device pe1 interface GigabitEthernet0/0/0/3 ip-address 192.168.1.6/30
   link-vlan 77
  !
  ```
* The second part lists devices for each role in the network and is in this example only used to dynamically render a network map in the Web UI.

  ```
  topology role ce
   device [ ce0 ce1 ce2 ce3 ce4 ce5 ]
  !
  topology role pe
   device [ pe0 pe1 pe2 pe3 ]
  !
  ```

The QOS configuration in service provider networks is complex and often requires a lot of different variations. It is also often desirable to be able to deliver different levels of QOS. This example shows how a QOS policy configuration can be stored in NSO and referenced from VPN service instances. Three different levels of QOS policies are defined; `GOLD`, `SILVER`, and `BRONZE` with different queuing parameters.

```
 qos qos-policy GOLD
 class BUSINESS-CRITICAL
  bandwidth-percentage 20
 !
 class MISSION-CRITICAL
  bandwidth-percentage 20
 !
 class REALTIME
  bandwidth-percentage 20
  priority
 !
!
qos qos-policy SILVER
 class BUSINESS-CRITICAL
  bandwidth-percentage 25
 !
 class MISSION-CRITICAL
  bandwidth-percentage 25
 !
 class REALTIME
  bandwidth-percentage 10
 !
```

Three different traffic classes are also defined with a DSCP value that will be used inside the MPLS core network as well as default rules that will match traffic to a class.

```
qos qos-class BUSINESS-CRITICAL
 dscp-value af21
 match-traffic ssh
  source-ip      any
  destination-ip any
  port-start     22
  port-end       22
  protocol       tcp
 !
!
qos qos-class MISSION-CRITICAL
 dscp-value af31
 match-traffic call-signaling
  source-ip      any
  destination-ip any
  port-start     5060
  port-end       5061
  protocol       tcp
 !
!
```

## Running the Example <a href="#d5e706" id="d5e706"></a>

Run the example as follows:

1. Make sure that you start clean, i.e. no old configuration data is present. If you have been running this or some other example before, make sure to stop any NSO or simulated network nodes (ncs-netsim) that you may have running. Output like 'connection refused (stop)' means no previous NSO was running and 'DEVICE ce0 connection refused (stop)...' no simulated network was running, which is good.

   ```
   Copy$
   ```

   \
   This will set up the environment and start the simulated network.
2. Before creating a new L3VPN service, we must sync the configuration from all network devices and then enter config mode. (A hint for this complete section is to have the `README` file from the example and cut and paste the CLI commands).

   ```
   Copyncs#
   ```
3. Add another VPN.

   ```
   top
   !
   vpn l3vpn ford
   as-number 65200
   endpoint main-office
   ce-device    ce2
   ce-interface GigabitEthernet0/5
   ip-network   192.168.1.0/24
   bandwidth    10000000
   !
   endpoint branch-office1
   ce-device    ce3
   ce-interface GigabitEthernet0/5
   ip-network   192.168.2.0/24
   bandwidth    5500000
   !
   endpoint branch-office2
   ce-device    ce5
   ce-interface GigabitEthernet0/5
   ip-network   192.168.7.0/24
   bandwidth    1500000
   !
   ```

   \
   The above sequence showed how NSO can be used to manipulate service abstractions on top of devices. Services can be defined for various purposes such as VPNs, Access Control Lists, firewall rules, etc. Support for services is added to NSO via a corresponding service package.

A service package in NSO comprises two parts:

1. **Service model:** the attributes of the service, and input parameters given when creating the service. In this example name, as-number, and end-points.
2. **Mapping**: what is the corresponding configuration of the devices when the service is applied. The result of the mapping can be inspected by the `commit dry-run outformat native` command.

We later show how to define this, for now, assume that the job is done.

## Service-Life Cycle Management <a href="#d5e757" id="d5e757"></a>

### Service Changes <a href="#d5e759" id="d5e759"></a>

When NSO applies services to the network, NSO stores the service configuration along with resulting device configuration changes. This is used as a base for the FASTMAP algorithm which automatically can derive device configuration changes from a service change.

**Example 1**

Going back to the example L3 VPN above, any part of `volvo` VPN instance can be modified.

A simple change like changing the `as-number` on the service results in many changes in the network. NSO does this automatically.

```
ncs(config)# vpn l3vpn volvo as-number 65102
ncs(config-l3vpn-volvo)# commit dry-run outformat native
native {
    device {
        name ce0
        data no router bgp 65101
             router bgp 65102
              neighbor 192.168.1.2 remote-as 100
              neighbor 192.168.1.2 activate
              network 10.10.1.0
             !
...
ncs(config-l3vpn-volvo)# commit
```

**Example 2**

Let us look at a more challenging modification.

A common use case is of course to add a new CE device and add that as an end-point to an existing VPN. Below is the sequence to add two new CE devices and add them to the VPNs. (In the CLI snippets below we omit the prompt to enhance readability).

First, we add them to the topology:

```
top
!
topology connection c7
endpoint-1 device ce7 interface GigabitEthernet0/1 ip-address 192.168.1.25/30
endpoint-2 device pe3 interface GigabitEthernet0/0/0/2 ip-address 192.168.1.26/30
link-vlan 103
!
topology connection c8
endpoint-1 device ce8 interface GigabitEthernet0/1 ip-address 192.168.1.29/30
endpoint-2 device pe3 interface GigabitEthernet0/0/0/2 ip-address 192.168.1.30/30
link-vlan 104
!
ncs(config)#commit
```

Note well that the above just updates NSO local information on topological links. It has no effect on the network. The mapping for the L3 VPN services does a look-up in the topology connections to find the corresponding `pe` router.

Next, we add them to the VPNs:

```
top
!
vpn l3vpn ford
endpoint new-branch-office
ce-device    ce7
ce-interface GigabitEthernet0/5
ip-network   192.168.9.0/24
bandwidth    4500000
!
vpn l3vpn volvo
endpoint new-branch-office
ce-device    ce8
ce-interface GigabitEthernet0/5
ip-network   10.8.9.0/24
bandwidth    4500000
!
```

Before we send anything to the network, let's look at the device configuration using a dry run. As you can see, both new CE devices are connected to the same PE router, but for different VPN customers.

```
ncs(config)# commit dry-run outformat native
```

Finally, commit the configuration to the network

```
(config)# commit
```

### Service Impacting Out-of-band Changes <a href="#d5e779" id="d5e779"></a>

Next, we will show how NSO can be used to check if the service configuration in the network is up to date.

In a new terminal window, we connect directly to the device `ce0` which is a Cisco device emulated by the tool `ncs-netsim`.

```bash
$ ncs-netsim cli-c ce0
```

We will now reconfigure an edge interface that we previously configured using NSO.

```
 enable
ce0# configure
Enter configuration commands, one per line. End with CNTL/Z.
ce0(config)# no policy-map volvo
ce0(config)# exit
ce0# exit
```

Going back to the terminal with NSO, check the status of the network configuration:

```cli
ncs# devices check-sync
sync-result {
    device ce0
    result out-of-sync
    info got: c5c75ee593246f41eaa9c496ce1051ea expected: c5288cc0b45662b4af88288d29be8667
...

ncs# vpn l3vpn * check-sync
vpn l3vpn ford check-sync
    in-sync true
vpn l3vpn volvo check-sync
    in-sync true

ncs# vpn l3vpn * deep-check-sync
vpn l3vpn ford deep-check-sync
    in-sync true
vpn l3vpn volvo deep-check-sync
    in-sync false
```

The CLI sequence above performs 3 different comparisons:

* Real device configuration versus device configuration copy in NSO CDB.
* Expected device configuration from the service perspective and device configuration copy in CDB.
* Expected device configuration from the service perspective and real device configuration.

Notice that the service `volvo` is out of sync with the service configuration. Use the `check-sync outformat cli` to see what the problem is:

```cli
ncs# vpn l3vpn volvo deep-check-sync outformat cli
cli  devices {
         devices {
             device ce0 {
                 config {
    +                ios:policy-map volvo {
    +                    class class-default {
    +                        shape {
    +                            average {
    +                                bit-rate 12000000;
    +                            }
    +                        }
    +                    }
    +                }
                 }
             }
         }
     }
```

Assume that a network engineer considers the real device configuration to be authoritative:

```cli
ncs# devices device ce0 sync-from
result true
```

Then they can restore the service:

```cli
ncs# vpn l3vpn volvo re-deploy dry-run { outformat native }
native {
    device {
        name ce0
        data policy-map volvo
               class class-default
                shape average 12000000
               !
              !

    }
}
ncs# vpn l3vpn volvo re-deploy
```

However, in some cases the device change was made with a good reason and it may be desirable to keep it. In that case, the engineer can either:

* Update the service configuration in NSO to reflect the new device configuration. This is the preferred way but requires the service to support the particular device configuration.
* Alternatively, accept the change as out-of-band service change, as described in [Out-of-band Interoperation](/guides/operation-and-usage/operations/out-of-band-interoperation).

### Service Deletion <a href="#d5e809" id="d5e809"></a>

In the same way, as NSO can calculate any service configuration change, it can also automatically delete the device configurations that resulted from creating services:

```cli
ncs(config)# no vpn l3vpn ford
ncs(config)# commit dry-run
cli  devices {
         device ce7
             config {
    -            ios:policy-map ford {
    -                class class-default {
    -                    shape {
    -                        average {
    -                            bit-rate 4500000;
    -                    }
    -                }
    -            }
    -        }
...
```

It is important to understand the two diffs shown above. The first diff as an output to `show configuration` shows the diff at the service level. The second diff shows the output generated by NSO to clean up the device configurations.

Finally, we commit the changes to delete the service.

```
(config)# commit
```

### Viewing Service Configurations <a href="#d5e819" id="d5e819"></a>

Service instances live in the NSO data store as well as a copy of the device configurations. NSO will maintain relationships between these two.

Show the configuration for a service

```cli
ncs(config)# show full-configuration vpn l3vpn
vpn l3vpn volvo
 as-number 65102
 endpoint branch-office1
  ce-device    ce1
  ce-interface GigabitEthernet0/11
  ip-network   10.7.7.0/24
  bandwidth    6000000
 !
...
```

You can ask NSO to list all devices that are touched by a service and vice versa:

```
ncs# show vpn l3vpn modified devices
NAME   DEVICES
------------------------------------
volvo  [ ce0 ce1 ce4 ce8 pe0 pe2 pe3 ]

ncs# show devices device services
NAME  ID
--------------------------------
ce0   /vpn/l3vpn[name='volvo']
ce1   /vpn/l3vpn[name='volvo']
ce2
ce3
ce4   /vpn/l3vpn[name='volvo']
ce5
ce6
ce7
ce8   /vpn/l3vpn[name='volvo']
p0
p1
p2
p3
pe0   /vpn/l3vpn[name='volvo']
pe1
pe2   /vpn/l3vpn[name='volvo']
pe3   /vpn/l3vpn[name='volvo']
```

Note that operational mode in the CLI was used above. Every service instance has an operational attribute that is maintained by the transaction manager and shows which device configuration it created. Furthermore, every device configuration has backward pointers to the corresponding service instances:

```cli
ncs(config)# show full-configuration devices device ce3 \
                    config | display service-meta-data
devices device ce3
 config
  ...
  /* Refcount: 1 */
  /* Backpointer: [ /l3vpn:vpn/l3vpn:l3vpn[l3vpn:name='ford'] ] */
  ios:interface GigabitEthernet0/2.100
   /* Refcount: 1 */
   description Link to PE / pe1 - GigabitEthernet0/0/0/5
   /* Refcount: 1 */
   encapsulation dot1Q 100
   /* Refcount: 1 */
   ip address 192.168.1.13 255.255.255.252
   /* Refcount: 1 */
   service-policy output ford
  exit

ncs(config)# show full-configuration devices device ce3 config \
                     | display curly-braces | display service-meta-data
...
ios:interface {
    GigabitEthernet 0/1;
    GigabitEthernet 0/10;
    GigabitEthernet 0/11;
    GigabitEthernet 0/12;
    GigabitEthernet 0/13;
    GigabitEthernet 0/14;
    GigabitEthernet 0/15;
    GigabitEthernet 0/16;
    GigabitEthernet 0/17;
    GigabitEthernet 0/18;
    GigabitEthernet 0/19;
    GigabitEthernet 0/2;
    /* Refcount: 1 */
    /* Backpointer: [ /l3vpn:vpn/l3vpn:l3vpn[l3vpn:name='ford'] ] */
    GigabitEthernet 0/2.100 {
        /* Refcount: 1 */
        description "Link to PE / pe1 - GigabitEthernet0/0/0/5";
        encapsulation {
            dot1Q {
                /* Refcount: 1 */
                vlan-id 100;
            }
        }
        ip {
            address {
                primary {
                    /* Refcount: 1 */
                    address 192.168.1.13;
                    /* Refcount: 1 */
                    mask    255.255.255.252;
                }
            }
        }
        service-policy {
            /* Refcount: 1 */
            output ford;
        }
    }

ncs(config)# show full-configuration devices device ce3 config \
                   | display service-meta-data | context-match Backpointer
devices device ce3
  /* Refcount: 1 */
  /* Backpointer: [ /l3vpn:vpn/l3vpn:l3vpn[l3vpn:name='ford'] ] */
  ios:interface GigabitEthernet0/2.100
devices device ce3
  /* Refcount: 2 */
  /* Backpointer: [ /l3vpn:vpn/l3vpn:l3vpn[l3vpn:name='ford'] ] */
  ios:interface GigabitEthernet0/5
```

The reference counter above makes sure that NSO will not delete shared resources until the last service instance is deleted. The context-match search is helpful, it displays the path to all matching configuration items.

### Using Commit Queues <a href="#d5e833" id="d5e833"></a>

As described in [Commit Queue](https://nso-docs.cisco.com/guides/operation-and-usage/operations/pages/auKQMOAF2p1jiGYJBweP#user_guide.devicemanager.commit-queue), the commit queue can be used to increase the transaction throughput. When the commit queue is for service activation, the services will have states reflecting outstanding commit queue items.

{% hint style="info" %}
When committing a service using the commit queue in *async* mode the northbound system can not rely on the service being fully activated in the network when the activation requests return.
{% endhint %}

We will now commit a VPN service using the commit queue and one device is down.

```bash
$ ncs-netsim stop ce0
DEVICE ce0 STOPPED
```

```cli
ncs(config)# show configuration
vpn l3vpn volvo
 as-number 65101
 endpoint branch-office1
  ce-device    ce1
  ce-interface GigabitEthernet0/11
  ip-network   10.7.7.0/24
  bandwidth    6000000
 !
 endpoint main-office
  ce-device    ce0
  ce-interface GigabitEthernet0/11
  ip-network   10.10.1.0/24
  bandwidth    12000000
 !
!

ncs# commit commit-queue async
commit-queue-id 10777927137
Commit complete.
ncs(config)# *** ALARM connection-failure: Failed to connect to device ce0: connection refused: Connection refused
```

This service is not provisioned fully in the network, since `ce0` was down. It will stay in the queue either until the device starts responding or when an action is taken to remove the service or remove the item. The commit queue can be inspected. As shown below we see that we are waiting for `ce0`. Inspecting the queue item shows the outstanding configuration.

```cli
ncs# show devices commit-queue | notab
devices commit-queue queue-item 10777927137
 age       1934
 status    executing
 devices   [ ce0 ce1 pe0 ]
 transient ce0
  reason "Failed to connect to device ce0: connection refused"
 is-atomic true

ncs# show vpn l3vpn volvo commit-queue | notab
commit-queue queue-item 1498812003922
```

The commit queue will constantly try to push the configuration towards the devices. The number of retry attempts and at what interval they occur can be configured.

```cli
ncs# show full-configuration devices global-settings commit-queue | details
devices global-settings commit-queue enabled-by-default false
devices global-settings commit-queue atomic true
devices global-settings commit-queue retry-timeout 30
devices global-settings commit-queue retry-attempts unlimited
```

If we start `ce0` and inspect the queue, we will see that the queue will finally be empty and that the `commit-queue` status for the service is empty.

```cli
ncs# show devices commit-queue | notab
devices commit-queue queue-item 10777927137
 age       3357
 status    executing
 devices   [ ce0 ce1 pe0 ]
 transient ce0
  reason "Failed to connect to device ce0: connection refused"
 is-atomic true

ncs# show devices commit-queue | notab
devices commit-queue queue-item 10777927137
 age       3359
 status    executing
 devices   [ ce0 ce1 pe0 ]
 is-atomic true

ncs# show devices commit-queue
% No entries found.

ncs# show vpn l3vpn volvo commit-queue
% No entries found.

ncs# show devices commit-queue completed | notab
devices commit-queue completed queue-item 10777927137
 when               2015-02-09T16:48:17.915+00:00
 succeeded          true
 devices            [ ce0 ce1 pe0 ]
 completed          [ ce0 ce1 pe0 ]
 completed-services [ /l3vpn:vpn/l3vpn:l3vpn[l3vpn:name='volvo'] ]
```

### Un-deploying Services <a href="#d5e863" id="d5e863"></a>

In some scenarios, it makes sense to remove the service configuration from the network but keep the representation of the service in NSO. This is called to `un-deploy` a service.

```cli
ncs# vpn l3vpn volvo check-sync
in-sync false
ncs# vpn l3vpn volvo re-deploy
ncs# vpn l3vpn volvo check-sync
in-sync true
```

## Defining Your Own Services <a href="#d5e871" id="d5e871"></a>

### Overview <a href="#d5e873" id="d5e873"></a>

To have NSO deploy services across devices, two pieces are needed:

1. A service model in YANG: the service model shall define the black-box view of a service; which are the input parameters given when creating the service? This YANG model will render an update of all NSO northbound interfaces, for example, the CLI.
2. Mapping, given the service input parameters, what is the resulting device configuration? This mapping can be defined in templates, code, or a combination of both.

### Defining the Service Model <a href="#d5e881" id="d5e881"></a>

The first step is to generate a skeleton package for a service (for details, see [Packages](/guides/administration/management/package-mgmt)). Create a directory under, for example, `~/my-sim-ios`similar to how it is done for the [examples.ncs/device-management/simulated-devices](https://github.com/NSO-developer/nso-examples/tree/6.7/device-management/simulated-devices) example. Make sure that you have stopped any running NSO and netsim.

Navigate to the simulated ios directory and create a new package for the VLAN service model:

```bash
$ cd examples.ncs/device-management/simulated-devices/packages
```

If the `packages` folder does not exist yet, such as when you have not run this example before, you will need to invoke the `ncs-setup` and `ncs-netsim create-network` commands as described in the `simulated-devices` `README` file.

The next step is to create the template skeleton by using the `ncs-make-package` utility:

```bash
$ ncs-make-package --service-skeleton template --root-container vlans --no-test  vlan
```

This results in a directory structure:

```
vlan
   package-meta-data.xml
   src
   templates
```

For now, let's focus on the `src/yang/vlan.yang` file.

```yang
            module vlan {
              namespace "http://com/example/vlan";
              prefix vlan;

              import ietf-inet-types {
                prefix inet;
              }
              import tailf-ncs {
                prefix ncs;
              }

              container vlans {
              list vlan {
                key name;

                uses ncs:service-data;
                ncs:servicepoint "vlan";

                leaf name {
                  type string;
                }

                // may replace this with other ways of refering to the devices.
                leaf-list device {
                  type leafref {
                    path "/ncs:devices/ncs:device/ncs:name";
                  }
                }

                // replace with your own stuff here
                leaf dummy {
                  type inet:ipv4-address;
                }
              }
              } // container vlans {
            }
```

If this is your first exposure to YANG, you can see that the modeling language is very straightforward and easy to understand. See [RFC 7950](https://www.ietf.org/rfc/rfc7950.txt) for more details and examples for YANG. The concept to understand in the above-generated skeleton is that the two lines of `uses ncs:service-data` and `ncs:servicepoint "vlan"` tells NSO that this is a service. The `ncs:service-data` grouping together with the `ncs:servicepoint` YANG extension provides the common definitions for a service. The two are implemented by the `$NCS_DIR/src/ncs/yang/tailf-ncs-services.yang`. So if a user wants to create a new VLAN in the network what should be the parameters? - A very simple service model would look like below (modify the `src/yang/vlan.yang` file):

```yang
  augment /ncs:services {
    container vlans {
      key name;

      uses ncs:service-data;
      ncs:servicepoint "vlan";
      leaf name {
        type string;
      }

      leaf vlan-id {
        type uint32 {
          range "1..4096";
        }
      }

      list device-if {
        key "device-name";
          leaf device-name {
            type leafref {
              path "/ncs:devices/ncs:device/ncs:name";
            }
          }
          leaf interface-type {
            type enumeration {
              enum FastEthernet;
              enum GigabitEthernet;
              enum TenGigabitEthernet;
            }
          }
          leaf interface {
            type string;
          }
      }
    }
}
```

This simple VLAN service model says:

1. We give a VLAN a name, for example, net-1, this must also be unique, it is specified as `key`.
2. The VLAN has an id from 1 to 4096.
3. The VLAN is attached to a list of devices and interfaces. To make this example as simple as possible the interface reference is selected by picking the type and then the name as a plain string.

The good thing with NSO is that already at this point you could load the service model to NSO and try if it works well in the CLI etc. Nothing would happen to the devices since we have not defined the mapping, but this is normally the way to iterate a model and test the CLI towards the network engineers.

To build this service model `cd` to the [examples.ncs/device-management/simulated-devices](https://github.com/NSO-developer/nso-examples/tree/6.7/device-management/simulated-devices) example `/packages/vlan/src` directory and type `make` (assuming you have the `make` build system installed).

```bash
$ make
```

Go to the root directory of the `simulated-ios` example:

```bash
$ cd $NCS_DIR/examples.ncs/device-management/simulated-devices
```

Start netsim, NSO, and the CLI:

```bash
$ ncs-netsim start
$ ncs --with-package-reload
$ ncs_cli -C -u admin
```

When starting NSO above we give NSO a parameter to reload all packages so that our newly added `vlan` package is included. Packages can also be reloaded without restart. At this point we have a service model for VLANs, but no mapping of VLAN to device configurations. This is fine, we can try the service model and see if it makes sense. Create a VLAN service:

```cli
admin@ncs(config)# services vlan net-0 vlan-id 1234 \
device-if c0 interface-type FastEthernet interface 1/0
admin@ncs(config-device-if-c0)# top
admin@ncs(config)# show configuration
services vlan net-0
 vlan-id 1234
 device-if c0
  interface-type FastEthernet
  interface      1/0
 !
!
admin@ncs(config)# services vlan net-0 vlan-id 1234 \
device-if c1 interface-type FastEthernet interface 1/0
admin@ncs(config-device-if-c1)# top
admin@ncs(config)# show configuration
services vlan net-0
 vlan-id 1234
 device-if c0
  interface-type FastEthernet
  interface      1/0
 !
 device-if c1
  interface-type FastEthernet
  interface      1/0
 !
!
admin@ncs(config)# commit dry-run outformat cli
cli {
    local-node {
        data  services {
             +    vlan net-0 {
             +        vlan-id 1234;
             +        device-if c0 {
             +            interface-type FastEthernet;
             +            interface 1/0;
             +        }
             +        device-if c1 {
             +            interface-type FastEthernet;
             +            interface 1/0;
             +        }
             +    }
              }
    }
}
admin@ncs(config)# commit
Commit complete.
admin@ncs(config)# no services vlan
admin@ncs(config)# commit
Commit complete.
```

Committing service changes does not affect the devices since we have not defined the mapping. The service instance data will just be stored in NSO CDB.

Note that you get tab completion on the devices since they are leafrefs to device names in CDB, the same for interface-type since the types are enumerated in the model. However the interface name is just a string, and you have to type the correct interface name. For service models where there is only one device type like in this simple example, we could have used a reference to the ios interface name according to the IOS model. However that makes the service model dependent on the underlying device types and if another type is added, the service model needs to be updated and this is most often not desired. There are techniques to get tab completion even when the data type is a string, but this is omitted here for simplicity.

Make sure you delete the `vlan` service instance as above before moving on with the example.

### Defining the Mapping <a href="#d5e954" id="d5e954"></a>

Now it is time to define the mapping from service configuration to actual device configuration. The first step is to understand the actual device configuration. Hard-wire the VLAN towards a device as example. This concrete device configuration is a boilerplate for the mapping, it shows the expected result of applying the service.

```cli
admin@ncs(config)# devices device c0 config ios:vlan 1234
admin@ncs(config-vlan)# top
admin@ncs(config)# devices device c0 config ios:interface \
                   FastEthernet 10/10 switchport trunk allowed vlan 1234
admin@ncs(config-if)# top
admin@ncs(config)# show configuration
devices device c0
 config
  ios:vlan 1234
  !
  ios:interface FastEthernet10/10
   switchport trunk allowed vlan 1234
  exit
 !
!
admin@ncs(config)# commit
```

The concrete configuration above has the interface and VLAN hard-wired. This is what we now will make into a template instead. It is always recommended to start like the above and create a concrete representation of the configuration the template shall create. Templates are device-configuration where parts of the config are represented as variables. These kinds of templates are represented as XML files. Show the above as XML:

```cli
admin@ncs(config)# show full-configuration devices device c0 \
                                 config ios:vlan | display xml
<config xmlns="http://tail-f.com/ns/config/1.0">
  <devices xmlns="http://tail-f.com/ns/ncs">
  <device>
    <name>c0</name>
      <config>
      <vlan xmlns="urn:ios">
        <vlan-list>
          <id>1234</id>
        </vlan-list>
      </vlan>
      </config>
  </device>
  </devices>
</config>

admin@ncs(config)# show full-configuration devices device c0 \
                                config ios:interface FastEthernet 10/10 | display xml
<config xmlns="http://tail-f.com/ns/config/1.0">
  <devices xmlns="http://tail-f.com/ns/ncs">
  <device>
    <name>c0</name>
      <config>
      <interface xmlns="urn:ios">
      <FastEthernet>
        <name>10/10</name>
        <switchport>
          <trunk>
            <allowed>
              <vlan>
                <vlans>1234</vlans>
              </vlan>
            </allowed>
          </trunk>
        </switchport>
      </FastEthernet>
      </interface>
      </config>
  </device>
  </devices>
</config>
admin@ncs(config)#
```

Now, we shall build that template. When the package was created a skeleton XML file was created in `packages/vlan/templates/vlan.xml`

```xml
<config-template xmlns="http://tail-f.com/ns/config/1.0"
                 servicepoint="vlan">
  <devices xmlns="http://tail-f.com/ns/ncs">
    <device>
      <!--
          Select the devices from some data structure in the service
          model. In this skeleton the devices are specified in a leaf-list.
          Select all devices in that leaf-list:
      -->
      <name>{/device}</name>
      <config>
        <!--
            Add device-specific parameters here.
            In this skeleton the service has a leaf "dummy"; use that
            to set something on the device e.g.:
            <ip-address-on-device>{/dummy}</ip-address-on-device>
        -->
      </config>
    </device>
  </devices>
</config-template>
```

We need to specify the right path to the devices. In our case, the devices are identified by `/device-if/device-name` (see the YANG service model).

For each of those devices, we need to add the VLAN and change the specified interface configuration. Copy the XML config from the CLI and replace it with variables:

```xml
<config-template xmlns="http://tail-f.com/ns/config/1.0"
                 servicepoint="vlan">
  <devices xmlns="http://tail-f.com/ns/ncs">
    <device>
      <name>{/device-if/device-name}</name>
      <config>
        <vlan xmlns="urn:ios">
          <vlan-list tags="merge">
            <id>{../vlan-id}</id>
          </vlan-list>
        </vlan>
        <interface xmlns="urn:ios">
          <?if {interface-type='FastEthernet'}?>
            <FastEthernet tags="nocreate">
              <name>{interface}</name>
              <switchport>
                <trunk>
                  <allowed>
                    <vlan tags="merge">
                      <vlans>{../vlan-id}</vlans>
                    </vlan>
                  </allowed>
                </trunk>
              </switchport>
            </FastEthernet>
          <?end?>
          <?if {interface-type='GigabitEthernet'}?>
            <GigabitEthernet tags="nocreate">
              <name>{interface}</name>
              <switchport>
                <trunk>
                  <allowed>
                    <vlan tags="merge">
                      <vlans>{../vlan-id}</vlans>
                    </vlan>
                  </allowed>
                </trunk>
              </switchport>
            </GigabitEthernet>
          <?end?>
          <?if {interface-type='TenGigabitEthernet'}?>
            <TenGigabitEthernet tags="nocreate">
              <name>{interface}</name>
              <switchport>
                <trunk>
                  <allowed>
                    <vlan tags="merge">
                      <vlans>{../vlan-id}</vlans>
                    </vlan>
                  </allowed>
                </trunk>
              </switchport>
            </TenGigabitEthernet>
          <?end?>
        </interface>
      </config>
    </device>
  </devices>
</config-template>
```

Walking through the template can give a better idea of how it works. For every `/device-if/device-name` from the service model do the following:

1. Add the VLAN to the VLAN list, the tag merge tells the template to merge the data into an existing list (the default is to replace).
2. For every interface within that device, add the VLAN to the allowed VLANs and set the mode to `trunk`. The tag `nocreate` tells the template to not create the named interface if it does not exist

It is important to understand that every path in the template above refers to paths from the service model in `vlan.yang`.

Request NSO to reload the packages:

```cli
admin@ncs# packages reload
reload-result {
    package cisco-ios
    result true
}
reload-result {
    package vlan
    result true
}
```

Previously we started NCS with a `reload` package option, the above shows how to do the same without starting and stopping NSO.

We can now create services that will make things happen in the network. (Delete any dummy service from the previous step first). Create a VLAN service:

```cli
admin@ncs(config)# services vlan net-0 vlan-id 1234 device-if c0 \
                                 interface-type FastEthernet interface 1/0
admin@ncs(config-device-if-c0)# top
admin@ncs(config)# services vlan net-0 device-if c1 \
                                 interface-type FastEthernet interface 1/0
admin@ncs(config-device-if-c1)# top
admin@ncs(config)# show configuration
services vlan net-0
 vlan-id 1234
 device-if c0
  interface-type FastEthernet
  interface      1/0
 !
 device-if c1
  interface-type FastEthernet
  interface      1/0
 !
!
admin@ncs(config)# commit dry-run outformat native
native {
    device {
        name c0
        data interface FastEthernet1/0
              switchport trunk allowed vlan 1234
             exit
    }
    device {
        name c1
        data vlan 1234
             !
             interface FastEthernet1/0
              switchport trunk allowed vlan 1234
             exit
    }
}
admin@ncs(config)# commit
Commit complete.
```

When working with services in templates, there is a useful debug option for commit which will show the template and XPATH evaluation.

```cli
admin@ncs(config)# commit | debug
Possible completions:
 template   Display template debug info
 xpath      Display XPath debug info
admin@ncs(config)# commit | debug template
```

We can change the VLAN service:

```cli
admin@ncs(config)# services vlan net-0 vlan-id 1222
admin@ncs(config-vlan-net-0)# top
admin@ncs(config)# show configuration
services vlan net-0
 vlan-id 1222
!
admin@ncs(config)# commit dry-run outformat native
native {
    device {
        name c0
        data no vlan 1234
             vlan 1222
             !
             interface FastEthernet1/0
              switchport trunk allowed vlan 1222
             exit
    }
    device {
        name c1
        data no vlan 1234
             vlan 1222
             !
             interface FastEthernet1/0
              switchport trunk allowed vlan 1222
             exit
    }
}
```

It is important to understand what happens above. When the VLAN ID is changed, NSO can calculate the minimal required changes to the configuration. The same situation holds true for changing elements in the configuration or even parameters of those elements. In this way, NSO does not need explicit mapping to define a VLAN change or deletion. NSO does not overwrite a new configuration on the old configuration. Adding an interface to the same service works the same:

```cli
admin@ncs(config)# services vlan net-0 device-if c2 interface-type FastEthernet interface 1/0
admin@ncs(config-device-if-c2)# top
admin@ncs(config)# commit dry-run outformat native
native {
    device {
        name c2
        data vlan 1222
             !
             interface FastEthernet1/0
              switchport trunk allowed vlan 1222
             exit
    }
}
admin@ncs(config)# commit
Commit complete.
```

To clean up the configuration on the devices, run the delete command as shown below:

```cli
admin@ncs(config)# no services vlan net-0
admin@ncs(config)# commit dry-run outformat native
native {
    device {
        name c0
        data no vlan 1222
             interface FastEthernet1/0
              no switchport trunk allowed vlan 1222
             exit
    }
    device {
        name c1
        data no vlan 1222
             interface FastEthernet1/0
              no switchport trunk allowed vlan 1222
             exit
    }
    device {
        name c2
        data no vlan 1222
             interface FastEthernet1/0
              no switchport trunk allowed vlan 1222
             exit
    }
}
admin@ncs(config)# commit
Commit complete.
```

To make the VLAN service package complete edit the package-meta-data.xml to reflect the service model purpose. This example showed how to use template-based mapping. NSO also allows for programmatic mapping and also a combination of the two approaches. The latter is very flexible if some logic needs to be attached to the service provisioning that is expressed as templates and the logic applies device agnostic templates.

### Reactive FASTMAP and Nano Services <a href="#d5e1023" id="d5e1023"></a>

FASTMAP is the NSO algorithm that renders any service change from the single definition of the `create` service. As seen above, the template or code only has to define how the service shall be created, NSO is then capable of defining *any* change from that single definition.

A limitation in the scenarios described so far is that the mapping definition could immediately do its work as a single atomic transaction. This is sometimes not possible. Typical examples are external allocation of resources such as IP addresses from an IPAM, spinning up VMs, and sequencing in general.

Nano services using Reactive FASTMAP handle these scenarios with an executable plan that the system can follow to provision the service. The general idea is to implement the service as several smaller (nano) steps or stages, by using reactive FASTMAP and provide a framework to safely execute actions with side effects.

The [examples.ncs/getting-started/netsim-sshkey](https://github.com/NSO-developer/nso-examples/tree/6.7/getting-started/netsim-sshkey) example implements key generation to files and service deployment of the key to set up network elements and NSO for public key authentication to illustrate this concept. The example is described in more detail in [Develop and Deploy a Nano Service](/guides/administration/installation-and-deployment/development-to-production-deployment/develop-and-deploy-a-nano-service).

## Reconciling Existing Services <a href="#d5e1032" id="d5e1032"></a>

A very common situation when we wish to deploy NSO in an existing network is that the network already has existing services implemented in the network. These services may have been deployed manually or through another provisioning system. The task is to introduce NSO and import the existing services into NSO. The goal is to use NSO to manage existing services, and to add additional instances of the same service type, using NSO. This is a non-trivial problem since existing services may have been introduced in various ways. The mapping operation is not necessarily reversible and it is therefore impossible, in general, to extract service parameters from the service instance.

A better approach is to start with a list of existing service instances. Maybe such a list exists in an inventory system, an external database, or maybe just an Excel spreadsheet. If the service configuration has been done consistently (but it rarely is), it may also be the case that we can:

1. Import all managed devices into NSO.
2. Execute a full `sync-from` on the entire network.
3. Write a program, using Python/Maapi or Java/Maapi that traverses the entire network configuration and computes the services list.

With the pre-existing services list, we also need to define the service YANG model and implement the service mapping logic in such a way that it results in a configuration that is already there in the existing network. Due to inconsistencies in actual configurations and different ways of configuring the same service, this usually requires significant effort. But it is required before a full service reconciliation is possible.

[Service Discovery and Import](https://nso-docs.cisco.com/guides/operation-and-usage/operations/pages/Bw8TviXCSEsM9XjXBC1d#ch_svcref.discovery) describes the necessary steps and procedures.

## Brownfield Networks <a href="#d5e1052" id="d5e1052"></a>

In contrast with service reconciliation, where the end goal is to manage the network services through NSO by incorporating existing configurations, there are also situations where NSO service activation solution is deployed in parallel with other solutions and these solutions must coexist in the network. By default, NSO expects to manage the full device configuration and can thus conflict with the configuration rendered from the other solutions.

For such situations NSO supports the `commit no-overwrite` operation. This commit flag restricts the device configuration to not overwrite data that NSO did not create. Since NSO 6.4, it also includes additional functionality for verifying device values that are required to compute the changes in the transaction (the values from the so-called transaction read-set) have not changed. This means `commit no-overwrite` in newer versions of NSO provides guarantees about correctness in the face of device changes that were not made through NSO.

## Advanced Services Orchestration <a href="#d5e1057" id="d5e1057"></a>

Some services need to be set up in stages where each stage can consist of setting up some device configuration and then waiting for this configuration to take effect before performing the next stage. In this scenario, each stage must be performed in a separate transaction which is committed separately. Most often an external notification or other event must be detected and trigger the next stage in the service activation.

NSO supports the implementation of such staged services with the use of Reactive FASTMAP patterns in nano services.

From the user's perspective, it is not important how a certain service is implemented. The implementation should not have an impact on how the user creates or modifies a service. However, knowledge about this can be necessary to explain the behavior of a certain service.

In short the life-cycle of an RFM nano service in not only controlled by the direct create/set/delete operations. Instead, there are one or many implicit `reactive-re-deploy` requests on the service that are triggered by external event detection. If the user examines an RFM service, e.g. using `get-modification`, the device impact will grow over time after the initial create.

### Nano Service Plans <a href="#d5e1065" id="d5e1065"></a>

Nano services autonomously will do `reactive-re-deploy` until all stages of the service are completed. This implies that a nano service normally is not completed when the initial create is committed. For the operator to understand that a nano service has run to completion there must typically be some service-specific operational data that can indicate this.

Plans are introduced to standardize the operational data that can show the progress of the nano service. This gives the user a standardized view of all nano services and can directly answer the question of whether a service instance has run to completion or not.

A plan consists of one or many component entries. Each component consists of two or many state entries where the state can be in status `not-reached`, `reached`, or `failed`. A plan must have a component named `self` and can have other components with arbitrary names that have meaning for the implementing nano service. A plan component must have a first state named `init` and a last state named `ready`. In between `init` and `ready`, a plan component can have additional state entries with arbitrary naming.

The purpose of the `self` component is to describe the main progress of the nano service as a whole. Most importantly the `self` component last state named `ready` must have the status `reached` if and only if the nano service as a whole has been completed. Other arbitrary components as well as states are added to the plan if they have meaning for the specific nano service i.e. more specific progress reporting.

A `plan` also defines an empty leaf `failed` which is set if and only if any *state* in any *component* has a status set to `failed`. As such this is an aggregation to make it easy to verify if a RFM service is progressing without problems or not.

The following is an illustration of using the plan to report the progress of a nano service:

```cli
ncs# show vpn l3vpn volvo plan
NAME                    TYPE   STATE              STATUS       WHEN
------------------------------------------------------------------------------------
self                    self   init               reached      2016-04-08T09:22:40
                               ready              not-reached  -
endpoint-branch-office  l3vpn  init               reached      2016-04-08T09:22:40
                               qos-configured     reached      2016-04-08T09:22:40
                               ready              reached      2016-04-08T09:22:40
endpoint-head-office    l3vpn  init               reached      2016-04-08T09:22:40
                               pe-created         not-reached  -
                               ce-vpe-topo-added  not-reached  -
                               vpe-p0-topo-added  not-reached  -
                               qos-configured     not-reached  -
                               ready              not-reached  -
```

### Service Progress Monitoring <a href="#d5e1109" id="d5e1109"></a>

Plans were introduced to standardize the operational data that show the progress of reactive fastmap (RFM) nano services. This gives the user a standardized view of all nano services and can answer the question of whether a service instance has run to completion or not. To keep track of the progress of plans, Service Progress Monitoring (SPM) is introduced. The idea with SPM is that time limits are put on the progress of plan states. To do so, a policy and a trigger are needed.

A policy defines what plan components and states need to be in what status for the policy to be true. A policy also defines how long time it can be false without being considered jeopardized and how long time it can be false without being considered violated. Further, it may define an action, that is called in case of a policy being jeopardized, violated, or successful.

A trigger is used to associate a policy with a service and a component.

The following is an illustration of using an SPM to track the progress of an RFM service, in this case, the policy specifies that the self-components ready state must be reached for the policy to be true:

```cli
ncs# show vpn l3vpn volvo service-progress-monitoring
                                                               JEOPARDY                       VIOLATION           SUCCESS
NAME  POLICY         START TIME           JEOPARDY TIME        RESULT    VIOLATION TIME       RESULT     STATUS   TIME
---------------------------------------------------------------------------------------------------------------------------
self  service-ready  2016-04-08T09:22:40  2016-04-08T09:22:40  -         2016-04-08T09:22:40  -          running  -
```

## Bulk Service Actions

In some scenarios, you may need to execute an action on multiple service instances at once. A typical example is after a NED migration where the NED upgrade touches the same paths that services manage. In such cases, all affected service instances must be re-deployed to reconcile the service view with the updated device model.

NSO provides the following actions:

* `/ncs:services/re-deploy`: `re-deploy` multiple service instances in a single operation.
* `/ncs:services/un-deploy`: `un-deploy` multiple service instances in a single operation.
* `/ncs:services/check-sync`: `check-sync` multiple service instances in a single operation.

You can filter by service type (service callpoint) and/or specify individual service instances using your choice of service filters. There are 3 types of service filter:

* `service-type`: filter on all services of a given type
* `service-id`: filter on specfic instances
* `select-services`: filter on services selected by an XPath expression.

The following is an illustration of excuting an action on mutiple service instances:

### Filtering on All Services of a Given Type

```cli
ncs# services re-deploy service-type [ /l3vpn:vpn/l3vpn:l3vpn ] dry-run { outformat native }
result {
    service-id /vpn/l3vpn[name='ford']
    native {
        device {
            name ce7
            data interface GigabitEthernet 0/1.103
                   description Link to PE / pe0 - GigabitEthernet0/0/0/5
                  !

        }
        device {
            name pe0
            data policy-map ford-ce7
                   class class-default
                    shape average 4500000 bps
                   !
                  !
                  interface GigabitEthernet 0/0/0/5.103
                   description Link to CE / ce7 - GigabitEthernet0/1
                   encapsulation dot1q 103
                   service-policy output ford-ce7
                   vrf ford
                   ipv4 address 192.168.1.26 255.255.255.252
                  !
                  router bgp 100
                   vrf ford
                    neighbor 192.168.1.25
                     remote-as 65204
                     address-family ipv4 unicast
                      route-policy ford in
                      route-policy ford out
                      as-override
                     !
                    !
                   !
                  !

        }
        device {
            name pe1
            data no policy-map ford-ce7
                  no interface GigabitEthernet 0/0/0/5.103
                  router bgp 100
                   vrf ford
                    no neighbor 192.168.1.25
                   !
                  !

        }
    }
}
result {
    service-id /vpn/l3vpn[name='volvo']
    native {
        device {
            name ce8
            data interface GigabitEthernet 0/1.104
                   description Link to PE / pe0 - GigabitEthernet0/0/0/5
                  !

        }
        device {
            name pe0
            data policy-map volvo-ce8
                   class class-default
                    shape average 4500000 bps
                   !
                  !
                  interface GigabitEthernet 0/0/0/5.104
                   description Link to CE / ce8 - GigabitEthernet0/1
                   encapsulation dot1q 104
                   service-policy output volvo-ce8
                   vrf volvo
                   ipv4 address 192.168.1.30 255.255.255.252
                  !
                  router bgp 100
                   vrf volvo
                    neighbor 192.168.1.29
                     remote-as 65104
                     address-family ipv4 unicast
                      route-policy volvo in
                      route-policy volvo out
                      as-override
                     !
                    !
                   !
                  !

        }
        device {
            name pe1
            data no vrf volvo
                  no policy-map volvo-ce8
                  no interface GigabitEthernet 0/0/0/5.104
                  no route-policy volvo
                  router bgp 100
                   no vrf volvo
                  !

        }
    }
}
```

### Filtering on Specific Service Instances

```cli
ncs# services re-deploy service-id [ /vpn/l3vpn[name='volvo'] ] dry-run { outformat native }
result {
    service-id /vpn/l3vpn[name='volvo']
    native {
        device {
            name ce8
            data interface GigabitEthernet 0/1.104
                   description Link to PE / pe0 - GigabitEthernet0/0/0/5
                  !

        }
        device {
            name pe0
            data policy-map volvo-ce8
                   class class-default
                    shape average 4500000 bps
                   !
                  !
                  interface GigabitEthernet 0/0/0/5.104
                   description Link to CE / ce8 - GigabitEthernet0/1
                   encapsulation dot1q 104
                   service-policy output volvo-ce8
                   vrf volvo
                   ipv4 address 192.168.1.30 255.255.255.252
                  !
                  router bgp 100
                   vrf volvo
                    neighbor 192.168.1.29
                     remote-as 65104
                     address-family ipv4 unicast
                      route-policy volvo in
                      route-policy volvo out
                      as-override
                     !
                    !
                   !
                  !

        }
        device {
            name pe1
            data no vrf volvo
                  no policy-map volvo-ce8
                  no interface GigabitEthernet 0/0/0/5.104
                  no route-policy volvo
                  router bgp 100
                   no vrf volvo
                  !

        }
    }
}
```

### Filtering on Services Evaluated by XPath Expression

```cli
ncs# services re-deploy select-services /vpn/l3vpn[name='ford'] dry-run { outformat native }
result {
    service-id /vpn/l3vpn[name='ford']
    native {
        device {
            name ce7
            data interface GigabitEthernet 0/1.103
                   description Link to PE / pe0 - GigabitEthernet0/0/0/5
                  !

        }
        device {
            name pe0
            data policy-map ford-ce7
                   class class-default
                    shape average 4500000 bps
                   !
                  !
                  interface GigabitEthernet 0/0/0/5.103
                   description Link to CE / ce7 - GigabitEthernet0/1
                   encapsulation dot1q 103
                   service-policy output ford-ce7
                   vrf ford
                   ipv4 address 192.168.1.26 255.255.255.252
                  !
                  router bgp 100
                   vrf ford
                    neighbor 192.168.1.25
                     remote-as 65204
                     address-family ipv4 unicast
                      route-policy ford in
                      route-policy ford out
                      as-override
                     !
                    !
                   !
                  !

        }
        device {
            name pe1
            data no policy-map ford-ce7
                  no interface GigabitEthernet 0/0/0/5.103
                  router bgp 100
                   vrf ford
                    no neighbor 192.168.1.25
                   !
                  !

        }
    }
}
```


# Device Manager

Learn the concepts of NSO device management.

The NSO device manager is the center of NSO. The device manager maintains a flat list of all managed devices. Normally NSO keeps the primary copy of the configuration for each managed device in the CDB. Whenever a configuration change is done to the list of device configuration primary copies, the device manager will partition this network configuration change into the corresponding changes for the managed devices. The device manager passes on the required changes to the NEDs (Network Element Drivers). A NED needs to be installed for every type of device OS, like Cisco IOS NED, Cisco XR NED, Juniper JUNOS NED, etc. The NEDs communicate through the native device protocol southbound.

The NEDs fall into the following categories:

* **NETCONF-capable device**: The Device Manager will produce NETCONF `edit-config` RPC operations for each participating device.
* **SNMP device**: The Device Manager translates the changes made to the configuration into the corresponding SNMP SET PDUs.
* **Device with Cisco CLI**: The device has a CLI with the same structure as Cisco IOS or XR routers. The Device Manager and a CLI NED are used to produce the correct sequence of CLI commands which reflects the changes made to the configuration.
* **Other devices**: For devices that do not fit into any of the above-mentioned categories, a corresponding Generic NED is invoked. Generic NEDs are used for proprietary protocols like REST and for CLI flavors that do not resemble IOS or XR. The Device Manager will inform the Generic NED about the made changes and the NED will translate these to the appropriate operations toward the device.

NSO orchestrates an atomic transaction that has the very desirable characteristic of either the transaction as a whole ending up on all participating devices and in the NSO primary copy, or alternatively, the whole transaction getting aborted and resultingly, all changes getting automatically rolled back.

The architecture of the NETCONF protocol is the enabling technology making it possible to push out configuration changes to managed devices and then in the case of other errors, roll back changes. Devices that do not support NETCONF, i.e., devices that do not have transactional capabilities can also participate, however depending on the device, error recovery may not be as good as it is for a proper NETCONF-enabled device.

To understand the main idea behind the NSO device manager it is necessary to understand the NSO data model and how NSO incorporates the YANG data models from the different managed devices.

The NEDs will publish YANG data models even for non-NETCONF devices. In the case of SNMP the YANG models are generated from the MIBs. For JunOS devices the JunOS NED generates a YANG from the JunOS XML Schema. For Schema-less devices like CLI devices, the NED developer writes YANG models corresponding to the CLI structure. The result of this is the device manager and NSO CDB has YANG data models for all devices independent of the underlying protocol.

Throughout this section, we will use the [examples.ncs/service-management/mpls-vpn-java](https://github.com/NSO-developer/nso-examples/tree/6.7/service-management/mpls-vpn-java) example. The example network consists of Cisco ASR 9k and Juniper core routers (P and PE) and Cisco IOS-based CE routers.

<div data-with-frame="true"><figure><img src="/files/PddnV8PVgxk8jLdRXaVv" alt="" width="563"><figcaption><p>NSO Example Network</p></figcaption></figure></div>

## Managed Device Tree <a href="#user_guide.devicemanager.device-tree" id="user_guide.devicemanager.device-tree"></a>

The central part of the NSO YANG model, in the file `tailf-ncs-devices.yang`, has the following structure:

{% code title="tailf-ncs-devices.yang" %}

```yang
submodule tailf-ncs-devices {
  belongs-to tailf-ncs {
    prefix ncs;
  }
  ...
  container devices {
    ......
    list device {
      key name;

      description
        "This list contains all devices managed by NCS.";

      leaf name {
        type string;
        description
          "A string uniquely identifying the managed device.";
      }

      leaf address {
        type inet:host;
        mandatory true;
        description
          "IP address or host name for the management interface on
           the device.";
      }
      leaf port {
        type inet:port-number;
        description
          "Port for the management interface on the device.  If this leaf
           is not configured, NCS will use a default value based on the
           type of device.  For example, a NETCONF device uses port 830,
           a CLI device over SSH uses port 22, and a SNMP device uses
           port 161.";
      }
      ....
      leaf authgroup {
        ....
      }
      container device-type {
      .......
      container config {
         ...
      }
  }
}
```

{% endcode %}

Each managed device is uniquely identified by its name, which is a free-form text string. This is typically the DNS name of the managed device but could equally well be the string format of the IP address of the managed device or anything else. Furthermore, each managed device has a mandatory address/port pair that together with the `authgroup` leaf provides information to NSO on how to connect and authenticate over SSH/NETCONF to the device. Each device also has a mandatory parameter `device-type` that specifies which southbound protocol to use for communication with the device.

The following device types are available:

* NETCONF
* CLI: A corresponding CLI NED is needed to communicate with the device. This requires YANG models with the appropriate annotations for the device CLI.
* SNMP: The device speaks SNMP, preferably in read-write mode.
* Generic NED: A corresponding Generic NED is needed to communicate with the device. This requires YANG models and Java code.

The NSO CLI command below lists the NED types for the devices in the example network.

```cli
ncs(config)# show full-configuration devices device device-type
devices device ce0
 device-type cli ned-id cisco-ios-cli-3.8
!
...
devices device p0
 device-type cli ned-id cisco-iosxr-cli-3.5
!
devices device p1
 device-type cli ned-id cisco-iosxr-cli-3.5
!
...
devices device pe2
 device-type netconf ned-id juniper-junos-nc-3.0
!
```

The empty container `/ncs:devices/device/config` is used as a mount point for the YANG models from the different managed devices.

As previously mentioned, NSO needs the following information to manage a device:

* The IP/Port of the device and authentication information.
* Some or all of the YANG data models for the device.

In the example setup, the address and authentication information are provided in the NSO database (CDB) initialization file. There are many different ways to add new managed devices. All of the NSO northbound interfaces can be used to manipulate the set of managed devices. This will be further described later.

Once NSO has started you can inspect the meta information for the managed devices through the NSO CLI. This is an example session:

{% code title="Example: Show Device Configuration in NSO CLI" %}

```cli
ncs(config)# show full-configuration devices device
devices device ce0
 address   127.0.0.1
 port      10022
 ssh host-key ssh-dss
 ...
 authgroup default
 device-type cli ned-id cisco-ios-cli-3.8
 state admin-state unlocked
 config
 ...
 !
!
devices device ce1
 address   127.0.0.1
 port      10023
 ssh host-key ssh-dss
...
 !
 authgroup default
 device-type cli ned-id cisco-ios-cli-3.8
 state admin-state unlocked
 config
 ...
 !
!
```

{% endcode %}

Alternatively, this information could be retrieved from the NSO northbound NETCONF interface by running the simple Python-based netconf-console program towards the NSO NETCONF server.

{% code title="Example: Show Device Configuration in NETCONF" %}

```bash
$ netconf-console --get-config -x "/devices/device[name='ce0']"
<?xml version="1.0" encoding="UTF-8"?>
<rpc-reply xmlns="urn:ietf:params:xml:ns:netconf:base:1.0" message-id="1">
  <data>
    <devices xmlns="http://tail-f.com/ns/ncs">
      <device>
        <name>ce0</name>
        <address>127.0.0.1</address>
        <port>10022</port>
        <ssh>
          <host-key>
            <algorithm>ssh-dss</algorithm>

            ...

        <authgroup>default</authgroup>
        <device-type>
          <cli>
          <ned-id xmlns:cisco-ios-cli-3.8="http://tail-f.com/ns/ned-id/cisco-ios-cli-3.8">
            cisco-ios-cli-3.8:cisco-ios-cli-3.8
          </ned-id>
          </cli>
        </device-type>
        <state>
          <admin-state>unlocked</admin-state>
        </state>
        <config>

        ...

        </config>
      </device>
    </devices>
  </data>
</rpc-reply>
```

{% endcode %}

All devices in the above two examples (Show Device Configuration in NSO CLI and Show Device Configuration in NETCONF) have `/devices/device/state/admin-state` set to `unlocked`, this will be described later in this section.

## The NED Packages <a href="#user_guide.devicemanager.ned-packages" id="user_guide.devicemanager.ned-packages"></a>

To communicate with a managed device, a NED for that device type needs to be loaded by NSO. A NED contains the YANG model for the device and corresponding driver code to talk CLI, REST, SNMP, etc. NEDs are distributed as packages.

{% code title="Example: Installed Packages" %}

```cli
ncs# show packages
packages package cisco-ios-cli-3.8
 package-version 3.8.0.1
 description     "NED package for Cisco IOS"
 ncs-min-version [ 3.2.2 3.3 3.4 ]
 directory       ./state/packages-in-use/1/cisco-ios-cli-3.8
 component IOSDp2
  callback java-class-name [ com.tailf.packages.ned.ios.IOSDp2 ]
 component IOSDp
  callback java-class-name [ com.tailf.packages.ned.ios.IOSDp ]
 component cisco-ios
  ned cli ned-id  cisco-ios-cli-3.8
  ned cli java-class-name com.tailf.packages.ned.ios.IOSNedCli
  ned device vendor Cisco
 ...
 oper-status up
packages package cisco-iosxr-cli-3.5
 package-version 3.5.0.7
 description     "NED package for Cisco IOS XR"
 ncs-min-version [ 3.2.2 3.3 ]
 directory       ./state/packages-in-use/1/cisco-iosxr-cli-3.5
 component cisco-ios-xr
  ned cli ned-id  cisco-iosxr-cli-3.5
  ned cli java-class-name com.tailf.packages.ned.iosxr.IosxrNedCli
  ned device vendor Cisco
 ...
 oper-status up
packages package juniper-junos-nc-3.0
 package-version 3.0.14.2
 description     "NED package for all JunOS based Juniper routers"
 ncs-min-version [ 3.0.0.1 3.1 3.2 3.3 3.4 ]
 directory       ./state/packages-in-use/1/juniper-junos-nc-3.0
 component junos
  ned netconf ned-id juniper-junos-nc-3.0
  ned device vendor Juniper
 oper-status up
 ...
```

{% endcode %}

The CLI command in the above example (Installed Packages) shows all the loaded packages. NSO loads packages at startup and can reload packages at run-time. By default, the packages reside in the `packages` directory in the NSO run-time directory.

<pre><code>$ ls -l $NCS_DIR/examples.ncs/service-management/mpls-vpn-java
total 160
...
drwxr-xr-x   8 stefan  staff    272 Oct  1 16:57 packages
...
<strong>$ ls -l $NCS_DIR/examples.ncs/service-management/mpls-vpn-java/packages
</strong>total 24
cisco-ios
cisco-iosxr
juniper-junos
...
</code></pre>

## Starting the NSO Daemon <a href="#user_guide.devicemanager.starting-ncs" id="user_guide.devicemanager.starting-ncs"></a>

Once you have access to the network information for a managed device, its IP address and authentication information, as well as the data models of the device, you can actually manage the device from NSO.

You start the `ncs` daemon in a terminal like:

```cli
% ncs
```

Which is the same as, NSO loads it config from a `ncs.conf` file

```cli
% ncs -c ./ncs.conf
```

During development, it is sometimes convenient to run `ncs` in the foreground as:

```cli
% ncs -c ./ncs.conf --foregound --verbose
```

Once the daemon is running, you can issue the command:

```cli
% ncs --status
vsn: 7.1
SMP support: yes, using 8 threads
Using epoll: yes
available modules: backplane,netconf,cdb,cli,snmp,webui
...
... lots of output
```

To get more information about options to `ncs` do:

```cli
% ncs --help
```

The `ncs --status` command produces a lengthy list describing for example which YANG modules are loaded in the system. This is a valuable debug tool.

The same information is also available in the NSO CLI (and thus through all available northbound interfaces, including Maapi for Java programmers)

```cli
ncs# show ncs-state
ncs-state version 7.1
ncs-state smp number-of-threads 8
ncs-state epoll true
ncs-state daemon-status started
...
```

## Synchronizing Devices <a href="#user_guide.devicemanager.sync" id="user_guide.devicemanager.sync"></a>

When the NSO daemon is running and has been initialized with IP/Port and authentication information, as well as imported all modules, you can start to manage devices through NSO.

NSO provides the ability to synchronize the configuration to or from the device. If you know that the device has the correct configuration you can choose to synchronize from a managed device whereas if you know NSO has the correct device configuration and the device is incorrect, you can choose to synchronize from NSO to the device.

In the normal case, the configuration on the device and the copy of the configuration inside NSO should be identical.

In a cold start situation like in the [examples.ncs/service-management/mpls-vpn-java](https://github.com/NSO-developer/nso-examples/tree/6.7/service-management/mpls-vpn-java) example, where NSO is empty and there are network devices to talk to, it makes sense to synchronize from the devices. You can choose to synchronize from one device at a time or from all devices at once. Here is a CLI session to illustrate this.

{% code title="Example: Synchronize From Devices" %}

```cli
ncs(config)# devices sync-from
sync-result {
    device ce0
    result true
}
sync-result {
    device ce1
    result true
}
sync-result {
    device ce2
    result true
...
ncs(config)# show full-configuration devices device ce0
devices device ce0
...
 config
  no ios:service pad
  no ios:ip domain-lookup
  no ios:ip http secure-server
  ios:ip source-route
  ios:interface GigabitEthernet0/1
  exit
  ios:interface GigabitEthernet0/10
  exit
  ios:interface GigabitEthernet0/11
  exit
...
[ok][2010-04-13 16:29:15]
```

{% endcode %}

The command `devices sync-from`, in example (Synchronize from Devices), is an action that is defined in the NSO data model. It is important to understand the model-driven nature of NSO. All devices are modeled in YANG, network services like MPLS VPN are also modeled in YANG, and the same is true for NSO itself. Anything that can be performed over the NSO CLI or any north-bound is defined in the YANG files. The NSO YANG files are located here:

```
$ls $NCS_DIR/src/ncs/yang/
```

Packages can add other YANG files as well. For example the directory `packages/cisco-ios/src/yang/` contains the YANG definition of an IOS device.

The `tailf-ncs.yang` file defines the main NSO YANG data model; it includes parts of the model from many different submodule files.

The actions `sync-from` and `sync-to` are modeled in the file `tailf-ncs-devices.yang`. The sync action(s) are defined as:

{% code title="Example: tailf-ncs-devices.yang sync actions" %}

```
  grouping sync-from-output {
    list sync-result {
      key device;
      leaf device {
        type leafref {
          path "/devices/device/name";
        }
      }
      uses sync-result;
    }
  }

  grouping sync-result {
    description
      "Common result data from a 'sync' action.";

    choice outformat {
      leaf result {
        type boolean;
      }
      anyxml result-xml;
      leaf cli {
        tailf:cli-preformatted;
        type string;
      }
    }
    leaf info {
      type string;
      description
        "If present, contains additional information about the result.";
    }
  }

  ...

  container devices {

    ...

    tailf:action sync-from {
      description
        "Synchronize the configuration by pulling from all unlocked
         devices.";
      tailf:info "Synchronize the config by pulling from the devices";
      tailf:actionpoint ncsinternal {
        tailf:internal;
      }
      input {
        leaf suppress-positive-result {
          type empty;
          description
            "Use this additional parameter to only return
             devices that failed to sync.";
        }
        container dry-run {
          presence "";
          leaf outformat {
            type outformat2;
            description
              "Report what would be done towards CDB, without
               actually doing anything.";
          }
        }
      }
      output {
        uses sync-from-output;
      }
    }

    ...

    tailf:action sync-to {
      ...
    }

    ...

    list device {
      description
        "This list contains all devices managed by NCS.";

      key name;

      leaf name {
        description "A string uniquely identifying the managed device";
        type string;
      }

      ...

      tailf:action sync-from {
        description
          "Synchronize the configuration by pulling from the device.";
        tailf:info "Synchronize the config by pulling from the device";
        tailf:actionpoint ncsinternal {
          tailf:internal;
        }
        input {
          container dry-run {
            presence "";
            leaf outformat {
              type outformat2;
              description
                "Report what would be done towards CDB, without
                 actually doing anything.";
            }
          }
        }
        output {
          uses sync-result;
        }
      }
      tailf:action sync-to {

      ...
```

{% endcode %}

Synchronizing from NSO to the device is common when a device has been configured out-of-band. NSO has no means to enforce that devices are not directly reconfigured behind the scenes of NSO; however, once an out-of-band configuration has been performed, NSO can detect the fact. When this happens, it may (or may not, depending on the situation at hand) make sense to synchronize from NSO to the device, i.e. undo the rogue reconfigurations.

The command to do that is:

```cli
ncs# devices device ce0 sync-to
result true
```

A `dry-run` option is available for the action `sync-to`.

```cli
ncs# devices device ce0 sync-to dry-run
data {
    ...
}
```

This makes it possible to investigate the changes before they are transmitted to the devices.

### Partial `sync-from` <a href="#d5e2872" id="d5e2872"></a>

It is possible to synchronize a part of the configuration (a certain subtree) from the device using the `partial-sync-from` action located under /devices. While it is primarily intended to be used by service developers as described in [Partial Sync](https://nso-docs.cisco.com/guides/operation-and-usage/operations/pages/Bw8TviXCSEsM9XjXBC1d#ch_svcref.partialsync), it is also possible to use directly from the NSO CLI (or any other northbound interface). The example below (Example of Running partial-sync-from Action via CLI) illustrates using this action via CLI, using a router device from the [examples.ncs/device-management/router-network](https://github.com/NSO-developer/nso-examples/tree/6.7/device-management/router-network) example.

{% code title="Example: Example of Running partial-sync-from Action via CLI" %}

```bash
$ ncs_cli -C -u admin
ncs# devices partial-sync-from path [ \
/devices/device[name='ex0']/config/r:sys/interfaces/interface[name='eth0'] \
/devices/device[name='ex1']/config/r:sys/dns/server ]
sync-result {
    device ex0
    result true
}
sync-result {
    device ex1
    result true
}
ncs# show running-config devices device ex0..1 config
devices device ex0
 config
  r:sys interfaces interface eth0
   unit 0
    enabled
   !
   unit 1
    enabled
   !
   unit 2
    enabled
    description "My Vlan"
    vlan-id     18
   !
  !
 !
!
devices device ex1
 config
  r:sys dns server 10.2.3.4
  !
 !
!
```

{% endcode %}

## Configuring Devices

It is now possible to configure several devices through the NSO inside the same network transaction. To illustrate this, start the NSO CLI from a terminal application.

{% code title="Example: Configure Devices" %}

```bash
$ ncs_cli -C -u admin
ncs# config
Entering configuration mode terminal
ncs(config)# devices device pe1 config cisco-ios-xr:snmp-server \
     community public RO
ncs(config-config)# top
ncs(config)# devices device ce0 config ios:snmp-server community public RO
ncs(config-config)# devices device pe2 config junos:configuration \
      snmp community public view RO
ncs(config-community-public)# top
ncs(config)# show configuration
devices device ce0
 config
  ios:snmp-server community public RO
 !
!
devices device pe1
 config
  cisco-ios-xr:snmp-server community public RO
 !
!
devices device pe2
 config
  ! first
  junos:configuration snmp community public
   view RO
  !
 !
!
ncs(config)# commit dry-run outformat native
native {
    device {
        name ce0
        data snmp-server community public RO
    }
    device {
        name pe1
        data snmp-server community public RO
    }
    device {
        name pe2
        data <rpc xmlns="urn:ietf:params:xml:ns:netconf:base:1.0"
                  message-id="1">
               <edit-config xmlns:nc="urn:ietf:params:xml:ns:netconf:base:1.0">
                 <target>
                   <candidate/>
                 </target>
                 <test-option>test-then-set</test-option>
                 <error-option>rollback-on-error</error-option>
                 <config>
                   <configuration xmlns="http://xml.juniper.net/xnm/1.1/xnm">
                     <snmp>
                       <community>
                         <name>public</name>
                         <view>RO</view>
                       </community>
                     </snmp>
                   </configuration>
                 </config>
               </edit-config>
             </rpc>
    }
}
ncs(config)# commit
```

{% endcode %}

The example above (Configure Devices) illustrates a multi-host transaction. In the same transaction, three hosts were re-configured. Had one of them failed, or been non-operational, the transaction as a whole would have failed.

As seen from the output of the command `commit dry-run outformat native`, NSO generates the native CLI and NETCONF commands which will be sent to each device when the transaction is committed.

Since the `/devices/device/config` path contains different models depending on the augmented device model NSO uses the data model prefix in the CLI names; `ios`, `cisco-ios-xr` and `junos`. Different data models might use the same name for elements and the prefix avoids name clashes.

NSO uses different underlying techniques to implement the atomic transactional behavior in case of any error. NETCONF devices are straightforward using confirmed commit. For CLI devices like IOS NSO calculates the reverse diff to restore the configuration to the state before the transaction was applied.

## Connection Management <a href="#user_guide.devicemanager.connection-mgmt" id="user_guide.devicemanager.connection-mgmt"></a>

Each managed device needs to be configured with the IP address and the port where the CLI, NETCONF server, etc. of the managed device listens for incoming requests.

Connections are established on demand as they are needed. It is possible to explicitly establish connections, but that functionality is mostly there for troubleshooting connection establishment. We can, for example, do:

```cli
ncs# devices connect
connect-result {
    device ce0
    result true
    info (admin) Connected to ce0 - 127.0.0.1:10022
}
connect-result {
    device ce1
    result true
    info (admin) Connected to ce1 - 127.0.0.1:10023
}
...
```

We were able to connect to all managed devices. It is also possible to explicitly attempt to test connections to individual managed devices:

```cli
ncs# devices device ce0 connect
result true
info (admin) Connected to ce0 - 127.0.0.1:10022
```

Established connections are typically not closed right away when not needed, but rather pooled according to the rules described in [Device Session Pooling](#user_guide.devicemanager.pooling). This applies to NETCONF sessions as well as sessions established by CLI or generic NEDs via a connection-oriented protocol. In addition to session pooling, underlying SSH connections for NETCONF devices are also reused. Note that a single NETCONF session occupies one SSH channel inside an SSH connection, so multiple NETCONF sessions can co-exist in a single connection. When an SSH connection has been idle (no SSH channels open) for 2 minutes, the SSH connection is closed. If a new connection is needed later, a connection is established on demand.

Three configuration parameters can be used to control the connection establishment: `connect-timeout`, `read-timeout`, and `write-timeout`. In the NSO data model file `tailf-ncs-devices.yang`, these timeouts are modeled as:

```yang
submodule tailf-ncs-devices {
  ...
  container devices {
    ...
    grouping timeouts {
      description
        "Timeouts used when communicating with a managed device.";

      leaf connect-timeout {
        type uint32;
        units "seconds";
        description
          "The timeout in seconds for new connections to managed
           devices.";
      }
      leaf read-timeout {
        type uint32;
        units "seconds";
        description
          "The timeout in seconds used when reading data from a
           managed device.";
      }
      leaf write-timeout {
        type uint32;
        units "seconds";
        description
          "The timeout in seconds used when writing data to a
           managed device.";
      }
    }
    ...
    container global-settings {
      ...
      uses timeouts {
        description
          "These timeouts can be overridden per device.";

        refine connect-timeout {
          default 20;
        }
        refine read-timeout {
          default 20;
        }
        refine write-timeout {
          default 20;
        }
      }
      ....
```

Thus, to change these parameters (globally for all managed devices) you do:

```cli
ncs(config)# devices global-settings connect-timeout 30
ncs(config)# devices global-settings read-timeout 30
ncs(config)# commit
```

Or, to use a profile:

```cli
ncs(config)# devices profiles profile slow-devices connect-timeout 60
ncs(config-profile-slow-devices)# read-timeout 60
ncs(config-profile-slow-devices)# write-timeout 60
ncs(config-profile-slow-devices)# commit

ncs(config)# devices device ce3 device-profile slow-devices
ncs(config-device-ce3)# commit
```

## Authentication Groups <a href="#user_guide.devicemanager.authgroups" id="user_guide.devicemanager.authgroups"></a>

When NSO connects to a managed device, it requires authentication information for that device. The `authgroups` are modeled in the NSO data model:

{% code title="Example: tailf-ncs-devices.yang - Authgroups" %}

```yang
submodule tailf-ncs-devices {
  ...
  container devices {
    ...

    container authgroups {
      description
        "Named authgroups are used to decide how to map a local NCS user to
         remote authentication credentials on a managed device.

         The list 'group' is used for NETCONF and CLI managed devices.

         The list 'snmp-group' is used for SNMP managed devices.";

      list group {
        key name;

        description
          "When NCS connects to a managed device, it locates the
           authgroup configured for that device.  Then NCS looks up
           the local NCS user name in the 'umap' list.  If an entry is
           found, the credentials configured is used when
           authenticating to the managed device.

           If no entry is found in the 'umap' list, the credentials
           configured in 'default-map' are used.

           If no 'default-map' has been configured, and the local NCS
           user name is not found in the 'umap' list, the connection
           to the managed device fails.";

        grouping remote-user-remote-auth {
          description
            "Remote authentication credentials.";

          choice login-credentials {
            mandatory true;
            case stored {
              choice remote-user {
                mandatory true;
                leaf same-user {
                  type empty;
                  description
                    "If this leaf exists, the name of the local NCS user is used
                     as the remote user name.";
                }
                leaf remote-name {
                  type string;
                  description
                    "Remote user name.";
                }
              }

              choice remote-auth {
                mandatory true;
                leaf same-pass {
                  type empty;
                  description
                    "If this leaf exists, the password used by the local user
                     when logging in to NCS is used as the remote password.";
                }
                leaf remote-password {
                  type tailf:aes-256-cfb-128-encrypted-string;
                  description
                    "Remote password.";
                }
                case public-key {
                  uses public-key-auth;
                }
              }
              leaf remote-secondary-password {
                type tailf:aes-256-cfb-128-encrypted-string;
                description
                  "Some CLI based devices require a second
                   additional password to enter config mode";
              }
            }
            case callback {
              leaf callback-node {
                description
                  "Invoke a standalone action to retrieve login credentials for
                  managed devices on the 'callback-node' instance.

                  The 'action-name' action is invoked on the callback node that
                  is specified by an instance identifer.";
                mandatory true;
                type instance-identifier;
              }
              leaf action-name {
                description
                  "The action to call when a notification is received.

                  The action must use 'authgroup-callback-input-params'
                  grouping for input and 'authgroup-callback-output-params'
                  grouping for output from tailf-ncs-devices.yang.";
                type yang:yang-identifier;
                mandatory true;
                tailf:validate ncs {
                   tailf:internal;
                   tailf:dependency "../callback-node";
                }
              }
            }
          }
        }

        grouping mfa-grouping {
          container mfa {
            presence "MFA";
            description
              "Settings for handling multi-factor authentication towards
               the device";
            leaf executable {
              description "Path to the external executable handling MFA";
              type string;
              mandatory true;
            }
            leaf opaque {
              description
                "Opaque data for the external MFA executable.
                 This string will be base64 encoded and passed to the MFA
                 executable along with other parameters";
              type string;
            }
          }
        }

        leaf name {
          type string;
          description
            "The name of the authgroup.";
        }

        container default-map {
          presence "Map unknown users";
          description
            "If an authgroup has a default-map, it is used if a local
             NCS user is not found in the umap list.";
          tailf:info "Remote authentication parameters for users not in umap";
          uses remote-user-remote-auth;
          uses mfa-grouping;
        }

        list umap {
          key local-user;
          description
            "The umap is a list with the local NCS user name as key.
             It maps the local NCS user name to remote authentication
             credentials.";
          tailf:info "Map NCS users to remote authentication parameters";
          leaf local-user {
            type string;
            description
              "The local NCS user name.";
          }
          uses remote-user-remote-auth;
          uses mfa-grouping;
        }
      }
```

{% endcode %}

Each managed device must refer to a named authgroup. The purpose of an authentication group is to map local users to remote users together with the relevant SSH authentication information.

Southbound authentication can be done in two ways. One is to configure the stored user and credential components as shown in the example below (Configured authgroup) and the next example (authgroup default-map). The other way is to configure a callback to retrieve user and credentials on demand as shown in the example below (authgroup-callback).

{% code title="Example: Configured authgroup" %}

```cli
ncs(config)# show full-configuration devices authgroups
devices authgroups group default
 umap admin
  remote-name     admin
  remote-password $4$wIo7Yd068FRwhYYI0d4IDw==
 !
 umap oper
  remote-name     oper
  remote-password $4$zp4zerM68FRwhYYI0d4IDw==
 !
!
devices authgroups snmp-group default
 umap admin
  usm remote-name admin
  usm security-level auth-priv
  usm auth sha remote-password $4$wIo7Yd068FRwhYYI0d4IDw==
  usm priv aes remote-password $4$wIo7Yd068FRwhYYI0d4IDw==
 !
!
```

{% endcode %}

In the example above (Configured authgroup) in the auth group named `default`, the two local users `oper` and `admin` shall use the remote users' name `oper` and `admin` respectively with identical passwords.

Inside an authgroup, all local users need to be enumerated. Each local user name must have credentials configured which should be used for the remote host. In centralized AAA environments, this is usually a bad strategy. You may also choose to instantiate a `default-map`. If you do that it probably only makes sense to specify the same user name/password pair should be used remotely as the pair that was used to log into NSO.

{% code title="Example: authgroup default-map" %}

```cli
ncs(config)# devices authgroups group default default-map same-user same-pass
ncs(config-group-default)# commit
Commit complete.
ncs(config-group-default)# top
ncs(config)# show full-configuration devices authgroups
devices authgroups group default
 default-map same-user
 default-map same-pass
 umap admin
  remote-name     admin
  remote-password $4$wIo7Yd068FRwhYYI0d4IDw==
 !
 umap oper
  remote-name     oper
  remote-password $4$zp4zerM68FRwhYYI0d4IDw==
 !
!
devices authgroups snmp-group default
 umap admin
  usm remote-name admin
  usm security-level auth-priv
  usm auth sha remote-password $4$wIo7Yd068FRwhYYI0d4IDw==
  usm priv aes remote-password $4$wIo7Yd068FRwhYYI0d4IDw==
 !
!
```

{% endcode %}

In the example (Configured authgroup), only two users `admin` and `oper` were configured. If the `default-map` in example (authgroup default-map) is configured, all local users not found in the `umap` list will end up in the `default-map`. For example, if the user `rocky` logs in to NSO with the password `secret`. Since NSO has a built-in SSH server and also a built-in HTTPS server, NSO will be able to pick up the clear text passwords and can then reuse the same password when NSO attempts to establish southbound SSH connections. The user `rocky` will end up in the `default-map` and when NSO attempts to propagate `rocky`'s changes towards the managed devices, NSO will use the remote user name `rocky` with whatever password `rocky` used to log into NSO.

Authenticating southbound using stored configuration has two main components to define remote user and remote credentials. This is defined by the authgroup. As for the southbound user, there exist two options, the same user logged in to NSO or another user, as specified in the authgroup. As for the credentials, there are three options.

1. Regular password.
2. Public key. This means that a private key, either from a file in the user's SSH key directory, or one that is configured in the /ssh/private-key list in the NSO configuration, is used for authentication. Refer to [Publickey Authentication](/guides/operation-and-usage/operations/ssh-key-management#d5e4113) for the details on how the private key is selected.
3. Finally, an interesting option is to use the 'same-pass' option. Since NSO runs its own SSH server and its own SSL server, NSO can pick up the password of a user in clear text. Hence, if the 'same-pass' option is chosen for an authgroup, NSO will reuse the same password when attempting to connect southbound to a managed device.

### Connecting Using SSH Keyboard-Interactive (Multi-Factor) Authentication

NSO can connect to a device that is using multi-factor authentication. For this, the `authgroup` must be configured with an executable for handling the keyboard-interactive part, and optionally some opaque data that is passed to the executable. i.e., the `/devices/authgroups/group/umap/mfa/executable` and `/devices/authgroups/group/umap/mfa/opaque` (or under `default-map` for users that are not in `umap`) must be configured.

The prompts from the SSH server (including the password prompt and any additional challenge prompts) are passed to the `stdin` of the executable along with some other relevant data. The executable must write a single line to its `stdout` as the reply to the prompt. This is the reply that NSO sends to the SSH server.

{% code title="Example: Configuring Authgroup For Keyboard-interactive Authentication" %}

```
admin@ncs(config)# devices authgroups group mfa umap admin
admin@ncs(config-umap-admin)# remote-name admin remote-password
(<AES encrypted string>): *********
admin@ncs(config-umap-admin)# mfa executable ./handle_mfa.py opaque foobar
admin@ncs(config-umap-admin)# commit
Commit complete.
```

{% endcode %}

For example, with the above configured for the authgroup, if the user `admin` is trying to log in to the device `dev0` with password `admin`, this is the line that is sent to the `stdin` of the `handle_mfa.py` script:

```
[ZGV2MA==;YWRtaW4=;YWRtaW4=;Zm9vYmFy;;;YWRtaW5AbG9jYWxob3N0J3MgcGFzc3dvcmQ6IA==;]
```

The input to the script is the device, username, password, opaque data, as well as the name, instruction, and prompt from the SSH server. All these fields are base64 encoded, and separated by a semi-colon (`;`). So, the above line in effect encodes the following:

```
[dev0;admin;admin;foobar;;;admin@localhost's password:;]
```

A small Python program can be used to implement the keyboard-interactive authentication towards a device, such as:

```python
#!/usr/bin/env python3
import base64
line = input()
(device, user, passwd, opaque, name, instr, prompt, _) = map(
        lambda x: base64.b64decode(x).decode('utf-8'),
        line.strip('[]').split(';'))
if prompt == "admin@localhost's password: ":
    print(passwd)
elif prompt == "Enter SMS passcode:":
    print("secretSMScode")
else:
    print("2")
```

This script will then be invoked with the above fields for every prompt from the server, and the corresponding output from the script will be sent as the reply to the server.

### Using a Callback to Provide Device Credentials

In the case of authenticating southbound using a callback, remote user and remote credentials are obtained by an action invocation. The action is defined by the `callback-node` and `action-name` as in the example below (authgroup-callback) and supported credentials are remote password and optionally a secondary password for the provided local user, authgroup, and device.

With remote passwords, you may encounter issues if you use special characters, such as quotes (`"`) and backslash (`\`) in your password. See [Configure Mode](/guides/operation-and-usage/cli/introduction-to-nso-cli#d5e2199) for recommendations on how to avoid running into password issues.

{% code title="Example: authgroup-callback" %}

```cli
ncs(config)# devices authgroups group default umap oper
ncs(config-umap-oper)# callback-node /callback action-name auth-cb
ncs(config-group-oper)# commit
Commit complete.
ncs(config-group-oper)# top
ncs(config)# show full-configuration devices authgroups
devices authgroups group default
 default-map same-user
 default-map same-pass
 umap admin
  remote-name     admin
  remote-password $4$wIo7Yd068FRwhYYI0d4IDw==
 !
 umap oper
  callback-node /callback
  action-name   auth-cb
 !
!
devices authgroups snmp-group default
 umap admin
  usm remote-name admin
  usm security-level auth-priv
  usm auth sha remote-password $4$wIo7Yd068FRwhYYI0d4IDw==
  usm priv aes remote-password $4$wIo7Yd068FRwhYYI0d4IDw==
 !
!
```

{% endcode %}

{% code title="Example: authgroup-callback.yang" %}

```yang
module authgroup-callback {
  namespace "http://com/example/authgroup-callback";
  prefix authgroup-callback;

  import tailf-common {
    prefix tailf;
  }

  import tailf-ncs {
    prefix ncs;
  }

  container callback {
    description
      "Example callback that defines an action to retrieve
       remote authentication credentials";
    tailf:action auth-cb {
      tailf:actionpoint auth-cb-point;
      input {
        uses ncs:authgroup-callback-input-params;
      }
      output {
        uses ncs:authgroup-callback-output-params;
      }
    }
  }
}
```

{% endcode %}

In the example above (`authgroup-callback`), the configuration for the `umap` entry of the `oper` user is changed to use a callback to retrieve southbound authentication credentials. Thus, NSO is going to invoke the action `auth-cb` defined in the callback-node `callback`. The callback node is of type `instance-identifier` and refers to the container called `callback` defined in the example, (`authgroup-callback.yang`), which includes an action defined by action-name `auth-cb` and uses groupings `authgroup-callback-input-params` and `authgroup-callback-output-params` for input and output parameters respectively. In the example, (authgroup-callback), `authgroup-callback` module was loaded in NSO within an example package. Package development and action callbacks are not described here but more can be read in [Package Development](/guides/development/advanced-development/developing-packages), the section called [DP API](https://nso-docs.cisco.com/guides/operation-and-usage/operations/pages/Uzy6qvKpLQF47FSwk0S2#ug.java_api_overview.dp) and [Python API Overview](/guides/development/core-concepts/api-overview/python-api-overview).

### Caveats <a href="#d5e3032" id="d5e3032"></a>

Authentication groups and the functionality they bring come with some limitations on where and how it is used.

* The callback option that enables `authgroup-callback` feature is not applicable for members of `snmp-group` list.
* Generic devices that implement their own authentication scheme do not use any mapping or callback functionality provided by Authgroups.
* Cluster nodes use their own authgroups and mapping model, thus functionality differs, e.g. callback option is not applicable.

## Device Session Pooling <a href="#user_guide.devicemanager.pooling" id="user_guide.devicemanager.pooling"></a>

Opening a session towards a managed device is potentially time and resource-consuming. Also, the probability that a recently accessed device is still subject to further requests is reasonably high. These are motives for having a managed devices session pool in NSO.

The NSO device session pool is by default active and normally needs no maintenance. However, under certain circumstances, it might be of interest to modify its behavior. Examples can be when some device type has characteristics that make session pooling undesired, or when connections to a specific device are very costly, and therefore the time that open sessions can stay in the pool should increase.

{% hint style="info" %}
Changes from the default configuration of the NSO session pool should only be performed when absolutely necessary and when all effects of the change are understood.
{% endhint %}

NSO presents operational data that represent the current state of the session pool. To visualize this, we use the CLI to connect to NSO and force connection to all known devices:

```bash
$ ncs_cli -C -u admin

admin connected from 127.0.0.1 using console on ncs
ncs# devices connect suppress-positive-result
```

We can now list all open sessions in the `session-pool`. But note that this is a live pool. Sessions will only remain open for a certain amount of time, the idle time.

```cli
ncs# show devices session-pool
        DEVICE            MAX        IDLE
DEVICE  TYPE    SESSIONS  SESSIONS   TIME
-------------------------------------------
ce0     cli     1         unlimited  30
ce1     cli     1         unlimited  30
ce2     cli     1         unlimited  30
ce3     cli     1         unlimited  30
ce4     cli     1         unlimited  30
ce5     cli     1         unlimited  30
pe0     cli     1         unlimited  30
pe1     cli     1         unlimited  30
pe2     cli     1         unlimited  30
```

In addition to the idle time for sessions, we can also see the type of device, current number of pooled sessions, and maximum number of pooled sessions.

We can close pooled sessions for specific devices.

```cli
ncs# devices session-pool pooled-device pe0 close
ncs# devices session-pool pooled-device pe1 close
ncs# devices session-pool pooled-device pe2 close
ncs# show devices session-pool
        DEVICE            MAX        IDLE
DEVICE  TYPE    SESSIONS  SESSIONS   TIME
-------------------------------------------
ce0     cli     1         unlimited  30
ce1     cli     1         unlimited  30
ce2     cli     1         unlimited  30
ce3     cli     1         unlimited  30
ce4     cli     1         unlimited  30
ce5     cli     1         unlimited  30
```

And we can close all pooled sessions in the session pool.

```cli
ncs# devices session-pool close
ncs# show devices session-pool
% No entries found.
```

The session pool configuration is found in the `tailf-ncs-devices.yang` submodel. The following part of the YANG device-profile-parameters grouping controls how the session pool is configured:

```
grouping device-profile-parameters {

  ...

    container session-pool {
      tailf:info "Control how sessions to related devices can be pooled.";
      description
        "NCS uses NED sessions when performing transactions, actions
         etc towards a device. When such a task is completed the NED
         session can either be closed or pooled.

         Pooling a NED session means that the session to the
         device is kept open for a configurable amount of
         time. During this time the session can be re-used for a new
         task. Thus the pooling concept exists to reduce the number
         of new connections needed towards a device that is often
         used.

         By default NCS uses pooling for all device types except
         SNMP. Normally there is no need to change the default
         values.";

      leaf max-sessions {
        type union {
          type enumeration {
            enum unlimited;
          }
          type uint32;
        }
        description
          "Controls the maximum number of open sessions in the pool for
           a specific device. When this threshold is exceeded the oldest
           session in the pool will be closed.
           A Zero value will imply that pooling is disabled for
           this specific device. The label 'unlimited' implies that no
           upper limit exists for this specific device";
      }

      leaf idle-time {
        tailf:info
          "The maximum time that a session is kept open in the pool";
        type uint32 {
          range "1 .. max";
        }
        units "seconds";
        description
          "The maximum time that a session is kept open in the pool.
           If the session is not requested and used before the
           idle-time has expired, the session is closed.
           If no idle-time is set the default is 30 seconds.";
      }
    }
  }
}
```

This grouping can be found in the NSO model under `/ncs:devices/global-settings/session-pool`, `/ncs:devices/profiles/profile/session-pool` and `/ncs:devices/device/session-pool` to be able to control session pooling for all devices, a group of devices, and a specific device respectively.

In addition under `/ncs:devices/global-settings/session-pool/default` it is possible to control the global max size of the session pool, as defined by the following yang snippet:

```yang
container global-settings {
  tailf:info "Global settings for all managed devices.";
  description
    "Global settings for all managed devices. Some of these
     settings can be overridden per managed device.";

  uses device-profile-parameters {

    ...

    augment session-pool {
      leaf pool-max-sessions {
        type union {
          type enumeration {
            enum unlimited;
          }
          type uint32;
        }
        description
          "Controls the grand total session count in the pool.
           Independently on how different devices are pooled the grand
           total session count can never exceed this value.
           A Zero value will imply that pooling is disabled for all devices.
           The label 'unlimited' implies that no upper limit exists for
           the number open sessions in the pool";
      }
    }
  }
}
```

Let's illustrate the possibilities with an example configuration of the session pool:

```cli
ncs# configure
ncs(config)# devices global-settings session-pool idle-time 100
ncs(config)# devices profiles profile small session-pool max-sessions 3
ncs(config-profile-small)# top
ncs(config)# devices device ce* device-profile small
ncs(config-device-ce*)# top
ncs(config)# devices device pe0 session-pool max-sessions 0
ncs(config-device-pe0)# top
ncs(config)# commit
Commit complete.
ncs(config)# exit
```

In the above configuration, the default idle time is set to 100 seconds for all devices. A device profile called `small` is defined which contains a max-session value of 3 sessions, this profile is set on all `ce*` devices. The devices `pe0` has a max-sessions 0 which implies that this device cannot be pooled. Let's connect all devices and see what happens in the session pool:

```cli
ncs# devices connect suppress-positive-result
ncs# show devices session-pool
        DEVICE            MAX        IDLE
DEVICE  TYPE    SESSIONS  SESSIONS   TIME
-------------------------------------------
ce0     cli     1         3          100
ce1     cli     1         3          100
ce2     cli     1         3          100
ce3     cli     1         3          100
ce4     cli     1         3          100
ce5     cli     1         3          100
pe1     cli     1         unlimited  100
pe2     cli     1         unlimited  100
```

Now, we set an upper limit to the maximum number of sessions in the pool. Setting the value to 4 is too small for a real situation but serves the purpose of illustration:

```cli
ncs# configure
ncs(config)# devices global-settings session-pool pool-max-sessions 4
ncs(config)# commit
Commit complete.
ncs(config)# exit
```

The number of open sessions in the pool will be adjusted accordingly:

```cli
ncs# show devices session-pool
        DEVICE            MAX        IDLE
DEVICE  TYPE    SESSIONS  SESSIONS   TIME
-------------------------------------------
ce4     cli     1         3          100
ce5     cli     1         3          100
pe1     cli     1         unlimited  100
pe2     cli     1         unlimited  100
```

## Device Session Limits <a href="#user_guide.devicemanager.session_limits" id="user_guide.devicemanager.session_limits"></a>

Some devices only allow a small number of concurrent sessions, in the extreme case it only allows one (for example through a terminal server). For this reason, NSO can limit the number of concurrent sessions to a device and make operations wait if the maximum number of sessions has been reached.

In other situations, we need to limit the number of concurrent connect attempts made by NSO. For example, the devices managed by NSO talk to the same server for authentication which can only handle a limited number of connections at a time.

The configuration for session limits is found in the `tailf-ncs-devices.yang` submodel. The following part of the YANG device-profile-parameters grouping controls how the session limits are configured:

```
grouping device-profile-parameters {

  ...

    container session-limits {
      tailf:info "Parameters for limiting concurrent access to the device.";
      leaf max-sessions {
        type union {
          type enumeration {
            enum unlimited;
          }
          type uint32 {
            range "1..max";
          }
        }
        default unlimited;
        description
          "Puts a limit to the total number of concurrent sessions
           allowed for the device. The label 'unlimited' implies that no
           upper limit exists for this device.";
      }
    }

  ...

  }
```

This grouping can be found in the NSO model under `/ncs:devices/global-settings/session-limits`, `/ncs:devices/profiles/profile/session-limits` and `/ncs:devices/device/session-limits` to be able to control session limits for all devices, a group of devices, and a specific device respectively.

In addition, under `/ncs:devices/global-settings/session-limits`, it is possible to control the number of concurrent connect attempts allowed and the maximum time to wait for a device to be available to connect.

```yang
container global-settings {
  tailf:info "Global settings for all managed devices.";
  description
    "Global settings for all managed devices. Some of these
     settings can be overridden per managed device.";

  uses device-profile-parameters {

  ...

    augment session-limits {
      description
        "Parameters for limiting concurrent access to devices.";
      container connect-rate {
        leaf burst {
          type union {
            type enumeration {
              enum unlimited;
            }
            type uint32 {
              range "1..max";
            }
          }
          default unlimited;
          description
            "The number of concurrent connect attempts allowed.
             For example, the devices managed by NSO talk to the same
             server for authentication which can only handle a limited
             number of connections at a time. Then we can limit
             the concurrency of connect attempts with this setting.";
        }
      }
      leaf max-wait-time {
        tailf:info
          "Max time in seconds to wait for device to be available.";
        type union {
          type enumeration {
            enum unlimited;
          }
          type uint32 {
            range "0..max";
          }
        }
        units "seconds";
        default 10;
        description
          "Max time in seconds to wait for a device being available
           to connect. When the maximum time is reached an error
           is returned. Setting this to 0 means that the error is
           returned immediately.";
      }
    }

  ...

}
```

## Tracing Device Communication <a href="#user_guide.devicemanager.tracing" id="user_guide.devicemanager.tracing"></a>

It is possible to turn on and off NED traffic tracing. This is often a good way to troubleshoot problems. To understand the trace output, a basic prerequisite is a good understanding of the native device interface. For NETCONF devices, an understanding of NETCONF RPC is a prerequisite. Similarly for CLI NEDs, a good understanding of the CLI capabilities of the managed devices is required.

To turn on southbound traffic tracing, we need to enable the feature and we must also configure a directory where we want the trace output to be written. It is possible to have the trace output in two different formats, `pretty` and `raw`. The format of the data depends on the type of the managed device. For NETCONF devices, the `pretty` mode indents all the XML data for enhanced readability and the `raw` mode does not. Sometimes when the XML is broken, `raw` mode is required to see all the data received. Tracing in `raw` mode will also signal to the corresponding NED to log more verbose tracing information.

To enable tracing, do:

```cli
ncs(config)# devices global-settings trace raw trace-dir .logs
ncs(config)# commit
```

The trace setting only affects new NED connections, so to ensure that we get any tracing data, we can do:

```cli
ncs(config)# devices disconnect
```

The above command terminates all existing connections.

At this point, if you execute a transaction towards one or several devices and then view the trace data.

```cli
ncs(config)# do file show logs/ned-cisco-ios-ce0.trace
>> 8-Oct-2014::18:23:18.512 CLI CONNECT to ce0-127.0.0.1:10022 as admin (Trace=true)

  *** output 8-Oct-2014::18:23:18.514 ***
-- SSH connecting to host: 127.0.0.1:10022 --
-- SSH initializing session --

  *** input 8-Oct-2014::18:23:18.547 ***

admin connected from 127.0.0.1 using ssh on ncs
...
ce0(config)#
  *** output 8-Oct-2014::18:23:19.428 ***
snmp-server community topsecret RW
```

It is possible to clear all existing trace files through the command

```cli
ncs(config)# devices clear-trace
```

Finally, it is worth mentioning the trace functionality does not come for free. It is fairly costly to have the trace turned on. Also, there exists no trace log wrapping functionality.

## Checking Device Configuration <a href="#user_guide.devicemanager.cheap-synch-check" id="user_guide.devicemanager.cheap-synch-check"></a>

When managing large networks with NSO, a good strategy is to consider the NSO copy of the network configuration to be the main primary copy. All device configuration changes must go through NSO and all other device re-configurations are considered rogue.

NSO does not contain any functionality which disallows rogue re-configurations of managed devices, however, it does contain a mechanism whereby it is a very cheap operation to discover if one or several devices have been configured out-of-band.

The underlying mechanism for cheap `check-sync` is to compare time stamps, transaction IDs, hash-sums, etc., depending on what the device supports. This is in order not to have to read the full configuration to check if the NSO copy is in sync.

The transaction IDs are stored in CDB and can be viewed as:

```cli
ncs# show devices device state last-transaction-id
NAME  LAST TRANSACTION ID
----------------------------------------
ce0   ef3bbd344ef94b3fecec5cb93ac7458c
ce1   48e91db163e294bf5c3978d154922c9
ce2   48e91db163e294bf5c3978d154922c9
ce3   48e91db163e294bf5c3978d154922c9
ce4   48e91db163e294bf5c3978d154922c9
ce5   48e91db163e294bf5c3978d154922c9
ce6   48e91db163e294bf5c3978d154922c9
ce7   48e91db163e294bf5c3978d154922c9
ce8   48e91db163e294bf5c3978d154922c9
p0    -
p1    -
p2    -
p3    -
pe0   -
pe1   -
pe2   1412-581909-661436
pe3   -
```

Some of the devices do not have a transaction ID, this is the case where the NED has not implemented the cheap `check-sync` mechanism. Although it is called transaction-id, the underlying value in the device can be anything to detect a config change, like for example a time-stamp.

To check for consistency, we execute:

```cli
ncs# devices check-sync
sync-result {
    device ce0
    result in-sync
}
...
sync-result {
    device p1
    result unsupported
}
...
```

Alternatively for all (or a subset) managed devices:

```cli
ncs# devices device ce0..3 check-sync
devices device ce0 check-sync
    result in-sync
devices device ce1 check-sync
    result in-sync
devices device ce2 check-sync
    result in-sync
devices device ce3 check-sync
    result in-sync
```

The following YANG grouping is used for the return value from the `check-sync` command:

```
grouping check-sync-result {
    description
      "Common result data from a 'check-sync' action.";

    leaf result {
      type enumeration {
        enum unknown {
          description
            "NCS have no record, probably because no
             sync actions have been executed towards the device.
             This is the initial state for a device.";
        }
        enum locked {
          tailf:code-name 'sync_locked';
          description
            "The device is administratively locked, meaning that NCS
             cannot talk to it.";
        }
        enum in-sync {
          tailf:code-name 'in-sync-result';
          description
            "The configuration on the device is in sync with NCS.";
        }
        enum out-of-sync {
          description
            "The device configuration is known to be out of sync, i.e.,
             it has been reconfigured out of band.";
        }
        enum unsupported {
          description
            "The device doesn't support the tailf-netconf-monitoring
             module.";
        }
        enum error {
          description
            "An error occurred when NCS tried to check the sync status.
             The leaf 'info' contains additional information.";
        }
      }
    }
  }
```

### Comparing Device Configurations <a href="#user_guide.devicemanager.comparing" id="user_guide.devicemanager.comparing"></a>

In the previous section, we described how we can easily check if a managed device is in sync. If the device is not in sync, we are interested to know what the difference is. The CLI sequence below shows how to modify `ce0` out-of-band using the ncs-netsim tool. Finally, the sequence shows how to do an explicit configuration comparison.

```bash
$ ncs-netsim cli-i ce0
admin connected from 127.0.0.1 using console on ncs
ce0> enable
ce0# configure
Enter configuration commands, one per line. End with CNTL/Z.
ce0(config)# snmp-server community foobar RW
ce0(config)# exit
ce0# exit
$ ncs_cli -C -u admin

admin connected from 127.0.0.1 using console on ncs
ncs# devices device ce0 check-sync
result out-of-sync
info got: 290fa2b49608df9975c9912e4306110 expected: ef3bbd344ef94b3fecec5cb93ac7458c

ncs# devices device ce0 compare-config
diff
 devices {
     device ce0 {
         config {
             ios:snmp-server {
+                community foobar {
+                    RW;
+                }
             }
         }
     }
 }
```

The diff in the above output should be interpreted as: what needs to be done in NSO to become in sync with the device.

Previously in the example (Synchronize from Devices), NSO was brought in sync with the devices by fetching configuration from the devices. In this case, where the device has a rogue re-configuration, NSO has the correct configuration. In such cases, you want to reset the device configuration to what is stored inside NSO.

When you decide to reset the configuration with the copy kept in NSO use the option `dry-run` in conjunction with `sync-to` and inspect what will be sent to the device:

```cli
ncs# devices device ce0 sync-to dry-run
data
      no snmp-server community foobar RW
ncs#
```

As this is the desired data to send to the device a `sync-to` can now safely be performed.

```cli
ncs# devices device ce0 sync-to
result true
ncs#
```

The device configuration should now be in sync with the copy in NSO and `compare-config` ought to yield an empty output:

```cli
ncs# devices device ce0 compare-config
ncs#
```

## Initialize Device <a href="#user_guide.devicemanager.initialize-device" id="user_guide.devicemanager.initialize-device"></a>

There exist several ways to initialize new devices. The two common ways are to initialize a device from another existing device or to use device templates.

### From Other <a href="#user_guide.devicemanager.initialize-from-other" id="user_guide.devicemanager.initialize-from-other"></a>

For example, another CE router has been added to our example network. You want to base the configuration of that host on the configuration of the managed device `ce0` which has a valid configuration:

```cli
ncs(config)# show full-configuration devices device ce0
devices device ce0
 address   127.0.0.1
 port      10022
 ssh host-key ssh-dss
  key-data "AAAAB3NzaC1kc3MAAACBAO9tkTdZgAqJMz8m...
 !
 authgroup default
 device-type cli ned-id cisco-ios-cli-3.8
 state admin-state unlocked
 config
  no ios:service pad
  no ios:ip domain-lookup
  no ios:ip http secure-server
  ios:ip source-route
  ios:interface GigabitEthernet0/1
  exit
  ios:interface GigabitEthernet0/10
  exit
  ios:interface GigabitEthernet0/11
  exit
  ios:interface GigabitEthernet0/12
  exit
  ios:interface GigabitEthernet0/13
  exit
  ios:interface GigabitEthernet0/14
  exit
....
```

If the configuration is accurate you can create a new managed device based on that configuration as:

{% code title="Example: Instantiate Device from Other" %}

```cli
ncs(config)# devices device ce9 address 127.0.0.1 port 10031
ncs(config-device-ce9)# device-type cli ned-id cisco-ios-cli-3.8
ncs(config-device-ce9)# authgroup default
ncs(config-device-ce9)# instantiate-from-other-device device-name ce0
ncs(config-device-ce9)# top
ncs(config)# show configuration
devices device ce9
 address   127.0.0.1
 port      10031
 authgroup default
 device-type cli ned-id cisco-ios-cli-3.8
 config
  no ios:service pad
  no ios:ip domain-lookup
  no ios:ip http secure-server
  ios:ip source-route
  ios:interface GigabitEthernet0/1
  exit
....
ncs(config)# commit
Commit complete.
```

{% endcode %}

In the example above (Instantiate Device from Other) the commands first create the new managed device, `ce9` and then populates the configuration of the new device based on the configuration of `ce0`.

This new configuration might not be entirely correct, you can modify any configuration before committing it.

The above concludes the instantiation of a new managed device. The new device configuration is committed and NSO returned OK without the device existing in the network (netsim). Try to force a sync to the device:

```cli
ncs(config)# devices device ce9 sync-to
result false
info Device ce9 is southbound locked
```

The device is `southbound locked`, this is a mode that is used where you can reconfigure a device, but any changes done to it are never sent to the managed device. This will be thoroughly described in the next section. Devices are by default created southbound locked. Default values are not shown if not explicitly requested:

```
(config)# show full-configuration devices device ce9 state | details
devices device ce9
 state admin-state southbound-locked
!
```

### By Template <a href="#user_guide.devicemanager.initialize-with-template" id="user_guide.devicemanager.initialize-with-template"></a>

Another alternative to instantiating a device from the actual working configuration of another device is to have a number of named device templates that manipulate the configuration.

The template tree looks like this:

```yang
submodule tailf-ncs-devices {
  namespace "http://tail-f.com/ns/ncs";
  ...
container devices {
    ........
    list template {
      description
        "This list is used to define named template configurations that
         can be used to either instantiate the configuration for new
         devices, or to apply snippets of configurations to existing
         devices.
         ...
         ";

      key name;
      leaf name {
        description "The name of a specific template configuration";
        type string;
      }
      list ned-id {
        key id;
        leaf id {
          type identityref {
            base ned:ned-id;
          }
        }
        container config {
          tailf:mount-point ncs-template-config;
          tailf:cli-add-mode;
          tailf:cli-expose-ns-prefix;
          description
            "This container is augmented with data models from the devices.";
        }
      }
    }
```

The tree for device templates is generated from all device YANG models. All constraints are removed and the data type of all leafs is changed to `string`. By default the schemas for device templates are not accessible from application client libraries such as MAAPI. This reduces the memory usage for large device data models. The schema can be made accessible with the `/ncs-config/enable-client-template-schemas` setting in `ncs.conf`.

A device template is created by setting the desired data in the configuration. The created device template is stored in NSO CDB.

{% code title="Example: Create ce-initialize Template" %}

```cli
ncs(config)# devices template ce-initialize ned-id cisco-ios-cli-3.8 config
ncs(config-config)# no ios:service pad
ncs(config-config)# no ios:ip domain-lookup
ncs(config-config)# ios:ip dns server
ncs(config-config)# no ios:ip http server
ncs(config-config)# no ios:ip http secure-server
ncs(config-config)# ios:ip source-route true
ncs(config-config)# ios:interface GigabitEthernet 0/1
ncs(config- GigabitEthernet-0/1)# exit
ncs(config-config)# ios:interface GigabitEthernet 0/2
ncs(config- GigabitEthernet-0/2)# exit
ncs(config-config)# ios:interface GigabitEthernet 0/3
ncs(config- GigabitEthernet-0/3)# exit
ncs(config-config)# ios:interface Loopback 0
ncs(config-Loopback-0)# exit
ncs(config-config)# ios:snmp-server community public RO
ncs(config-community-public)# exit
ncs(config-config)# ios:snmp-server trap-source GigabitEthernet 0/2
ncs(config-config)# top
ncs(config)# commit
```

{% endcode %}

The device template created in the example above (Create ce-initialize template) can now be used to initialize single devices or device groups, see [Device Groups](#user_guide.devicemanager.device_groups).

In the following CLI session, a new device `ce10` is created:

```cli
ncs(config)# devices device ce10 address 127.0.0.1 port 10032
ncs(config-device-ce10)# device-type cli ned-id cisco-ios-cli-3.8
ncs(config-device-ce10)# authgroup default
ncs(config-device-ce10)# top
ncs(config)# commit
```

Initialize the newly created device `ce10` with the device template `ce-initialize`:

```cli
ncs(config)# devices device ce10 apply-template template-name ce-initialize
apply-template-result {
    device ce10
    result no-capabilities
    info No capabilities found for device: ce10. Has a sync-from the device
         been performed?
}
```

When initializing devices, NSO does not have any knowledge about the capabilities of the device, no connect has been done. This can be overridden by the option `accept-empty-capabilities`

```cli
ncs(config)# devices device ce10 \
apply-template template-name ce-initialize accept-empty-capabilities
apply-template-result {
    device ce10
    result ok
}
```

Inspect the changes made by the template `ce-initialize`:

```cli
ncs(config)# show configuration
devices device ce10
 config
  ios:ip dns server
  ios:interface GigabitEthernet0/1
  exit
  ios:interface GigabitEthernet0/2
  exit
  ios:interface GigabitEthernet0/3
  exit
  ios:interface Loopback0
  exit
  ios:snmp-server community public RO
  ios:snmp-server trap-source GigabitEthernet0/2
 !
!
```

## Device Templates <a href="#ncs.user_guide.devicemanager.device.templates" id="ncs.user_guide.devicemanager.device.templates"></a>

This section shows how device templates can be used to create and change device configurations. See [Introduction](/guides/development/core-concepts/templates#introduction) in Templates for other ways of using templates.

Device templates are part of the NSO configuration. Device templates are created and changed in the tree `/devices/template/ned-id/config` the same way as any other configuration data and are affected by rollbacks and upgrades. Device templates can only manipulate configuration data in the `/devices/device/config` tree i.e., only device data.

The [examples.ncs/service-management/mpls-vpn-java](https://github.com/NSO-developer/nso-examples/tree/6.7/service-management/mpls-vpn-java) example comes with a pre-populated template for SNMP settings.

```cli
ncs(config)# show full-configuration devices template
devices template snmp1
 ned-id cisco-ios-cli-3.8
  config
   ios:snmp-server community {$COMMUNITY}
    RO
   !
  !
 !
 ned-id cisco-iosxr-cli-3.5
  config
   cisco-ios-xr:snmp-server community {$COMMUNITY}
    RO
   !
  !
 !
 ned-id juniper-junos-nc-3.0
  config
   junos:configuration snmp community {$COMMUNITY}
    authorization read-only
   !
  !
 !
!
```

{% hint style="info" %}
The variable `$DEVICE` is used internally by NSO and can not be used in a template.
{% endhint %}

Templates can be created like any configuration data and use the CLI tab completion to navigate. Variables can be used instead of hard-coded values. In the template above the community string is a variable. The template can cover several device types/NEDs, by making use of the namespace information. This will make sure that only devices modeled with this particular namespace will be affected by this part of the template. Hence, it is possible for one template to handle a multitude of devices from various manufacturers.

A template can be applied to a device, a device group, and a range of devices. It can be used as shown in [By Template](#user_guide.devicemanager.initialize-with-template) to create the day-zero config for a newly created device.

Applying the `snmp1` template, providing a value for the `COMMUNITY` template variable:

```cli
ncs(config)# devices device ce2 apply-template template-name \
      snmp1 variable { name COMMUNITY value 'FUZBAR' }
ncs(config)# show configuration
devices device ce2
 config
  ios:snmp-server community FUZBAR RO
 !
!
ncs(config)# commit dry-run outformat native
native {
    device {
        name ce2
        data snmp-server community FUZBAR RO
    }
}
ncs(config)# commit
Commit complete.
```

The result of applying the template:

```cli
ncs(config)# show full-configuration devices device ce2 config\
   ios:snmp-server
devices device ce2
 config
  ios:snmp-server community FUZBAR RO
 !
!
```

### Tags <a href="#d5e3344" id="d5e3344"></a>

The default operation for templates is to merge the configuration. Tags can be added to templates to have the template `merge`, `replace`, `delete`, `create` or `nocreate` configuration. A tag is inherited to its sub-nodes until a new tag is introduced.

* `merge`*:* Merge with a node if it exists, otherwise create the node. This is the default operation if no operation is explicitly set.
* `replace`*:* Replace a node if it exists, otherwise create the node.
* `create`*:* Creates a node. The node can not already exist.
* `nocreate`*:* Merge with a node if it exists. If it does not exist, it will *not* be created.

Example of how to set a tag:

```cli
ncs(config)# tag add devices template snmp1 ned-id cisco-ios-cli-3.8 config\
 ios:snmp-server community {$COMMUNITY} replace
```

Displaying Tags information::

```cli
ncs(config)# show configuration
devices template snmp1
 ned-id cisco-ios-cli-3.8
  config
   ! Tags: replace
   ios:snmp-server community {$COMMUNITY}
   !
  !
 !
!
```

### Debug <a href="#d5e3374" id="d5e3374"></a>

By adding the CLI pipe flag `debug template` when applying a template, the CLI will output detailed information on what is happening when the template is being applied:

```cli
ncs(config)# devices device ce2 apply-template template-name \
      snmp1 variable { name COMMUNITY value 'FUZBAR' } | debug template
Operation 'merge' on existing node: /devices/device[name='ce2']
The device /devices/device[name='ce2'] does not support
namespace 'http://tail-f.com/ned/cisco-ios-xr' for node "'snmp-server'"
Skipping...
The device /devices/device[name='ce2'] does not support
namespace 'http://xml.juniper.net/xnm/1.1/xnm' for node "configuration"
Skipping...
Variable $COMMUNITY is set to "FUZBAR"
Operation 'merge' on non-existing node:
/devices/device[name='ce2']/config/ios:snmp-server/community[name='FUZBAR']
Operation 'merge' on non-existing node:
/devices/device[name='ce2']/config/ios:snmp-server/community[name='FUZBAR']/RO
```

## Generating Device Templates From Configuration

To simplify template creation, NSO features the `/devices/create-template` action that can initiate a template from a set of device configurations by finding common structural patterns. The resulting template can be used as as-is or as a starting point for further refinement.

In addition to extracting patterns from configuration already present in NSO, the action can also consume configuration snippets directly. Snippets can be supplied either from a file on the NSO server filesystem or as inline payload data. Supported formats are NETCONF-style XML wrapped in a `<config>` element, Cisco XR style CLI (`cli-c`), Juniper curly-brace CLI (`cli-j`), and Juniper set commands (`cli-j-cmd`). Delete operations in the input, such as Cisco-style `no` commands or XML `operation="remove"` attributes, are translated into `delete` tags in the generated device template.

The algorithm works by traversing the data depth-first, keeping track of the rate of occurrence of configuration nodes, and any values that compare equal. Values that do not compare equal are parameterized. For example:

{% code overflow="wrap" %}

```bash
admin@ncs(config)# devices create-template name syslog path [ /devices/device[device-type/netconf/ned-id='router-nc-1.0:router-nc-1.0']/config/sys/syslog ]
admin@ncs(config)# show configuration                                                                   devices template syslog
 ned-id router-nc-1.0
  config
   sys syslog server 10.3.4.5
    enabled
    selector 8
     facility [ "{$server-selector-facility}" ]
    !
   !
  !
 !
!
admin@ncs(config)# commit
Commit complete.
```

{% endcode %}

The action takes a number of arguments to control how the resulting template looks:

* `path` - A list of XPath 1.0 expressions pointing into `/devices/device/config` to create the template from. The template is only created from the paths that are common in the node-set.
* `match-rate` - Device configuration is included in the resulting template based on the rate of occurrence given by this setting.
* `exclude-service-config` - Exclude configuration that is already under service management.
* `collapse-list-keys` - Decides what lists to make variables of, either `all`, `automatic` (default), or those specified by the `list-path` parameter. The default is to find those lists that differ among the device configurations.

## Renaming Devices in NSO

The usual way to rename an instance in a list is to delete it and create a new instance. Aside from having to explicitly create all its children, an obvious problem with this method is the dependencies - if there is a leafref that refers to this instance, this method of deleting and recreating will fail unless the leafref is also explicitly reset to the value of the new instance.

The `/devices/device/rename` action renames an existing device and fixes the node/data dependencies in CDB. When renaming a device, the action fixes the following dependencies:

* Leafrefs and instance-identifiers (both config true and config false).
* Monitor and kick-node of kickers, if they refer to this device.
* Diff-sets and forward-diff-sets of services that touch this device (This includes nano-services and also zombies).

NSO maintains a history of past renames at `/devices/device/rename-history`.

### Examples <a href="#d5e3500" id="d5e3500"></a>

```
admin@ncs> request devices device ex0 rename new-name foo
result true
[ok][2024-04-16 20:51:51]
admin@ncs> show devices device foo rename-history | tab
FROM  TO   WHEN                              USER
----------------------------------------------------
ex0   foo  2024-04-16T18:51:51.578439+00:00  admin

[ok][2024-04-16 20:52:07]
admin@ncs> show configuration devices device ex0
---------------------------------------------^
syntax error: element does not exist
[error][2024-04-16 20:52:09]
admin@ncs> show configuration devices device foo
address   127.0.0.1;
port      12022;
...
```

The `rename` action takes a device lock to prevent modifications to the device while renaming it. Depending on the input parameters, the action will either immediately fail if it cannot get the device lock, or wait wait a specified amount of seconds before timing out.

```
admin@ncs> request devices commit-queue add-lock device [ ex1 ]
commit-queue-id 1713297244546
[ok][2024-04-16 21:54:04]
admin@ncs> request devices device ex1 rename new-name foo wait-for-lock { timeout 5 }
result false
info ex1: A timeout occured when trying to add device lock to the commit queue
[ok][2024-04-16 21:54:26]
```

The parameter `no-wait-for-lock` makes the action fail immediately if the device lock is unavailable, while a timeout of `infinity` can be used to make it wait indefinitely for the lock.

### Limitations <a href="#d5e3516" id="d5e3516"></a>

If a nano-service has components whose names are derived from the device name, and that device is renamed, the corresponding service components in its plan are not automatically renamed.

For example, let's say the nano-service has components with names matching device names.

```cli
admin@ncs% run show vlan-state test plan | tab
                                                                           POST
              BACK                                                         ACTION
TYPE  NAME  TRACK  GOAL  STATE        STATUS   WHEN                 ref  STATUS
---------------------------------------------------------------------------------
self  self  false  -     init         reached  2024-04-16T21:38:34  -    -
                         ready        reached  2024-04-16T21:38:34  -    -
vlan  ex1   false  -     init         reached  2024-04-16T21:38:34  -    -
                         router-init  reached  2024-04-16T21:38:34  -    -
                         ready        reached  2024-04-16T21:38:34  -    -
vlan  ex2   false  -     init         reached  2024-04-16T21:38:34  -    -
                         router-init  reached  2024-04-16T21:38:34  -    -
                         ready        reached  2024-04-16T21:38:34  -    -

[ok][2024-04-16 21:38:44]
```

If this device is renamed, the corresponding nano-service component is not renamed.

```cli
admin@ncs% request devices device ex1 rename new-name newex1
result true
[ok][2024-04-16 21:39:21]

[edit]
admin@ncs% run show vlan-state test plan | tab
                                                                           POST
              BACK                                                         ACTION
TYPE  NAME  TRACK  GOAL  STATE        STATUS   WHEN                 ref  STATUS
---------------------------------------------------------------------------------
self  self  false  -     init         reached  2024-04-16T21:38:34  -    -
                         ready        reached  2024-04-16T21:38:34  -    -
vlan  ex1   false  -     init         reached  2024-04-16T21:38:34  -    -
                         router-init  reached  2024-04-16T21:38:34  -    -
                         ready        reached  2024-04-16T21:38:34  -    -
vlan  ex2   false  -     init         reached  2024-04-16T21:38:34  -    -
                         router-init  reached  2024-04-16T21:38:34  -    -
                         ready        reached  2024-04-16T21:38:34  -    -

[ok][2024-04-16 21:39:24]
```

To handle this, the component with the old name must be force-back-tracked and the service re-deployed.

```cli
admin@ncs% request vlan-state test plan component vlan ex1 force-back-track
result true
[ok][2024-04-16 21:39:51]

[edit]
admin@ncs% run show vlan-state test plan | tab
                                                                         POST
            BACK                                                         ACTION
TYPE  NAME  TRACK  GOAL  STATE        STATUS   WHEN                 ref  STATUS
---------------------------------------------------------------------------------
self  self  false  -     init         reached  2024-04-16T21:38:34  -    -
                         ready        reached  2024-04-16T21:38:34  -    -
vlan  ex2   false  -     init         reached  2024-04-16T21:38:34  -    -
                         router-init  reached  2024-04-16T21:38:34  -    -
                         ready        reached  2024-04-16T21:38:34  -    -

[ok][2024-04-16 21:39:54]

[edit]
admin@ncs% request vlan test re-deploy
[ok][2024-04-16 21:40:02]

[edit]
admin@ncs% run show vlan-state test plan | tab
                                                                           POST
              BACK                                                         ACTION
TYPE  NAME    TRACK  GOAL  STATE        STATUS   WHEN                 ref  STATUS
-----------------------------------------------------------------------------------
self  self    false  -     init         reached  2024-04-16T21:38:34  -    -
                           ready        reached  2024-04-16T21:40:02  -    -
vlan  ex2     false  -     init         reached  2024-04-16T21:38:34  -    -
                           router-init  reached  2024-04-16T21:38:34  -    -
                           ready        reached  2024-04-16T21:38:34  -    -
vlan  newex1  false  -     init         reached  2024-04-16T21:40:02  -    -
                           router-init  reached  2024-04-16T21:40:02  -    -
                           ready        reached  2024-04-16T21:40:02  -    -

[ok][2024-04-17 08:40:05]
```

When a device is renamed, all components that derive their name from that device's name in all the service instances must be force-back-tracked.

## Auto-configuring Devices <a href="#user_guide.devicemanager.auto-configuring-devices" id="user_guide.devicemanager.auto-configuring-devices"></a>

Provisioning new devices in NSO requires the user to be familiar with the concept of Network Element Drivers and the unique ned-id they use to distinguish their schema. For an end user interacting with a northbound client of NSO, the concept of a ned-id might feel too abstract. It could be challenging to know what device type and ned-id to select when configuring a device for the first time in NSO. After initial configuration, there are also additional steps required before the device can be operated from NSO.

NSO can auto-configure devices during initial provisioning. Under `/devices/device/auto-configure`, a user can specify either the ned-id explicitly or a combination of the device vendor and `product-family` or `operating-system`. These are meta-data specified in the `package-meta-data.xml` file in the NED package. Based on the combination of this meta-data or using the ned-id explicitly configured, a ned-id from a matching NED package is selected from the currently loaded packages. If multiple packages match the given combination, the package with the latest version is selected.

When a transaction with a newly auto-configured device gets committed, NSO fetches the device host keys (if required) and synchronizes the configuration from the device. Depending on the NED used, additional transactions may be required. Also, if the device is unreachable, NSO will retry the operation at intervals, specified in the settings under `/devices/global-settings/auto-configure`. The `oper-state` leaf indicates when the device becomes `enabled`. Once the device is in sync, the auto-configuration stops. If the configured retry attempts are exhausted, NSO raises an `auto-configure-failed` alarm.

If several devices are committed simultaneously in the transaction with `auto-configure`, NSO will retry these immediately in separate transactions. This ensures that auto-configuration for a single device is not dependent on the success of the other devices.

### Examples <a href="#d5e3539" id="d5e3539"></a>

NSO will auto-configure a new device in a transaction if either `/devices/device/auto-configure/vendor` or `/devices/device/auto-configure/ned-id` is set in that transaction.

```cli
admin@ncs% show packages package component ned device
packages package router-nc-1.0
 component router
  ned device vendor "Acme Inc."
  ned device product-family [ "Acme Netconf router 1.0" ]
  ned device operating-system [ AcmeOS "AcmeOS 2.0" ]
[ok][2024-04-16 19:53:20]
admin@ncs% set devices device mydev address 127.0.0.1 port 12022 authgroup default
[ok][2024-04-16 19:53:34]

[edit]
admin@ncs% set devices device mydev auto-configure vendor "Acme Inc." operating-system AcmeOs
[ok][2024-04-16 19:53:36]

[edit]
admin@ncs% commit | details
...
 2024-04-16T19:53:37.655 device mydev: auto-configuring...
 2024-04-16T19:53:37.659 device mydev: configuring admin state... ok (0.000 s)
 2024-04-16T19:53:37.659 device mydev: fetching ssh host keys... ok (0.011 s)
 2024-04-16T19:53:37.671 device mydev: copying configuration from device... ok (0.054 s)
 2024-04-16T19:53:37.726 device mydev: auto-configuring: ok (0.070 s)
...
```

One can configure either `vendor` and `product-family`, or `vendor` and `operating-system` or just the `ned-id` explicitly.

```cli
admin@ncs% set devices device d1 auto-configure vendor "Acme Inc." product-family "Acme router"

admin@ncs% set devices device d2 auto-configure vendor "Acme Inc." operating-system AcmeOS

admin@ncs% set devices device d3 auto-configure ned-id router-nc-1.0
```

The `admin-state` for the device, if configured, will be honored. I.e., while auto-configuring a new device, if the `admin-state` is set to be southbound-locked, NSO will only pick the ned-id automatically. NSO will not fetch host keys and synchronize config from the device. NSO will not try again, even if the `admin-state` is changed.

```cli
admin@ncs% set devices device mydev2 auto-configure vendor "Acme Inc." operating-system AcmeOS
[ok][2024-04-16 20:03:05]

[edit]
admin@ncs% set devices device mydev2 state admin-state southbound-locked
[ok][2024-04-16 20:03:05]

[edit]
admin@ncs% commit | details
...
 2024-04-16T20:03:08.604 device mydev2: auto-configuring...
 2024-04-16T20:03:08.606 device mydev2: configuring admin state... ok (0.000 s)
 2024-04-16T20:03:08.606 device mydev2: fetching ssh host keys... skipped - 'southbound-locked' configured (0.001 s)
 2024-04-16T20:03:08.608 device mydev2: auto-configuring: ok (0.003 s)
...
```

Many NEDs require additional custom configuration to be operational. This applied in particular to Generic NEDs. Information about such additional configuration can be found in the files `README.md` and `README-ned-settings.md` bundled with the NED package.

## `oper-state` and `admin-state` <a href="#user_guide.devicemanager.state" id="user_guide.devicemanager.state"></a>

NSO differentiates between `oper-state` and `admin-state` for a managed device. `oper-state` is the actual state of the device. We have chosen to implement a very simple `oper-state` model. A managed device `oper-state` is either enabled or disabled. `oper-state` can be mapped to an alarm for the device. If the device is disabled, we may have additional error information. For example, the `ce9` device created from another device and `ce10` created with a device template in the previous section is disabled, and no connection has been established with the device, so its state is completely unknown:

```cli
ncs# show devices device ce9 state oper-state
state oper-state disabled
```

Or, a slightly more interesting CLI usage:

```cli
ncs# show devices device state oper-state
      OPER
NAME  STATE
----------------
ce0   enabled
ce1   enabled
ce10  disabled
ce2   enabled
ce3   enabled
ce4   enabled
ce5   enabled
ce6   enabled
ce7   enabled
ce8   enabled
ce9   disabled
p0    enabled
p1    enabled
p2    enabled
p3    enabled
pe0   enabled
pe1   enabled
pe2   enabled
pe3   enabled

ncs# show devices device ce0..9 state oper-state
      OPER
NAME  STATE
----------------
ce0   enabled
ce1   enabled
ce2   enabled
ce3   enabled
ce4   enabled
ce5   enabled
ce6   enabled
ce7   enabled
ce8   enabled
ce9   disabled
```

If you manually stop a managed device, for example `ce0`, NSO doesn't immediately indicate that. NSO may have an active SSH connection to the device, but the device may voluntarily choose to close its end of that (idle) SSH connection. Thus the fact that a socket from the device to NSO is closed by the managed device doesn't indicate anything. The only certain method NSO has to decide a managed device is non-operational - from the point of view of NSO - is NSO cannot SSH connect to it. If you manually stop managed device `ce0`, you still have:

```bash
$ ncs-netsim stop ce0
DEVICE ce0 STOPPED
$ ncs_cli -C -u admin
ncs# show devices device ce0 state oper-state
state oper-state enabled
```

NSO cannot draw any conclusions from the fact that a managed device closed its end of the SSH connection. It may have done so because it decided to time out an idle SSH connection. Whereas if NSO tried to initiate any operations towards the dead device, the device would be marked as `oper-state` `disabled`:

```cli
ncs(config)# devices device ce0 config ios:snmp-server contact joe@acme.com
ncs(config-config)# commit
Aborted: Failed to connect to device ce0: connection refused: Connection refused
ncs(config-config)# *** ALARM connection-failure: Failed to
connect to device ce0: connection refused: Connection refused
```

Now, NSO has failed to connect to it, NSO knows that `ce0` is dead:

```cli
ncs# show devices device ce0 state oper-state
state oper-state disabled
```

This concludes the `oper-state` discussion. The next state to be illustrated is the `admin-state`. The `admin-state` is what the operator configures, this is the desired state of the managed device.

In `tailf-ncs.yang` we have the following configuration definition for `admin-state`:

{% code title="Example: tailf-ncs-devices.yang - admin-state" %}

```yang
submodule tailf-ncs-devices {
  ....

  typedef admin-state {
    type enumeration {
      enum locked {
        description
          "When a device is administratively locked, it is not possible
           to modify its configuration, and no changes are ever
           pushed to the device.";
      }
      enum unlocked {
        description
          "Device is assumed to be operational.
           All changes are attempted to be sent southbound.";
      }
      enum southbound-locked {
        description
          "It is possible to configure the device, but
           no changes are sent to the device. Useful admin mode
           when pre provisioning devices. This is the default
           when a new device is created.";
      }
      enum config-locked {
        description
          "It is possible to send live-status commands or RPCs
           but it is not possible to modify the configuration
           of the device.";
      }
    }
  }

  ....
  container devices {
     ....
     container state {
        ....
        leaf admin-state {
          type admin-state;
          default southbound-locked;
        }

        leaf admin-state-description {
          type string;
          description
            "Reason for the admin state.";

        }
```

{% endcode %}

In the example above (tailf-ncs-devices.yang - admin-state), you can see the four different admin states for a managed device as defined in the YANG model.

* `locked` - This means that all changes to the device are forbidden. Any transaction which attempts to manipulate the configuration of the device will fail. It is still possible to read the configuration of the device.
* `unlocked` -This is the state a device is set into when the device is operational. All changes to the device are attempted to be sent southbound.
* `southbound-locked` - This is the default value. It means that it is possible to manipulate the configuration of the device but changes done to the device configuration are never pushed to the device. This mode is useful during e.g. pre-provisioning, or when we instantiate new devices.
* `config-locked` - This means that any transaction which attempts to manipulate the configuration of the device will fail. It is still possible to read the configuration of the device and send live-status commands or RPCs.

## Configuration Source <a href="#user_guide.devicemanager.source" id="user_guide.devicemanager.source"></a>

NSO manages a set of devices that are given to NSO through any means like CLI, inventory system integration through XML APIs, or configuration files at startup. The list of devices to manage in an overall integrated network management solution is shared between different tools and therefore it is important to keep an authoritative database of this and share it between different tools including NSO. The purpose of this part is to identify the source of the population of managed devices. The `source` attribute should indicate the source of the managed device like "inventory", "manual", or "EMS".

{% code title="Example: tailf-ncs-devices.yang - source" %}

```yang
submodule tailf-ncs-devices {
  ...
      container source {
        tailf:info "How the device was added to NCS";
        leaf added-by-user {
          type string;
        }
        leaf context {
          type string;
        }
        leaf when {
          type yang:date-and-time;
        }
        leaf from-ip {
          type inet:ip-address;
        }
        leaf source {
          type string;
          reference "TMF518 NRB Network Resource Basics";
        }
      }
```

{% endcode %}

These attributes should be automatically set by the integration towards the inventory source, rather than manipulated manually.

* `added-by-user`: Identify the user who loaded the managed device.
* `context`: In what context was the device loaded.
* `when`: When the device was added to NSO.
* `from-ip`: From which IP the load activity was run.
* `source`: Identify the source of the managed device such as the inventory system name or the name of the source file.

### Capabilities, Modules, and Revision Management <a href="#user_guide.devicemanager.capas" id="user_guide.devicemanager.capas"></a>

The NETCONF protocol mandates that the first thing both the server and the client have to do is to send its list of NETCONF capabilities in the `<hello>` message. A capability indicates what the peer can do. For example the `validate:1.0` indicates that the server can validate a proposed configuration change, whereas the capability `http://acme.com/if` indicates the device implements the `http://acme.com` proprietary capability.

The NEDs report the capabilities of the devices at connection time. The NEDs also load the YANG modules for NSO. For a NETCONF/YANG device, all this is straightforward, for non-NETCONF devices the NEDs do the translation.

The capabilities announced by a device also contain the YANG version 1 modules supported. In addition to this, YANG version 1.1 modules are advertised in the YANG library module on the device. NSO checks both the capabilities and the YANG library to find out which YANG modules a device supports.

The capabilities and modules detected by NSO are available in two different lists, `/devices/device/capability` and `devices/device/module`. The `capability` list contains all capabilities announced and all YANG modules in the YANG library. The `module` list contains all YANG modules announced that are also supported by the NED in NSO.

```cli
ncs# show devices device ce0 capability
capability urn:ietf:params:netconf:capability:with-defaults:1.0?basic-mode=trim
capability urn:ios
 revision 2015-03-16
 module   tailf-ned-cisco-ios
capability urn:ios-stats
 revision 2015-03-16
 module   tailf-ned-cisco-ios-stats

ncs#  show devices device ce0 capability module
NAME                       REVISION    FEATURE  DEVIATION
-----------------------------------------------------------
tailf-ned-cisco-ios        2015-03-16  -        -
tailf-ned-cisco-ios-stats  2015-03-16  -        -
```

NSO can be used to handle all or some of the YANG configuration modules for a device. A device may announce several modules through its capability list which NSO ignores. NSO will only handle the YANG modules for a device which are loaded (and compiled through `ncsc --ncs-compile-bundle`) or `ncsc --ncs-compile-module`) all other modules for the device are ignored. If you require a situation where NSO is entirely responsible for a device so that complete device backup/configurations are stored in NSO you must ensure NSO indeed has support for all modules for the device. It is not possible to automate this process since a capability URI doesn't necessarily indicate actual configuration.

### Discovery of a NETCONF Device <a href="#d5e3479" id="d5e3479"></a>

When a device is added to NSO its NED ID must be set. For a NETCONF device, it is possible to configure the generic NETCONF NED id `netconf` (defined in the YANG module `tailf-ncs-ned`). If this NED ID is configured, we can then ask NSO to connect to the device and then check the `capability` list to see which modules this device implements.

```cli
ncs(config)# devices device foo address 127.0.0.1 port 12033 authgroup default
ncs(config-device-foo)# device-type netconf ned-id netconf
ncs(config-device-foo)# state admin-state unlocked
ncs(config-device-foo)# commit
Commit complete.
ncs(config-device-foo)# exit
ncs(config)# exit
ncs# devices fetch-ssh-host-keys device foo
fetch-result {
    device foo
    result updated
    fingerprint {
        algorithm ssh-rsa
        value 14:3c:79:87:69:8e:e2:f0:6d:43:07:8c:89:41:fd:7f
    }
}
ncs# devices device foo connect
result true
info (admin) Connected to foo - 127.0.0.1:12033
ncs# show devices device foo capability
capability :candidate:1.0
capability :confirmed-commit:1.0
...
capability http://xml.juniper.net/xnm/1.1/xnm
 module junos
capability urn:ietf:params:xml:ns:yang:ietf-yang-types
 revision 2013-07-15
 module   ietf-yang-types
capability urn:juniper-rpc
 module junos-rpc
...
```

We can also check which modules the loaded NEDs support. Then we can pick the most suitable NED and configure the device with this NED ID.

```cli
ncs# show devices ned-ids
ID                    NAME                          REVISION
--------------------------------------------------------------
cisco-ios-xr-v2       tailf-ned-cisco-ios-xr        -
                      tailf-ned-cisco-ios-xr-stats  -
lsa-netconf
netconf
snmp
alu-sr-cli-3.4        tailf-ned-alu-sr              -
                      tailf-ned-alu-sr-stats        -
cisco-ios-cli-3.8     tailf-ned-cisco-ios           -
                      tailf-ned-cisco-ios-stats     -
cisco-iosxr-cli-3.5   tailf-ned-cisco-ios-xr        -
                      tailf-ned-cisco-ios-xr-stats  -
juniper-junos-nc-3.0  junos                         -
                      junos-rpc                     -
ncs# config
Entering configuration mode terminal
ncs(config)# devices device foo device-type netconf ned-id juniper-junos-nc-3.0
ncs(config-device-foo)# commit
Commit complete.
```

## Configuration Datastore Support <a href="#user_guide.devicemanager.candidate" id="user_guide.devicemanager.candidate"></a>

NSO works best if the managed devices support the NETCONF candidate configuration datastore. However, NSO reads the capabilities of each managed device and executes different sequences of NETCONF commands towards different types of devices.

For implementations of the NETCONF protocol that do not support the candidate datastore, and in particular, devices that do not support NETCONF commit with a timeout, NSO tries to do the best of the situation.

NSO divides devices into the following groups.

* `start_trans_running`: This mode is used for devices that support the Tail-f proprietary transaction extension defined by `http://tail-f.com/ns/netconf/transactions/1.0`. Read more on this in the Tail-f ConfD user guide. In principle it's a means to - over the NETCONF interface - control transaction processing towards the running data store. This may be more efficient than going through the candidate data store. The downside is that it is Tail-f proprietary non-standardized technology.
* `lock_candidate`: This mode is used for devices that support the candidate data store but disallow direct writes to the running data store.
* `lock_reset_candidate`: This mode is used for devices that support the candidate data and also allow direct writes to the running data store. This is the default mode for Tail-f ConfD NETCONF server. Since the running data store is configurable, we must, before each configuration attempt, copy all of the running to the candidate. (ConfD has optimized this particular usage pattern, so this is a very cheap operation for ConfD)
* `startup`: This mode is used for devices that have writable running, no candidate but do support the startup data store. This is the typical mode for Cisco-like devices.
* `running-only`: This mode is used for devices that only support writable running.
* `NED`: The transaction is controlled by a Network Element Driver. The exact transaction mode depends on the type of the NED.

Which category NSO chooses for a managed device depends on which NETCONF capabilities the device sends to NSO in its NETCONF hello message. You can see in the CLI what NSO has decided for a device as in:

```cli
ncs# show devices device ce0 state transaction-mode
state transaction-mode ned
ncs# show devices device pe2 state transaction-mode
state transaction-mode lock-candidate
```

NSO talking to ConfD device running in its standard configuration, thus `lock-reset-candidate`.

Another important discriminator between managed devices is whether they support the confirmed commit with a timeout capability, i.e., the `confirmed-commit:1.0` standard NETCONF capability. If a device supports this capability, NSO utilizes it. This is the case with for example Juniper routers.

If a managed device does not support this capability, NSO attempts to do the best it can.

This is how NSO handles common failure scenarios:

* The operator aborts the transaction, or the NSO loses the SSH connection to another managed device which is also participating in the same network transaction. If the device does support the `confirmed-commit` capability, NSO aborts the outstanding yet-uncommitted transaction simply by closing the SSH connection. When the device does not support the `confirmed-commit` capability, NSO has the reverse diff and simply sends the precise undo information to the device instead.
* The device rejects the transaction in the first place, i.e. the NSO attempts to modify its running data store. This is an easy case since NSO then simply aborts the transaction as a whole in the initial `commit confirmed [time]` attempt.
* NSO loses SSH connectivity to the device during the timeout period. This is a real error case and the configuration is now in an unknown state. NSO will abort the entire transaction, but the configuration of the failing managed device is now probably in error. The correct procedure once network connectivity has been restored to the device is to sync it in the direction from NSO to the device. The NSO copy of the device configuration will be what was configured before the failed transaction.

Thus, even if not all participating devices have first-class NETCONF server implementations, NSO will attempt to fake the `confirmed-commit` capability.

## Action Proxy <a href="#user_guide.devicemanager.action_proxy" id="user_guide.devicemanager.action_proxy"></a>

When the managed device defines top-level NETCONF RPCs or alternatively, define `tailf:action` points inside the YANG model, these RPCs and actions are also imported into the data model that resides in NSO.

For example, the Juniper NED comes with a set of JunOS RPCs defined in: `$NCS_DIR/packages/neds/juniper-junos/src/yang/junos-rpc.yang`

```yang
module junos-rpc {
  ...
  rpc request-package-add {
  ...
  rpc request-reboot {
  ...
  rpc get-software-information {
  ...
  rpc ping {
```

Thus, since all RPCs and actions from the devices are accessible through the NSO data model, these actions are also accessible through all NSO northbound APIs, REST, JAVA MAAPI, etc. Hence it is possible to - from user scripts/code - invoke actions and RPCs on all managed devices. The RPCs are augmented below an RPC container:

```cli
ncs(config)# devices device pe2 rpc rpc-
Possible completions:
  rpc-get-software-information  rpc-idle-timeout  rpc-ping \
  rpc-request-package-add  rpc-request-reboot

ncs(config)# devices device pe2 rpc \
rpc-get-software-information get-software-information brief
```

In the simulated environment of the [examples.ncs/service-management/mpls-vpn-java](https://github.com/NSO-developer/nso-examples/tree/6.7/service-management/mpls-vpn-java) example, these RPCs might not have been implemented.

## Device Groups <a href="#user_guide.devicemanager.device_groups" id="user_guide.devicemanager.device_groups"></a>

The NSO device manager has a concept of groups of devices. A group is nothing more than a named group of devices. What makes this interesting is that we can invoke several different actions in the group, thus implicitly invoking the action on all members in the group. This is especially interesting for the `apply-template` action.

The definition of device groups resides at the same layer in the NSO data model as the device list, thus we have:

{% code title="Example: Device Groups" %}

```yang
submodule tailf-ncs-devices {
  namespace "http://tail-f.com/ns/ncs";
  ...
  container devices {
     .....
    list device {
     ...
     }
    list device-group {
      key name;
      leaf name {
        type string;
      }
      description
        "A named group of devices, some actions can be
         applied to an entire  group of devices, for example
         apply-template, and the sync actions.";
      leaf-list device-name {
        type leafref {
          path "/devices/device/name";
        }
      }
      leaf-list device-group {
        type leafref {
          path "/devices/device-group/name";
        }
        description
          "A list of device groups contained in this device group.

           Recursive definitions are not valid.";
      }
      leaf-list member {
        type leafref {
          path "/devices/device/name";
        }
        config false;
        description
          "The current members of the device-group.  This is a flat list
           of all the devices in the group.";
      }
      uses connect-grouping ;
      uses sync-grouping;
      uses check-sync-grouping;
      uses apply-template-grouping;
    }
  }
}
```

{% endcode %}

The MPLS VPN example comes with a couple of pre-defined device-groups:

```cli
ncs(config)# show full-configuration devices device-group
devices device-group C
 device-name [ ce0 ce1 ce3 ce4 ce5 ce6 ce7 ce8 ]
!
devices device-group P
 device-name [ p0 p1 p2 p3 ]
!
devices device-group PE
 device-name [ pe0 pe1 pe2 pe3 ]
!
```

Device groups are created like below:

{% code title="Example: Create Device Group" %}

```cli
ncs(config)# devices device-group my-group device-name ce0
ncs(config-device-group-my-group)# device-name pe
Possible completions:
  pe0  pe1  pe2  pe3
ncs(config-device-group-my-group)# device-name pe0
ncs(config-device-group-my-group)# device-name p0
ncs(config-device-group-my-group)# commit
```

{% endcode %}

Device groups can reference other device groups. There is an operational attribute that flattens all members in the group. The CLI sequence below adds the `PE` group to `my-group`. Then it shows the configuration of that group followed by the status of this group. The status for the group contains a `members` attribute that lists all device members.

```
ncs(config-device-group-my-group)# device-group PE
ncs(config-device-group-my-group)# commit

ncs(config)# show full-configuration devices device-group my-group
devices device-group my-group
 device-name  [ ce0 p0 pe0 ]
 device-group [ PE ]
!
ncs(config)# exit

ncs# show devices device-group my-group
NAME      MEMBER                      INDETERMINATES  CRITICALS  MAJORS  MINORS  WARNINGS
-------------------------------------------------------------------------------------------
my-group  [ ce0 p0 pe0 pe1 pe2 pe3 ]  0               0          1       0       0
```

Once you have a group, you can sync and check-sync the entire group.

```cli
ncs# devices device-group C sync-to
```

However, what makes device groups really interesting is the ability to apply a template to a group. You can use the pre-populated templates to apply SNMP settings to device groups.

```cli
ncs(config)# devices device-group C apply-template \
template-name snmp1 variable { name COMMUNITY value 'cinderella' }
ncs(config)# show configuration
devices device ce0
 config
  ios:snmp-server community cinderella RO
 !
!
devices device ce1
 config
  ios:snmp-server community cinderella RO
 !
!
...
ncs(config)# commit
```

## Policies <a href="#user_guide.devicemanager.policies" id="user_guide.devicemanager.policies"></a>

Policies allow you to specify network-wide constraints that always must be true. If someone tries to apply a configuration change over any northbound interface that would be evaluated to false, the configuration change is rejected by NSO. Policies can be of type warning means that it is possible to override them, or error which cannot be overridden.

Assume you would like to enforce all CE routers to have a Gigabit interface `0/1`.

<pre data-title="Example: Policies"><code>ncs(config)# policy rule gb-one-zero
ncs(config-rule-gb-one-zero)# foreach /ncs:devices/device[starts-with(name,'ce')]/config
ncs(config-rule-gb-one-zero)# expr ios:interface/ios:GigabitEthernet[ios:name='0/1']
<strong>ncs(config-rule-gb-one-zero)# warning-message "{../name} should have 0/1 interface"
</strong>ncs(config-rule-gb-one-zero)# commit
zork(config-rule-gb-one-zero)# top
zork(config)# !
ncs(config)# show full-configuration policy
policy rule gb-one-zero
 foreach         /ncs:devices/device[starts-with(name,'ce')]/config
 expr            ios:interface/ios:GigabitEthernet[ios:name='0/1']
 warning-message "{../name} should have 0/1 interface"
!
ncs(config)# no devices device ce0 config ios:interface GigabitEthernet 0/1
ncs(config)# validate
Validation completed with warnings:
  ce0 should have 0/1 interface
ncs(config)# no devices device ce1 config ios:interface GigabitEthernet 0/1
ncs(config)# validate
Validation completed with warnings:
  ce1 should have 0/1 interface
  ce0 should have 0/1 interface
ncs(config)# commit
The following warnings were generated:
  ce1 should have 0/1 interface
  ce0 should have 0/1 interface
Proceed? [yes,no] yes
Commit complete.
</code></pre>

As seen in the example above (Policies) , a policy rule has (an optional) for each statement and a mandatory expression and error message. The `foreach` statement evaluates to a node set, and the expression is then evaluated on each node. So in this example, the expression would be evaluated for every device in NSO which begins with ce. The name variable in the warning message refers to a leaf available from the for-each node set.

Validation is always performed at commit but can also be requested interactively.

Note any configuration can be activated or deactivated. This means that to temporarily turn off a certain policy you can deactivate it. Note also that if the configuration was changed by any other means than NSO by local tools to the device like a CLI, a `devices sync-from` operation might fail if the device configuration violates the policy.

## Commit Queue <a href="#user_guide.devicemanager.commit-queue" id="user_guide.devicemanager.commit-queue"></a>

One of the strengths of NSO is the concept of network-wide transactions. When you commit data to NSO that spans multiple devices in the `/ncs:devices/device` tree, NSO will - within the NSO transaction - commit the data on all devices or none, keeping the network consistent with CDB. The NSO transaction doesn't return until all participants have acknowledged the proposed configuration change. The downside of this is that the slowest device in each transaction limits the overall transactional throughput in NSO. Such things as out-of-sync checks, network latency, calculation of changes sent southbound, or device deficiencies all affect the throughput.

Typically when automation software north of NSO generates network change requests it may very well be the case more requests arrive than what can be handled. In NSO deployment scenarios where you wish to have higher transactional throughput than what is possible using network-wide transactions, you can use the commit queue instead. The goal of the commit queue is to increase the transactional throughput of NSO while keeping an eventual consistency view of the database. With the commit queue, NSO will compute the configuration change for each participating device, put it in an outbound queue item, and immediately return. The queue is then independently run.

Another use case where you can use the commit queue is when you wish to push a configuration change to a set of devices and don't care about whether all devices accept the change or not. You do not want the default behavior for transactions which is to reject the transaction as a whole if one or more participating devices fail to process its part of the transaction.

An example of the above could be if you wish to set a new NTP server on all managed devices in our entire network, if one or more devices currently are non-operational, you still want to push out the change. You also want the change automatically pushed to the non-operational devices once they go live again.

The big upside of this scheme is that the transactional throughput through NSO is considerably higher. Also, transient devices are handled better. The downsides are:

1. If a device rejects the proposed change, NSO and the device are now *out of sync* until any error recovery is performed. Whenever this happens, an NSO alarm (called commit-through-queue-failed) is generated.
2. While a transaction remains in the queue, i.e., it has been accepted for delivery by NSO but is not yet delivered, the view of the network in NSO is not (yet) correct. Eventually, though, the queued item will be delivered, thus achieving eventual consistency.

To facilitate the two use cases of the commit queue the outbound queue item can be either in an atomic or non-atomic mode.

In atomic mode the outbound queue item will push all configuration changes concurrently once there are no intersecting devices ahead in the queue. If any device rejects the proposed change, all device configuration changes in the queue item will be rejected as a whole, leaving the network in a consistent state. The atomic mode also allows for automatic error recovery to be performed by NSO.

In the non-atomic mode, the outbound queue item will push configuration changes for a device whenever all occurrences of it are completed or it doesn't exist ahead in the queue. The drawback to this mode is that there is no automatic error recovery that can be performed by NSO.

In the following sequences, the simulated device `ce0` is stopped to illustrate the commit queue. This can be achieved by the following sequence including returning to the NSO CLI config mode:

```bash
$ ncs-netsim stop ce0
DEVICE ce0 STOPPED
$ ncs_cli -C -u admin

admin connected from 127.0.0.1 using console on ncs
ncs# config
```

By default, the commit queue is turned off. You can configure NSO to run a transaction, device, or device group through the commit queue in a number of different ways, either by providing a flag to the `commit` command as:

```cli
ncs(config)# commit commit-queue
Possible completions:
  async    Commit through commit queue and return immediately
  bypass   Bypass commit-queue when queue is enabled by default
  sync     Commit through commit queue and wait for reply
ncs(config)# commit commit-queue async
```

Or, by configuring NSO to always run all transactions through the commit queue as in:

```cli
ncs(config)# devices global-settings commit-queue enabled-by-default
[false,true] (false): true
ncs(config)# commit
```

Or, by configuring a number of devices to run through the commit queue as default:

```cli
ncs(config)# devices device ce0..2 commit-queue enabled-by-default
[false,true] (false): true
ncs(config)# commit
```

When enabling the commit queue as default on a per device/device group basis, an NSO transaction will compute the configuration change for each participating device, put the devices enabled for the commit queue in the outbound queue, and then proceed with the normal transaction behavior for those devices not commit queue enabled. The transaction will still be successfully committed even if some of the devices added to the outbound queue will fail. If the transaction fails in the validation phase the entire transaction will be aborted, including the configuration change for those devices added to the commit queue. If the transaction fails after the validation phase, the configuration change for the devices in the commit queue will still be delivered.

Do some changes and commit through the commit queue:

{% code title="Example: Commit through Commit Queue" %}

```cli
ncs(config)# devices device ce0..2 config ios:snmp-server \
    trap-source GigabitEthernet 0/1
ncs(config-config)# commit
commit-queue-id 9494446997
Commit complete.
ncs(config-config)# *** ALARM connection-failure: Failed to
connect to device ce0: connection refused: Connection refused
```

{% endcode %}

### Commit Queue Scheduling

In the example above (Commit through Commit Queue), the commit affected three devices, `ce0`, `ce1` and `ce2`. If you immediately would have launched yet another transaction, as in the second one (see example below), manipulating an interface of `ce2`, that transaction would have been queued instead of immediately launched. The idea here is to queue entire transactions that touch any device that has anything queued ahead in the queue.

```cli
ncs(config)# devices device ce0 config ios:interface GigabitEthernet 0/25
ncs(config-if)# commit
commit-queue-id 9494530158
Commit complete.
ncs(config-if)# *** ALARM commit-through-queue-blocked:
Commit Queue item 9494530158 is blocked because qitem 9494446997
cannot connect to ce0
```

Each transaction committed through the queues becomes a queue item. A queue item has an ID number. A bigger number means that it's scheduled later. Each queue item waits for something to happen. A queue item is in either of three states.

1. `waiting`: The queue item is waiting for other queue items to finish. This is because the *waiting* queue item has participating devices that are part of other queue items, ahead in the queue. It is waiting for a set of devices, to not occur ahead of itself in the queue.
2. `executing`: The queue item is currently being processed. Multiple queue items can run concurrently as long as they don't share any managed devices or if the atomic behaviour of the queue items are set to `false`. If NSO fails to connect to a device or the change is being rejected due to the device being locked, it is shown as a transient error in the `transient` list. NSO will retry aginast the device at intervals specified in `/ncs:devices/global-settings/commit-queue/retry-timeout`. Transient errors are potentially bad since the queue might grow if new items are added, waiting for the same device.
3. `locked`: This queue item is locked and will not be processed until it has been unlocked, see the action `/ncs:devices/commit-queue/queue-item/unlock`. A locked queue item will block all subsequent queue items that are using any device in the locked queue item.

### Viewing and Manipulating the Commit Queue <a href="#d5e3710" id="d5e3710"></a>

You can view the queue in the CLI. There are three different view modes, `summary`, `normal`*,* and `detailed`. Depending on the output, both the `summary` and the `normal` look good:

{% code title="Example: Viewing Queue Items" %}

```cli
ncs# show devices commit-queue | notab
devices commit-queue queue-item 9494446997
 age       144
 status    executing
 devices   [ ce0 ce1 ce2 ]
 transient ce0
  reason "Failed to connect to device ce0: connection refused"
 is-atomic true
devices commit-queue queue-item 9494530158
 age         61
 status      blocked
 devices     [ ce0 ]
 waiting-for [ ce0 ]
 is-atomic   true
```

{% endcode %}

The `age` field indicated how many seconds a queue item has been in the queue.

You can also view the queue items in detailed mode:

```cli
ncs# show devices commit-queue queue-item 9494530158 details | notab
devices commit-queue queue-item 9494530158
 age         278
 status      blocked
 devices     [ ce0 ]
 waiting-for [ ce0 ]
 is-atomic   true
 modification ce0
  data       <interface xmlns="urn:ios">
               <GigabitEthernet>
                 <name>0/25</name>
               </GigabitEthernet>
             </interface>

  local-user admin
```

The queue items are stored persistently, thus if NSO is stopped and restarted, the queue remains the same. Similarly, if NSO runs in HA (High Availability) mode, the queue items are replicated, ensuring the queue is processed even in case of failover.

{% hint style="info" %}
The commit queue is disabled when both HA is enabled, and its HA role is `none`, i.e., not `primary` or `secondary`. See [Mode of Operation](https://nso-docs.cisco.com/guides/operation-and-usage/operations/pages/qG3CMifhI63daJ1BZfmB#ha.moo).
{% endhint %}

A number of useful actions are available to manipulate the queue:

1. `devices commit-queue add-lock device [ ... ]`. This adds a fictive queue item to the commit queue. Any queue item, affecting the same devices, which is entering the commit queue will have to wait for this lock item to be unlocked or deleted. If no devices are specified, all devices in NSO are locked.
2. `devices commit-queue clear`. This action clears the entire queue. All devices present in the commit queue will, after this action, have executed be out of sync. The `clear` action is a rather blunt tool and is not recommended to be used in any normal use case.
3. `devices commit-queue prune device [ ... ]` . This action prunes all specified devices from all queue items in the commit queue. The affected devices will, after this action has been executed, be out of sync. Devices that are currently being committed to will not be pruned unless the `force` option is used. Atomic queue items will not be affected, unless all devices in it are pruned. The `force` option will brutally kill an ongoing commit. This could leave the device in a bad state. It is not recommended in any normal use case.
4. `devices commit-queue set-atomic-behaviour atomic [ true,false ]`. This action sets the atomic behavior of all queue items. If these are set to false, the devices contained in these queue items can start executing if the same devices in other non-atomic queue items ahead of it in the queue are completed. If set to true, the atomic integrity of these queue items is preserved.
5. `devices commit-queue wait-until-empty`. This action waits until the commit queue is empty. The default is to wait `infinity`. A `timeout` can be specified to wait for a number of seconds. The result is `empty` if the queue is empty or `timeout` if there are still items in the queue to be processed.
6. `devices commit-queue queue-item [ id ] lock`. This action puts a lock on an existing queue item. A locked queue item will not start executing until it has been unlocked.
7. `devices commit-queue queue-item [ id ] unlock`. This action unlocks a locked queue item. Unlocking a queue item that is not locked is silently ignored.
8. `devices commit-queue queue-item [ id ] delete`. This action deletes a queue item from the queue. If other queue items are waiting for this (deleted) item, they will all automatically start to run. The devices of the deleted queue item will, after the action has been executed, be out of sync if they haven't started executing. Any error option set for the queue item will also be disregarded. The `force` option will brutally kill an ongoing commit. This could leave the device in a bad state. It is not recommended in any normal use case.
9. `devices commit-queue queue-item [ id ] prune device [ ... ]`. This action prunes the specified devices from the queue item. Devices that are currently being committed to will not be pruned unless the `force` option is used. Atomic queue items will not be affected, unless all devices in it are pruned. The `force` option will brutally kill an ongoing commit. This could leave the device in a bad state. It is not recommended in any normal use case.
10. `devices commit-queue queue-item [ id ] set-atomic-behaviour atomic [ true,false ]`. This action sets the atomic behavior of this queue item. If this is set to false, the devices contained in this queue item can start executing if the same devices in other non-atomic queue items ahead of it in the queue are completed. If set to true, the atomic integrity of the queue item is preserved.
11. `devices commit-queue queue-item [ id ] wait-until-completed`. This action waits until the queue item is completed. The default is to wait `infinity`. A `timeout` can be specified to wait for a number of seconds. The result is `completed` if the queue item is completed or `timeout` if the timer expired before the queue item was completed.
12. `devices commit-queue queue-item [ id ] retry`. This action retries devices with transient errors instead of waiting for the automatic retry attempt. The `device` option will let you specify the devices to retry.

A typical use scenario is where one or more devices are not operational. In the example above (Viewing Queue Items), there are two queue items, waiting for the device `ce0` to come alive. `ce0` is listed as a transient error, and this is blocking the entire queue. Whenever a queue item is blocked because another item ahead of it cannot connect to a specific managed device, an alarm is generated:

```cli
ncs# show alarms alarm-list alarm ce0 commit-through-queue-blocked
alarms alarm-list alarm ce0 commit-through-queue-blocked /devices/device[name='ce0'] 9494530158
 is-cleared              false
 last-status-change      2015-02-09T16:48:17.915+00:00
 last-perceived-severity warning
 last-alarm-text         "Commit queue item 9494530158 is blocked because item 9494446997 cannot connect to ce0"
 status-change 2015-02-09T16:48:17.915+00:00
  received-time      2015-02-09T16:48:17.915+00:00
  perceived-severity warning
  alarm-text         "Commit queue item 9494530158 is blocked because item 9494446997 cannot connect to ce0"
```

1. Block other affecting device `ce0` from entering the commit queue:

   ```cli
   ncs(config)# devices commit-queue add-lock device [ ce0 ] block-others
   commit-queue-id 9577950918
   ncs# show devices commit-queue | notab
   devices commit-queue queue-item 9494446997
    age       1444
    status    executing
    devices   [ ce0 ce1 ce2 ]
    transient ce0
     reason "Failed to connect to device ce0: connection refused"
    is-atomic true
   devices commit-queue queue-item 9494530158
    age         1361
    status      blocked
    devices     [ ce0 ]
    waiting-for [ ce0 ]
    is-atomic   true
   devices commit-queue queue-item 9577950918
    age         55
    status      locked
    devices     [ ce0 ]
    waiting-for [ ce0 ]
    is-atomic   true
   ```

   \
   Now queue item `9577950918` is blocking other items using `ce0` from entering the queue.
2. Prune the usage of the device `ce0` from all queue items in the commit queue:

   ```cli
   ncs(config)# devices commit-queue set-atomic-behaviour atomic false
   ncs(config)# devices commit-queue prune device [ ce0 ]
   num-affected-queue-items 2
   num-deleted-queue-items 1
   ncs(config)# show devices commit-queue | notab
   devices commit-queue queue-item 9577950918
    age              102
    status           locked
    kilo-bytes-size  1
    devices          [ ce0 ]
    is-atomic        true
   ```

   \
   The lock will be in the queue until it has been deleted or unlocked. Queue items affecting other devices are still allowed to enter the queue.
3. Fix the problem with the device `ce0`, remove the lock item and sync from the device:

   ```cli
   ncs(config)# devices commit-queue queue-item 9577950918 delete
   ncs(config)# devices device ce0 sync-from
   result true
   ```

### Commit Queue in a Cluster Environment <a href="#d5e3826" id="d5e3826"></a>

In an LSA cluster, each remote NSO has its own commit queue. When committing through the commit queue on the upper node NSO will automatically create queue items on the lower nodes where the devices in the transaction reside. The progress of the lower node queue items is monitored through a queue item on the upper node. The remote NSO is treated as a device in the queue item and the remote queue items and devices are opaque to the user of the upper node.

{% code title="Example: Commit Queue in an LSA Cluster" %}

```cli
ncs(config)# show configuration
vpn l3vpn volvo
 as-number 65101
 endpoint branch-office1
  ce-device    ce1
  ce-interface GigabitEthernet0/11
  ip-network   10.7.7.0/24
  bandwidth    6000000
 !
 endpoint main-office
  ce-device    ce0
  ce-interface GigabitEthernet0/11
  ip-network   10.10.1.0/24
  bandwidth    12000000
 !
!

ncs(config-if)# commit commit-queue async
commit-queue-id 9494530158

ncs# show devices commit-queue | notab
devices commit-queue queue-item 9494446997
 age       60
 status    executing
 devices   [ lsa-nso2 lsa-nso3 ]
 is-atomic true

ncs# show devices commit-queue | notab
devices commit-queue queue-item 9494446997
 age       66
 status    executing
 devices   [ lsa-nso2 ]
 completed [ lsa-nso3 ]
 is-atomic true

ncs# show devices commit-queue
% No entries found.
```

{% endcode %}

{% hint style="danger" %}
Generally, it is not recommended to interfere with the queue items of the lower nodes that have been created by an upper NSO. This can cause the upper queue item to not synchronize with the lower ones correctly.
{% endhint %}

### Configuring Commit Queue in a Cluster Environment <a href="#d5e3839" id="d5e3839"></a>

To be able to track the commit queue on the lower cluster nodes, NSO uses the built-in stream `ncs-events` that generates northbound notifications for internal events. This stream is required if running the commit queue in a clustered scenario. It is enabled in `ncs.conf`:

{% code title="Example: Enabling the ncs-events Stream" %}

```xml
<stream>
  <name>ncs-events</name>
  <description>NCS event according to tailf-ncs-devices.yang</description>
  <replay-support>true</replay-support>
  <builtin-replay-store>
    <enabled>true</enabled>
    <dir>./state</dir>
    <max-size>S10M</max-size>
    <max-files>50</max-files>
  </builtin-replay-store>
</stream>
```

{% endcode %}

In addition, the commit queue needs to be enabled in the cluster configuration.

```cli
ncs(config)# cluster commit-queue enabled
ncs(config)# commit
```

For more detailed information on how to set up clustering, see [LSA Overview](/guides/administration/advanced-topics/layered-service-architecture).

### Error Recovery with Commit Queue <a href="#d5e3852" id="d5e3852"></a>

The goal of the commit queue is to increase the transactional throughput of NSO while keeping an eventual consistency view of the database. This means no matter if changes committed through the commit queue originate as pure device changes or as the effect of service manipulations the effects on the network should eventually be the same as if performed without a commit queue no matter if they succeed or not. This should apply to a single NSO node as well as NSO nodes in an LSA cluster.

Depending on the selected `error-option` NSO will store the reverse of the original transaction to be able to undo the transaction changes and get back to the previous state. This data is stored in the `/ncs:devices/commit-queue/completed` tree from where it can be viewed and invoked with the `rollback` action. When invoked the data will be removed.

{% code title="Example: Viewing Completed Queue items" %}

```cli
ncs# show devices commit-queue completed | notab
devices commit-queue completed queue-item 9494446997
 when      2015-02-09T16:48:17.915+00:00
 succeeded false
 devices   [ ce0 ce1 ce2 ]
 failed ce0
  reason "Failed to connect to device ce0: closed"
devices commit-queue completed queue-item 9494530158
 when      2015-02-09T16:48:17.915+00:00
 succeeded false
 devices   [ ce0 ]
 failed ce0
  reason "Deleted by user"
```

{% endcode %}

The error option can be configured under `/ncs:devices/global-settings/commit-queue/error-option`. Possible values are: `continue-on-error`, `rollback-on-error`*,* and `stop-on-error`. The `continue-on-error` value means that the commit queue will continue on errors. No rollback data will be created. The `rollback-on-error` value means that the commit queue item will roll back on errors. The commit queue will place a lock on the failed queue item, thus blocking other queue items with overlapping devices from being executed. The `rollback` action will then automatically be invoked when the queue item has finished its execution. The lock will be removed as part of the rollback. The `stop-on-error` means that the commit queue will place a lock on the failed queue item, thus blocking other queue items with overlapping devices from being executed. The lock must then either manually be released when the error is fixed or the `rollback` action under `/devices/commit-queue/completed` be invoked. The `rollback` action is as:

{% code title="Example: Execute Rollback Action" %}

```cli
ncs(config)# devices commit-queue completed queue-item 9494446997 rollback
```

{% endcode %}

The error option can also be given as a commit parameter through the shared `commit-queue/error-option` model described in [Commit Parameters](/guides/operation-and-usage/operations/lifecycle-operations#d5e5048).

{% hint style="info" %}
To guarantee service integrity NSO checks for overlapping service or device modifications against the items in the commit queue and returns an error if such exists. If a service instance does a shared set on the same data as a service instance in the queue actually changed, the reference count will be increased but no actual change is pushed to the device(s). This will give a false positive that the change is actually deployed in the network. The `rollback-on-error` and `stop-on-error` error options will automatically create a queue lock on the involved services and devices to prevent such a case.
{% endhint %}

In a clustered environment, different parts of the resulting configuration change set will end up on different lower nodes. This means on some nodes the queue item could succeed and on others, it could not.

The error option in a cluster environment will originate on the upper node. The reverse of the original transaction will be committed on this node and propagated through the cluster down to the lower nodes. The net effect of this is the state of the network will be the same as before the original change.

{% hint style="info" %}
As the error option in a cluster environment will originate on the upper node, any configuration on the lower nodes will be meaningless.
{% endhint %}

When NSO is recovering from a failed commit, the rollback data of the failed queue items in the cluster is applied and committed through the commit queue. In the rollback, the no-networking flag will be set on the commits towards the failed lower nodes or devices to get CDB consistent with the network. Towards the successful nodes or devices, the commit is done as before. This is what the `rollback` action in `/ncs:devices/commit-queue/completed/queue-item` does.

<div data-with-frame="true"><figure><img src="/files/2zq6LCR4kyOGYAG7U5Rn" alt="" width="563"><figcaption><p>Error Recovery in a Single Node Deployment</p></figcaption></figure></div>

1. TR1; service `s1` creates `ce0:a` and `ce1:b`. The nodes `a` and `b` are created in CDB. In the changes of the queue item, `CQ1`, `a` and `b` are created.
2. TR2; service `s2` creates `ce1:c` and `ce2:d`. The nodes `c` and `d` are created in CDB. In the changes of the queue item, `CQ2`, `c`*,* and `d` are created.
3. The queue item from `TR1`, `CQ1`, starts to execute. The node `a` cannot be created on the device. The node `b` was created on the device but that change is reverted as `a` failed to be created.

<div data-with-frame="true"><figure><img src="/files/ElhcOfEYtznrYA8Q6JZm" alt="" width="563"><figcaption></figcaption></figure></div>

4. The reverse of `TR1`, the rollback of `CQ1`, `TR3`, is committed.
5. `TR3`; service `s1` is applied with the old parameters. Thus the effect of `TR1` is reverted. Nothing needs to be pushed towards the network, so no queue item is created.
6. `TR2`; as the queue item from `TR2`, `CQ2`, is not the same service instance and has no overlapping data on the `ce1` device, this queue item executes as normal.

<div data-with-frame="true"><figure><img src="/files/yNtx1c2wwyh8ohNEmXWI" alt="" width="563"><figcaption><p>Error Recovery in an LSA Cluster</p></figcaption></figure></div>

1. `NSO1`:`TR1`; service `s1` dispatches the service to `NSO2` and `NSO3` through the queue item `NSO1`:`CQ1`. In the changes of `NSO1`:`CQ1`, `NSO2:s1` and `NSO3:s1` are created.
2. `NSO1`:`TR2`; service `s2` dispatches the service to `NSO2` through the queue item `NSO1`:`CQ2`. In the changes of `NSO1`:`CQ2`, `NSO2:s2` is created.
3. The queue item from `NSO2`:`TR1`, `NSO2`:`CQ1`, starts to execute. The node `a` cannot be created on the device. The node `b` was created on the device, but that change is reverted as `a` failed to be created.
4. The queue item from `NSO3`:`TR1`, `NSO3`:`CQ1`, starts to execute. The changes in the queue item are committed successfully to the network.

<div data-with-frame="true"><figure><img src="/files/zSoHKj9spwkIPDzc8wa6" alt="" width="563"><figcaption></figcaption></figure></div>

5. The reverse of `TR1`, rollback of `CQ1`, `TR3`, is committed on all nodes part of `TR1` that failed.
6. `NSO2`:`TR3`; service `s1` is applied with the old parameters. Thus the effect of `NSO2`:`TR1` is reverted. Nothing needs to be pushed towards the network, so no queue item is created.
7. `NSO1`:`TR3`; service `s1` is applied with the old parameters. Thus the effect of `NSO1`:`TR1` is reverted. A queue item is created to push the transaction changes to the lower nodes that didn't fail.
8. `NSO3`:`TR3`; service `s1` is applied with the old parameters. Thus the effect of `NSO3`:`TR1` is reverted. Since the changes in the queue item `NSO3`:`CQ1` was successfully committed to the network a new queue item `NSO3`:`CQ3` is created to revert those changes.

If for some reason the rollback transaction fails there are, depending on the failure, different techniques to reconcile the services involved:

* Make sure that the commit queue is blocked to not interfere with the error recovery procedure. Do a sync-from on the non-completed device(s) and then re-deploy the failed service(s) with the `reconcile` option to reconcile original data, i.e., take control of that data. This option acknowledges other services controlling the same data. The reference count will indicate how many services control the data. Release any queue lock that was created.
* Make sure that the commit queue is blocked to not interfere with the error recovery procedure. Use un-deploy with the no-networking option on the service and then do sync-from on the non-completed device(s). Make sure the error is fixed and then re-deploy the failed service(s) with the `reconcile` option. Release any queue lock that was created.

### Commit Queue Tuning <a href="#d5e3976" id="d5e3976"></a>

As the goal of the commit queue is to increase the transactional throughput of NSO, it means that we need to calculate the configuration change towards the device(s) outside of the transaction lock. To calculate a configuration change, NSO needs a pre-commit running and a running view of the database. The key enabler to support this in the commit queue is to allow different views of the database to live beyond the commit. In NSO, this is implemented by keeping a snapshot database of the configuration tree for devices and storing configuration changes towards this snapshot database on a per-device basis. The snapshot database is updated when a device in the queue has been processed. This snapshot database is stored on disk for persistence (the CDB `S` files in the `ncs-cdb` directory).

The snapshot database can be populated either eagerly or lazily. This is controlled by the `/ncs-config/cdb/snapshot/pre-populate` setting in `ncs.conf`.

* When `pre-populate` is `true`, NSO pre-populates the snapshot during CDB upgrade. This mode is optimized for systems that use the commit queue extensively. The benefit is better commit queue performance with no extra first-commit penalty per device. The trade-offs are longer upgrade time and almost twofold memory consumption.
* When `pre-populate` is `false`, NSO is optimized for the default transaction behavior and uses lazy population. If there is no existing snapshot, each device snapshot is created when that device is committed through the commit queue for the first time. That first commit can be noticeably slower, especially for large device configurations. Subsequent commit queue commits on the same device do not have this penalty.

After a snapshot has been populated, NSO does not re-populate it on every startup or commit. Re-population is done only when there is no existing snapshot, for example if `S.cdb` has been removed or re-initialized. If the snapshot is missing, NSO falls back to populating it on demand when a device is committed through the commit queue.

## NETCONF Call Home <a href="#d5e3986" id="d5e3986"></a>

The NSO device manager has built-in support for the NETCONF Call Home client protocol operations over SSH as defined in [RFC 8071](https://www.ietf.org/rfc/rfc8071.txt).

With NETCONF SSH Call Home, the NETCONF client listens for TCP connection requests from NETCONF servers. The SSH client protocol is started when the connection is accepted. The SSH client validates the server's presented host key with credentials stored in NSO. If no matching host key is found the TCP connection is closed immediately. Otherwise, the SSH connection is established, and NSO is enabled to communicate with the device. The SSH connection is kept open until the device itself terminates the connection, an NSO user disconnects the device, or the idle connection timeout is triggered (configurable in the `ncs.conf` file).

NSO will generate an asynchronous notification event whenever there is a connection request. An application can subscribe to these events and, for example, add an unknown device to the device tree with the information provided, or invoke actions on the device if it is known.

If an SSH connection is established, any outstanding configuration in the commit queue for the device will be pushed. Any notification stream for the device will also be reconnected.

NETCONF Call Home is enabled and configured under `/ncs-config/netconf-call-home` in the `ncs.conf` file. By default NETCONF Call Home is disabled.

A device can be connected through the NETCONF Call Home client only if `/devices/device/state/admin-state` is set to `call-home`. This state prevents any southbound communication to the device unless the connection has already been established through the NETCONF Call Home client protocol.

See [examples.ncs/device-management/netconf-call-home](https://github.com/NSO-developer/nso-examples/tree/6.7/device-management/netconf-call-home) for an example.

## Notifications <a href="#d5e4000" id="d5e4000"></a>

The NSO device manager has built-in support for device notifications. Notifications are a means for the managed devices to send structured data asynchronously to the manager. NSO has native support for NETCONF event notifications (see RFC 5277) but could also receive notifications from other protocols implemented by the Network Element Drivers.

Notifications can be utilized in various use-case scenarios. It can be used to populate alarms in the Alarm manager, collect certain types of errors over time, build a network-wide audit log, react to configuration changes, etc.

The basic mode of operation is the manager subscribes to one or more *named* notification channels which are announced by the managed device. The manager keeps an open SSH channel towards the managed device, and then, the managed device may asynchronously send structured XML data on the SSH channel.

The notification support in NSO is usable as is without any further programming. However, NSO cannot understand any semantics contained inside the received XML messages, thus for example a notification with a content of "Clear Alarm 456" cannot be processed by NSO without any additional programming.

When you add programs to interpret and act upon notifications, make sure that resulting operations are idempotent. This means that they should be able to be called any number of times while guaranteeing that side effects only occur once. The reason for this is that, for example, replaying notifications can sometimes mean that your program will handle the same notifications multiple times.

In the `tailf-ncs.yang` data model, you find a YANG data model that can be used to:

* Setup subscriptions. A subscription is configuration data from the point of view of NSO, thus if NSO is restarted, all configured subscriptions are automatically resumed.
* Inspect which named streams a managed device publishes.
* View all received notifications.

{% hint style="info" %}
Notifications must be defined at the top level of a YANG module. NSO does currently not support defining notifications inside lists or containers as specified in section 7.16 in [RFC 7950](https://www.ietf.org/rfc/rfc7950.txt).
{% endhint %}

### An Example Session <a href="#d5e4020" id="d5e4020"></a>

In this section, we will use the [examples.ncs/device-management/web-server-basic](https://github.com/NSO-developer/nso-examples/tree/6.7/device-management/web-server-basic) example.

Let's dive into an example session with the NSO CLI. In the NSO example collection, the webserver publishes two NETCONF notification structures, indicating what they intend to send to any interested listeners. They all have the YANG module:

{% code title="Example: notif.yang" %}

```yang
module notif {
  namespace "http://router.com/notif";
  prefix notif;

  import ietf-inet-types {
    prefix inet;
  }


  notification startUp {
    leaf node-id {
      type string;
    }
  }

  notification linkUp {
    leaf ifName {
      type string;
      mandatory true;
    }
    leaf extraId {
      type string;
    }
    list linkProperty {
      max-elements 64;
      leaf newlyAdded {
        type empty;
      }
      leaf flags {
        type uint32;
        default 0;
      }
      list extensions {
        max-elements 64;
        leaf name {
          type uint32;
          mandatory true;
        }
        leaf value {
          type uint32;
          mandatory true;
        }
      }
    }

    list address {
      key ip;
      leaf ip {
        type inet:ipv4-address;
      }
      leaf mask {
        type inet:ipv4-address;
      }
    }

    leaf-list iface-flags {
      type enumeration {
        enum UP;
        enum DOWN;
        enum BROADCAST;
        enum RUNNING;
        enum MULTICAST;
        enum LOOPBACK;
      }
    }
  }


  notification linkDown {
    leaf ifName {
      type string;
      mandatory true;
    }
  }
}
```

{% endcode %}

Follow the instructions in the README file if you want to run the example: build the example, start netsim, and start NCS.

```cli
admin@ncs# show devices device pe2 notifications stream | notab
notifications stream NETCONF
 description    "default NETCONF event stream"
 replay-support false
notifications stream tailf-audit
 description    "Tailf Commit Audit events"
 replay-support true
notifications stream interface
 description              "Example notifications"
 replay-support           true
 replay-log-creation-time 2014-10-14T11:21:12+00:00
 replay-log-aged-time     2014-10-14T11:53:19.649207+00:00
```

The above shows how we can inspect - as status data - which named streams the managed device publishes. Each stream also has some associated data. The data model for that looks like this:

{% code title="Example: tailf-ncs.yang Notification Streams" %}

```yang
module tailf-ncs {
  namespace "http://tail-f.com/ns/ncs";
  ...
  container devices {
     list device {
       ....
       container notifications {
          ....

          list stream {
             description "A list of the notification streams
                          provided by the device. NCS reads this list in
                          real time";

             config false;
             key name;
             leaf name {
               description "The name of the the stream";
               type string;
             }
             leaf description {
               description "A textual description of the stream";
               type string;
             }
             leaf replay-support {
               description "An indication of whether or not event replay
                            is available on this stream.";
               type boolean;
             }
             leaf replay-log-creation-time {
               description "The timestamp of the creation of the log
                           used to support the replay function on
                           this stream.
                           Note that this might be earlier then
                           the earliest available
                           notification in the log.  This object
                           is updated if the log resets
                           for some reason.";

               type yang:date-and-time;
             }
             leaf replay-log-aged-time {
               description "The timestamp of the last notification
                            aged out of the log";
               type yang:date-and-time;
             }
           }
```

{% endcode %}

Let's set up a subscription for the stream called `interface`. The subscriptions are NSO configuration data, thus to create a subscription we need to enter configuration mode:

{% code title="Example: Configuring a Subscription" %}

```cli
admin@ncs(config)# devices device www0..2 notifications \
      subscription mysub stream interface
admin@ncs(config-subscription-mysub)# commit
```

{% endcode %}

The above example created subscriptions for the `interface` stream on all web servers, i.e. managed devices, `www0`, `www1`, and `www2`. Each subscription must have an associated stream to it, this is however not the key for an NSO notification, the key is a free-form text string. This is because we can have multiple subscriptions to the same stream. More on this later when we describe the filter that can be associated with a subscription. Once the notifications start to arrive, they are read by NSO and stored in stable storage as CDB operational data. they are stored under each managed device - and we can view them as:

{% code title="Example: Viewing the Received Notifications" %}

```cli
admin@ncs# show devices device notifications | notab
devices device www0
 notifications subscription mysub
  local-user admin
  status     running
 notifications stream NETCONF
  description    "default NETCONF event stream"
  replay-support false
 notifications stream tailf-audit
  description    "Tailf Commit Audit events"
  replay-support true
 notifications stream interface
  description              "Example notifications"
  replay-support           true
  replay-log-creation-time 2014-10-14T11:21:12+00:00
  replay-log-aged-time     2014-10-14T11:56:45.755964+00:00
 notifications notification-name startUp
  uri http://router.com/notif
 notifications notification-name linkUp
  uri http://router.com/notif
 notifications notification-name linkDown
  uri http://router.com/notif
 notifications received-notifications notification 2014-10-14T11:54:43.692371+00:00 0
  user          admin
  subscription  mysub
  stream        interface
  received-time 2014-10-14T11:54:43.695191+00:00
  data linkUp ifName eth2
  data linkUp linkProperty
   newlyAdded
   flags      42
   extensions
    name  1
    value 3
   extensions
    name  2
    value 4668
  data linkUp address 192.168.128.55
   mask 255.255.255.0
```

{% endcode %}

Each received notification has some associated metadata, such as the time the event was received by NSO, which subscription and which stream is associated with the notification, and also which user created the subscription.

It is fairly instructive to inspect the XML that goes on the wire when we create a subscription and then also receive the first notification. We can do:

```cli
ncs(config)# devices global-settings trace pretty trace-dir ./logs
ncs(config)# commit

ncs(config)# devices disconnect

ncs(config)# devices device pe2 notifications \
     subscription foo stream interface
ncs(config-subscription-foo)# top
ncs(config)# exit

ncs# file show ./logs/netconf-pe2.trace
<<<<in 14-Oct-2014::13:59:52.295 device=pe2 session-id=14
<notification xmlns="urn:ietf:params:xml:ns:netconf:notification:1.0">
  <eventTime>2014-10-14T11:58:51.816077+00:00</eventTime>
  <linkUp xmlns="http://router.com/notif">
    <ifName>eth2</ifName>
    <linkProperty>
      <newlyAdded/>
      <flags>42</flags>
      <extensions>
        <name>1</name>
        <value>3</value>
      </extensions>
      <extensions>
        <name>2</name>
        <value>4668</value>
      </extensions>
    </linkProperty>
    <address>
      <ip>192.168.128.55</ip>
      <mask>255.255.255.0</mask>
    </address>
  </linkUp>
</notification>
 .........
```

Thus, once the subscription has been configured, NSO continuously receives, and stores in CDB oper persistent storage, the notifications sent from the managed device. The notifications are stored in a circular buffer, to set the size of the buffer, we can do:

```cli
ncs(config)# devices device www0 notifications \
   received-notifications max-size 100
admin@ncs(config-device-www0)# commit
```

The default value is 200. Once the size of the circular buffer is exceeded, the older notification is removed.

### Subscription Status <a href="#d5e4063" id="d5e4063"></a>

A running subscription can be in either of three states. The YANG model has:

```yang
module tailf-ncs {
  namespace "http://tail-f.com/ns/ncs";
  ...
  container devices {
     list device {
       ....
       container notifications {
          ....
          list subscription {
             .....
            leaf status {
            description "Is this subscription currently running";
            config false;
            type enumeration {
              enum running {
                description "The subscription is established and we should
                             be receiving notifications";
              }
              enum connecting {
                description "Attempting to establish the subscription";
              }
              enum failed {
                description
                "The subscription has failed, unless the failure is
                 in the connection establishing, i.e connect() failed
                 there will be no automatic re-connect";
              }
            }
          }
```

If a subscription is in the *failed* state, an optional *failure-reason* field indicates the reason for the failure. If a subscription fails due to, not being able to connect to the managed device or if the managed device closed its end of the SSH socket, NSO will attempt to automatically reconnect. The re-connect attempt interval is configurable.

```cli
ncs# show devices device notifications subscription
             LOCAL           FAILURE  ERROR
NAME  NAME   USER   STATUS   REASON   INFO
---------------------------------------------
www0  foo    admin  running  -        -
      mysub  admin  running  -        -
www1  mysub  admin  running  -        -
www2  mysub  admin  running  -        -
```

## SNMP Notifications <a href="#d5e4074" id="d5e4074"></a>

SNMPv3 `auth-priv` notifications can be received by NSO and acted upon. The SNMP receiver is a stand-alone process and by default, all notifications are ignored. IP addresses must be opted in and a handler must be defined to take actions on certain notifications. This can be used to for example listen to configuration change notifications and trigger a log action or a resync for example

These actions are programmed in Java, see the [SNMP Notification Receiver](/guides/development/connected-topics/snmp-notification-receiver) for how to do this.

## Southbound Datastore Subscriptions

The NSO device manager has built-in support for southbound datastore subscriptions. It is a means for the managed devices to send datastore updates asynchronously to the manager. NSO has native support for NETCONF YANG-Push ([RFC 8641](https://www.rfc-editor.org/rfc/rfc8641.html)) but could also receive data from other telemetry protocols implemented by the Network Element Drivers.

Telemetry data can be utilized to drive provisioning scenarios. In particular, it can be used to keep devices in sync, handle out-of-band changes automatically, or react to (configuration) changes in general. The telemetry support in NSO is designed for such use cases, not as a general purpose telemetry collector.

The basic mode of operation is where the manager subscribes to updates to a device datastore. The manager keeps an open SSH channel towards the managed device, and then, the managed device may asynchronously send structured XML data on the SSH channel.

The telemetry support in NSO allows NSO to subscribe to and then receive updates out of the box. However, additional programming is necessary for NSO to understand and react to the received datastore updates. See [Telemetry Kickers](/guides/development/advanced-development/kicker#telemetry-kicker-concepts) on how to process the data.

In the `tailf-ncs.yang` data model, you find a YANG data model that can be used to:

* Setup telemetry subscriptions. A subscription is configuration data from the point of view of NSO, thus if NSO is restarted, all configured subscriptions are automatically resumed.
* View the status of all configured telemetry subscriptions.

When setting up the subscriptions, be mindful of the fact that every subscription likely requires an active device connection. Too many subscriptions, or subscriptions to the whole configuration/operational data store could overwhelm NSO or the device. It is best to use telemetry for specific parts of the data tree and for specific devices only.

Additionally, there is a short time window after a subscription is configured and before it is running. For example, configuring a subscription on a device, and changing configuration of this device in the same transaction, may or may not trigger the new subscription right away.

The [examples.ncs/service-management/implement-a-service/iface-v6](https://github.com/NSO-developer/nso-examples/tree/6.7/service-management/implement-a-service/iface-v6) example configures a subscription to device link operational state changes. It also allows you to inspect the datastore update messages generated by the device.

## Inactive Configuration <a href="#d5e4079" id="d5e4079"></a>

NSO can configure inactive parameters on the devices that support inactive configuration. Currently, these devices include Juniper devices and devices that announce `http://tail-f.com/ns/netconf/inactive/1.0` capability. NSO itself implements `http://tail-f.com/ns/netconf/inactive/1.0` capability which is formally defined in `tailf-netconf-inactive` YANG module.

To recap, a node that is marked as inactive exists in the data store but is not used by the server. The nodes announced as inactive by the device will also be inactive in the device's configuration in NSO, and activating/deactivating a node in NSO will push the corresponding change to the device. This also means that for NSO to be able to manage inactive configuration, both `/ncs-config/enable-inactive` and `/ncs-config/netconf-north-bound/capabilities/inactive` need to be enabled in `ncs.conf`.

If the inactive feature is disabled in `ncs.conf`, NSO will still be able to manage devices that have inactive configuration in their datastore, but the inactive attribute will be ignored, so the data will appear as active in NSO and it would not be possible for NSO to activate/deactivate such nodes in the device.


# Out-of-band Interoperation

Manage out-of-band changes.

The preferred way of making changes in the network is to perform all changes through NSO, which keeps the NSO copy of device configurations up-to-date (in sync) at all times. This approach has many benefits, as it allows NSO to:

* Avoid making provisioning decisions based on stale data
* Provide a single pane of glass to network configuration
* Act as a network source of truth
* Better aid in troubleshooting scenarios
* Provide improved performance, and
* Expose advanced compliance and reporting capabilities

However, in some situations, such setup is undesirable or not possible due to historic, organizational, or other reasons. While an organization may decide to forgo most of these benefits by managing the network through multiple systems, it is essential for NSO provisioning code to work with current data.

<div data-with-frame="true"><img src="/files/qY2rG6IBIHGnwUC74Ek0" alt="Out-of-band Changes" width="375"></div>

To better allow coexistence with other systems and processes that manage the same devices, NSO 6.5 introduces an innovative, patent-pending approach to the so-called "out-of-band" changes. Out-of-band changes are changes to NSO-managed devices not done through NSO. From a high-level perspective, this approach consists of:

* "Ships passing in the night" handling of configuration not relevant to NSO-managed parts
* Verification of data used in provisioning decisions prior to being pushed out to the network, and
* Policy-based retention of changes by other systems and agents on NSO-managed configuration

It now becomes possible to manage a network device by never doing a sync-from/sync-to operation (in practice the first sync-from may still be desirable to allow reading from NSO). At the same time, special-purpose pre-provisioning checks become unnecessary for the majority of cases, as NSO verifies the correctness of data used in the transaction.

<div data-with-frame="true"><img src="/files/l1JQ1MXwaDwv94SOUBEu" alt="Handling Out-of-band Changes" width="375"></div>

Such an approach allows NSO to use targeted correctness checks that have another benefit when used with devices which have huge configurations, such as various controllers. If only small parts of the configuration are relevant to NSO, the checks can be optimized. Limiting the checks to only the required parts allows the system to scale with the extent of the change, not the size or time-complexity of producing the full device configurations.

## Introducing `confirm-network-state`

Handling out-of-band changes requires NSO to make additional checks and perform additional processing when provisioning network changes, so this functionality is opt-in. The first option to invoke the out-of-band processing machinery is to use the `commit confirm-network-state` commit variant, which takes effect for the current commit only.

{% hint style="info" %}
`confirm-network-state` validates device state during commit and therefore requires device reachability. NSO must be able to discover the device capabilities (either already known, or fetched by connecting to the device as part of the operation). A prior `sync-from` is not required, but in brownfield deployments it is recommended (full or partial) to seed CDB with baseline data and reduce excessive/false out-of-band markings.
{% endhint %}

This option is great for testing out different scenarios and getting familiar with the out-of-band features of NSO. In addition to `commit`, there are other commands that can also be `confirm-network-state` enabled, such as device `sync-from` and service `re-deploy`.

However, the recommended way for normal, day-to-day use is to enable`confirm-network-state` for a set of devices through device settings. For example:

```bash
admin@ncs(config)# devices device c1 confirm-network-state enabled-by-default true
```

Or:

```bash
admin@ncs(config)# devices profiles profile prod confirm-network-state enabled-by-default true
```

Or:

```bash
admin@ncs(config)# devices global-settings confirm-network-state enabled-by-default true
```

Commit and other operations then no longer require using the`confirm-network-state` option explicitly; it is enabled automatically for those devices.

By default, `confirm-network-state` now keeps the resulting service impact scoped to the current transaction. If NSO discovers out-of-band data that affects other services, those additional services are not re-deployed automatically unless you opt in with `re-deploy-all`. You can use `re-deploy-all` as part of the commit parameters, configure it under device `confirm-network-state` settings to make it the default for selected devices, or pass it to `confirm-network-state`-enabled device actions such as `sync-from` and `partial-sync-from`.

Once NSO uses `confirm-network-state` for a device change, it no longer checks device sync status, so the commit may go through even if parts of device configuration are out-of-sync. To find out if the device configuration is out-of-sync before committing, use `dry-run` together with `confirm-network-state`.

NSO keeps track of all reads in a given transaction and then verifies that these values (which were presumably used to influence the provisioning decisions) remain the same on the device. Behind the scenes, this mechanism uses the same transaction read-set that is also used for [concurrency checks](/guides/development/core-concepts/nso-concurrency-model).

For example, let's say you want to set interface MTU to at least 1520 with a Python script:

```python
import ncs
with ncs.maapi.single_write_trans('admin', 'python') as t:
    root = ncs.maagic.get_root(t)
    intf = root.devices.device['c1'].config.interface.GigabitEthernet['0/1']
    if intf.mtu is None or intf.mtu < 1520:
        intf.mtu = 1520
    params = t.get_params()
    params.confirm_network_state()
    t.apply_params(True, params)
```

Using the `confirm-network-state`-enabled commit ensures that the script does not overwrite values previously set out of band:

```bash
$ python3 update-mtu.py || echo 'inconsistency detected!'
...
inconsistency detected!
$ ncs_cli -Cu admin
admin@ncs# devices device c1 compare-config
diff
 devices {
     device c1 {
         config {
             interface {
                 GigabitEthernet 0/1 {
+                    mtu 9000;
                 }
             }
         }
     }
 }
```

Note that the failure of the script in this scenario is expected; the script should retry the operation once the out-of-band changes are inspected and either accepted with a (partial) sync-from, or rejected with a sync-to. If, instead, this was service code, NSO would retry the operation automatically with the updated data.

If the commit is successful, it includes the out-of-band changes that NSO found. For example, setting a value may trigger validating a YANG `must` expression, requiring NSO to read additional configuration from the device for the purpose of verification. If this configuration has changed out of band, NSO will validate the commit with the new data (the `must` expression must be satisfied) and include the change with the commit.

Including the out-of-band changes with the commit allows you to revert the whole operation, if necessary, and ensures CDB consistency. After commit, the CDB contains the updated configuration for the parts that affected provisioning, while safely ignoring other out-of-band changes.

It also enables you to preview the out-band-changes you are bringing in as part of the `commit dry-run`, illustrated in the following output.

```bash
admin@ncs(config)# no devices device c1 config interface GigabitEthernet 0/1\
 ip dhcp snooping trust
admin@ncs(config)# commit dry-run outformat cli-c confirm-network-state
...
        confirm-network-state {
            device {
                name c1
                out-of-band devices device c1
                             config
                              interface GigabitEthernet0/1
                               mtu 9000
                              exit
                             !
                            !
                data devices device c1
                      config
                       interface GigabitEthernet0/1
                        no ip dhcp snooping trust
                       exit
                      !
                     !
            }
        }
```

The not-overwriting functionality is shared with `commit no-overwrite` and ensures provisioning code in NSO works with up-to-date data. The difference between the two is that `confirm-network-state` also updates the CDB while evaluating service out-of-band policies for the relevant services.

## Service Out-of-band Policies

The `confirm-network-state` mode of operation shows its true power when used in combination with services. Services in NSO, through service mapping code and templates, manage the required network configuration. NSO knows what device configuration belongs to which service through the [backpointer references](/guides/development/advanced-development/developing-services/services-deep-dive) and can therefore detect when out-of-band changes are made to a configuration that belongs to a service.

When NSO detects such a change, the question becomes what to do with it. The answer depends on the service and on the kind of change; some changes need to be accepted and others rejected. The service out-of-band policy specifies how the change is to be handled for a specific case.

The service policy is defined per service type (servicepoint) and contains a set of rules. For example:

```
services out-of-band policy iface-servicepoint
 rule allow-mtu
  path         ios:interface/GigabitEthernet/mtu
  at-create    sync-from-device
  at-delete    sync-from-device
  at-value-set sync-from-device
 !
 rule reject-ip-address
  path         ios:interface/GigabitEthernet/ip/address
  at-create    sync-to-device
  at-delete    sync-to-device
  at-value-set sync-to-device
 !
!
```

Each rule defines an action NSO should take when encountering an out-of-band change at the given device path. The paths in the preceding printout are relative to `/devices/device/config` and tell NSO:

* We allow other systems or operators to change MTU for interfaces managed by the `iface` service; by specifying `sync-from-device`, NSO copies the new device value to CDB. This is a good choice for values that are mostly unrelated to the service and unlikely to break it.
* We reject changes to the IP address on the interface with `sync-to-device`, making NSO revert the change from the other system back to what is in the CDB (usually generated by the service mapping). This is a good choice for values that are vital for the correct operation of the service.

The example also shows how to differentiate between the type of change (operation); is the changed configuration node newly introduced (`at-create`), removed (`at-delete`), or has it gotten a new value (`at-value-set`)? The`at-create` operation makes little sense for configuration that is provisioned by the service (the configuration obviously already exists) but is useful when additional configuration parameters are introduced under service-created ones. Additionally, `at-create` might be used when a service deletes device configuration which is then introduced back out of band.

Using the type of change allows you to express more complicated policies. For example, suppose the `iface` service really requires just some IP address on the interface, not necessarily the one it initially provisioned. As it does not matter what particular IP address is used, it can be changed out of band, as long as there is one. You can describe this with a rule, such as:

```
 rule reject-no-ip-address
  path         ios:interface/GigabitEthernet/ip/address
  at-delete    sync-to-device
  at-value-set sync-from-device
 !
```

A rule can specify a default action that is used when no operation-specific action has been specified. If a rule contains both a default action and an operation-specific action, then the operation-specific action takes precedence. The following rule is functionally equivalent to the `allow-mtu` rule in the service policy above:

```
 rule allow-mtu
  path           ios:interface/GigabitEthernet/mtu
  default-action sync-from-device
 !
```

<div data-with-frame="true"><img src="/files/fyaGT9bXZgo4QY1sBXfC" alt="Out-of-band Policy" width="563"></div>

This, however, brings up another question: what should happen if you redeploy the service? Should NSO use the service-provided IP or should the out-of-band configured value be used instead? With the `sync-from-device` policy action, NSO overwrites the out-of-band value with the service-provided one. Instead, if the service should keep the out-of-band value, use the `manage-by-service` policy action, for example:

```
 rule reject-no-ip-address
  path         ios:interface/GigabitEthernet/ip/address
  at-delete    sync-to-device
  at-value-set manage-by-service
 !
```

Specifying `manage-by-service` not only updates device configuration in the CDB with the out-of-band value, it also adds the value under service instance's out-of-band changes (also called extra operations). NSO takes these changes into account when calculating service configuration after mapping code runs. It allows the service to preserve an out-of-band value during a redeploy. Additionally, it ties the value to the lifecycle of the service; if the service is deleted, so is the out-of-band configuration.

It may be desirable to abort out-of-band handling entirely and fail the transaction with an out-of-sync error if certain out-of-band changes are detected on a device. This can be achieved using the `abort` action, for example:

```
 rule abort-if-mtu-is-set
  path         ios:interface/GigabitEthernet/mtu
  at-value-set abort
 !
```

The rule above will cause out-of-band handling to be aborted if the `mtu` leaf has been set out-of-band.

### Rule Behavior Example

Consider a setup from [examples.ncs/service-management/confirm-network-state](https://github.com/NSO-developer/nso-examples/tree/6.7/service-management/confirm-network-state), started by `make demo`, with the following out-of-band policy:

```
services out-of-band policy iface-servicepoint
 rule allow-mtu
  path         ios:interface/GigabitEthernet/mtu
  at-create    sync-from-device
  at-delete    sync-from-device
  at-value-set sync-from-device
 !
 rule reject-no-ip-address
  path         ios:interface/GigabitEthernet/ip/address
  at-delete    sync-to-device
  at-value-set manage-by-service
 !
!
```

Initially, the service provides some device configuration:

```bash
admin@ncs# iface instance1 get-modifications outformat cli-c
cli-c {
    local-node {
        data devices device c1
               config
                interface GigabitEthernet0/1
                 ip address 10.1.2.3 255.255.255.240
                exit
               !
              !
    }
}
```

At some later point in time, perhaps after a support call from a customer, a technician changes a number of things either directly on the device, or through some other system:

```bash
admin@ncs# devices device c1 compare-config
diff
 devices {
     device c1 {
         config {
             interface {
                 GigabitEthernet 0/1 {
                     ip {
                         address {
                             primary {
-                                address 10.1.2.3;
-                                mask 255.255.255.240;
                             }
                         }
                     }
+                    mtu 1520;
                 }
                 GigabitEthernet 0/2 {
                     ip {
                         address {
                             primary {
-                                address 10.2.2.3;
+                                address 10.2.2.8;
                             }
                         }
                     }
                 }
             }
         }
     }
 }
```

If you now perform sync-from, the out-of-band policy will get processed and do the following:

* Re-provision the GigabitEthernet0/1 IP address due to rule #2 `at-delete: sync-to-device`.
* Keep the MTU at 1520 due to rule #1 `at-create: sync-from-device`.
* Keep the GigabitEthernet0/2 IP address due to rule #2 `at-value-set: manage-by-service`.
* Tie the GigabitEthernet0/2 IP to the service lifecycle.

To see the last part take effect, you can inspect the service modifications:

```bash
admin@ncs# iface instance2 get-modifications forward { only-out-of-band }
cli {
    local-node {
        data  devices {
                   device c1 {
                       config {
                           interface {
              +                GigabitEthernet 0/2 {
              +                    ip {
              +                        address {
              +                            primary {
              +                                address 10.2.2.8;
              +                            }
              +                        }
              +                    }
              +                }
                           }
                       }
                   }
               }
    }
}
```

The difference between `sync-to-device` and `manage-by-service` is also pronounced when you remove the service:

```bash
admin@ncs(config)# no iface
admin@ncs(config)# commit and-quit
admin@ncs# show running-config devices device c1 config interface GigabitEthernet
devices device c1
 config
  interface GigabitEthernet0/1
   mtu 1520
  exit
 !
!
```

The MTU setting is left behind since it is not tied to the lifecycle of the service. But note that, if the service had initially created the container in which it is configured, it would get removed as well when the container would be removed.

On the other hand, the new IP address for instance2 GigabitEthernet0/2 is gone with the service, since it is tied to the service lifecycle according to the policy.

### Default Policy

The service out-of-band policy is part of the NSO dynamic configuration, allowing an operator to tailor it to their needs. However, a service designer may already foresee some common scenarios where out-of-band handling of data is beneficial and provide a default out-of-band policy for their service.

NSO populates the service point entry under `/services/out-of-band/policy` with the service package defined default policy unless an entry is already present. The operator is then free to change this policy as they see fit. (But note that policy changes take effect after the policy is committed, not during the same transaction.)

To revert back to the default service-provided policy, an operator must delete the whole service point entry from `/services/out-of-band/policy`. Note that this is different from deleting all the rules from policy for a service point, which actually represents an empty policy (effectively `sync-from-device`).

A service developer defines the default policy for their service type in YANG. It has almost the same structure as the policy configuration in NSO, but uses YANG statements and is defined on the top level of a YANG (sub)module. For example:

```yang
module iface-service {
  // ...

  ncs:out-of-band iface-servicepoint {
    ncs:policy {
      ncs:rule "reject-no-ip-address" {
        ncs:path "ios:interface/GigabitEthernet/ip/address";
        ncs:at-delete sync-to-device;
        ncs:at-value-set manage-by-service;
      }
    }
  }
}
```

## Policy Rule Evaluation

NSO processes the out-of-band data with the service policy when:

* NSO performs a device operation that is `confirm-network-state`-enabled (either through the command itself or participating device setting), and
* NSO finds out-of-band data that is related to the service.

This is an optimization that allows NSO to no longer request or process parts of device configuration which are not related to the current operation. To ensure all current out-of-band data for a device is processed, you can invoke a`confirm-network-state`-enabled sync-from for this device, such as:

```bash
admin@ncs# devices device c1 sync-from confirm-network-state
```

When NSO encounters out-of-band data, it checks if this data resides in a part of configuration that is managed by one or more services. If that is the case, NSO uses backpointer references to identify individual service instances and the corresponding servicepoints. For each service, NSO searches the out-of-band policy rules for that servicepoint and handles the change according to specified action.

In particular, NSO compares rules and checks for the best match rule, where:

* Rule's `path` matches node or one of its parents; longer matches are checked first. For example, path `ios:interface/GigabitEthernet/ip` is tested before its parent `ios:interface/GigabitEthernet`.
* Rule must define an action for the type of change (operation) to match. For example, a rule without `at-delete` does not match an out-of-band delete.
* If multiple rules are found, NSO checks their priority value; numerically lower values are matched first.
* If `filter-expr` of a rule is set, it must evaluate true to match.

If no matching rule is found at all, `sync-from-device` is used as a fallback.

For example, consider the following rule set:

```
services out-of-band policy iface-servicepoint
 rule 1-no-delete-address
  path         ios:interface/GigabitEthernet/ip/address
  at-delete    sync-to-device
 !
 rule 2-specific-address
  path         ios:interface/GigabitEthernet/ip/address
  filter-expr  ". = '10.1.1.1'"
  at-create    sync-to-device
  at-delete    sync-to-device
  at-value-set sync-to-device
 !
 rule 3-ip-for-specific-interface
  path         ios:interface/GigabitEthernet[name='0/2']/ip
  priority     2
  at-create    sync-from-device
  at-delete    sync-from-device
  at-value-set sync-from-device
 !
 rule 4-ip
  path         ios:interface/GigabitEthernet/ip
  priority     1
  at-create    manage-by-service
  at-delete    manage-by-service
  at-value-set manage-by-service
 !
!
```

When a device's GigabitEthernet0/2 IP address is changed (value-set), say from 10.2.2.3 to 10.2.2.5, and NSO starts processing the rule set, it selects the`manage-by-service` action for this change because:

* Rule 1 has a matching path but no `at-value-set` action, so it does not match.
* Rule 2 also has a matching path but `filter-expr` does not match.
* Rule 3 matches a parent path with priority 2.
* Rule 4 matches the same parent path with priority 1 and is selected over priority 2 rule.

To get more detailed information about how the running system processes out-of-band changes, you can enable and set the level for`out-of-band-policy-log` in `ncs.conf`.

### Policy Rule Filter Expression

Note that `path` in the policy rule definition is a special variant of YANG`instance-identifier` that may be absolute or relative to`/devices/device/config`. As such, it is limited to selecting data nodes, and predicates can only be used for selecting keys.

On the other hand, `filter-expr` can specify a full XPath 1.0 expression that allows fine-grained selection of which rule applies where. It also supports`filter-expr`-specific extensions to the XPath language in the form of predefined variables and additional functions.

The expression is evaluated with the path of the out-of-band-changed node as the current XPath context and NSO data root as XPath root. Note that the expression operates on the values in the current transaction, that is, values that NSO sees. If you wish to access the changed values, that is "new" device values, you need to use a special XPath function `oob:context()`.

Say an IP address changes from 10.2.2.3 to 10.2.2.8 out-of-band on the device. Then:

<table><thead><tr><th valign="top">Expression</th><th valign="top">Result</th></tr></thead><tbody><tr><td valign="top"><code>.</code></td><td valign="top">10.2.2.3</td></tr><tr><td valign="top"><code>oob:context()</code></td><td valign="top">10.2.2.8</td></tr></tbody></table>

Also note that `.` refers to the currently-processing changed node, which may be a sub-node of the rule's `path` value, for example it could be`ip/address/primary/address` even though `path` points to `ip/address`.

Additional variables supported by the filter expression:

<table><thead><tr><th valign="top">Variable</th><th valign="top">Description</th></tr></thead><tbody><tr><td valign="top"><code>SERVICE</code></td><td valign="top">Path to the service instance the rule is evaluating for. Example use: <code>$SERVICE/name = 'instance1'</code>.</td></tr><tr><td valign="top"><code>RULE_PATH</code></td><td valign="top"><code>path</code> value of the rule that is evaluating.</td></tr></tbody></table>

Additional functions supported by the filter expression:

<table><thead><tr><th valign="top">Function</th><th valign="top">Description</th></tr></thead><tbody><tr><td valign="top"><code>oob:is-leaf([nodeset])</code></td><td valign="top">Check if specified nodes, or the current node when <em><code>nodeset</code></em> is not specified, are leaves. Returns boolean.</td></tr><tr><td valign="top"><code>oob:is-service-data([nodeset])</code></td><td valign="top">Check if specified nodes, or the current node when <em><code>nodeset</code></em> is not specified, are configured by service. Allows to easily differentiate nodes that are in addition to what service provisions. Returns boolean.</td></tr><tr><td valign="top"><code>oob:rule-paths()</code></td><td valign="top">Rule's path selector evaluated for the current change. Useful for lists, where the rule's path typically refers to all list items, but <code>oob:rule-paths()</code> selects the one with the change. Returns nodeset with one node.</td></tr><tr><td valign="top"><code>oob:context([nodeset])</code></td><td valign="top">Use the out-of-band version of data when evaluating the specified nodes or the current node when <em><code>nodeset</code></em> is not specified.</td></tr></tbody></table>

`oob:rule-paths()` perhaps requires an example to explain fully. A typical use case for this function is to more easily reference one specific parent of a changed node. Suppose a service configures BGP routing and you want to distinguish between BGP neighbors that are owned by the service versus those that are added out of band. The following rule would match all kinds of out-of-band changes but only for service-provisioned BGP neighbors:

```
 rule service-owned-neighbors
  path         ios:router/bgp/neighbor
  filter-expr  "oob:is-service-data(oob:rule-paths())"
  at-create    manage-by-service
  at-delete    manage-by-service
  at-value-set manage-by-service
 !
```

For a change of `.../bgp[as-no='65000']/neighbor[id='192.168.1.1']/remote-as`, the `oob:rule-paths()` would produce a node`.../bgp[as-no='65000']/neighbor[id='192.168.1.1']`. The filter expression is similar to `oob:is-service-data(current()/..)` but also works for nested nodes under the BGP neighbor, not just direct children.

## Service-managed Out-of-band Data

Using `manage-by-service` in an out-of-band policy ties an out-of-band change to a service instance and instructs NSO FASTMAP algorithm to take the change into account. FASTMAP treats the out-of-band changes as additional configuration, applied on top of service mapping logic.

For example, when a service-defined value is changed out-of-band and the policy specifies `manage-by-service`, the change is preserved across service redeploys. To differentiate between data from service mapping and out-of-band data, additional parameters can be used with service `get-modifications forward` action:

* `only-out-of-band`: display service-managed out-of-band configuration only.
* `only-service`: display configuration produced by service mapping only.
* `with-out-of-band`: display complete configuration, combined with out-of-band part.

For a service, where DHCP snooping rate-limit was configured out of band, the combined configuration might be:

```bash
admin@ncs# iface instance1 get-modifications forward { with-out-of-band } outformat cli-c
cli-c {
    local-node {
        data devices device c1
               config
                interface GigabitEthernet0/1
                 ip address 10.1.2.3 255.255.255.240
                 ip dhcp snooping limit rate 10
                 ip dhcp snooping trust
                exit
               !
              !
    }
}
```

The same is reflected in the service-meta-data of device configuration, which also shows origin of each part (note the `Out-of-band:` reference):

```bash
admin@ncs# show running-config devices device c1 config interface GigabitEthernet 0/1\
 | display service-meta-data
devices device c1
 config
  ! Refcount: 2
  ! Backpointer: [ /iface:iface[iface:name='instance1'] ]
  interface GigabitEthernet0/1
   ! Refcount: 1
   ip address 10.1.2.3 255.255.255.240
   ! Refcount: 1
   ! Out-of-band: [ /iface:iface[iface:name='instance1'] ]
   ip dhcp snooping limit rate 10
   ! Refcount: 1
   ! Backpointer: [ /iface:iface[iface:name='instance1'] ]
   ip dhcp snooping trust
   mtu 1520
  exit
 !
!
```

The lifecycle of the out-of-band parts is tied to the service lifecycle and the change is deleted when the service instance is deleted. But it is not truly managed in the sense of how mapping-generated configuration is managed.

For example, observe what happens when service interface parameter changes:

```bash
admin@ncs(config)# iface instance1 interface 0/4
admin@ncs(config-iface-instance1)# commit dry-run outformat cli-c
cli-c {
    local-node {
        data iface instance1
              interface 0/4
             !
             devices device c1
              config
               interface GigabitEthernet0/4
                ip address 10.1.2.3 255.255.255.240
                ip dhcp snooping trust
               exit
               interface GigabitEthernet0/1
                no ip address 10.1.2.3 255.255.255.240
                no ip dhcp snooping trust
               exit
              !
             !
    }
}
```

The configuration produced by service mapping uses the new interface, however, the out-of-band configuration is not migrated along with it:

```bash
admin@ncs(config-iface-instance1)# commit and-quit
Commit complete.
admin@ncs# iface instance1 get-modifications forward { with-out-of-band } outformat cli-c
cli-c {
    local-node {
        data devices device c1
               config
                interface GigabitEthernet0/1
                 ip dhcp snooping limit rate 10
                exit
                interface GigabitEthernet0/4
                 ip address 10.1.2.3 255.255.255.240
                 ip dhcp snooping trust
                exit
               !
              !
    }
}
```

In general, NSO cannot migrate the out-of-band changes on its own, since they may be inapplicable or even break the new service configuration. But in this specific case, the rest of the service configuration is removed from the interface, and DHCP snooping part would not be picked up by the service out-of band policy (if the change was done after the service update). While you can manually remove the residual GigabitEthernet0/1 configuration, a service re-deploy would reprovision it (unless you also [detach](#attach-and-detach-out-of-band-data) out-of-band data). A simpler approach is to reevaluate out-of-band policy.

### Reevaluating Policy

To avoid leftover configuration, or catch up with an updated out-of-band policy, you can instruct NSO to recompute service out-of-band changes according to the policy.

NSO will reapply the relevant policies if you use the`commit confirm-network-state re-evaluate-policies` commit variant when updating the service instance. Continuing the previous example:

```bash
admin@ncs(config)# show configuration
iface instance1
 interface    0/4
!
admin@ncs(config-iface-instance1)# commit and-quit confirm-network-state re-evaluate-policies
admin@ncs# iface instance1 get-modifications forward { with-out-of-band } outformat cli-c
cli-c {
    local-node {
        data devices device c1
               config
                interface GigabitEthernet0/4
                 ip address 10.1.2.3 255.255.255.240
                 ip dhcp snooping trust
                exit
               !
              !
    }
}
```

Since the out-of-band policy was reapplied, and the old interface is no longer part of the configuration that is provisioned by the service, its out-of-band configuration is gone.

If you have already committed the updated service instance without`confirm-network-state re evaluate-policies`, or have just updated the out-of-band policy, you can perform the same through a service redeploy:

```bash
admin@ncs# iface instance1 re-deploy confirm-network-state { re-evaluate-policies }\
 dry-run { outformat cli-c }
cli-c {
    local-node {
        data devices device c1
               config
                interface GigabitEthernet0/1
                 no ip dhcp snooping limit rate 10
                exit
               !
              !
    }
}
admin@ncs# iface instance1 re-deploy confirm-network-state { re-evaluate-policies }
```

Another potential effect using `re-evaluate-policies` has, is bringing in existing configuration. Suppose the above service instance, instead of GigabitEthernet0/4, uses GigabitEthernet0/3 interface, which already has some pre-existing configuration (configuration before being provisioned for this service).

```bash
admin@ncs(config)# show full-configuration devices device c1 config\
 interface GigabitEthernet 0/3
devices device c1
 config
  interface GigabitEthernet0/3
   ip address 10.2.2.10 255.255.255.240
  exit
 !
!
admin@ncs(config)# ! Change the interface:
admin@ncs(config)# iface instance1 interface 0/3
admin@ncs(config-iface-instance1)# commit confirm-network-state re-evaluate-policies and-quit
Commit complete.
admin@ncs# iface instance1 get-modifications forward { only-out-of-band } outformat cli-c
cli-c {
    local-node {
        data devices device c1
               config
                interface GigabitEthernet0/3
                 ip address 10.2.2.10 255.255.255.240
                exit
               !
              !
    }
}
```

Observe that the IP address is not the one configured by the service mapping (10.1.2.3); the existing value is instead being treated as a service-managed out-of-band change (as defined by the policy).

Therefore, if you wish to retain out-of-band parts that are no longer under service-managed configuration, you need to migrate them manually first. But consider that updating the service to support this kind of configuration natively is a much better choice that will save you a lot of time and trouble in the future.

### Attach and Detach Out-of-band Data

An alternative to `confirm-network-state re-evaluate-policies` for updating service out-of-band data are two service `re-deploy reconcile` actions:`attach-non-service-config` and `detach-non-service-config`.

Detach makes all current service-managed out-of-band data unmanaged. That is, it keeps the out-of band data but removes the references to the service from it. The out-of-band data behaves like it had a policy `sync-from-device` instead of`manage-by-service`.

You can use detach, for example, before removing a service in order to keep out-of-band changes.

On the other hand, attach performs similarly as confirm-network-state-enabled commit would for detected out-of-band data. It looks at all the service-owned configuration and finds parts that would stay if the service was removed (they have non-service refcounts). Then it makes these parts service managed out-of-band data.

You should use attach, instead of a regular `re-deploy reconcile`, when [importing existing services to NSO](/guides/development/advanced-development/developing-services/services-deep-dive). Using attach ensures the service also picks up out-of-band data according to policy.

Likewise, you can use attach to reattach configuration that you have previously, perhaps mistakenly, detached.

In addition, reconcile also supports `discard-non-service-config`, which allows you to discard all non-service-managed out-of-band changes.

To drop all out-of-band changes, not just unmanaged ones, and return the service to its pristine state, with only service-mapping-generated configuration, first detach out-of-band data, followed by a discard. For example:

```bash
admin@ncs# show running-config devices device c1 config interface GigabitEthernet 0/1\
 | display service-meta-data
devices device c1
 config
  ! Refcount: 2
  ! Backpointer: [ /iface:iface[iface:name='instance1'] ]
  interface GigabitEthernet0/1
   ! Refcount: 1
   ip address 10.1.2.3 255.255.255.240
   ! Refcount: 1
   ! Out-of-band: [ /iface:iface[iface:name='instance1'] ]
   ip dhcp snooping limit rate 10
   ! Refcount: 1
   ! Backpointer: [ /iface:iface[iface:name='instance1'] ]
   ip dhcp snooping trust
   mtu 1520
  exit
 !
!
admin@ncs# iface instance1 re-deploy reconcile { detach-non-service-config }
admin@ncs# iface instance1 re-deploy reconcile { discard-non-service-config }\
 dry-run { outformat cli-c }
cli-c {
    local-node {
        data devices device c1
               config
                interface GigabitEthernet0/1
                 no ip dhcp snooping limit rate 10
                 no mtu 1520
                exit
               !
              !

    }
}
```

In the example, both managed (`ip dhcp snooping limit rate 10`) and unmanaged (`mtu 1520`) out of-band changes are going to be dropped.

## Configuration and Command Reference

To globally [enable out-of-band data processing](#introducing-confirm-network-state) described in this section, configure:

```bash
admin@ncs(config)# devices global-settings confirm-network-state enabled-by-default true
```

To enable it for a set of devices, use device profiles:

```bash
admin@ncs(config)# devices profiles profile <PROFILE> confirm-network-state\
 enabled-by-default true
```

To enable it per individual device, configure:

```bash
admin@ncs(config)# devices device <DEVICE> confirm-network-state enabled-by-default true
```

Inspect out-of-band changes on a device without updating the CDB configuration:

```bash
admin@ncs# devices device <DEVICE> compare-config
```

CDB is updated automatically with referenced out-of-band data during a confirm-network-state-enabled commit. To manually update the CDB and process all out-of-band changes for a device, use device `sync-from`.

```bash
admin@ncs# devices device <DEVICE> sync-from
```

Inspect [out-of-band policy](#service-out-of-band-policies) for a service:

```bash
admin@ncs# show running-config services out-of-band policy <SERVICEPOINT>
```

Inspect [service-managed out-of-band changes](#service-managed-out-of-band-data) for a service:

```bash
admin@ncs# <SERVICE INSTANCE> get-modifications forward { only-out-of-band }
```

[Reevaluate out-of-band policy](#reevaluating-policy) (when updating service instance):

```bash
admin@ncs(config)# commit confirm-network-state re-evaluate-policies
```

Reevaluate out-of-band policy during service redeploy:

```bash
admin@ncs# <SERVICE INSTANCE> re-deploy confirm-network-state { re-evaluate-policies }
```

[Attach, detach, and discard](#attach-and-detach-out-of-band-data) out-of-band changes for a service:

```bash
admin@ncs# <SERVICE INSTANCE> re-deploy reconcile { attach-non-service-config }
admin@ncs# <SERVICE INSTANCE> re-deploy reconcile { detach-non-service-config }
admin@ncs# <SERVICE INSTANCE> re-deploy reconcile { discard-non-service-config }
```


# SSH Key Management

Learn about NSO SSH key management.

The SSH protocol uses public key technology for two distinct purposes:

1. **Server Authentication**: This use is a mandatory part of the protocol. It allows an SSH client to authenticate the server, i.e. verify that it is really talking to the intended server and not some man-in-the-middle intruder. This requires that the client has prior knowledge of the server's public keys, and the server proves its possession of one of the corresponding private keys by using it to sign some data. These keys are normally called 'host keys', and the authentication procedure is typically referred to as 'host key verification' or 'host key checking'.
2. **Client Authentication**: This use is one of several possible client authentication methods, i.e. it is an alternative to the commonly used password authentication. The server is configured with one or more public keys which are authorized for authentication of a user. The client proves possession of one of the corresponding private keys by using it to sign some data - i.e. the exact reverse of the server authentication provided by host keys. The method is called 'public key authentication' in SSH terminology.

These two usages are fundamentally independent, i.e., host key verification is done regardless of whether the client authentication is via public key, password, or some other method. However host key verification is of particular importance when client authentication is done via password, since failure to detect a man-in-the-middle attack in this case will result in the cleartext password being divulged to the attacker.

## NSO as SSH Server <a href="#ug.ssh_keys.server" id="ug.ssh_keys.server"></a>

NSO can act as an SSH server for northbound connections to the CLI or the NETCONF agent, and for connections from other nodes in an NSO cluster - cluster connections use NETCONF, and the server side setup used is the same as for northbound connections to the NETCONF agent. It is possible to use either the NSO built-in SSH server or an external server such as OpenSSH, for all of these cases. When using an external SSH server, host keys for server authentication and authorized keys for client/user authentication need to be set up per the documentation for that server, and there is no NSO-specific key management in this case.

When the NSO built-in SSH server is used, the setup is very similar to the one OpenSSH uses:

### Host Keys <a href="#d5e4103" id="d5e4103"></a>

The private host key(s) must be placed in the directory specified by `/ncs-config/aaa/ssh-server-key-dir` in `ncs.conf`, and named either `ssh_host_dsa_key` (for a DSA key) or `ssh_host_rsa_key` (for a RSA key). The key(s) must be in PEM format (e.g. as generated by the OpenSSH **ssh-keygen** command), and must not be encrypted - protection can be achieved by file system permissions (not enforced by NSO). The corresponding public key(s) is/are typically stored in the same directory with a `.pub` extension to the file name, but they are not used by NSO. The NSO installation creates a DSA private/public key pair in the directory specified by the default `ncs.conf`.

### Public Key Authentication <a href="#d5e4113" id="d5e4113"></a>

The public keys that are authorized for authentication of a given user must be placed in the user's SSH directory. Refer to [Public Key Login](https://nso-docs.cisco.com/guides/operation-and-usage/operations/pages/oUchYxSeOKkwhPOOefVU#ug.aaa.public_key_login) for details on how NSO searches for the keys to use.

## NSO as SSH Client <a href="#ug.ssh_keys.client" id="ug.ssh_keys.client"></a>

NSO can act as an SSH client for connections to managed devices that use SSH (this is always the case for devices accessed via NETCONF, typically also for devices accessed via CLI), and for connections to other nodes in an NSO cluster. In all cases, a built-in SSH client is used. The [examples.ncs/aaa/ssh-keys](https://github.com/NSO-developer/nso-examples/tree/6.7/aaa/ssh-keys) example in the NSO example collection has a detailed walk-through of the NSO functionality that is described in this section.

### Host Key Verification <a href="#ug.ssh_keys.client.host_keys" id="ug.ssh_keys.client.host_keys"></a>

#### **Verification Level**

The level of host key verification can be set globally via `/ssh/host-key-verification`. The possible values are:

* `reject-unknown`: The host key provided by the device or cluster node must be known by NSO for the connection to succeed.
* `reject-mismatch`: The host key provided by the device or cluster node may be unknown, but it must not be different from the "known" key for the same key algorithm, for the connection to succeed.
* `none`: No host key verification is done - the connection will never fail due to the host key provided by the device or cluster node.

The default is `reject-unknown`, and it is not recommended to use a different value, although it can be useful or needed in certain circumstances. E.g. `none` maybe useful in a development scenario, and temporary use of `reject-mismatch` maybe motivated until host keys have been configured for a set of existing managed devices.

{% code title="Allowing SSH Connections With Unknown Host Keys" %}

```bash
admin@ncs(config)# ssh host-key-verification reject-mismatch
admin@ncs(config)# commit
Commit complete.
```

{% endcode %}

#### **Connection to a Managed Device**

The public host keys for a device that is accessed via SSH are stored in the `/devices/device/ssh/host-key` list. There can be several keys in this list, one each for the `ssh-ed25519` (ED25519 key), `ssh-dss` (DSA key) and `ssh-rsa` (RSA key) key algorithms. In case a device has entries in its `live-status-protocol` list that use SSH, the host keys for those can be stored in the `/devices/device/live-status-protocol/ssh/host-key` list, in the same way as the device keys - however if `/devices/device/live-status-protocol/ssh` does not exist, the keys from `/devices/device/ssh/host-key` are used for that protocol. The keys can be configured e.g. via input directly in the CLI, but in most cases, it will be preferable to use the actions described below to retrieve keys from the devices. These actions will also retrieve any `live-status-protocol` keys for a device.

The level of host key verification can also be set per device, via `/devices/device/ssh/host-key-verification`. The default is to use the global value (or default) for `/ssh/host-key-verification`, but any explicitly set value will override the global value. The possible values are the same as for `/ssh/host-key-verification`.

There are several actions that can be used to retrieve the host keys from a device and store them in the NSO configuration:

* `/devices/fetch-ssh-host-keys`: Retrieve the host keys for all devices. Successfully retrieved keys are committed to the configuration.
* `/devices/device-group/fetch-ssh-host-keys`: Retrieve the host keys for all devices in a device group. Successfully retrieved keys are committed to the configuration.
* `/devices/device/ssh/fetch-host-keys`: Retrieve the host keys for one or more devices. In the CLI, range expressions can be used for the device name, e.g. using '\*' will retrieve keys for all devices, etc. The action will commit the retrieved keys if possible, i.e. if the device entry is already committed, otherwise (i.e., if the action is invoked from "configure mode" when the device entry has been created but not committed), the keys will be written to the current transaction, but not committed.

The fingerprints of the retrieved keys will be reported as part of the result from these actions, but it is also possible to ask for the fingerprints of already retrieved keys by invoking the `/devices/device/ssh/host-key/show-fingerprint` action (`/devices/device/live-status-protocol/ssh/host-key/show-fingerprint` for live-status protocols that use SSH).

{% code title="Retrieving SSH Host Keys for All Configured Devices" %}

```bash
admin@ncs# devices fetch-ssh-host-keys
fetch-result {
    device c0
    result unchanged
    fingerprint {
        algorithm ssh-dss
        value 03:64:fc:b7:87:bd:34:5e:3b:6e:d8:71:4d:3f:46:76
    }
}
fetch-result {
    device h0
    result unchanged
    fingerprint {
        algorithm ssh-dss
        value 03:64:fc:b7:87:bd:34:5e:3b:6e:d8:71:4d:3f:46:76
    }
}
```

{% endcode %}

#### **Connection to an NSO Cluster Node**

This is very similar to the case of a connection to a managed device, it differs mainly in locations - and in the fact that SSH is always used for connection to a cluster node. The public host keys for a cluster node are stored in the `/cluster/remote-node/ssh/host-key` list, in the same way as the host keys for a device. The keys can be configured e.g. via input directly in the CLI, but in most cases, it will be preferable to use the action described below to retrieve keys from the cluster node.

The level of host key verification can also be set per cluster node, via `/cluster/remote-node/ssh/host-key-verification`. The default is to use the global value (or default) for `/ssh/host-key-verification`, but any explicitly set value will override the global value. The possible values are the same as for `/ssh/host-key-verification`.

The `/cluster/remote-node/ssh/fetch-host-keys` action can be used to retrieve the host keys for one or more cluster nodes. In the CLI, range expressions can be used for the node name, e.g. using '\*' will retrieve keys for all nodes, etc. The action will commit the retrieved keys if possible, but if it is invoked from "configure mode" when the node entry has been created but not committed, the keys will be written to the current transaction, but not committed.

The fingerprints of the retrieved keys will be reported as part of the result from this action, but it is also possible to ask for the fingerprints of already retrieved keys by invoking the `/cluster/remote-node/ssh/host-key/show-fingerprint` action.

{% code title="Retrieving SSH Host Keys for All Cluster Nodes" %}

```bash
admin@ncs# cluster remote-node * ssh fetch-host-keys
cluster remote-node ncs1 ssh fetch-host-keys
    result updated
    fingerprint {
        algorithm ssh-dss
        value 03:64:fc:b7:87:bd:34:5e:3b:6e:d8:71:4d:3f:46:76
    }
cluster remote-node ncs2 ssh fetch-host-keys
    result updated
    fingerprint {
        algorithm ssh-dss
        value 03:64:fc:b7:87:bd:34:5e:3b:6e:d8:71:4d:3f:46:76
    }
cluster remote-node ncs3 ssh fetch-host-keys
    result updated
    fingerprint {
        algorithm ssh-dss
        value 03:64:fc:b7:87:bd:34:5e:3b:6e:d8:71:4d:3f:46:76
    }
```

{% endcode %}

### Public Key Authentication

#### **Private Key Selection**

The private key used for public key authentication can be taken either from the SSH directory for the local user or from a list of private keys in the NSO configuration. The user's SSH directory is determined according to the same logic as for the server-side public keys that are authorized for authentication of a given user, see [Public Key Login](https://nso-docs.cisco.com/guides/operation-and-usage/operations/pages/oUchYxSeOKkwhPOOefVU#ug.aaa.public_key_login), but of course, different files in this directory are used, see below. Alternatively, the key can be configured in the `/ssh/private-key` list, using an arbitrary name for the list key. In both cases, the key must be in PEM format (e.g. as generated by the OpenSSH **ssh-keygen** command), and it may be encrypted or not. Encrypted keys configured in `/ssh/private-key` must have the passphrase for the key configured via `/ssh/private-key/passphrase`.

#### **Connection to a Managed Device**

The specific private key to use is configured via the `authgroup` indirection and the `umap` selection mechanisms as for password authentication, just a different alternative. Setting `/devices/authgroups/group/umap/public-key` (or `default-map` instead of `umap` for users that are not in `umap`) without any additional parameters will select the default of using a file called `id_dsa` in the local user's SSH directory, which must have an unencrypted key. A different file name can be set via `/devices/authgroups/group/umap/public-key/private-key/file/name`. For an encrypted key, the passphrase can be set via `/devices/authgroups/group/umap/public-key/private-key/file/passphrase`, or `/devices/authgroups/group/umap/public-key/private-key/file/use-password` can be set to indicate that the password used (if any) by the local user when authenticating to NSO should also be used as a passphrase for the key. To instead select a private key from the `/ssh/private-key` list, the name of the key is set via `/devices/authgroups/group/umap/public-key/private-key/name`.

{% code title="Configuring a Private Key File for Publickey Authentication to Devices" %}

```bash
admin@ncs(config)# devices authgroups group default umap admin
admin@ncs(config-umap-admin)# public-key private-key file name /home/admin/.ssh/id-dsa
admin@ncs(config-umap-admin)# public-key private-key file passphrase
(<AES encrypted string>): *********
admin@ncs(config-umap-admin)# commit
Commit complete.
```

{% endcode %}

#### **Connection to an NSO Cluster Node**

This is again very similar to the case of a connection to a managed device, since the same `authgroup`/`umap` scheme is used. Setting `/cluster/authgroup/umap/public-key` (or `default-map` instead of `umap` for users that are not in `umap`) without any additional parameters will select the default of using a file called `id_dsa` in the local user's SSH directory, which must have an unencrypted key. A different file name can be set via `/cluster/authgroup/umap/public-key/private-key/file/name`. For an encrypted key, the passphrase can be set via `/cluster/authgroup/umap/public-key/private-key/file/passphrase`, or `/cluster/authgroup/umap/public-key/private-key/file/use-password` can be set to indicate that the password used (if any) by the local user when authenticating to NSO should also be used as a passphrase for the key. To instead select a private key from the `/ssh/private-key` list, the name of the key is set via `/cluster/authgroup/umap/public-key/private-key/name`.

{% code title="Configuring a Private Key File for Publickey Authentication in Cluster" %}

```bash
admin@ncs(config)# cluster authgroup default umap admin
admin@ncs(config-umap-admin)# public-key private-key file name /home/admin/.ssh/id-dsa
admin@ncs(config-umap-admin)# public-key private-key file passphrase
(<AES encrypted string>): *********
admin@ncs(config-umap-admin)# commit
Commit complete.
```

{% endcode %}


# Alarm Manager

Manage NSO alarms with native alarm manager.

NSO embeds a generic alarm manager. It manages NSO native alarms and can easily be extended with application-specific alarms. Alarm sources can be notifications from devices, undesired states on services detected or anything provided via the Java API.

The Alarm Manager has three main components:

* **Alarm List**: A list of alarms in NSO. Each list entry represents an alarm state for a specific device, an object within the device, and an alarm type.
* **Alarm Model**: For each alarm type, you can configure the mapping to for example X.733 alarm standard parameters that are sent as notifications northbound.
* **Operator Actions**: Actions to set operator states on alarms such as acknowledgement, and also actions to administratively manage the alarm list such as deleting alarms.

<div data-with-frame="true"><figure><img src="/files/gnvWh8KAC5wVxfxJd1DH" alt="" width="375"><figcaption><p>The Alarm Manager</p></figcaption></figure></div>

The alarm manager is accessible over all northbound interfaces. A read-only view including an SNMP alarm table and alarm notifications is available in an SNMP Alarm MIB. This MIB is suitable for integration with SNMP-based alarm systems.

To populate the alarm list there is a dedicated Java API. This API lets a developer add alarms, change states on alarms, etc. A common usage pattern is to use the SNMP notification receiver to map a subset of the device traps into alarms.

## Alarm Concepts <a href="#ug.alarmmgr.alarms" id="ug.alarmmgr.alarms"></a>

First of all, it is important to clearly define what an alarm means: "An alarm denotes an undesirable state in a resource for which an operator action is required". Alarms are often confused with general logging and event mechanisms, thereby overflooding the operator with alarms. In NSO, the alarm manager shows undesired resource states that an operator should investigate. NSO contains other mechanisms for logging in general. Therefore, NSO does not naively populate the alarm list with traps received in the SNMP notification receiver.

Before looking into how NSO handles alarms, it is important to define the fundamental concepts. We make a clear distinction between alarms and events in general. Alarms should be taken seriously and be investigated. Alarms have states; they go active with a specific severity, they change severity, and they are cleared by the resource. The same alarm may become active again. A common mistake is to confuse the operator view with the resource view. The model described so far is the resource view. The resource itself may consider the alarm cleared. The alarm manager does not automatically delete cleared alarms. An alarm that has existed in the network may still need investigation. There are dedicated actions an operator can use to manage the alarm list, for example, delete the alarms based on criteria such as cleared and date. These actions can be performed over all northbound interfaces.

Rather than viewing alarms as a list of alarm notifications, NSO defines alarms as states on objects. The NSO alarm list uses four keys for alarms: the alarming object within a device, the alarm type, and an optional specific problem.

Alarm types are normally unique identifiers for a specific alarm state and are defined statically. An alarm type corresponds to the well-known X.733 alarm standard tuple event type and probable cause. A specific problem is an optional key that is string-based and can further redefine an alarm type at run-time. This is needed for alarms that are not known before a system is deployed.

Imagine a system with general digital inputs. A MIB might specify traps called `input-high`, or `input-low`. When defining the SNMP notification reception, an integrator might define an alarm type called "External-Alarm". `input-high` might imply a major alarm and `input-low` might imply clear.

At installation, some detectors report "fire-alarm" and some "door-open" alarms. This is configured at the device and sent as free text in the SNMP var-binds. This is then managed by using the specific problem field of the NSO alarm manager to separate these different alarm types.

The data model for the alarm manager is outlined below.

<div data-with-frame="true"><figure><img src="/files/0oigSJ69OAXEjGSeFXK0" alt="" width="563"><figcaption><p>Alarm Model</p></figcaption></figure></div>

This means that we have a list with key: (managed device, managed object, alarm type, specific problem). In the example above, we might have the following different alarms:

* Device : House1; Managed Object : Detector1; Alarm-Type : External Alarm; Specific Problem = Smoke;
* Device : House1; Managed Object : Detector2; Alarm-Type : External Alarm; Specific Problem = Door Open;

Each alarm entry shows the last status change for the alarm and also a child list with all status changes sorted in chronological order.

* `is-cleared`: was the last state change clear?
* `last-status-change`: timestamp for the last status change.
* `last-perceived-severity`: last severity (not equal to clear).
* `last-alarm-text`: the last alarm text (not equal to clear).
* `status-change`, `event-time`: the time reported by the device.
* `status-change`, `received-time`: the time the state change was received by NSO.
* `status-change`, `perceived-severity`: the new perceived severity.
* `status-change`, `alarm-text`: descriptive text associated with the new alarm status.

It is fundamental to define alarm types (specific problem) and the managed objects with a fine-grained mechanism that still is extensible. For objects we allow YANG instance-identifiers to refer to a YANG instance identifier, an SNMP OID, or a string. Strings can be used when the underlying object is not modeled. We use YANG identities to define alarm types. This has the benefit that alarm types can be defined in a named hierarchy and thereby provide an extensible mechanism. To support "dynamic alarm types" so that alarms can be separated by information only available at run-time, the string-based field-specific problem can also be used.

So far we have described the model based on the resource view. It is common practice to let operators manipulate the alarms corresponding to the operator's investigation. We clearly separate the resource and the operator view, for example, there is no such thing as an operator "clearing an alarm". Rather the alarm entries can have a corresponding alarm handling state. Operators may want to acknowledge an alarm and set the alarm state to closed or similar.

### Alarm List Administrative Actions

We also support some alarm list administrative actions:

* **Synchronize alarms***:* try to read the alarm states in the underlying resources and update the alarm list accordingly (this action needs to be implemented by user code for specific applications).
* **Purge alarms***:* delete entries in the alarm list based on several different filter criteria.
* **Filter alarms***:* with an XPATH as filter input, this action returns all alarms fulfilling the filter.
* **Compress alarms***:* since every entry may contain a large amount of state change entries this action compresses the history to the latest state change.

Alarms can be forwarded over NSO northbound interfaces. In many telecom environments, alarms need to be mapped to X.733 parameters. We provide an alarm model where every alarm type is mapped to the corresponding X.733 parameters such as event type and probable cause. In this way, it is easy to integrate NSO alarms into whatever X.733 enumerated values the upper fault management system requires.

### Filtering Outgoing Alarm Notifications by Type

The configuration leaf-list `/alarms/control/filter-types` suppresses outbound alarm notifications for selected alarm types. For each generated alarm notification, NSO compares the alarm’s type with the configured filter-types entries. If the alarm type exactly matches a configured entry, NSO does not emit the corresponding SNMP trap, NETCONF alarm notification, or RESTCONF alarm notification.

This feature filters notification delivery only. The alarm is still created and maintained in the NSO alarm list, where it can be viewed, acknowledged, cleared, and otherwise handled using the normal alarm management workflows.

This means `filter-types` is useful when some alarm types should remain visible inside NSO but should not be forwarded to external monitoring systems. A common use case is reducing noise from known or operationally uninteresting alarm types.

Note that this mechanism filters by alarm type, not by severity level. Also note that the filtering takes effect for notifications generated after the configuration is applied; it does not retroactively remove existing alarms or old log entries.

For example, if `tailf-ncs-alarms:connection-failure` is added to `/alarms/control/filter-types`, NSO still stores `connection-failure` alarms in the alarm list, but it does not send matching SNMP, NETCONF, or RESTCONF alarm notifications.

#### Configuration Example

Configure `/alarms/control/filter-types` with one or more alarm types for which outbound alarm notifications should be suppressed. Each entry is an alarm type identity, for example `tailf-ncs-alarms:connection-failure`.

{% code title="Example: Configure Alarm Notification Filtering" overflow="wrap" %}

```bash
admin@ncs# configure
admin@ncs(config)# set alarms control filter-types tailf-ncs-alarms:connection-failure
admin@ncs(config)# commit
```

{% endcode %}

The filtering takes effect for notifications generated after the configuration is committed.

If `tailf-ncs-alarms:connection-failure` is configured under `/alarms/control/filter-types`, NSO still creates and maintains `connection-failure` alarms in `/alarms/alarm-list`:

{% code overflow="wrap" %}

```bash
admin@ncs# show alarms alarm-list
```

{% endcode %}

## The Alarm Model <a href="#ug.alarmmgr.model" id="ug.alarmmgr.model"></a>

The central part of the YANG Alarm model `tailf-ncs-alarms.yang` has the following structure.

{% code title=" tailf-ncs-alarms.yang" %}

```yang
module tailf-ncs-alarms {

  namespace "http://tail-f.com/ns/ncs-alarms";
  prefix "al";
  ...
 typedef managed-object-t {
    type union {
      type instance-identifier {
        require-instance false;
        }
      type yang:object-identifier;
      type string;
    }


  ...
  typedef event-type  {
    type enumeration {
      enum other {value 1;}
      enum communicationsAlarm {value 2;}
      enum qualityOfServiceAlarm {value 3;}
      enum processingErrorAlarm {value 4;}
      enum equipmentAlarm {value 5;}
      ...
    }
    description
    "...";
    reference
    "ITU Recommendation X.736, 'Information Technology - Open
     Systems Interconnection - System Management: Security
     Alarm Reporting Function', 1992";
  }

  typedef severity-t  {
    type enumeration {
      enum cleared {value 1;}
      enum indeterminate {value 2;}
      enum critical {value 3;}
      enum major {value 4;}
      enum minor {value 5;}
      enum warning {value 6;}
    }
    description
      "...";
  }
  ...
  identity alarm-type {
    description
    "Base identity for alarm types."
    ...
  }

  identity ncs-dev-manager-alarm {
    base alarm-type;
  }

  identity ncs-service-manager-alarm {
    base alarm-type;
  }

  identity connection-failure {
    base ncs-dev-manager-alarm;
    description
      "NCS failed to connect to a device";
  }
  ....
  container alarm-model {
    list alarm-type {
      key "type";
        leaf type {
          type alarm-type-t;
        }

        uses alarm-model-parameters;
     }
  }

      ...


    container alarm-list {
      config false;
      leaf number-of-alarms {
        type yang:gauge32;
      }

      leaf last-changed {
        type yang:date-and-time;
      }

      list alarm {
        key "device type managed-object specific-problem";
        uses common-alarm-parameters;
        leaf is-cleared {
          type boolean;
          mandatory true;
        }

        leaf last-status-change {
          type yang:date-and-time;
          mandatory true;
        }

        leaf last-perceived-severity {
          type severity-t;
        }

        leaf last-alarm-text {
          type alarm-text-t;
        }

        list status-change {
          key event-time;
          min-elements 1;
          uses alarm-state-change-parameters;
        }

        leaf last-alarm-handling-change {
          type yang:date-and-time;
        }

        list alarm-handling {
          key time;
          leaf time {
            tailf:info "Time stamp for operator action";
            type yang:date-and-time;
          }
          leaf state {
            tailf:info "The operators view of the alarm state";
            type alarm-handling-state-t;
            mandatory true;
            description
              "The operators view of the alarm state.";
          }
          ...
        }
        ...
        notification alarm-notification {
        ...
        rpc synchronize-alarms {
        ...
        rpc compress-alarms {
        ...
        rpc purge-alarms {
```

{% endcode %}

The first part of the YANG listing above shows the definition for `managed-object` type in order for alarms to refer to YANG, SNMP, and other resources. We also see basic definitions from the X.733 standard for severity levels.

Note well the definition of alarm type using YANG identities. In this way, we can create a structured alarm-type hierarchy all rooted at `alarm-type`. For you to add your specific alarm types, define your own alarm types YANG file and add identities using `alarm-type` as a base.

The `alarm-model` container contains the mapping from alarm types to X.733 parameters used for north-bound interfaces.

The `alarm-list` container is the actual alarm list where we maintain a list mapping (device, managed-object, alarm-type, specific-problem) to the corresponding alarm state changes \[(time, severity, text)].

Finally, we see the northbound alarm notification and alarm administrative actions.

## Alarm Handling <a href="#d5e4392" id="d5e4392"></a>

The NSO alarm manager has support for the operator to acknowledge alarms. We call this alarm handling. Each alarm has an associated list of alarm handling entries as:

```yang
container alarms {
  ....
  container alarm-list {
    config false;
    ....
    list alarm {
      key "device type managed-object specific-problem";

      .....

      list alarm-handling {
        key time;
        leaf time {
          type yang:date-and-time;
          description
            "Time-stamp for operator action on alarm.";
        }
        leaf state {
          mandatory true;
          type alarm-handling-state-t;
          description
            "The operators view of the alarm state";
        }
        leaf user {
          description "Which user has acknowledged this alarm";
          mandatory true;
          type string;
        }
        leaf description {
          description "Additional optional textual information regarding
            this new alarm-handling entry";
          type string;
        }
      }

        tailf:action handle-alarm {
          tailf:info "Set the operator state of this alarm";
          description
            "An action to allow the operator to add an entry to the
             alarm-handling list. This is a means for the operator to indicate
             the level of human intervention on an alarm.";
          input {
            leaf state {
              type alarm-handling-state-t;
              mandatory true;
            }
          }
        }
      }
```

The following typedef defines the different states an alarm can be set into.

{% code title="Alarm state" %}

```
  typedef alarm-handling-state-t  {
    type enumeration {
      enum none {
        value 1;
      }
      enum ack {
        value 2;
      }
      enum investigation {
        value 3;
      }
      enum observation {
        value 4;
      }
      enum closed {
        value 5;
      }
    }
    description
      "Operator actions on alarms";
  }
```

{% endcode %}

It is of course also possible to manipulate the alarm handling list from either Java code or Javascript code running in the web browser using the `js_maapi` library.

Below is a simple scenario to illustrate the alarm concepts. The example can be found in [examples.ncs/service-management/mpls-vpn-simple](https://github.com/NSO-developer/nso-examples/tree/6.7/service-management/mpls-vpn-simple).

```bash
$ make stop clean all start
$ ncs-netsim stop pe0
$ ncs-netsim stop pe1
$ ncs_cli -u admin -C
admin connected from 127.0.0.1 using console on host
admin@ncs# devices connect
...
connect-result {
    device pe0
    result false
    info Failed to connect to device pe0: connection refused
}
connect-result {
    device pe1
    result false
    info Failed to connect to device pe1: connection refused
}
...
admin@ncs# show alarms alarm-list
alarms alarm-list number-of-alarms 2
alarms alarm-list last-changed 2015-02-18T08:02:49.162436+00:00
alarms alarm-list alarm pe0 connection-failure /devices/device[name='pe0'] ""
 is-cleared              false
 last-status-change      2015-02-18T08:02:49.162734+00:00
 last-perceived-severity major
 last-alarm-text         "Failed to connect to device pe0: connection refused"
 status-change 2015-02-18T08:02:49.162734+00:00
  received-time      2015-02-18T08:02:49.162734+00:00
  perceived-severity major
  alarm-text         "Failed to connect to device pe0: connection refused"
alarms alarm-list alarm pe1 connection-failure /devices/device[name='pe1'] ""
 is-cleared              false
 last-status-change      2015-02-18T08:02:49.162436+00:00
 last-perceived-severity major
 last-alarm-text         "Failed to connect to device pe1: connection refused"
 status-change 2015-02-18T08:02:49.162436+00:00
  received-time      2015-02-18T08:02:49.162436+00:00
  perceived-severity major
  alarm-text         "Failed to connect to device pe1: connection refused"
```

In the above scenario, we stop two of the devices and then ask NSO to connect to all devices. This results in two alarms for `pe0` and `pe1`. Note that the key for the alarm is the device name, the alarm type, the full path to the object (in this case, the device and not an object within the device), and finally an empty string for the specific problem.

In the next command sequence, we start the device and request NSO to connect. This will clear the alarms.

```bash
admin@ncs# exit
$ ncs-netsim start pe0
DEVICE pe0 OK STARTED
$ ncs-netsim start pe1
DEVICE pe1 OK STARTED
$ ncs_cli -u admin -C
$ admin@ncs# devices connect
...
connect-result {
    device pe0
    result true
    info (admin) Connected to pe0 - 127.0.0.1:10028
}
connect-result {
    device pe1
    result true
    info (admin) Connected to pe1 - 127.0.0.1:10029
}
...
admin@ncs# show alarms alarm-list
alarms alarm-list number-of-alarms 2
alarms alarm-list last-changed 2015-02-18T08:05:04.942637+00:00
alarms alarm-list alarm pe0 connection-failure /devices/device[name='pe0'] ""
 is-cleared              true
 last-status-change      2015-02-18T08:05:04.942637+00:00
 last-perceived-severity major
 last-alarm-text         "Failed to connect to device pe0: connection refused"
 status-change 2015-02-18T08:02:49.162734+00:00
  received-time      2015-02-18T08:02:49.162734+00:00
  perceived-severity major
  alarm-text         "Failed to connect to device pe0: connection refused"
 status-change 2015-02-18T08:05:04.942637+00:00
  received-time      2015-02-18T08:05:04.942637+00:00
  perceived-severity cleared
  alarm-text         "Connected as admin"
alarms alarm-list alarm pe1 connection-failure /devices/device[name='pe1'] ""
 is-cleared              true
 last-status-change      2015-02-18T08:05:04.84115+00:00
 last-perceived-severity major
 last-alarm-text         "Failed to connect to device pe1: connection refused"
 status-change 2015-02-18T08:02:49.162436+00:00
  received-time      2015-02-18T08:02:49.162436+00:00
  perceived-severity major
  alarm-text         "Failed to connect to device pe1: connection refused"
 status-change 2015-02-18T08:05:04.84115+00:00
  received-time      2015-02-18T08:05:04.84115+00:00
  perceived-severity cleared
  alarm-text         "Connected as admin"
```

Note that there are two status-change entries for the alarm and that the alarm is cleared. In the following scenario, we will state that the alarm is closed and finally purge (delete) all alarms that are cleared and closed (Again, note the distinction between operator states and the states from the underlying resources).

```bash
admin@ncs# alarms alarm-list alarm pe0 connection-failure /devices/device[name='pe0']
          "" handle-alarm state closed description Fixed

admin@ncs# show alarms alarm-list alarm alarm-handling

DEVICE  TYPE                 STATE   USER   DESCRIPTION
---------------------------------------------------------
pe0     connection-failure   closed  admin  Fixed

admin@ncs# alarms purge-alarms alarm-handling-state-filter { state closed }
Value for 'alarm-status' [any,cleared,not-cleared]: cleared
purged-alarms 1
```

Assume that you need to configure the northbound parameters. This is done using the alarm model. A logical mapping of the connection problem above is to map it to X.733 probable cause `connectionEstablishmentError (22)` . This is done in the NSO CLI in the following way:

```bash
admin@ncs# config
Entering configuration mode terminal
admin@ncs(config)# alarms alarm-model alarm-type connection-failure probable-cause 22
admin@ncs(config-alarm-type-connection-failure/*)# commit
Commit complete.
admin@ncs(config-alarm-type-connection-failure/*)# show full-configuration
alarms alarm-model alarm-type connection-failure *
 event-type     communicationsAlarm
 has-clear      true
 kind-of-alarm  root-cause
 probable-cause 22
```


# Plug-and-Play Scripting

Use NSO's plug-and-play scripting mechanism to add new functionality to NSO.

A scripting mechanism can be used together with the CLI (scripting is not available for any other northbound interfaces). This section is intended for users who are familiar with UNIX shell scripting and/or programming. With the scripting mechanism, an end-user can add new functionality to NSO in a plug-and-play-like manner. No special tools are needed.

There are three categories of scripts:

* `command` scripts: Used to add new commands to the CLI.
* `policy` scripts: Invoked at validation time and may control the outcome of a transaction. Policy scripts have the mandate to cause a transaction to abort.
* `post-commit` scripts: Invoked when a transaction has been committed. Post-commit scripts can for example be used for logging, sending external events etc.

The terms 'script' and 'scripting' used throughout this description refer to how functionality can be added without a requirement for integration using the NSO programming APIs. NSO will only run the scripts as UNIX executables. Thus they may be written as shell scripts, or by using another scripting language that is supported by the OS, e.g., Python, or even as compiled code. The scripts are run with the same user ID as NSO.

The examples in this section are written using shell scripts as the least common denominator, but they can be written in another suitable language, e.g., Python or C.

## Script Storage <a href="#d5e4460" id="d5e4460"></a>

Scripts are stored in a directory tree with a predefined structure where there is a sub-directory for each script category:

```
scripts/
        command/
        policy/
        post-commit/
```

For all script categories, it suffices to just add a valid script in the correct sub-directory to enable the script. See the details for each script category for how a valid script of that category is defined. Scripts with a name beginning with a dot character ('.') are ignored.

The directory path to the location of the scripts is configured with the `/ncs-config/scripts/dir` configuration parameter. It is possible to have several script directories. The sample `ncs.conf` file that comes with the NSO release specifies two script directories: `./scripts` and `${NCS_DIR}/scripts`.

## Script Interface <a href="#d5e4471" id="d5e4471"></a>

All scripts are required to provide a formal description of their interface. When the scripts are loaded, NSO will invoke the scripts with (one of) the following as an argument depending on the script category.

* `--command`
* `--policy`
* `--post-commit`

The script must respond by writing its formal interface description on `stdout` and exit normally. Such a description consists of one or more sections. Which sections are required, depends on the category of the script.

The sections do however have a common syntax. Each section begins with the keyword `begin` followed by the type of section. After that one or more lines of settings follow. Each such setting begins with a name, followed by a colon character (`:`), and after that the value is stated. The section ends with the keyword `end`. Empty lines and spaces may be used to improve readability.

For examples see each corresponding section below.

## Script Loading <a href="#d5e4494" id="d5e4494"></a>

Scripts are automatically loaded at startup and may also be manually reloaded with the CLI command `script reload`. The command takes an optional `verbosity` parameter which may have one of the following values:

* `diff`: Shows info about those scripts that have been changed since the latest (re)load. This is the default.
* `all`: Shows info about all scripts regardless of whether they have been changed or not.
* `errors`: Shows info about those scripts that are erroneous, regardless of whether they have been changed or not. Typical errors are invalid file permissions and syntax errors in the interface description.

Yet another parameter may be useful when debugging the reload of scripts:

* `debug`: Shows additional debug info about the scripts.

An example session reloading scripts using the [examples.ncs/sdk-api/scripting](https://github.com/NSO-developer/nso-examples/tree/6.7/sdk-api/scripting) example:

```cli
admin@ncs# script reload all
$NCS_DIR/examples.ncs/sdk-api/scripting/scripts:
ok
command:
    add_user.sh: unchanged
    echo.sh: unchanged
policy:
    check_dir.sh: unchanged
post-commit:
    show_diff.sh: unchanged
/opt/ncs/scripts: ok
command:
    device_brief.sh: unchanged
    device_brief_c.sh: unchanged
    device_list.sh: unchanged
    device_list_c.sh: unchanged
    device_save.sh: unchanged
```

## Command Scripts <a href="#d5e4525" id="d5e4525"></a>

Command scripts are used to add new commands to the CLI. The scripts are executed in the context of a transaction. When the script is run in `oper` mode, this is a read-only transaction, when it is run in `config` mode, it is a read-write transaction. In that context, the script may make use of the environment variables `NCS_MAAPI_USID` and `NCS_MAAPI_THANDLE` in order to attach to the active transaction. This makes it simple to make use of the `ncs-maapi` command (see the [ncs-maapi(1)](/guides/resources/man/ncs-maapi.1) in Manual Pages manual page) for various purposes.

Each command script must be able to handle the argument `--command` and, when invoked, write a `command` section to `stdout`. If the CLI command is intended to take parameters, one `param` section per CLI parameter must also be emitted.

The command is not paginated by default in the CLI and will only do so if it is piped to `more`.

```
joe@io> example_command_script | more
```

### `command` Section <a href="#d5e4544" id="d5e4544"></a>

The following settings can be used to define a command:

* `modes`: Defines in which CLI mode(s) that the command should be available. The value can be `oper`, `config` or both (separated with space).
* `styles`: Defines in which CLI styles the command should be available. The value can be one or more of `c`, `i` and `j` (separated with space). `c` means Cisco style, `i`means Cisco IOS, and `j` J-style.
* `cmdpath`: Is the full CLI command path. For example, the `command` path `my script echo` implies that the command will be called `my script echo` in the CLI.
* `help`: Command help text.

An example of a `command` section is:

```
begin command
  modes: oper
  styles: c i j
  cmdpath: my script echo
  help: Display a line of text
end
```

### `param` Section <a href="#d5e4581" id="d5e4581"></a>

Now let's look at various aspects of a parameter. This may both affect the parameter syntax for the end-user in the CLI as well as what the command script will get as arguments.

The following settings can be used to customize each CLI parameter:

* `name`: Optional name of the parameter. If provided, the CLI will prompt for this name before the value. By default, the name is not forwarded to the script. See `flag` and `prefix`.
* `type`: The type of the parameter. By default each parameter has a value, but by setting the type to `void` the CLI will not prompt for a value. To be useful the `void` type must be combined with `name` and either `flag` or `prefix`.
* `presence`: Controls whether the parameter must be present in the CLI input or not. Can be set to `optional` or `mandatory`.
* `words`: Controls the number of words that the parameter value may consist of. By default, the value must consist of just one word (possibly quoted if it contains spaces). If set to `any`, the parameter may consist of any number of words. This setting is only valid for the last parameter.
* `flag`: Extra argument added before the parameter value. For example, if set to `-f` and the user enters `logfile`, the script will get `-f logfile` as arguments.
* `prefix`: Extra string prepended to the parameter value (as a single word). For example, if set to `--file=` and the user enters `logfile`, the script will get `--file=logfile` as argument.
* `help`: Parameter help text.

If the command takes a parameter to redirect the output to a file, a `param` section might look like this:

```
begin param
 name: file
 presence: optional
 flag: -f
 help: Redirect output to file
end
```

### Full `command` Example <a href="#d5e4639" id="d5e4639"></a>

A command denying changes the configured `trace-dir` for a set of devices, it can use the `check_dir.sh` script.

```bash
#!/bin/bash

set -e

while [ $# -gt 0 ]; do
    case "$1" in
        --command)
            # Configuration of the command
            #
            # modes   - CLI mode (oper config)
            # styles  - CLI style (c i j)
            # cmdpath - Full CLI command path
            # help    - Command help text
            #
            # Configuration of each parameter
            #
            # name     - (optional) name of the parameter
            # more     - (optional) true or false
            # presence - optional or mandatory
            # type     - void - A parameter without a value
            # words    - any - Multi word param. Only valid for the last param
            # flag     - Extra word added before the parameter value
            # prefix   - Extra string prepended to the parameter value
            # help     - Command help text
            cat << EOF

begin command
  modes: config
  styles: c i j
  cmdpath: user-wizard
  help: Add a new user
end
EOF
            exit
            ;;
        *)
            break
            ;;
    esac
    shift
done

## Ask for user name
while true; do
    echo -n "Enter user name: "
    read user

    if [ ! -n "${user}" ]; then
        echo "You failed to supply a user name."
    elif ncs-maapi --exists "/aaa:aaa/authentication/users/user{${user}}"; then
        echo "The user already exists."
    else
        break
    fi
done

## Ask for password
while true; do
    echo -n "Enter password: "
    read -s pass1
    echo

    if [ "${pass1:0:1}" == "$" ]; then
        echo -n "The password must not start with $. Please choose a "
        echo    "different password."
    else
        echo -n "Confirm password: "
        read -s pass2
        echo

        if [ "${pass1}" != "${pass2}" ]; then
            echo "Passwords do not match."
        else
            break
        fi
    fi
done

groups=`ncs-maapi --keys "/nacm/groups/group"`
while true; do
    echo "Choose a group for the user."
    echo -n "Available groups are: "
    for i in ${groups}; do echo -n "${i} "; done
    echo
    echo -n "Enter group for user: "
    read group

    if [ ! -n "${group}" ]; then
        echo "You must enter a valid group."
    else
        for i in ${groups}; do
            if [ "${i}" == "${group}" ]; then
                # valid group found
                break 2;
            fi
        done
        echo "You entered an invalid group."
    fi
    echo
done

echo "Creating user"

ncs-maapi --create "/aaa:aaa/authentication/users/user{${user}}"
ncs-maapi --set "/aaa:aaa/authentication/users/user{${user}}/password" \
                "${pass1}"

echo "Setting home directory to: /homes/${user}"
ncs-maapi --set "/aaa:aaa/authentication/users/user{${user}}/homedir" \
            "/homes/${user}"

echo "Setting ssh key directory to: /homes/${user}/ssh_keydir"
ncs-maapi --set "/aaa:aaa/authentication/users/user{${user}}/ssh_keydir" \
            "/homes/${user}/ssh_keydir"

ncs-maapi --set "/aaa:aaa/authentication/users/user{${user}}/uid" "1000"
ncs-maapi --set "/aaa:aaa/authentication/users/user{${user}}/gid" "100"

echo "Adding user to the ${group} group."
gusers=`ncs-maapi --get "/nacm/groups/group{${group}}/user-name"`

for i in ${gusers}; do
    if [ "${i}" == "${user}" ]; then
        echo "User already in group"
        exit 0
    fi
done

ncs-maapi --set "/nacm/groups/group{${group}}/user-name" "${gusers} ${user}"
```

Running the [examples.ncs/sdk-api/scripting](https://github.com/NSO-developer/nso-examples/tree/6.7/sdk-api/scripting) `/scripts/command/echo.sh` script with the argument `--command` argument produces a `command` section and a couple of `param` sections:

```bash
$ ./echo.sh --command
begin command
  modes: oper
  styles: c i j
  cmdpath: my script echo
  help: Display a line of text
end

begin param
 name: nolf
 type: void
 presence: optional
 flag: -n
 help: Do not output the trailing newline
end

begin param
 name: file
 presence: optional
 flag: -f
 help: Redirect output to file
end

begin param
 presence: mandatory
 words: any
 help: String to be displayed
end
```

In the complete example, [examples.ncs/sdk-api/scripting](https://github.com/NSO-developer/nso-examples/tree/6.7/sdk-api/scripting), there is a `README` file and a simple command script `scripts/command/echo.sh`.

## Policy Scripts <a href="#d5e4655" id="d5e4655"></a>

Policy scripts are invoked at validation time before a change is committed. A policy script can reject the data, accept it, or accept it with a warning. If a warning is produced, it will be displayed for interactive users (e.g. through the CLI or Web UI). The user may choose to abort or continue to commit the transaction.

Policy scripts are typically assigned to individual leafs or containers. In some cases, it may be feasible to use a single policy script, e.g. on the top-level node of the configuration. In such a case, this script is responsible for the validation of all values and their relationships throughout the configuration.

All policy scripts are invoked on every configuration change. The policy scripts can be configured to depend on certain subtrees of the configuration, which can save time but it is very important that all dependencies are stated and also updated when the validation logic of the policy script is updated. Otherwise, an update may be accepted even though a dependency should have denied it.

There can be multiple dependency declarations for a policy script. Each declaration consists of a dependency element specifying a configuration subtree that the validation code is dependent upon. If any element in any of the subtrees is modified, the policy script is invoked. A subtree is specified as an absolute path.

If there are no declared dependencies, the root of the configuration tree (/) is used, which means that the validation code is executed when any configuration element is modified. If dependencies are declared on a leaf element, an implicit dependency on the leaf itself is added.

Each policy script must handle the argument `--policy` and, when invoked, write a `policy` section to `stdout`. The script must also perform the actual validation when invoked with the argument `--keypath`.

### `policy` Section <a href="#d5e4667" id="d5e4667"></a>

The following settings can be used to configure a policy script:

* `keypath`: Mandatory. The keypath is the path to a node in the configuration data tree. The policy script will be associated with this node. The path must be absolute. A keypath can for example be `/devices/device/c0`. The script will be invoked if the configuration node, referred to by the keypath, is changed or if any node in the subtree under the node (if the node is a container or list) is changed.
* `dependency`: Declaration of a dependency. The dependency must be an absolute key path. Multiple dependency settings can be declared. Default is `/`.
* `priority`: An optional integer parameter specifying the order policy scripts will be evaluated, in order of increasing priority, where a lower value is higher priority. The default priority is `0`.
* `call`: This optional setting can only be used if the associated node, declared as `keypath`, is a list. If set to `once`, the policy script is only called once even though there exists many list entries in the data store. This is useful if we have a huge amount of instances or if values assigned to each instance have to be validated in comparison with its siblings. Default is `each`.

A policy that will be run for every change on or under `/devices/device`.

```
begin policy
 keypath: /devices/device
 dependency: /devices/global-settings
 priority: 4
 call: each
end
```

### Validation <a href="#d5e4700" id="d5e4700"></a>

When NSO has concluded that the policy script should be invoked to perform its validation logic, the script is invoked with the option `--keypath`. If the registered node is a leaf, its value will be given with the `--value` option. For example `--keypath /devices/device/c0` or if the node is a leaf `--keypath /devices/device/c0/address --value 127.0.0.1`.

Once the script has performed its validation logic it must exit with a proper status.

The following exit statuses are valid:

* `0`: Validation ok. Vote for commit.
* `1`: When the outcome of the validation is dubious, it is possible for the script to issue a warning message. The message is extracted from the script output on stdout. An interactive user can choose to abort or continue to commit the transaction. Non-interactive users automatically vote for commit.
* `2`: When the validation fails, it is possible for the script to issue an error message. The message is extracted from the script output on stdout. The transaction will be aborted.

### Full `policy` Example <a href="#d5e4724" id="d5e4724"></a>

A policy denying changes the configured `trace-dir` for a set of devices, it can use the `check_dir.sh` script.

```bash
#!/bin/sh

usage_and_exit() {
    cat << EOF
Usage: $0 -h
       $0 --policy
       $0 --keypath <keypath> [--value <value>]

  -h                    display this help and exit
  --policy              display policy configuration and exit
  --keypath <keypath>   path to node
  --value <value>       value of leaf

Return codes:

  0 - ok
  1 - warning message is printed on stdout
  2 - error message   is printed on stdout
EOF
    exit 1
}

while [ $# -gt 0 ]; do
    case "$1" in
        -h)
            usage_and_exit
            ;;
        --policy)
            cat << EOF
begin policy
  keypath: /devices/global-settings/trace-dir
  dependency: /devices/global-settings
  priority: 2
  call: each
end
EOF
            exit 0
            ;;
        --keypath)
            if [ $# -lt 2 ]; then
                echo "<ERROR> --keypath <keypath> - path omitted"
                usage_and_exit
            else
                keypath=$2
                shift
            fi
            ;;
        --value)
            if [ $# -lt 2 ]; then
                echo "<ERROR> --value <value> - leaf value omitted"
                usage_and_exit
            else
                value=$2
                shift
            fi
            ;;
        *)
            usage_and_exit
            ;;
    esac
    shift
done

if [ -z "${keypath}" ]; then
    echo "<ERROR> --keypath <keypath> is mandatory"
    usage_and_exit
fi

if [ -z "${value}" ]; then
    echo "<ERROR> --value <value> is mandatory"
    usage_and_exit
fi

orig="./logs"
dir=${value}
# dir=`ncs-maapi --get /devices/global-settings/trace-dir`
if [ "${dir}" != "${orig}" ] ; then
    echo "/devices/global-settings/trace-dir: must retain it original value (${orig})"
    exit 2
fi
```

Trying to change that parameter would result in an aborted transaction

```bash
admin@ncs(config)# devices global-settings trace-dir ./testing
admin@ncs(config)# commit
Aborted: /devices/global-settings/trace-dir: must retain it original
value (./logs)
```

In the complete example, [examples.ncs/sdk-api/scripting](https://github.com/NSO-developer/nso-examples/tree/6.7/sdk-api/scripting) there is a `README` file and a simple policy script `scripts/policy/check_dir.sh`.

## Post-commit Scripts <a href="#d5e4739" id="d5e4739"></a>

Post-commit scripts are run when a transaction has been committed, but before any locks have been released. The transaction hangs until the script has returned. The script cannot change the outcome of the transaction. Post-commit scripts can for example be used for logging, sending external events etc. The scripts run as the same user ID as NSO.

The script is invoked with `--post-commit` at script (re)load. In future releases, it is possible that the `post-commit` section will be used for control of the post-commit scripts behavior.

At post-commit, the script is invoked without parameters. In that context, the script may make use of the environment variables `NCS_MAAPI_USID` and `NCS_MAAPI_THANDLE` in order to attach to the active (read-only) transaction.

This makes it simple to make use of the `ncs-maapi` command. Especially the command `ncs-maapi --keypath-diff /` may turn out to be useful, as it provides a listing of all updates within the transaction on a format that is easy to parse.

### `post-commit` Section <a href="#d5e4754" id="d5e4754"></a>

All post-commit scripts must be able to handle the argument `--post-commit` and, when invoked, write an empty `post-commit` section to `stdout`:

```
begin post-commit
end
```

### Full `post-commit` Example <a href="#d5e4762" id="d5e4762"></a>

Assume the administrator of a system would want to have a mail each time a change is performed on the system, a script such as `mail_admin.sh`:

```bash
#!/bin/bash

set -e

if [ $# -gt 0 ]; then
    case "$1" in
        --post-commit)
            cat <<EOF
begin post-commit
end
EOF
            exit 0
            ;;
        *)
            echo
            echo "Usage: $0 [--post-commit]"
            echo
            echo "  --post-commit Mandatory for post-commit scripts"
            exit 1
            ;;
    esac
else
    file="mail_admin.log"
    NCS_DIFF=$(ncs-maapi --keypath-diff /)
    mail -s "NCS Mailer" admin@example.com <<EOF
AutoGenerated mail from NCS

$NCS_DIFF
EOF
fi
```

If the `admin` then loads this script:

```bash
admin@ncs# script reload debug
$NCS_DIR/examples.ncs/device-management/simulated-devices/scripts:
ok
    post-commit:
        mail_admin.sh: new
--- Output from
$NCS_DIR/examples.ncs/device-management/simulated-devices/scripts/post-commit/mail_admin.sh
--post-commit ---
1: begin post-commit
2: end
3:
---
admin@ncs# config
Entering configuration mode terminal
admin@ncs(config)# devices global-settings trace-dir ./again
admin@ncs(config)# commit
Commit complete.
```

This configuration change will produce an email to `admin@example.com` with subject `NCS Mailer` and body.

```
AutoGenerated mail from NCS
value set  : /devices/global-settings/trace-dir
```

In the complete example, [examples.ncs/sdk-api/scripting](https://github.com/NSO-developer/nso-examples/tree/6.7/sdk-api/scripting) , there is a `README` file and a simple post-commit script `scripts/post-commit/show_diff.sh`.


# Compliance Reporting

Audit and verify your network for configuration compliance.

When the network configuration is broken, there is a need to gather information and verify the network. NSO has numerous functions to show different aspects of such a network configuration verification. However, to simplify this task, compliance reporting can assemble information using a selection of these NSO functions and present the resulting information in one report. This report aims to answer two fundamental questions:

* Who has done what?
* Is the network correctly configured?

What defines a correctly configured network? Where is the authoritative configuration kept? Naturally, NSO, with the configurations stored in CDB, is the authority. Checking the live devices against the NSO-stored device configuration is a fundamental part of compliance reporting. Compliance reporting can also be based on one or a number of stored templates which the live devices are compared against. The compliance reports can also be a combination of both approaches.

Compliance reporting can be configured to check the current situation, check historical events, or both. To assemble historical events, rollback files are used. Therefore this functionality must be enabled in NSO before report execution, otherwise, the history view cannot be presented.

The reports are stored in a SQLite database file and can be exported to plain text, HTML or DocBook XML format. The report results can be re-exported to a new format at any time. The DocBook XML format allows you to use the report in further post-processing, such as creating a PDF using Apache FOP and your own custom styling. Every consecutive run of the report is stored in the same SQLite database. This allows for comparing the report results over time and one such comparison is available in the Web UI. The previous behavior before NSO 6.5 of getting one SQLite file per report run is available by setting `common-db` under the report definition to `false.`

{% hint style="info" %}
Reports can be generated using either the CLI or Web UI. The suggested and favored way of generating compliance reports is via the Web UI, which provides a convenient way of creating, configuring, and consuming compliance reports. In the NSO Web UI, compliance reporting options are accessible from the **Tools** menu (see [Web User Interface](/guides/operation-and-usage/webui) for more information). The CLI options are described in the sections below.
{% endhint %}

## Creating Compliance Report Definitions <a href="#d5e4797" id="d5e4797"></a>

It is possible to create several named compliance report definitions. Each named report defines the devices, services, and/or templates that should be part of the network configuration verification.

Let us walk through a simple compliance report definition. This example is based on the [examples.ncs/service-management/mpls-vpn-java](https://github.com/NSO-developer/nso-examples/tree/6.7/service-management/mpls-vpn-java) example. For the details of the included services and devices in this example, see the `README` file.

Each report definition has a name and can specify device and service checks. Device checks are further classified into sync and configuration checks. Device sync checks verify the in-sync status of the devices included in the report, while device configuration checks verify individual device configuration against a compliance template (see [Device Configuration Checks](#device-configuration-checks)).

For device checks, you can select the devices to be checked in four different ways:

* `all-devices` - Check all defined devices.
* `device-group` - Specified list of device groups.
* `device` - Specified list of devices.
* `select-devices` - Specified by an XPath expression.

Consider the following example report definition named `gold-check`:

```bash
ncs(config)# compliance reports report gold-check
ncs(config-report-gold-check)# device-check all-devices
```

This report definition, when executed, checks whether all devices known to NSO are in sync.

For such a check, the behavior of the verification can be specified:

* To request a check-sync action to verify that the device is currently in sync. This behavior is controlled by the leaf `current-out-of-sync` (default `true`).
* To scan the commit log (i.e. rollback files) for changes on the devices and report these. This behavior is controlled by the leaf `historic-changes` (default `true`).

```bash
ncs(config-report-gold-check)# device-check ?
Possible completions:
  all-devices            Report on all devices
  current-out-of-sync    Should current check-sync action be performed?
  device                 Report on specific devices
  device-group           Report on specific device groups
  historic-changes       Include commit log events from within the report
                         interval
  select-devices         Report on devices selected by an XPath expression
  <cr>
```

For the example `gold-check`, you can also use service checks. This type of check verifies if the specified service instances are in sync, that is if the network devices contain configuration as defined by these services. You can select the services to be checked in four different ways:

* `all-services` - Check all known service instances.
* `service` - Specified list of service instances.
* `select-services` - Specified list of service instances through an XPath expression.
* `service-type` - Specified list of service types.

For service checks, the verification behavior can be specified as well:

* To request a check-sync action to verify that the service is currently in sync. This behavior is controlled by the leaf `current-out-of-sync` (default `true`).
* To scan the commit log (i.e., rollback files) for changes on the services and report these. This behavior is controlled by the leaf `historic-changes` (default `true`).

```bash
ncs(config-report-gold-check)# service-check ?
Possible completions:
  all-services          Report on all services
  current-out-of-sync   Should current check-sync action be performed?
  historic-changes      Include commit log events from within the report
                        interval
  select-services       Report on services selected by an XPath expression
  service               Report on specific services
  service-type          The type of service.
  <cr>
```

In the example report, you might choose the default behavior and check all instances of the `l3vpn` service:

```bash
ncs(config-report-gold-check)# service-check service-type /l3vpn:vpn/l3vpn:l3vpn
ncs(config-report-gold-check)# commit
Commit complete.
ncs(config-report-gold-check)# show full-configuration
compliance reports report gold-check
 device-check all-devices
 service-check service-type /l3vpn:vpn/l3vpn:l3vpn
!
```

You can also use the web UI to define compliance reports. See the section [Compliance Reporting](https://nso-docs.cisco.com/guides/operation-and-usage/operations/pages/YbD21HK8ZwE5HK5p6bay#sec.webui_compliance) for more information.

## Running Compliance Reports <a href="#d5e4911" id="d5e4911"></a>

Compliance reporting is a read-only operation. When running a compliance report, the result is stored in a file located in a sub-directory `compliance-reports` under the NSO `state` directory. NSO has operational data for managing this report storage which makes it possible to list existing reports.

Here is an example of such a report listing:

```bash
ncs# show compliance report-results
compliance report-results report 1
 name              gold-check
 title             "GOLD NW 1"
 time              2015-02-04T18:48:57+00:00
 who               admin
 compliance-status violations
 location          http://.../report_1_admin_1_2015-2-4T18:48:57:0.xml
compliance report-results report 2
 name              gold-check
 title             "GOLD NW 2"
 time              2015-02-04T18:51:48+00:00
 who               admin
 compliance-status violations
 location          http://.../report_2_admin_1_2015-2-4T18:51:48:0.text
compliance report-results report 3
 name              gold-check
 title             "GOLD NW 3"
 time              2015-02-04T19:11:43+00:00
 who               admin
 compliance-status violations
 location          http://.../report_3_admin_1_2015-2-4T19:11:43:0.text
```

There is also a `remove` action to remove report results (and the corresponding file):

```bash
ncs# compliance report-results report 2..3 remove
ncs# show compliance report-results
compliance report-results report 1
 name              gold-check
 title             "GOLD NW 1"
 time              2015-02-04T18:48:57+00:00
 who               admin
 compliance-status violations
 location          http://.../report_1_admin_1_2015-2-4T18:48:57:0.xml
```

When running the report, there are a number of parameters that can be specified with the specific `run` action.

The parameters that are possible to specify for a report `run` action are:

* `title`: The title in the resulting report.
* `from`: The date and time from which the report should start the information gathering. If not set, the oldest available information is implied.
* `to`: The date and time when the information gathering should stop. If not set, the current date and time are implied. If set, no new check-syncs of devices and/or services will be attempted.
* `outformat`: One of the formats from `xml`, `html`, `text`, or `sqlite`. If `xml` is specified, the report will be formatted using the DocBook schema. The generated file can be [downloaded](#downloading-compliance-reports), for example, using standard CLI tools like `curl` or using Python requests via the URL returned by NSO.

We will request a report run with a `title` and formatted as `text`.

```bash
ncs# compliance reports report gold-check run \
> title "My First Report" outformat text
```

In the above command, the report was run without a `from` or a `to` argument. This implies that historical information gathering will be based on all available information. This includes information gathered from rollback files.

When a `from` argument is supplied to a compliance report run action, this implies that only historical information younger than the `from` date and time is checked.

```bash
ncs# compliance reports report gold-check run \
> title "First check" from 2015-02-04T00:00:00
```

When a `to` argument is supplied, this implies that historical information will be gathered for all logged information up to the date and time of the `to` argument.

```bash
ncs# compliance reports report gold-check run \
> title "Second check" to 2015-02-05T00:00:00
```

The `from` and a `to` arguments can be combined to specify a fixed historic time interval.

```bash
ncs# compliance reports report gold-check run \
> title "Third check" from 2015-02-04T00:00:00 to 2015-02-05T00:00:00
```

When a compliance report is run, the action will respond with a flag indicating if any discrepancies were found. Also, it reports how many devices and services have been verified in total by the report.

```bash
ncs# compliance reports report gold-check run \
> title "Fourth check" outformat text
time 2015-2-4T20:42:45.019012+00:00
compliance-status violations
info Checking 17 devices and 2 services
location http://.../report_7_admin_1_2015-2-4T20:42:45.019012+00:00.text
```

Below is an example of a compliance report result (in `text` format):

{% code title="Compliance Report Result" %}

```bash
$ cat ./state/compliance-reports/report_7_admin_1_2015-2-4T20\:42\:45.019012+00\:00.text
reportcookie : g2gCbQAAAAtGaWZ0aCBjaGVja20AAAAKZ29sZC1jaGVjaw==

Compliance report : Fourth check

        Publication date : 2015-2-4 20:42:45
        Produced by user : admin

Chapter : Summary

        Compliance result titled "Fourth check" defined by report "gold-check"
        Resulting in violations
        Checking 17 devices and 2 services
        Produced 2015-2-4 20:42:45
        From : Oldest available information
        To : 2015-2-4 20:42:45

Devices out of sync

p0

        check-sync unsupported for device

p1

        check-sync unsupported for device

p2

        check-sync unsupported for device

p3

        check-sync unsupported for device

pe0

        check-sync unsupported for device

pe1

        check-sync unsupported for device

pe3

        check-sync unsupported for device



Template discrepancies

gold-conf

        Discrepancies in device
        ce0
        ce1
        ce2
        ce3


Chapter : Details


Commit list

        SeqNo   ID      User    Client  Timestamp            Label  Comment
        0       10031   admin   cli     2015-02-04 20:31:42
        1       10030   admin   cli     2015-02-04 20:03:41
        2       10029   admin   cli     2015-02-04 19:54:40
        3       10028   admin   cli     2015-02-04 19:45:20
        4       10027   admin   cli     2015-02-04 18:38:05


Service commit changes

        No service data commits saved for the time interval


Device commit changes

        No device data commits saved for the time interval


Service differences

        No service data diffs found


Template discrepancies details

gold-conf

Device ce0

 config {
     ios:snmp-server {
+        community public {
+        }
     }
 }

Device ce1

 config {
     ios:snmp-server {
+        community public {
+        }
     }
 }

Device ce2

 config {
     ios:snmp-server {
+        community public {
+        }
     }
 }

Device ce3

 config {
     ios:snmp-server {
+        community public {
+        }
     }
 }
```

{% endcode %}

### Downloading Compliance Reports

NSO generates a report file and returns a `location` URL pointing to it after running a compliance report using the command `compliance reports <report-name> run outformat <format>` . This URL is a direct HTTP(S) link to the report, which can be downloaded, for example, using a standard tool like `curl` or using Python requests. With basic authentication, the tools authenticate with NSO using a username and password, and allow users to retrieve and save the report file locally for further processing, automation, or archiving. You must first establish a JSON-RPC session before downloading the report. If the connection is closed before requesting the file, as is typically done with `curl`, use the returned session cookie to download the report.

The examples below clarify how to make requests.

{% tabs %}
{% tab title="curl" %}
**Session-based authentication using the provided cookie to identify the session**

{% code title="Example" overflow="wrap" fullWidth="false" %}

```bash
# 1. Start a session and save the cookie
$ curl -X POST -H 'Content-Type: application/json' --cookie-jar cookie.txt -d '{"jsonrpc": "2.0", "id": 1, "method": "login", "params": {"user": "admin", "passwd": "admin"}}' http://localhost:8080/jsonrpc

# 2. Use the cookie to identify the session and download the report
$ curl --cookie cookie.txt --output report.txt "http://localhost:8080/compliance-reports/report_2025-10-09T13:48:32.663282+00:00.txt"
```

{% endcode %}
{% endtab %}

{% tab title="Python requests" %}
**Session-based authentication**

{% code title="Example" overflow="wrap" %}

```python
import requests

url = "http://localhost:8080/jsonrpc"

# 1. Start a session
session = requests.Session()
headers = {
    "Content-Type": "application/json"
}
data = {
    "jsonrpc": "2.0",
    "id": 1,
    "method": "login",
    "params": {
        "user": "admin",
        "passwd": "admin"
    }
}

response = session.post(url, json=data, headers=headers,verify=False)
print("Status code:", response.status_code)
print("Response:", response.text)
file_url = "http://localhost:8080/compliance-reports/report_2025-10-09T13:48:32.663282+00:00.txt"
filename = file_url.split("/")[-1]

# 2. Use the session to download the report
file_response = session.get(file_url, stream=True)

if file_response.status_code == 200:
    with open("report.txt", "wb") as f:
        for chunk in file_response.iter_content(chunk_size=8192):
            if chunk:
                f.write(chunk)
else:
    print(file_response.text)
```

{% endcode %}
{% endtab %}
{% endtabs %}

## Device Configuration Checks

Services are the preferred way to manage device configuration in NSO as they provide numerous benefits (see [Why services?](/guides/development/core-concepts/services#d5e536) in Development). However, on your journey to full automation, perhaps you only use NSO to configure a subset of all the services (configuration) on the devices. In this case, you can still perform generic configuration validation on other parts with the help of device configuration checks.

Often, each device will have a somewhat different configuration, such as its own set of IP addresses, which makes checking against a static template impossible. For this reason, NSO supports compliance templates.

These templates are similar to but separate from, device templates. With compliance templates, you use regular expressions to check compliance, instead of simple fixed values. You can also define and reference variables that get their values when a report is run. All selected devices are then checked against the compliance template and the differences (if any) are reported as a compliance violation.

You can create a compliance template from scratch. For example, to check that the router uses only internal DNS servers from the 10.0.0.0/8 range, you might create a compliance template such as:

```bash
admin@ncs(config)# compliance template internal-dns
admin@ncs(config-template-internal-dns)# ned-id router-nc-1.0 config sys dns server 10\\\\..+
```

Here, the value of the `/sys/dns/server` must start with `10.`, followed by any string (the regular expression `.+`). Since a dot has a special meaning with regular expressions (any character), it must be escaped with a backslash to match only the actual dot character. But note the required multiple escaping (`\\\\`) in this case.

Compliance-template values support W3C XML Schema regular expressions. The expression must match the entire configuration value. Supported constructs include `.`, character classes such as `[0-9]`, grouping such as `(admin|root)`, alternation (`|`), and the quantifiers `?`, `*`, `+`, and `{m,n}`. Do not use `^` or `$` as start or end anchors; they are treated as literal characters.

As these expressions can be non-trivial to construct, the templates have a `check` command that allows you to quickly check compliance for a set of devices, which is a great development aid.

{% code overflow="wrap" %}

```bash
admin@ncs(config)# show full-configuration devices device ex0 config sys dns server
devices device ex0
 config
  sys dns server 10.2.3.4
  !
  sys dns server 192.168.100.10
  !
 !
!
admin@ncs(config)# compliance template internal-dns
admin@ncs(config-template-internal-dns)# check device ex0
check-result {
    device ex0
    result violations
    diff  config {
     sys {
         dns {
+            # after server 10.2.3.4
+            /* No match of 10\\..+ */
+            server 192.168.100.10;
         }
     }
 }

}
```

{% endcode %}

To simplify template creation, NSO features the `/compliance/create-template` action that can initiate a compliance template from a set of device configurations or an existing device template. The resulting template can be used as-is or as a starting point for further refinement. For example:

In addition to extracting patterns from configuration already present in NSO or from an existing device template, the action can also consume configuration snippets directly. Snippets can be supplied either from a file on the NSO server filesystem or as inline payload data. Supported formats are NETCONF-style XML wrapped in a `<config>` element, Cisco XR style CLI (`cli-c`), Juniper curly-brace CLI (`cli-j`), and Juniper set commands (`cli-j-cmd`). Delete operations in the input, such as Cisco-style `no` commands or XML `operation="remove"` attributes, are translated into `absent` tags in the generated compliance template.

{% code overflow="wrap" %}

```bash
admin@ncs(config)# show full-configuration devices template use-internal-dns
devices template use-internal-dns
 ned-id router-nc-1.0
  config
   ! Tags: replace (/devices/template{use-internal-dns}/ned-id{router-nc-1.0:router-nc-1.0}/config/r:sys/dns)
   sys dns server 10.8.8.8
   !
  !
 !
!
admin@ncs(config)# compliance create-template name internal-dns device-template use-internal-dns
admin@ncs(config)# show configuration
compliance template internal-dns
 ned-id router-nc-1.0
  config
   ! Tags: replace (/compliance/template{internal-dns}/ned-id{router-nc-1.0:router-nc-1.0}/config/r:sys/dns)
   sys dns server 10.8.8.8
   !
  !
 !
!
admin@ncs(config)# compliance template internal-dns
admin@ncs(config-template-internal-dns)# ned-id router-nc-1.0 config sys dns server 10\\\\..+
```

{% endcode %}

By providing a list of device configuration paths, the `create-template` action can find common structural patterns in the device configurations and create a compliance template based on it.

The algorithm works by traversing the data depth-first, keeping track of the rate of occurrence of configuration nodes, and any values that compare equal. Values that do not compare equal are made into regex match-all expressions. For example:

{% code overflow="wrap" %}

```bash
admin@ncs(config)# compliance create-template name syslog path [ /devices/device[device-type/netconf/ned-id='router-nc-1.0:router-nc-1.0']/config/sys/syslog ]
admin@ncs(config)# show configuration                                                                   compliance template syslog
 ned-id router-nc-1.0
  config
   sys syslog server 10.3.4.5
    enabled
    selector 8
     facility [ .* ]
    !
   !
  !
 !
!
admin@ncs(config)# commit
Commit complete.
```

{% endcode %}

The action takes a number of arguments to control how the resulting template looks:

* `path` - A list of XPath 1.0 expressions pointing into `/devices/device/config` to create the template from. The template is only created from the paths that are common in the node-set.
* `match-rate` - Device configuration is included in the resulting template based on the rate of occurrence given by this setting.
* `exclude-service-config` - Exclude configuration that is already under service management.
* `collapse-list-keys` - Decides what lists to do matching on, either `all`, `automatic` (default), or those specified by the `list-path` parameter. The default is to find those lists that differ among the device configurations.

Finally, to use compliance templates in a report, reference them from `device-check/template`:

```bash
admin@ncs(config-report-gold-check)# device-check template internal-dns
admin@ncs(config-template-internal-dns)# exit
admin@ncs(config-report-gold-check)# device-check template syslog
```

{% hint style="info" %}
By default the schemas for compliance templates are not accessible from application client libraries such as MAAPI. This reduces the memory usage for large device data models. The schema can be made accessible with the `/ncs-config/enable-client-template-schemas` setting in `ncs.conf`.
{% endhint %}

## Device Live-Status Checks

In addition to configuration, compliance templates can also check for operational data. This can be used, for example, to check device interface statuses and device software versions.

This feature is opt-in and requires the NEDs to be re-compiled with the `--ncs-with-operational-compliance` [ncsc(1)](/guides/resources/man/ncsc.1) flag. Instructions on how to re-compile a NED is included in each NED package.

```bash
admin@ncs(config)# compliance template interface-up
admin@ncs(config-template-interface-up)# ned-id router-nc-1.0 live-status sys interfaces interface eth0
admin@ncs(config-interface-eth0)# status link up
admin@ncs(config-interface-eth0)# commit
Commit complete.
```

When running a check against a device, the result will show a violation if the status of the interface link is not up.

```bash
admin@ncs(config-template-interface-up)# check device [ ex0 ]
check-result {
    device ex0
    result violations
    diff  live-status {
     sys {
         interfaces {
             interface eth0 {
                 status {
-                    link up;
+                    link down;
                 }
             }
         }
     }
 }

}
```

{% hint style="info" %}
Running checks on live-status data is slower than configuration data since it requires connecting to devices to read the data. In comparison, configuration data is checked against data in CDB.
{% endhint %}

## Additional Template Functionality

In some cases, it is insufficient to only check that the required configuration is present, as other configurations on the device can interfere with the desired functionality. For example, a service may configure a routing table entry for the 198.51.100.0/24 network. If someone also configures a more specific entry, say 198.51.100.0/28, that entry will take precedence and may interfere with the way the service requires the traffic to be routed. In effect, this additional configuration can render the service inoperable.

### `strict` Checks

To help operators ensure there is no such extraneous configuration on the managed devices, the compliance reporting feature supports the so-called `strict` mode. This mode not only checks whether the required configuration is present but also reports any configuration present on the device that is not part of the template.

You can enable `strict` mode in a report definition or when running a compliance-template check directly, for example:

```bash
ncs(config)# compliance template interfaces check device ios0 strict
```

Consider the following template and device configuration:

```bash
compliance template interfaces
 ned-id cisco-ios-cli-3.8
  config
   interface GigabitEthernet 0/0
    ip address 192.168.1.1
    ip address 255.255.255.0
   !
  !
 !
!
devices device ios0
 config
  interface GigabitEthernet0/0
   duplex full
   ip address 192.168.1.1 255.255.255.0
   no shutdown
  exit
 !
!
```

The device will be compliant with a regular template check. When using `strict`, all unexpected configuration will be shown in the diff.

```bash
check-result {
    device ios0
    result violations
    diff  config {
     interface {
+        FastEthernet 0/0 {
+        }
+        FastEthernet 1/0 {
+        }
         GigabitEthernet 0/0 {
+            duplex full;
         }
     }
 }

}
```

### `strict` Sub-Tree Tag

A `strict` tag enables strict checking for the tagged configuration statement and its subtree. For a list, the tag can only be applied to a specific list instance, including its key values. It cannot be applied to the list node without selecting an instance.\
\
For example, a `strict` tag on `username admin` checks for unexpected configuration below the `admin` entry, but it does not detect other `username` entries. To report usernames other than those represented in the template, such as `admin` and `root`, run the compliance check in whole-template `strict` mode.

The following example applies the `strict` tag to a specific interface list instance:

```bash
ncs(config)# tag add compliance template interfaces ned-id cisco-ios-cli-3.8 config interface GigabitEthernet 0/0 strict
```

This example adds the `strict` tag to the specific `GigabitEthernet 0/0` list instance, so strict checking applies only to that instance’s subtree.

```bash
admin@ncs(config)# compliance template interfaces check device ios0                                       check-result {
    device ios0
    result violations
    diff  config {
     interface {
         GigabitEthernet 0/0 {
+            duplex full;
         }
     }
 }

}
```

### `allow-empty` Tag

A compliance template can be used on many different devices. The configuration on the devices, however, is not always identical. The following template checks that interfaces are set to be reachable.

```bash
compliance template no-unreachables
 ned-id cisco-ios-cli-3.8
  config
   interface FastEthernet *
    ip unreachables false
   !
   interface GigabitEthernet *
    ip unreachables false
   !
  !
 !
!
devices device ios0
 config
  interface GigabitEthernet0/0
   duplex full
   ip address 192.168.1.1 255.255.255.0
   no ip unreachables
   no shutdown
  exit
 !
!
```

The device in this example only has a `GigabitEthernet` interface which will result in a violation.

```bash
ncs(config)# compliance template no-unreachables check device ios0
check-result {
    device ios0
    result violations
    diff  config {
     interface {
-        FastEthernet .* {
-        }
     }
 }

}
```

In this case, we are only interested in interfaces that are actually configured on the device. This is where the `allow-empty` tag comes in. By setting this tag on each interface, the check will only be run if there are interfaces configured of that type.

```bash
ncs(config)# tag add compliance template no-unreachables ned-id cisco-ios-cli-3.8 config interface FastEthernet .* allow-empty
ncs(config)# tag add compliance template no-unreachables ned-id cisco-ios-cli-3.8 config interface GigabitEthernet .* allow-empty
```

With this tag, the device will no longer have any violations.

```bash
ncs(config)# compliance template no-unreachables check device ios0                                  check-result {
    device ios0
    result no-violation
}
```

It will still result in violations if the configuration is incorrect, but not if it's empty.

### `absent` and `delete` Tags

To ensure that configuration does not exist on a device, add the `absent` tag to the corresponding node in a compliance template. The `delete` tag can also be used and is synonymous with `absent` in compliance templates.

Both tags assert that the tagged configuration must be absent. If neither tag is specified, the compliance template asserts that the configuration is present.

```bash
devices device ios0
 config
  no service password-encryption
  service finger
 !
!
compliance template no-finger
 ned-id cisco-ios-cli-3.8
  config
   ! Tags: absent
   service finger
  !
 !
!
```

Using the `delete` tag instead of `absent` in this template produces the same compliance result.

This template will result in a violation if `service finger` is configured on the device.

```bash
ncs(config)# compliance template no-finger check device ios0
check-result {
    device ios0
    result violations
    diff  config {
     service {
+        finger;
     }
 }

}
```

## XML Compliance Templates

In addition to CDB-based compliance templates (configured under `/compliance/template`), NSO supports **XML compliance templates**: standalone `.xml` files that are loaded from the file system at startup, in the same way service templates are loaded. This makes it possible to version-control compliance checks alongside service templates and ship them as part of a package.

### XML Template File Format

An XML compliance template is an `.xml` file whose root element is `<compliance-template>` with the namespace `http://tail-f.com/ns/config/1.0`. The `name` attribute gives the template its unique name:

```xml
<compliance-template xmlns="http://tail-f.com/ns/config/1.0"
                     name="my-template">
  <devices xmlns="http://tail-f.com/ns/ncs">
    <device>
      <name>{$DEVICE}</name>
      <config>
        <!-- device configuration to verify against, values as regexes -->
      </config>
    </device>
  </devices>
</compliance-template>
```

Values in the template are treated as regular expressions, and variables use the `{$VAR_NAME}` syntax and are substituted at check time.

### Loading XML Compliance Templates

**Via `ncs.conf` (file-system directories)**

Add one or more directory paths to `/ncs-config/load-path/templates-dir/compliance` in `ncs.conf`:

```xml
<load-path>
  <templates-dir>
    <compliance>/opt/ncs/compliance-templates</compliance>
    <compliance>/etc/ncs/site-compliance</compliance>
  </templates-dir>
</load-path>
```

NSO scans each configured directory recursively for `.xml` files at startup and loads them as compliance templates.

**Via packages**

XML compliance templates can be placed in the `templates/` subdirectory of a package alongside service templates. NSO automatically discovers and loads compliance templates from packages when they are loaded. Templates provided by a package are listed under `/packages/package/compliance-templates`.

### Processing Instructions

XML compliance templates support all processing instructions available for [service templates](/guides/development/core-concepts/templates), **except** `for`, `foreach`, and `copy-tree`. The supported instructions are:

| Instruction                                    | Description                               |
| ---------------------------------------------- | ----------------------------------------- |
| `<?if {xpath-expr}?>` … `<?else?>` … `<?end?>` | Conditionally check parts of the template |
| `<?set var = {xpath-expr}?>`                   | Assign a variable                         |
| `<?if-ned-id ned-id?>` … `<?end?>`             | Conditionally check based on the NED ID   |
| `<?macro name args?>` … `<?endmacro?>`         | Define a reusable template fragment       |
| `<?expand macro-name args?>`                   | Expand a macro                            |

#### `<?assert?>` Processing Instruction

The `assert` instruction is new and specific to XML compliance templates. It evaluates an XPath boolean expression and, if the expression evaluates to false, records a failure with the provided error message. Failed assertions are returned in the output of the `check` action.

Syntax:

```xml
<?assert {xpath-boolean-expression} error message text?>
```

For example, to assert that exactly two ethernet interfaces exist on a device:

```xml
<?set num = {count(./interface[starts-with(name, 'eth')])}?>
...
<?assert {$num = 2}
 Expected 2 interfaces starting with 'eth' - found {$num}?>
```

When an assertion fails, the `check` action output includes a `failed-assertions` list containing the `location`, `context`, and `message` of each failure.

### Inspecting Loaded XML Compliance Templates

All loaded XML compliance templates are visible under `/compliance/xml-templates/template`:

```
ncs# show compliance xml-templates template
NAME         FILENAME                                   PACKAGE  VARIABLES
----------------------------------------------------------------------------
compliance1  /opt/ncs/compliance-templates/comp1.xml    -        [DEVICE]
dns_check    packages/dns/templates/dns_compliance.xml  dns      [DEVICE]
```

Each entry shows the template `name`, `filename` (path on disk), `package` (empty for templates loaded from the load-path rather than a package), and the list of `variables` used in the template.

Any loading errors — such as duplicate template names, malformed XML, or unsupported processing instructions — are reported in `/compliance/xml-templates/error-info`:

```
ncs# show compliance xml-templates error-info
compliance xml-templates error-info "wrong_assert.xml:10 failed to compile: 'foo()' reason: Undefined function foo/0"
compliance xml-templates error-info "pi_for.xml:9 The processing instruction 'for' is unsupported."
```

### Running a Compliance Check Against an XML Template

Use the `check` action under `/compliance/xml-templates`, supplying the template name and the target devices:

```
ncs# compliance xml-templates check template-name dns_check \
     device [ router0 router1 ]
check-result {
    device router0
    result no-violation
}
check-result {
    device router1
    result violations
    diff devices device router1
 config
  no sys dns options timeout 30
  no sys dns options attempts 2
 !
!
}
```

You can also target all devices or a device group:

```
ncs# compliance xml-templates check template-name dns_check all-devices
ncs# compliance xml-templates check template-name dns_check device-group core-routers
```

When the template uses `assert` instructions, any failures appear in a `failed-assertions` list within the check result:

```
check-result {
    device core-rtr-1
    result violations
    failed-assertions {
        location ./templates/interface-check.xml:10
        context /devices/device[name='core-rtr-1']/config/r:sys/interfaces
        message Expected 2 interfaces starting with 'eth' - found 1
    }
}
```

### Reloading XML Compliance Templates

To reload all templates (both from the load-path and from packages), use:

```
ncs# packages reload
```

To reload only the templates loaded from the load-path (i.e., not from packages), use the dedicated reload action:

```
ncs# compliance xml-templates reload
```

{% hint style="info" %}
As of NSO 6.7, XML compliance templates cannot be configured in compliance reports. Support for referencing XML templates from compliance report definitions (via `device-check/xml-template`) is planned for the next maintenance release.
{% endhint %}

## Re-run existing compliance report results

When a compliance report result includes violations, you typically want to run that compliance report again after fixing the violations. What if you could run only the violating items from the old report results again, skipping everything else?

Re-running only the violating items is possible by using the `/compliance/report-results/report/re-run` action.

### Example of re-run

First, we produce a compliance report result. This example assumes that you already know how to create compliance reports. Here, we use a compliance report called `scale-test`.

```bash
ncs(config)# compliance reports report scale-test run outformat text
time 2026-04-09T20:42:45.019012+00:00
compliance-status violations
info Checking 4 devices and 4 services
```

We now inspect this compliance report result and find what causes these violations. Maybe some leaf values are incorrectly configured, maybe services are out of sync. There are six violations in this run, on two devices and four services.

We re-run the report result and notice that only the devices and services involved in violations are now part of the new re-run report result.

```bash
ncs(config)# compliance report-results report 2026-04-09T20:42:45.019012+00:00 re-run outformat text
time 2026-04-09T20:44:45.019012+00:00
compliance-status violations
info Checking 2 devices and 4 services
```

We fix these violations, and re-run the compliance report based on the new report result.

```bash
ncs(config)# compliance report-results report 2026-04-09T20:44:45.019012+00:00 re-run outformat text
time 2026-04-09T20:46:48.019012+00:00
compliance-status no-violation
info Checking 2 devices and 4 services
```

Note that `compliance-status` is now `no-violation`. We have managed to fix all violations and we did not need to re-run the full compliance report in order to achieve this. This saves time and makes the resulting report smaller and easier to read.

If we re-run the successful report result, we see that this final re-run does not include any devices nor any services.

```bash
ncs(config)# compliance report-results report 2026-04-09T20:46:48.019012+00:00 re-run outformat text
time 2026-04-09T20:49:48.019012+00:00
compliance-status no-violation
info Checking no devices and no services
```

### Renamed devices and removed services

The `/compliance/report-results/report/re-run` action will ignore any renamed devices and removed services, these will be dropped from the resulting new compliance report result.


# Listing Packages

View currently loaded packages.

NSO packages contain data models and code for a specific function. It might be a NED for a specific device, a service application like MPLS VPN, a WebUI customization package, etc. Packages can be added, removed, and upgraded in run time.

The currently loaded packages can be viewed with the following command:

{% code title="Show Currently Loaded Packages" %}

```bash
admin@ncs# show packages
packages package cisco-ios
 package-version 3.0
 description     "NED package for Cisco IOS"
 ncs-min-version [ 3.0.2 ]
 directory       ./state/packages-in-use/1/cisco-ios
 component upgrade-ned-id
  upgrade java-class-name com.tailf.packages.ned.ios.UpgradeNedId
 component cisco-ios
  ned cli ned-id  cisco-ios
  ned cli java-class-name com.tailf.packages.ned.ios.IOSNedCli
  ned device vendor Cisco
NAME      VALUE
---------------------
show-tag  interface

 build-info date "2015-01-29 23:40:12"
 build-info file ncs-3.4_HEAD-cisco-ios-3.0.tar.gz
 build-info arch linux.x86_64
 build-info java "compiled Java class data, version 50.0 (Java 1.6)"
 build-info package name cisco-ios
 build-info package version 3.0
 build-info package ref 3.0
 build-info package sha1 a8f1329
 build-info ncs version 3.4_HEAD
 build-info ncs sha1 81a1e4c
 build-info dev-support version 0.99
 build-info dev-support branch e4d3fa7
 build-info dev-support sha1 e4d3fa7
 oper-status up
```

{% endcode %}

Thus, the above command shows that NSO currently has only one package loaded, the NED package for Cisco IOS. The output includes the name and version of the package, the minimum required NSO version, the Java components included, package build details, and finally the operational status of the package. The operational status is of particular importance—if it is anything other than `up`, it indicates that there was a problem with the loading or the initialization of the package. In this case, an item `error-info` may also be present, giving additional information about the problem. To show only the operational status for all loaded packages, this command can be used:

```bash
admin@ncs# show packages package * oper-status
packages package cisco-ios
 oper-status up
```


# Lifecycle Operations

Manipulate and manage existing services and devices.

Devices and services are the most important entities in NSO. Once created, they may be manipulated in several different ways. The three main categories of operations that affect the state of services and devices are:

* **Commit Parameters:** Commit parameters modify the transaction semantics. In the CLI they are exposed as commit flags.
* **Device Actions:** Explicit actions that modify the devices.
* **Service Actions:** Explicit actions that modify the services.

The purpose of this section is more of a quick reference guide, an enumeration of available commands. The context in which these commands should be used is found in other parts of the documentation.

## Commit Parameters <a href="#d5e5048" id="d5e5048"></a>

NSO uses the shared YANG module `tailf-ncs-commit-params.yang` to define commit parameters, dry-run parameters, and commit results for northbound interfaces and SDK APIs. The groupings and `sx:structure` definitions in this module are reused by NSO modules such as `tailf-netconf-ncs.yang`, `tailf-restconf-ncs.yang`, `tailf-yang-patch-ncs.yang`, `tailf-ncs-rollback.yang`, `tailf-ncs-services.yang`, and `tailf-ncs-devices.yang`.

In the CLI, these parameters are exposed as commit flags:

```cli
commit label nightly-batch no-networking
```

The same shared model is used by the main northbound interfaces:

* JSON-RPC passes a structured `params` object that matches `tailf-ncs-commit-params:commit-params`.
* RESTCONF passes the same structure as base64-encoded JSON in the `params` query parameter or the `X-Cisco-NSO-Commit-Params` header. The header form is an alternative to the `params` query parameter when the client prefers not to place the often long, URL-encoded base64 commit-parameter payload in the request URI, for example when other query parameters are also used.
* NETCONF augments `commit`, `edit-config`, `copy-config`, and `prepare-transaction` with the same parameters.
* MAAPI and the language SDKs expose helpers for the built-in parameters and generic tagged-value or XML access for augmented parameters.

Some of these parameters may be configured to apply globally for all commits under `/devices/global-settings`, or per device profile under `/devices/profiles`.

The shared commit parameters are:

* `label`, `comment`: Add user-defined metadata to the transaction. The data is visible in rollback files, compliance reports, notifications, and events. If supported, it is also propagated to participating devices.
* `dry-run`: Validate and return the resulting changes without updating CDB or devices. The `outformat` leaf selects `xml`, `cli`, `native`, or `cli-c`. The `reverse` leaf can be used with `native` or `cli-c` output to show the commands needed to return to the current running state. The `with-service-meta-data` leaf includes FASTMAP service metadata in the diff.
* `confirm-network-state`: Check device state as part of the commit and process out-of-band changes according to policy. The `re-evaluate-policies` leaf also reprocesses out-of-band policies for services touched by the commit. The `re-deploy-all` leaf expands the operation to also re-deploy all services affected by discovered out-of-band data; without it, the impact stays scoped to the original transaction.
* `no-networking`: Update CDB but do not send configuration southbound. The affected devices become out of sync until the change is pushed later, for example with `sync-to`.
* `no-out-of-sync-check`: Continue even if NSO detects that a device is out of sync.
* `no-overwrite`: Perform a fine-grained check that the data NSO is about to modify has not changed on the device compared to NSO's view.
* `no-revision-drop`: Fail instead of silently dropping configuration that is not supported by an older device model revision.
* `no-deploy`: Write service data without invoking the service create callback.
* `reconcile`: Reconcile service ownership for existing configuration. The `keep-non-service-config`, `discard-non-service-config`, `attach-non-service-config`, and `detach-non-service-config` leafs control how non-service-owned data is handled. The `include` and `exclude` leaf-lists limit reconciliation to selected service configuration paths.
* `use-lsa`, `no-lsa`: Control whether LSA nodes are handled as LSA nodes or as ordinary devices.
* `commit-queue`: Commit through the commit queue instead of pushing configuration transactionally in the same operation. The `async`, `sync`, and `bypass` leafs select the queue behavior. The `sync` container accepts either `timeout` or `infinity`. The `lock`, `block-others`, `atomic`, and `error-option` leafs control the resulting queue item. See [Commit Queue](https://nso-docs.cisco.com/guides/operation-and-usage/operations/pages/auKQMOAF2p1jiGYJBweP#user_guide.devicemanager.commit-queue) for details and error-recovery behavior.

Some combinations of parameters are not allowed. For example, `dry-run` cannot be combined with `no-overwrite`, `no-out-of-sync-check`, or `commit-queue`, and `use-lsa` cannot be combined with `no-lsa`.

The CLI also has local commit modifiers such as `check` and `and-quit`. These affect CLI behavior but are not part of `tailf-ncs-commit-params.yang`.

### Augmenting Commit Parameters

Developers can augment the `sx:structure commit-params` structure in `tailf-ncs-commit-params.yang` with their own parameters. Once augmented, the new parameters become available wherever the shared commit-parameter model is used, including propagation to lower LSA nodes.

```yang
module example-commit-params {
  yang-version 1.1;
  namespace "http://example.com/example-commit-params";
  prefix ecp;

  import ietf-yang-structure-ext {
    prefix sx;
  }
  import tailf-ncs-commit-params {
    prefix ncp;
  }

  sx:augment-structure "/ncp:commit-params" {
    container audit-context {
      presence "Attach extra audit information to the commit";
      leaf ticket-id {
        type string;
      }
    }
  }
}
```

### Accessing From User Code

Augmented commit parameters are accessible through MAAPI together with the built-in ones:

* Python: Use `trans.get_params()` or `trans.get_trans_params()`. Built-in parameters have helper methods on `CommitParams`, while augmented parameters can be read through `CommitParams.root` or `CommitParams.get_tagvalues()`.
* Java: Use `maapi.getTransParams(th)`. Built-in parameters have helper methods on `CommitParams`, while augmented parameters can be inspected through `CommitParams.getConfXMLParam()`.
* Lower-level MAAPI APIs expose the same data as tagged values or XML parameter arrays.

The [examples.ncs/sdk-api/maapi-commit-parameters](https://github.com/NSO-developer/nso-examples/tree/6.7/sdk-api/maapi-commit-parameters) example shows how to augment the shared commit-parameter model, how to access built-in commit parameters from Python and Java user code, and how to access augmented parameters from Python and Java user code. The Java package also illustrates the lower-level `CommitParams.getConfXMLParam()` access pattern for augmented parameters. A northbound example covering CLI, RESTCONF, NETCONF, and JSON-RPC is available in [examples.ncs/northbound-interfaces/commit-parameters](https://github.com/NSO-developer/nso-examples/tree/6.7/northbound-interfaces/commit-parameters).

All commands in NSO can also have pipe commands. A useful pipe command for commit is `details`:

```cli
ncs% commit | details
```

This will give feedback on the steps performed in the commit.

When working with templates, there is a pipe command `debug` which can be used to troubleshoot templates. To enable debugging on all templates use:

```cli
ncs% commit | debug template
```

When configuring using many templates the debug output can be overwhelming. For this reason, there is an option to only get debug information for one template, in this example, a template named `l3vpn`:

```cli
ncs% commit | debug template l3vpn
```

## Device Actions <a href="#d5e5227" id="d5e5227"></a>

Actions for devices can be performed globally on the `/devices` path and for individual devices on `/devices/device/name`. Many actions are also available on device groups as well as device ranges.

<details>

<summary><code>add-capability</code></summary>

This action adds a capability to the list of capabilities. If `uri` is specified, then it is parsed as a YANG capability string and `module`, `revision`, `feature` and `deviation` parameters are derived from the string. If `module` is specified, then the namespace is looked up in the list of loaded namespaces, and the capability string is constructed automatically. If the `module` is specified and the attempt to look it up fails, then the action does nothing. If `module` is specified or can be derived from the capability string, then the `module` is also added/replaced in the list of modules. This action is only intended to be used for pre-provisioning; it is not possible to override capabilities and modules provided by the NED implementation using this action.

</details>

<details>

<summary><code>apply-template</code></summary>

Take a named template and apply its configuration here.

If the `accept-empty-capabilities` parameter is included, the template is applied to devices even if the capability of the device is unknown.

This action will behave differently depending on whether it is invoked with a transaction or not. When invoked with a transaction (such as via the CLI) it will apply the template to it and leave it to the user to commit or revert the resulting changes. If invoked without a transaction (for example when invoked via RESTCONF), the action will automatically create one and commit the resulting changes. An error will be returned and the transaction aborted if the template failed to apply on any of the devices.

The `device-select` option takes an XPath 1.0 expression that applies the action to the selected devices. The XPath expression can be a location path or an expression evaluated as a predicate to the `/devices/device` list. The `device-group` option takes a list of group names that expand to their group members. The `device`, `device-select`, and `device-group` options can be combined.

</details>

<details>

<summary><code>check-sync</code></summary>

Check if the NSO copy of the device configuration is in sync with the actual device configuration, using device-specific mechanisms. This operation is usually cheap as it only compares a signature of the configuration from the device rather than comparing the entire configuration.

Depending on the device the signature is implemented as a transaction-id, timestamp, hash-sum, or not at all. The capability must be supported by the corresponding NED. The output might say unsupported, and then the only way to perform this would be to do a full `compare-config` command.

As some NEDs implement the signature as a hash-sum of the entire configuration, this operation might for some devices be just as expensive as performing a full `compare-config` command.

The `device-select` option takes an XPath 1.0 expression that applies the action to the selected devices. The XPath expression can be a location path or an expression evaluated as a predicate to the `/devices/device` list. The `device-group` option takes a list of group names that expand to their group members. The `device`, `device-select`, and `device-group` options can be combined.

</details>

<details>

<summary><code>check-yang-modules</code></summary>

Check if the device YANG modules loaded by NSO have revisions that are compatible with the ones reported by the device.

This can indicate for example that the device has a YANG module of later revision than the corresponding NED.

The `device-select` option takes an XPath 1.0 expression that applies the action to the selected devices. The XPath expression can be a location path or an expression evaluated as a predicate to the `/devices/device` list. The `device-group` option takes a list of group names that expand to their group members. The `device`, `device-select`, and `device-group` options can be combined.

</details>

<details>

<summary><code>clear-trace</code></summary>

Clear all trace files for all active traces for all managed devices.

The `device-select` option takes an XPath 1.0 expression that applies the action to the selected devices. The XPath expression can be a location path or an expression evaluated as a predicate to the `/devices/device` list. The `device-group` option takes a list of group names that expand to their group members. The `device`, `device-select`, and `device-group` options can be combined.

</details>

<details>

<summary><code>compare-config</code></summary>

Retrieve the config from the device and compare it to the NSO locally stored copy.

</details>

<details>

<summary><code>connect</code></summary>

Set up a session to the unlocked device. This is not used in real operational scenarios. NSO automatically establishes connections on demand. However, it is useful for test purposes when installing new NEDs, adding devices, etc.

When a device is southbound locked, all southbound communication is turned off. The `override-southbound-locked` flag overrides the southbound lock for connection attempts. Thus, this is a way to update the capabilities including revision information for a managed device although the device is southbound locked.

The `device-select` option takes an XPath 1.0 expression that applies the action to the selected devices. The XPath expression can be a location path or an expression evaluated as a predicate to the `/devices/device` list. The `device-group` option takes a list of group names that expand to their group members. The `device`, `device-select`, and `device-group` options can be combined.oup members. The `device`, `device-select`, and `device-group` options can be combined.

</details>

<details>

<summary><code>copy-capabilities</code></summary>

This action copies the list of capabilities and the list of modules from another device or profile. When used on a device, this action is only intended to be used for pre-provisioning: it is not possible to override capabilities and modules provided by the NED implementation using this action.

Note that this action overwrites the existing list of capabilities.

</details>

<details>

<summary><code>delete-config</code></summary>

Delete the device configuration in NSO without executing the corresponding delete on the managed device.

</details>

<details>

<summary><code>disconnect</code></summary>

Close all sessions to the device.

The `device-select` option takes an XPath 1.0 expression that applies the action to the selected devices. The XPath expression can be a location path or an expression evaluated as a predicate to the `/devices/device` list. The `device-group` option takes a list of group names that expand to their group members. The `device`, `device-select`, and `device-group` options can be combined.

</details>

<details>

<summary><code>fetch-ssh-host-keys</code></summary>

Retrieve the SSH host keys from all devices, or all devices in the given device group, and store them in each device's `ssh/host-key` list. Successfully retrieved new or updated keys are always committed by the action.

The `device-select` option takes an XPath 1.0 expression that applies the action to the selected devices. The XPath expression can be a location path or an expression evaluated as a predicate to the `/devices/device` list. The `device-group` option takes a list of group names that expand to their group members. The `device`, `device-select`, and `device-group` options can be combined.

</details>

<details>

<summary><code>find-capabilities</code></summary>

This action populates the list of capabilities based on the configured ned-id for the device, if possible. NSO will look up the package corresponding to the ned-id and add all the modules from these packages to the list of device capabilities and list of modules. It is the responsibility of the caller to verify that the automatically populated list of capabilities matches the actual device's capabilities. The list of capabilities can then be fine-tuned using `add-capability` and `capability/remove` actions. Currently, this approach will only work for CLI and generic devices. This action is only intended to be used for pre-provisioning: it is not possible to override capabilities and modules provided by the NED implementation using this action.

Note that this action overwrites the existing list of capabilities.

</details>

<details>

<summary><code>instantiate-from-other-device</code></summary>

Instantiate the configuration for the device as a copy of the configuration of some other already working device.

</details>

<details>

<summary><code>load-native-config</code></summary>

Load configuration data in native format into the transaction. This action is only applicable to devices with NETCONF, CLI, and generic NEDs.

The action can load the configuration data either from a file in the local filesystem or as a string through the northbound client. If loading XML the data must be a valid XML document, either with a single namespace or wrapped in a config node with the <http://tail-f.com/ns/config/1.0> namespace.

The `verbose` option can be used to show additional parse information reported by the NED. By default, the behavior is to merge the configuration that is applied. This can be changed by setting the `mode` option to replace. This will replace the entire device configuration.

This action will behave differently depending on if it is invoked with a transaction or not. When invoked with a transaction (such as via the CLI), it will load the configuration into it and leave it to the user to commit or revert the resulting changes. If invoked without a transaction (for example, when invoked via RESTCONF), the action will automatically create one and commit the resulting changes.

Since NSO 6.4 the `load-native-config` will create list entries with sharedCreate() and set leafs with sharedSet() if invoked inside a service for refcounters and backpointers to be created or updated.

</details>

<details>

<summary><code>migrate</code></summary>

Change the NED identity and migrate all data. As a side-effect reads and commits the actual device configuration.

The action reports what paths have been modified and the services affected by those changes. If the `verbose` option is used, all service instances are reported instead of just the service points. If the `dry-run` option is used, the action simply reports what it would do.

If the `no-networking` option is used, no southbound traffic is generated toward the devices. Only the device configuration in CDB is used for the migration. If used, NSO can not know if the device is in sync. To determine this, the **compare-config** or the **sync-from** action must be used.

The `device-select` option takes an XPath 1.0 expression that applies the action to the selected devices. The XPath expression can be a location path or an expression evaluated as a predicate to the `/devices/device` list. The `device-group` option takes a list of group names that expand to their group members. The `device`, `device-select`, and `device-group` options can be combined.

</details>

<details>

<summary><code>partial-sync-from</code></summary>

Synchronize parts of the devices' configuration by pulling from the network.

</details>

<details>

<summary><code>ping</code></summary>

ICMP pings the device.

</details>

<details>

<summary><code>scp-from</code></summary>

Securely copy the file from the device.

The `port` option specifies the port to connect to on the device. If this leaf is not configured, NSO will use the port for the management interface of the device.

The `preserve` option preserves modification times, access times, and modes from the original file. This is not always supported by the device.

The `protocol` option selects which protocol to use for the file transfer. SCP (default) or SFTP.

</details>

<details>

<summary><code>scp-to</code></summary>

Securely copy the file to the device.

The `port` option specifies the port to connect to on the device. If this leaf is not configured, NSO will use the port for the management interface of the device.

The `preserve` option preserves modification times, access times, and modes from the original file. This is not always supported by the device.

The `protocol` option selects which protocol to use for the file transfer. SCP (default) or SFTP.

The `device-select` option takes an XPath 1.0 expression that applies the action to the selected devices. The XPath expression can be a location path or an expression evaluated as a predicate to the `/devices/device` list. The `device-group` option takes a list of group names that expand to their group members. The `device`, `device-select`, and `device-group` options can be combined.

</details>

<details>

<summary><code>sync-from</code></summary>

Synchronize the NSO copy of the device configuration by reading the actual device configuration. The change will be immediately committed to NSO.

If the `dry-run` option is used, the action simply reports (in different formats) what it would do. The `verbose` option can be used to show additional parse information reported by the NED.

If you have any services that have created a configuration on the device, the corresponding service might be out of sync. Use the commands `check-sync` and `re-deploy` to reconcile this.

</details>

<details>

<summary><code>sync-to</code></summary>

Synchronize the device configuration by pushing the NSO copy to the device.

NSO pushes a minimal diff to the device. The diff is calculated by reading the configuration from the device and comparing it with the configuration in NSO.

If the `dry-run` option is used, the action simply reports (in different formats) what it would do.

Some of the operations above can't be performed while the device is being committed to (or waiting in the commit queue). This is to avoid getting inconsistent data when reading the configuration. The `wait-for-lock` option in these specifies a timeout to wait for a device lock to be placed in the commit queue. The lock will be automatically released once the action has been executed. If the `no-wait-for-lock` option is specified, the action will fail immediately for the device if the lock is taken for the device or if the device is placed in the commit queue. The `wait-for-lock` and the `no-wait-for-lock` options are device settings as well; they can be set as a device profile, device, and global setting. The `no-wait-for-lock` option is set in the global settings by default. If neither `wait-for-lock` and the `no-wait-for-lock` options are provided together with the action, the device setting is used.

The `device-select` option takes an XPath 1.0 expression that applies the action to the selected devices. The XPath expression can be a location path or an expression evaluated as a predicate to the `/devices/device` list. The `device-group` option takes a list of group names that expand to their group members. The `device`, `device-select`, and `device-group` options can be combined.

</details>

## Service Actions <a href="#d5e5403" id="d5e5403"></a>

Service actions are performed on the service instance.

<details>

<summary><code>check-sync</code></summary>

Check if the service has been undermined, i.e., if the service was to be redeployed, would it do anything? This action will invoke the FASTMAP code to create the change set that is compared to the existing data in CDB locally.

If `outformat` is a boolean, `true` is returned if the service is in sync, i.e., a re-deploy would do nothing. If `outformat` is `cli`, `xml` or `native`, the changes that the service would do to the network if re-deployed are returned.

If configuration changes have been made out-of-band, then `deep-check-sync` is needed to detect an out-of-sync condition.

The `deep` option is used to recursively `check-sync` stacked services. The `shallow` option only `check-sync` the topmost service.

If the parameter `with-service-meta-data` is given, service meta-data will also be considered when determining if the service is in sync. This provides a more comprehensive check that includes both configuration data and service meta-data.

</details>

<details>

<summary><code>deep-check-sync</code></summary>

Check if the service has been undermined on the device itself. The action `check-sync` compares the output of the service code to what is stored in CDB locally. This action retrieves the configuration from the devices touched by the service and compares the forward diff set of the service to the retrieved data. This is thus a fairly heavyweight operation. As opposed to the `check-sync` action that invokes the FASTMAP code, this action re-applies the forward diff-set. This is the same output you see when inspecting the `get-modifications` operational field in the service instance.

If the device is in sync with CDB, the output of this action is identical to the output of the cheaper `check-sync` action.

</details>

<details>

<summary><code>get-modifications</code></summary>

Returns the data the service modified, either in CLI curly bracket format or NETCONF XML edit-config format. The modifications are shown as if the service instance was the only instance that modifies the data. This data is only available if the parameter `/services/global-settings/collect-forward-diff` is set to true.

If the parameter `reverse` is given, the modifications needed to reverse the effect of the service is shown. The modifications are shown as if this service instance was the last service instance. This will be applied if the service is deleted. This data is always available.

The `deep` option is used to recursively `get-modifications` for stacked services. The `shallow` option only `get-modifications` for the topmost service.

</details>

<details>

<summary><code>re-deploy</code></summary>

Run the service code again, possibly writing the changes of the service to the network once again. There are several reasons for performing this operation, such as:

* a `device sync-from` action has been performed to incorporate an out-of-band change.
* data referenced by the service has changed such as topology information, QoS policy definitions, etc.

The `deep` option is used to recursively `re-deploy` stacked services. The `shallow` option only `re-deploy` the topmost service.

If the `dry-run` option is used, the action simply reports (in different formats) what it would do. When the parameters `dry-run` and `with-service-meta-data` are used together with `outformat cli` or `outformat cli-c`, any changes to service meta-data that would be affected by the re-deploy operation will be included in the diff output.

Use the option `reconcile` if the service should reconcile original data, i.e., take control of that data. This option acknowledges other services controlling the same data. All data that existed before the service was created will now be owned by the service. When the service is removed, that data will also be removed. In technical terms, the reference count will be decreased by one for everything that existed prior to the service. If manually configured data exists below in the configuration tree that data is kept unless the option `discard-non-service-config` is used.

To control which configurations are to be reconciled, use option `include` / `exclude` together with option `reconcile`. Both options `include` and `exclude` will accept a list of instance identifiers as parameters.

* Option `include` will specify which configurations are to be reconciled in the scope of the redeployed service.
* Option `exclude` will specify which configurations are to be ignored during reconciliation. Note that `exclude` option can only be used with configuration subtrees whose root is not a child of any node created by the same service. This restriction ensures that when the service is removed, the parent of the excluded subtree is not deleted, which would otherwise result in the unintended removal of all child nodes, including those in the excluded subtree.

**Note**: The action is idempotent. If no configuration diff exists, then nothing needs to be done.

**Note**: The NSO general principle of minimum change applies.

</details>

<details>

<summary><code>reactive-re-deploy</code></summary>

This is a tailored `re-deploy` intended to be used in the reactive FASTMAP scenario. It differs from the ordinary `re-deploy` in that this action does not take any commit parameters.

This action will `re-deploy` the services as a shallow depth `re-deploy`. It will be performed with the same user as the original commit. Also, the commit parameters will be identical to the latest commit involving this service.

By default, this action is asynchronous and returns nothing. Use the `sync` leaf to get synchronous behavior and block until the service `re-deploy` transaction is committed. The `sync` leaf also means that the action will possibly return a commit result, such as a commit queue ID if any, or an error if the transaction failed.

</details>

<details>

<summary><code>touch</code></summary>

This action marks the service as changed.

Executing the action `touch` followed by a commit is the same as executing the action `re-deploy shallow`.

By using the action `touch`, several re-deploys can be performed in the same transaction.

</details>

<details>

<summary><code>un-deploy</code></summary>

Undo the effects of the service instance but keep the service itself. The service can later be re-deployed. This is a means to deactivate a service while keeping it in the system.

</details>

## Bulk Service Actions

Service actions are performed on the multiple service instances filtered by services type (service callpoint).

<details>

<summary><code>re-deploy</code></summary>

This is a tailored `re-deploy` intended to be used on multiple services. This action takes the same input parameters as `re-deploy` service instance action.

There are 3 choices of service filter that can be used as input.

* Use `service-type` to filter on services of a certain type.
* Use `service-id` to filter on specific service instances.
* Use `select-services` to filter on services evaluated by an XPath expression.

The output of this action is a list of service ids and the results of action `re-deploy`.

Use the option `suppress-positive-result` to suppress results with empty diffs.

</details>

<details>

<summary><code>un-deploy</code></summary>

This is a tailored `un-deploy` intended to be used on multiple services. This action takes the same input parameters as `un-deploy` service instance action.

There are 3 choices of service filter can be used as input.

* Use `service-type` to filter on services of a certain type.
* Use `service-id` to filter on specific services instances.
* Use `select-services` to filter on services evaluated by an XPath expression.

The output of this action is a list of service ids and the results of action `un-deploy`.

Use the option `suppress-positive-result` to suppress results with empty diffs.

</details>

<details>

<summary><code>check-sync</code></summary>

This is a tailored `check-sync` intended to be used on multiple services. This action takes the same input parameters as `check-sync` service instance action.

There are 3 choices of service filter that can be used as input.

* Use `service-type` to filter on services of a certain type.
* Use `service-id` to filter on specific service instances.
* Use `select-services` to filter on services evaluated by an XPath expression.

The output of this action is a list of service ids and the results of action `check-sync`.

Use the option `suppress-positive-result` to suppress results with empty diffs.

</details>

## Dry-run Drift Detection

Dry-run drift detection compares the changeset from a dry-run with the changeset at commit time to determine whether anything has changed in between. This helps prevent users from accidentally committing unintended changes. When a dry-run is performed, a checksum of the changeset is saved. During commit, a new checksum is calculated from the current changeset and compared with the saved checksum. Both checksums are calculated after the validation phase and before the prepare phase. If no dry-run has been performed in the transaction, no comparison is made. This feature applies to the CLI, JSON-RPC, and Web UI.

This behavior is configured through the `dry-run-drift-detection` container in `tailf-ncs-aaa.yang`, which has two leaves: `enabled` and `mode`. The `mode` leaf can be set to `warn` or `strict`. In `warn` mode, the user receives a validation warning if the changeset differs between dry-run and commit, and can choose whether to continue. This is handled the same way as other validation warnings. In `strict` mode, an error is returned instead, the transaction is aborted, and a new dry-run is required before the commit can proceed. In `tailf-ncs-aaa.yang`, `enabled` defaults to `false` and `mode` defaults to `warn`. However, `enabled` is set to `true` in `ncs_defaults.xml.in`, so the feature is enabled by default for new NSO deployments. The setting can be configured globally by setting the `dry-run-drift-detection` leaves in `tailf-ncs-aaa.yang`, for example in the CLI:

```cli
admin@ncs% set session dry-run-drift-detection enabled
admin@ncs% set session dry-run-drift-detection disabled
admin@ncs% set session dry-run-drift-detection mode strict
admin@ncs% set session dry-run-drift-detection mode warn
```

It can also be set per user, for example:

```cli
admin@ncs% set user admin session dry-run-drift-detection enabled
admin@ncs% set user admin session dry-run-drift-detection disabled
admin@ncs% set user admin session dry-run-drift-detection mode strict
admin@ncs% set user admin session dry-run-drift-detection mode warn
```

In the CLI, it is also possible to set this per session in operational mode:

```cli
admin@ncs> dry-run-drift-detection true
admin@ncs> dry-run-drift-detection false
admin@ncs> dry-run-drift-detection mode warn
admin@ncs> dry-run-drift-detection mode strict
```

This does not require a commit and will only last for this CLI session, meaning the changes are not persistent. This works in the same way as the existing commit-prompt setting.

### Examples

Consider this small service model:

```yang
module topology-service {
  namespace "http://example.com/topology-service";
  prefix topo;

  container topology {
    list connection {
      key "name";
      leaf name {
        type string;
      }
      leaf endpoint-1 {
        type leafref {
          path "/ncs:devices/ncs:device/ncs:name";
        }
      }
      leaf endpoint-2 {
        type leafref {
          path "/ncs:devices/ncs:device/ncs:name";
        }
      }
      leaf link-type {
        type enumeration {
          enum ethernet;
        }
        default ethernet;
      }
    }
  }

  augment "/ncs:services" {
    list link-config {
      key id;
      uses ncs:service-data;
      ncs:servicepoint link-config-servicepoint;
      leaf connection-name {
        type leafref {
          path "/topo:topology/topo:connection/topo:name";
        }
        mandatory true;
      }
      leaf timeout {
        type uint32; default 60;
      }
      leaf description {
        type string;
      }
    }
  }
}
```

#### Example A - create a changeset mismatch scenario with two transactions

Transaction A sets up the topology connections and then creates a service that depends on this topology, and then performs a dry-run:

```cli
admin@ncs% set topology connection link1 endpoint-1 ex0 endpoint-2 ex1 link-type ethernet
admin@ncs% commit
admin@ncs% set services link-config svc1 connection-name link1 timeout 10
admin@ncs% commit dry-run
cli {
    local-node {
        data  devices {
                  device ex0 {
                      config {
                          sys {
                              dns {
                                  options {
             +                        timeout 10;
             +                        attempts 3;
                                  }
                              }
                          }
                      }
                  }
                  device ex1 {
                      config {
                          sys {
                              dns {
                                  options {
             +                        timeout 10;
             +                        attempts 3;
                                  }
                              }
                          }
                      }
                  }
              }
              services {
             +    link-config svc1 {
             +        connection-name link1;
             +        timeout 10;
             +    }
              }
    }
}
```

Another transaction, transaction B, then changes endpoint-2 to point to ex2 instead of ex1 before transaction A commits:

```cli
admin@ncs% set topology connection link1 endpoint-2 ex2
admin@ncs% commit
Commit complete.
```

Transaction A then tries to commit its changes:

```cli
admin@ncs% commit
The following warnings were generated:
  Commit changeset does not match dry-run changeset
Proceed? [yes,no] no
Aborted: by user
```

When committing transaction A, validation runs again and the service re-evaluates against the updated topology, so the resulting transaction changeset is no longer the same as the one that was saved at dry-run. Because the changeset of transaction A was changed by the changes committed by transaction B, the dry-run-drift-detection validation warning is returned and the user can choose to continue or abort. If the user aborts, a new dry-run has to be performed to commit:

```cli
admin@ncs% commit dry-run
cli {
    local-node {
        data  devices {
                  device ex0 {
                      config {
                          sys {
                              dns {
                                  options {
             +                        timeout 10;
             +                        attempts 3;
                                  }
                              }
                          }
                      }
                  }
                  device ex2 {
                      config {
                          sys {
                              dns {
                                  options {
             +                        timeout 10;
             +                        attempts 3;
                                  }
                              }
                          }
                      }
                  }
              }
              services {
             +    link-config svc1 {
             +        connection-name link1;
             +        timeout 10;
             +    }
              }
    }
}
admin@ncs% commit
Commit complete.
```

The new dry-run output demonstrates that the original commit would have updated ex0 and ex1, but after the commit made by transaction B, transaction A would have actually updated ex0 and ex2, which would have been unintentional if the user was not aware of the changes made in transaction B.

#### Example B

This example demonstrates a scenario that does not cause a mismatch. Two transactions edit topology only, so no service is consuming topology in the open transaction. The transaction changeset of transaction A is not changed by the changes made in transaction B.

Transaction A:

```cli
admin@ncs% set topology connection link1 endpoint-2 ex1
admin@ncs% commit dry-run
cli {
    local-node {
        data  topology {
                  connection link1 {
             -        endpoint-2 ex2;
             +        endpoint-2 ex1;
                  }
              }
    }
}
```

Transaction B:

```cli
admin@ncs% set topology connection link1 endpoint-2 ex0
admin@ncs% commit
Commit complete.
```

Transaction A:

```cli
admin@ncs% commit
Commit complete.
```

No warning appears. If a new dry-run had been made in transaction A after the commit of transaction B, it would have displayed the following output:

```cli
admin@ncs% commit dry-run
cli {
    local-node {
        data  topology {
                  connection link1 {
             -        endpoint-2 ex0;
             +        endpoint-2 ex1;
                  }
              }
    }
}
```

The difference between the two dry-run outputs show that instead of changing endpoint-2 from ex2 to ex1, it was changed from ex0 to ex1 because of the change by transaction B. However, this does not affect the changeset of transaction A or the final result of the commit, which is that endpoint 2 now points to ex1. So even though the dry-run output differs, the intent of the commit is the same.

#### Example C

This example demonstrates a mismatch occurring when a change has been made in the same transaction. This also demonstrates "strict" mode

```cli
admin@ncs% set topology connection link1 endpoint-1 ex1
admin@ncs% commit dry-run
cli {
    local-node {
        data  topology {
                  connection link1 {
             -        endpoint-1 ex0;
             +        endpoint-1 ex1;
                  }
              }
    }
}
admin@ncs% set topology connection link1 endpoint-1 ex2
admin@ncs% commit
Aborted: Commit changeset does not match dry-run changeset
admin@ncs% commit dry-run
cli {
    local-node {
        data  topology {
                  connection link1 {
             -        endpoint-1 ex0;
             +        endpoint-1 ex2;
                  }
              }
    }
}
admin@ncs% commit
Commit complete.
```

In this scenario, transaction A first updated endpoint-1 to point to ex1 and performed a dry-run. Then transaction A updates endpoint-1 again to point to ex2 instead and tries to commit. The commit is aborted because this change in the same transaction has caused the transaction changeset to be updated, and dry-run-drift-detection is now in strict mode.

### Interaction with commit-prompt

Commit-prompt is a CLI behavior where each ordinary commit first runs an automatic dry-run and then prompts the user on whether or not to continue before the commit actually proceeds. This reduces the risk of committing unintended changes that dry-run drift detection is meant to catch. It is still possible to use commit-prompt with dry-run-drift-detection though. When used with dry-run-drift-detection, the changeset will be saved during the automatic dry-run, just as with any other dry-run. Then, if another transaction makes changes that affect the first transaction's changeset while the user is still at the prompt, a warning/error will be returned if the user confirms the commit with "yes", since confirming with "yes" runs validation again. So dry-run drift detection applies in the usual way if the changeset at commit time (after confirming the commit) differs from the one that was checksummed at the automatic dry-run.

### Known limitations

In some nano-service situations, NSO does not compare the dry-run and commit changeset checksums. This is during deletion of a nano-service as well as when a nano-service is created with converge-on-re-deploy. Instead, it emits a dedicated validation warning that explains why the comparison was skipped. The reasoning behind this warning is that the transaction might contain several other changes as well, and the user might want to have the comparison done for these changes. So when the warning appears, the user can choose to either continue with the commit with no comparison, or do a new commit that excludes the changes that cannot be compared, and later do those in a separate commit. The message is shown as a normal validation warning.

#### Nano-service delete

If the transaction deletes one or more nano-services, NSO skips the dry-run versus commit changeset comparison for that commit. This is the warning that will be returned during commit in the CLI:

```cli
The following warnings were generated:
  Cannot perform changeset comparison for this commit because service(s) were deleted: my-nano-servicepoint
Proceed? [yes,no]
```

If several service points are involved, their names appear separated by commas.

The changeset comparison is not possible in this case because the delete paths are different for a dry-run and a commit. When a nano-service is removed, NSO cannot run the full, multi-step teardown that a real delete performs. Internally, dry-run therefore follows a shortened delete path — essentially unwinding planned nano changes in a way that produces a useful preview without executing a complete delete. A real commit, however, uses the full nano delete path, which includes zombie handling etc. Because of this, the changesets will often differ between a dry-run and commit even without other changes being made, and therefore performing a comparison adds no value.

#### Nano-services using "converge-on-re-deploy"

For a nano-service declared with converge-on-re-deploy, NSO does not fully converge the service inside the same transaction where the service instance is created. Instead, in that creation commit NSO records the service intent, establishes the plan as required for this mode, and schedules a reactive re-deploy so the heavy convergence runs after that commit. During a dry-run however, NSO simulates what updates the nano-service would apply if it converged in that pass, so it executes the convergence path so that the dry-run can display a preview. Because of this, the changesets that are created during dry-run and commit will be different even if no other changes have been made in the transaction, so performing a comparison of the changesets has no purpose.

If the transaction involves nano-service work under a service point that has converge-on-re-deploy set, NSO skips the dry-run versus commit changeset checksum comparison for that commit and emits the warning below.

```cli
The following warnings were generated:
  Cannot perform changeset comparison for this commit because service(s) my-servicepoint use converge-on-re-deploy.
Proceed? [yes,no]
```

If several service points are involved, their names appear separated by commas.


# Network Simulator

Use NSO's network simulator to simulate your network and test functionality.

The `ncs-netsim` program is a useful tool to simulate a network of devices to be managed by NSO. It makes it easy to test NSO packages towards simulated devices. All you need is the NSO NED packages for the devices that you need to simulate. The devices are simulated with the Tail-f ConfD product.

Many NSO examples use `ncs-netsim` to simulate the devices. A good way to learn how to use `ncs-netsim` is to study them.

## Using Netsim <a href="#ug.netsim.using" id="ug.netsim.using"></a>

The `ncs-netsim` tool takes any number of NED packages as input. The user can specify the number of device instances per package (device type) and a string that is used as a prefix for the name of the devices. The command takes the following parameters:

```bash
admin$ ncs-netsim --help
Usage ncs-netsim  [--dir <NetsimDir>]
            create-network <NcsPackage> <NumDevices> <Prefix> |
            create-device <NcsPackage> <DeviceName> |
            add-to-network <NcsPackage> <NumDevices> <Prefix> |
            add-device <NcsPackage> <DeviceName> |
            delete-network                     |
            [-a | --async]  start [devname]    |
            [-a | --async ] stop [devname]     |
            [-a | --async ] reset [devname]    |
            [-a | --async ] restart [devname]  |
            list                      |
            is-alive [devname]        |
            status [devname]          |
            whichdir                  |
            ncs-xml-init [devname]    |
            ncs-xml-init-remote <RemoteNodeName> [devname] |
            [--force-generic]         |
            packages                  |
            netconf-console devname [XpathFilter] |
            [-w | --window] [cli | cli-c | cli-i] devname
```

Assume that you have prepared an NSO package for a device called `router`. (See the [examples.ncs/device-management/router-network](https://github.com/NSO-developer/nso-examples/tree/6.7/device-management/router-network) example). Also, assume the package is in `./packages/router`. At this point, you can create the simulated network by:

```bash
$ ncs-netsim create-network ./packages/router 3 device --dir ./netsim
```

This creates three devices; `device0`, `device1`, and `device2`. The simulated network is stored in the `./netsim` directory. The output structure is:

```
          ./netsim/device/
               device0/<ConfD files>, <log files>
               device1/
               ....
```

There is one separate directory for every ConfD simulating the devices.

The network can be started with:

```bash
$ ncs-netsim start
```

You can add more devices to the network in a similar way as it was created. E.g., if you created a network with some Juniper devices and want to add some Cisco IOS devices. Point to the NED you want to use (see `{NCS_DIR}/packages/neds/`) and run the command. Remember to start the new devices after they have been added to the network.

```bash
$ ncs-netsim add-to-network ${NCS_DIR}/packages/neds/cisco-ios 2 c-device --dir ./netsim
```

To extract the device data from the simulated network to a file in XML format:

```bash
$ ncs-netsim ncs-xml-init > devices.xml
```

This data is usually used to load the simulated network into NSO. Putting the XML file in the `./ncs-cdb` folder will load it when NSO starts. If NSO is already started, it can be reloaded while running.

```bash
$ ncs_load -l -m devices.xml
```

The generated device data creates devices of the same type as the device being simulated. This is true for NETCONF, CLI, and SNMP devices. When simulating generic devices, the simulated device will run as a NETCONF device.

Under very special circumstances, one can choose to force running the simulation as a generic device with the option `--force-generic`.

The simulated network device info can be shown with:

```bash
 $ ncs-netsim list
...
 name=device0 netconf=12022 snmp=11022 ipc=5010 cli=10022 dir=examples.ncs/device-management/router-network/netsim/device/device0
...
```

Here you can see the device name, the working directory, and the port number for different services to be accessed on the simulated device (NETCONF SSH, SNMP, IPC, and direct access to the CLI).

You can reach the CLI of individual devices with:

```bash
$ ncs-netsim cli-c device0
```

The simulated devices actually provide three different styles of CLI:

* `cli`: J-Style
* `cli-c`: Cisco XR Style
* `cli-i`: Cisco IOS Style

Individual devices can be started and stopped with:

```bash
$ ncs-netsim start device0
$ ncs-netsim stop device0
```

You can check the status of the simulated network. Either a short version just to see if the device is running or a more verbose with all the information.

```bash
$ ncs-netsim is-alive device0
$ ncs-netsim status device0
```

View which packages are used in the simulated network:

```bash
$ ncs-netsim packages
```

It is also possible to reset the network back to the state of initialization:

```bash
$ ncs-netsim reset
```

When you are done, remove the network:

```bash
$ ncs-netsim delete-network
```

### Using ConfD Tools with Netsim <a href="#ug.netsim.using_confd_tools" id="ug.netsim.using_confd_tools"></a>

The netsim tool includes a standard ConfD distribution and the ConfD C API library (libconfd) that the ConfD tools use. The library is built with default settings where the values for MAXDEPTH and MAXKEYLEN are 20 and 9, respectively. These values define the size of `confd_hkeypath_t` struct and this size is related to the size of data models in terms of depth and key lengths. Default values should be big enough even for very large and complex data models. But in some rare cases, one or both of these values might not be large enough for a given data model.

One might observe a limitation when the data models that are used by simulated devices exceed these limits. Then it would not be possible to use the ConfD tools that are provided with the netsim. To overcome this limitation, it is advised to use the corresponding NSO tools to perform desired tasks on devices.

NSO and ConfD tools and Python APIs are basically the same except for naming, the default IPC port and the MAXDEPTH and MAXKEYLEN values, where for NSO tools, the values are set to 60 and 18, respectively. Thus, the advised solution is to use the NSO tools and NSO Python API with netsim.

E.g., instead of using the below command:

```bash
$ CONFD_IPC_PORT=5010 ${NCS_DIR}/netsim/confd/bin/confd_load -m -l *.xml
```

One may use:

```bash
$ NCS_IPC_PORT=5010 ncs_load -m -l *.xml
```

### Learn More <a href="#ug.netsim.learnmore" id="ug.netsim.learnmore"></a>

The README file in [examples.ncs/device-management/router-network](https://github.com/NSO-developer/nso-examples/tree/6.7/device-management/router-network) example gives a good introduction on how to use `ncs-netsim`.

The


# Get Started

Develop services and more in NSO.

## Introduction to Automation

<table data-view="cards"><thead><tr><th></th><th></th><th data-hidden data-card-target data-type="content-ref"></th></tr></thead><tbody><tr><td><strong>CDB and YANG</strong></td><td>Learn about NSO's configuration DB &#x26; YANG.</td><td><a href="/pages/pZUZmkxLAnUONrqDb5PR">/pages/pZUZmkxLAnUONrqDb5PR</a></td></tr><tr><td><strong>Basic Python Automation</strong></td><td>Learn basics of NSO automation with Python.</td><td><a href="/pages/TxJY7WsTvmyw37b0w6Nm">/pages/TxJY7WsTvmyw37b0w6Nm</a></td></tr><tr><td><strong>Develop a Simple Service</strong></td><td>Take first steps to develop a simple NSO service.</td><td><a href="/pages/8ue8V7101TddJ17baqbw">/pages/8ue8V7101TddJ17baqbw</a></td></tr><tr><td><strong>Applications in NSO</strong></td><td>Automate NSO with applications.</td><td><a href="/pages/sXgWFsCfe9tG1xRpMdkO">/pages/sXgWFsCfe9tG1xRpMdkO</a></td></tr></tbody></table>

## Core Concepts

<table data-view="cards"><thead><tr><th></th><th></th><th data-hidden data-card-target data-type="content-ref"></th></tr></thead><tbody><tr><td><strong>Services</strong></td><td>Learn the concepts of NSO services and automation.</td><td><a href="/pages/IAQGSOrRFtEfgMpsPElu">/pages/IAQGSOrRFtEfgMpsPElu</a></td></tr><tr><td><strong>Implementing Services</strong></td><td>Learn NSO service development in detail.</td><td><a href="/pages/1ErWKzeM15TwAei4yXWm">/pages/1ErWKzeM15TwAei4yXWm</a></td></tr><tr><td><strong>Templates</strong></td><td>Develop and deploy NSO templates.</td><td><a href="/pages/Whkl286OHBVy52qnpzLh">/pages/Whkl286OHBVy52qnpzLh</a></td></tr><tr><td><strong>Nano Services</strong></td><td>Learn about nano services for staged provisioning.</td><td><a href="/pages/vJ0F1vOS4MeZ3S4VAOGS">/pages/vJ0F1vOS4MeZ3S4VAOGS</a></td></tr><tr><td><strong>Packages</strong></td><td>Learn about NSO packages and how they work.</td><td><a href="/pages/Co2LbpwYJk0phPiX0Gdp">/pages/Co2LbpwYJk0phPiX0Gdp</a></td></tr><tr><td><strong>Using CDB</strong></td><td>Concepts of importance in usage of the CDB.</td><td><a href="/pages/FxpCNgv5QKnfWrJw4nXf">/pages/FxpCNgv5QKnfWrJw4nXf</a></td></tr><tr><td><strong>YANG</strong></td><td>Explore YANG data modeling and its use.</td><td><a href="/pages/wD3v9o26MqWM8gLSMCJU">/pages/wD3v9o26MqWM8gLSMCJU</a></td></tr><tr><td><strong>NSO Concurrency Model</strong></td><td>Understand NSO's concurrency model.</td><td><a href="/pages/8PJW3J1uuhcu5ttTWYkl">/pages/8PJW3J1uuhcu5ttTWYkl</a></td></tr><tr><td><strong>Service Handling of ADMs</strong></td><td>Perform Handling of ambiguous device models.</td><td><a href="/pages/KXtxbH7UIdS0MakRJBfx">/pages/KXtxbH7UIdS0MakRJBfx</a></td></tr><tr><td><strong>NSO Virtual Machines</strong></td><td>Learn about Java and Python virtual machines.</td><td><a href="/pages/yGfkdnr1VpEtgrxAU0w0">/pages/yGfkdnr1VpEtgrxAU0w0</a></td></tr><tr><td><strong>API Overview</strong></td><td>Learn concepts and usage of Java and Python APIs.</td><td><a href="/pages/UVNTYYEFNC1zzWA5XYMP">/pages/UVNTYYEFNC1zzWA5XYMP</a></td></tr><tr><td><strong>Northbound APIs</strong></td><td>Learn working mechanism of northbound APIs.</td><td><a href="/pages/TItWBhukD9D6FJkD3eWB">/pages/TItWBhukD9D6FJkD3eWB</a></td></tr></tbody></table>

## Advanced Development

<table data-view="cards"><thead><tr><th></th><th></th><th data-hidden data-card-target data-type="content-ref"></th></tr></thead><tbody><tr><td><strong>Dev Env &#x26; Resources</strong></td><td>Useful info to get started with NSO development.</td><td><a href="/pages/UI92G83zmOpE8suAEtJj">/pages/UI92G83zmOpE8suAEtJj</a></td></tr><tr><td><strong>Developing Services</strong></td><td>Develop and deploy NSO services/nano services.</td><td><a href="/pages/3IPqmFAvtcCLGCnYJvvZ">/pages/3IPqmFAvtcCLGCnYJvvZ</a></td></tr><tr><td><strong>Developing Packages</strong></td><td>Develop and deploy NSO packages.</td><td><a href="/pages/Yb0PdnXkYaOCtAjNU3ch">/pages/Yb0PdnXkYaOCtAjNU3ch</a></td></tr><tr><td><strong>Developing NEDs</strong></td><td>Develop and deploy NSO NEDs.</td><td><a href="/pages/8ejgwCG15ge0KL57wXne">/pages/8ejgwCG15ge0KL57wXne</a></td></tr><tr><td><strong>Developing Alarm Apps</strong></td><td>Develop and deploy NSO alarm applications.</td><td><a href="/pages/x0g4ATc1LLQtq5WGL5ri">/pages/x0g4ATc1LLQtq5WGL5ri</a></td></tr><tr><td><strong>Kicker</strong></td><td>Trigger declarative notification actions in NSO.</td><td><a href="/pages/l1veKT5BYCJEPHqk6ytW">/pages/l1veKT5BYCJEPHqk6ytW</a></td></tr><tr><td><strong>Scaling and Performance</strong></td><td>Optimize your NSO automation solution.</td><td><a href="/pages/INyCuLO2h06pPPZrhcFY">/pages/INyCuLO2h06pPPZrhcFY</a></td></tr><tr><td><strong>Progress Trace</strong></td><td>Debug, diagnose, and profile events in NSO.</td><td><a href="/pages/kpsJEQm3pOtFku8CHOm1">/pages/kpsJEQm3pOtFku8CHOm1</a></td></tr><tr><td><strong>Web UI Development</strong></td><td>Develop enhancements for NSO Web UI.</td><td><a href="/pages/CaqxNy7kRXnmVOt6aIiJ">/pages/CaqxNy7kRXnmVOt6aIiJ</a></td></tr></tbody></table>

## Connected Topics

<table data-view="cards"><thead><tr><th></th><th></th><th data-hidden data-card-target data-type="content-ref"></th></tr></thead><tbody><tr><td><strong>SNMP Notifications</strong></td><td>Configure NSO as SNMP notification receiver.</td><td><a href="/pages/7JdgJxfER9fTkxUSsVi5">/pages/7JdgJxfER9fTkxUSsVi5</a></td></tr><tr><td><strong>Web Server</strong></td><td>Use embedded server to deliver static/CGI content.</td><td><a href="/pages/2Q779klNQ2CNgat78Ncg">/pages/2Q779klNQ2CNgat78Ncg</a></td></tr><tr><td><strong>Scheduler</strong></td><td>Schedule time-based jobs for background tasks.</td><td><a href="/pages/CjjXi3Jz3h1ieY3Cbm5a">/pages/CjjXi3Jz3h1ieY3Cbm5a</a></td></tr><tr><td><strong>External Logging</strong></td><td>Send log data to external commands.</td><td><a href="/pages/YmgQfKSxYe7Kdi8QDXYY">/pages/YmgQfKSxYe7Kdi8QDXYY</a></td></tr><tr><td><strong>Encryption Strings</strong></td><td>Store encrypted values in NSO.</td><td><a href="/pages/Mjhh9xTOj3hWijpPDr9o">/pages/Mjhh9xTOj3hWijpPDr9o</a></td></tr></tbody></table>


# Introduction to Automation

Get started with NSO automation by understanding fundamental concepts.


# CDB and YANG

Learn how NSO keeps a record of its managed devices using CDB.

Cisco NSO is a network automation platform that supports a variety of uses. This can be as simple as a configuration of a standard-format hostname, which can be implemented in minutes. Or it could be an advanced MPLS VPN with custom traffic-engineered paths in a Service Provider network, which might take weeks to design and code.

Regardless of complexity, any network automation solution must keep track of two things: intent and network state.

The Configuration Database (CDB) built into NSO was designed for this exact purpose:

* Firstly, the CDB will store the intent, which describes what you want from the network. Traditionally we call this intent a network service since this is what the network ultimately provides to its users.
* Secondly, the CDB also stores a copy of the configuration of the managed devices, that is, the network state. Knowledge of the network state is essential to correctly provision new services. It also enables faster diagnosis of problems and is required for advanced functionality, such as self-healing.

This section describes the main features of the CDB and explains how NSO stores data there. To help you better understand the structure of the CDB, you will also learn how to add your data to it.

## Key Features of the CDB <a href="#d5e47" id="d5e47"></a>

The CDB is a dedicated built-in storage for data in NSO. It was built from the ground up to efficiently store and access network configuration data, such as device configurations, service parameters, and even configuration for NSO itself. Unlike traditional SQL databases that store data as rows in a table, the CDB is a hierarchical database, with a structure resembling a tree. You could think of it as somewhat like a big XML document that can store all kinds of data.

There are a number of other features that make the CDB an excellent choice for a configuration store:

* Fast lightweight database access through a well-defined API.
* Subscription (“push”) mechanism for change notification.
* Transaction support for ensuring data consistency.
* Rich and extensible schema based on YANG.
* Built-in support for schema and associated data upgrade.
* Close integration with NSO for low-maintenance operation.

To speed up operations, CDB keeps a configurable amount of configuration data in RAM, in addition to persisting it to disk (see [CDB Persistence](/guides/administration/advanced-topics/cdb-persistence) for details). The CDB also stores transient operational data, such as alarms and traffic statistics. By default, this operational data is only kept in RAM and is reset during restarts, however, the CDB can be instructed to persist it if required.

{% hint style="info" %}
The automatic schema update feature is useful not only when performing an actual upgrade of NSO itself, it also simplifies the development process. It allows individual developers to add and delete items in the configuration independently.

Additionally, the schema for data in the CDB is defined with a standard modeling language called YANG. YANG (RFC 7950, <https://tools.ietf.org/html/rfc7950>) describes constraints on the data and allows the CDB to store values more efficiently.
{% endhint %}

## Compilation and Loading of YANG Modules <a href="#d5e74" id="d5e74"></a>

All of the data stored in the CDB follows the data model provided by various YANG modules. Each module usually comes as one or more files with a `.yang` extension and declares a part of the overall model.

NSO provides a base set of YANG modules out of the box. They are located in `$NCS_DIR/src/ncs/yang` if you wish to inspect them. These modules are required for proper system operation.

All other YANG modules are provided by packages and extend the base NSO data model. For example, each Network Element Driver (NED) package adds the required nodes to store the configuration for that particular type of device. In the same way, you can store your custom data in the CDB by providing a package with your own YANG module.

However, the CDB can't use the YANG files directly. The bundled compiler, `ncsc`, must first transform a YANG module into a final schema (`.fxs`) file. The reason is that internally and in the programming APIs NSO refers to YANG nodes with integer values instead of names. This conserves space and allows for more efficient operations, such as switch statements in the application code. The `.fxs` file contains this mapping and needs to be recreated if any part of the YANG model changes. The compilation process is usually started from the package Makefile by the `make` command.

## Showcase: Extending the CDB with Packages <a href="#d5e87" id="d5e87"></a>

{% hint style="info" %}
See [examples.ncs/getting-started/cdb-yang](https://github.com/NSO-developer/nso-examples/tree/6.7/getting-started/cdb-yang) for an example implementation.
{% endhint %}

### Prerequisites

Ensure that:

* No previous NSO or netsim processes are running. Use the `ncs --stop` and `ncs-netsim stop` commands to stop them if necessary.
* NSO Local Install with a fresh runtime directory has been created by the `ncs-setup --dest ~/nso-lab-rundir` or similar command.
* The environment variable `NSO_RUNDIR` points to this runtime directory, such as set by the `export NSO_RUNDIR=~/nso-lab-rundir` command. It enables the below commands to work as-is, without additional substitution needed.

### Step 1 - Create a Package <a href="#d5e102" id="d5e102"></a>

The easiest way to add your data fields to the CDB is by creating a service package. The package includes a YANG file for the service-specific data, which you can customize. You can create the initial package by simply invoking the `ncs-make-package` command. This command also sets up a `Makefile` with the code for compiling the YANG model.

Use the following command to create a new package:

```bash
$ ncs-make-package --service-skeleton python --build \
    --dest $NSO_RUNDIR/packages/my-data-entries my-data-entries
mkdir -p ../load-dir
mkdir -p java/src//
/nso/bin/ncsc  `ls my-data-entries-ann.yang  > /dev/null 2>&1 && echo "-a my-data-entries-ann.yang"` \
              -c -o ../load-dir/my-data-entries.fxs yang/my-data-entries.yang
$
```

The command line switches instruct the command to compile the YANG file and place the package in the right location.

### Step 2 - Add Package to NSO <a href="#d5e111" id="d5e111"></a>

Now start the NSO process if it is not running already and connect to the CLI:

```bash
$ cd $NSO_RUNDIR ; ncs ; ncs_cli -Cu admin

admin connected from 127.0.0.1 using console on nso
admin@ncs#
```

Next, instruct NSO to load the newly created package:

```bash
admin@ncs# packages reload

>>> System upgrade is starting.
>>> Sessions in configure mode must exit to operational mode.
>>> No configuration changes can be performed until upgrade has completed.
>>> System upgrade has completed successfully.
reload-result {
    package my-data-entries
    result true
}
```

Once the package loading process is completed, you can verify the data model from your package was incorporated into NSO. Use the `show` command, which now supports an additional parameter:

```bash
admin@ncs# show my-data-entries
% No entries found.
admin@ncs#
```

This command tells you that NSO knows about the extended data model but there is no actual data configured for it yet.

### Step 3 - Set Data <a href="#d5e124" id="d5e124"></a>

More interestingly, you are now able to add custom entries to the configuration. First, enter the CLI configuration mode:

```bash
admin@ncs# config
Entering configuration mode terminal
admin@ncs(config)#
```

Then add an arbitrary entry under my-data-entries:

```bash
admin@ncs(config)# my-data-entries "entry number 1"
admin@ncs(config-my-data-entries-entry number 1)#
```

What is more, you can also set a dummy IP address:

```bash
admin@ncs(config-my-data-entries-entry number 1)# dummy 0.0.0.0
admin@ncs(config-my-data-entries-entry number 1)#
```

However, if you try to use something different from a dummy, you will get an error. Likewise, if you try to assign a dummy a value that is not an IP address. How did NSO learn about this dummy value?

If you assumed from the YANG file, you are correct. YANG files provide the schema for the CDB and that dummy value comes from the YANG model in your package. Let's take a closer look.

### Step 4 - Inspect the YANG Module <a href="#d5e137" id="d5e137"></a>

Exit the configuration mode and discard the changes by typing `abort`:

```bash
admin@ncs(config-my-data-entries-entry number 1)# abort
admin@ncs#
```

Open the YANG file in an editor or list its contents from the CLI with the following command:

```bash
admin@ncs# file show packages/my-data-entries/src/yang/my-data-entries.yang
module my-data-entries {
< ... output omitted ... >
  list my-data-entries {
    < ... output omitted ... >
    leaf dummy {
      type inet:ipv4-address;
    }
  }
}
```

At the start of the output, you can see the module `my-data-entries`, which contains your data model. By default, the `ncs-make-package` gives it the same name as the package. You can check that this module is indeed loaded:

```bash
admin@ncs# show ncs-state loaded-data-models data-model my-data-entries

                                                                              EXPORTED  EXPORTED
NAME             REVISION  NAMESPACE                         PREFIX           TO ALL    TO
--------------------------------------------------------------------------------------------------
my-data-entries  -         http://com/example/mydataentries  my-data-entries  X         -

admin@ncs#
```

The `list my-data-entries` statement, located a bit further down in the YANG file, allowed you to add custom entries before. And near the end of the output, you can find the `leaf dummy` definition, with IPv4 as the type. This is the source of information that enables NSO to enforce a valid IP address as the value.

## Data Modeling Basics <a href="#d5e154" id="d5e154"></a>

NSO uses YANG to structure and enforce constraints on data that it stores in the CDB. YANG was designed to be extensible and handle all kinds of data modeling, which resulted in a number of language features that helped achieve this goal. However, there are only four fundamental elements (node types) for describing data:

* leaf nodes
* leaf-list nodes
* container nodes
* list nodes

You can then combine these elements into a complex, tree-like structure, which is why we refer to individual elements as nodes (of the data tree). In general, YANG separates nodes into those that hold data (`leaf`, `leaf-list`) and those that hold other nodes (container, list).

A `leaf` contains simple data such as an integer or a string. It has one value of a particular type and no child nodes. For example:

```yang
leaf host-name {
    type string;
    description "Hostname for this system";
}
```

This code describes the structure that can hold a value of a hostname (of some device). A `leaf` node is used because the hostname only has a single value, that is, the device has one (canonical) hostname. In the NSO CLI, you set a value of a `leaf` simply as:

```cli
admin@ncs(config)# host-name "server-NY-01"
```

A `leaf-list` is a sequence of leaf nodes of the same type. It can hold multiple values, very much like an array. For example:

```
leaf-list domains {
    type string;
    description "My favourite internet domains";
}
```

This code describes a data structure that can hold many values, such as a number of domain names. In the CLI, you can assign multiple values to a `leaf-list` with the help of square bracket syntax:

```bash
admin@ncs(config)# domains [ cisco.com tail-f.com ]
```

`leaf` and `leaf-list` describe nodes that hold simple values. As a model keeps expanding, having all data nodes on the same (top) level can quickly become unwieldy. A container node is used to group related nodes into a subtree. It has only child nodes and no value. A container may contain any number of child nodes of any type (including leafs, lists, containers, and leaf-lists). For example:

```yang
container server-admin {
    description "Administrator contact for this system";
    leaf name {
        type string;
    }
}
```

This code defines the concept of a server administrator. In the CLI, you first select the container before you access the child nodes:

```bash
admin@ncs(config)# server-admin name "Ingrid"
```

Similarly, a `list` defines a collection of container-like list entries that share the same structure. Each entry is like a record or a row in a table. It is uniquely identified by the value of its key leaf (or leaves). A list definition may contain any number of child nodes of any type (leafs, containers, other lists, and so on). For example:

```yang
list user-info {
    description "Information about team members";
    key "name";
    leaf name {
        type string;
    }
    leaf expertise {
        type string;
    }
}
```

This code defines a list of users (of which there can be many), where each user is uniquely identified by their name. In the CLI, lists take an additional parameter, the key value, to select a single entry:

```bash
admin@ncs(config)# user-info "Ingrid"
```

To set a value of a particular list entry, first specify the entry, then the child node, like so:

```bash
admin@ncs(config)# user-info "Ingrid" expertise "Linux"
```

Combining just these four fundamental YANG node types, you can build a very complex model that describes your data. As an example, the model for the configuration of a Cisco IOS-based network device, with its myriad features, is created with YANG. However, it makes sense to start with some simple models, to learn what kind of data they can represent and how to alter that data with the CLI.

## Showcase: Building and Testing a Model <a href="#d5e195" id="d5e195"></a>

{% hint style="info" %}
See [examples.ncs/getting-started/cdb-yang](https://github.com/NSO-developer/nso-examples/tree/6.7/getting-started/cdb-yang) for an example implementation.
{% endhint %}

### Prerequisites

Ensure that:

* No previous NSO or netsim processes are running. Use the `ncs --stop` and `ncs-netsim stop` commands to stop them if necessary.
* NSO Local Install with a fresh runtime directory has been created by the `ncs-setup --dest ~/nso-lab-rundir` or similar command.
* The environment variable `NSO_RUNDIR` points to this runtime directory, such as set by the `export NSO_RUNDIR=~/nso-lab-rundir` command. It enables the below commands to work as-is, without additional substitution needed.

### Step 1 - Create a Model Skeleton <a href="#d5e210" id="d5e210"></a>

You can add custom data models to NSO by using packages. So, you will build a package to hold the YANG module that represents your model. Use the following command to create a package (if you are building on top of the previous showcase, the package may already exist and will be updated):

```bash
$ ncs-make-package --service-skeleton python \
    --dest $NSO_RUNDIR/packages/my-data-entries my-data-entries
$
```

Change the working directory to the directory of your package:

```bash
$ cd $NSO_RUNDIR/packages/my-data-entries
```

You will place the YANG model into the `src/yang/my-test-model.yang` file. In a text editor, create a new file and add the following text at the start:

```yang
module my-test-model {
    namespace "http://example.tail-f.com/my-test-model";
    prefix "t";
```

The first line defines a new module and gives it a name. In addition, there are two more statements required: the `namespace` and `prefix`. Their purpose is to help avoid name collisions.

### Step 2 - Fill Out the Model <a href="#d5e224" id="d5e224"></a>

Add a statement for each of the four fundamental YANG node types (leaf, leaf-list, container, list) to the `my-test-model.yang` model.

```yang
    leaf host-name {
        type string;
        description "Hostname for this system";
    }
    leaf-list domains {
        type string;
        description "My favourite internet domains";
    }
    container server-admin {
        description "Administrator contact for this system";
        leaf name {
            type string;
        }
    }
    list user-info {
        description "Information about team members";
        key "name";
        leaf name {
            type string;
        }
        leaf expertise {
            type string;
        }
    }
```

Also, add the closing bracket for the module at the end:

```
}
```

Remember to finally save the file as `my-test-model.yang` in the `src/yang/` directory of your package. It is a best practice for the name of the file to match the name of the module.

### Step 3 - Compile and Load the Model <a href="#d5e234" id="d5e234"></a>

Having completed the model, you must compile it into an appropriate (`.fxs`) format. From the text editor first, return to the shell and then run the `make` command in the `src/` subdirectory of your package:

```bash
$ make -C src/
make: Entering directory 'nso-run/packages/my-data-entries/src'
/nso/bin/ncsc  `ls my-test-model-ann.yang  > /dev/null 2>&1 && echo "-a my-test-model-ann.yang"` \
              -c -o ../load-dir/my-test-model.fxs yang/my-test-model.yang
make: Leaving directory 'nso-run/packages/my-data-entries/src'
$
```

The compiler will report if there are errors in your YANG file, and you must fix them before continuing.

Next, start the NSO process and connect to the CLI:

```bash
$ cd $NSO_RUNDIR && ncs && ncs_cli -C -u admin

admin connected from 127.0.0.1 using console on nso
admin@ncs#
```

Finally, instruct NSO to reload the packages:

```bash
admin@ncs# packages reload

>>> System upgrade is starting.
>>> Sessions in configure mode must exit to operational mode.
>>> No configuration changes can be performed until upgrade has completed.
>>> System upgrade has completed successfully.
reload-result {
    package my-data-entries
    result true
}
admin@ncs#
```

### Step 4 - Test the Model <a href="#d5e248" id="d5e248"></a>

Enter the configuration mode by using the `config` command and test out how to set values for the data nodes you have defined in the YANG model:

* `host-name` leaf
* `domains` leaf-list
* `server-admin` container
* `user-info` list

Use the `?` and `TAB` keys to see the possible completions.

Now feel free to go back and experiment with the YANG file to see how your changes affect the data model. Just remember to rebuild and reload the package after you make any changes.

## Initialization Files <a href="#d5e268" id="d5e268"></a>

Adding a new YANG module to the CDB enables it to store additional data, however, there is nothing in the CDB for this module yet. While you can add configuration with the CLI, for example, there are situations where it makes sense to start with some initial data in the CDB already. This is especially true when a new instance starts for the first time and the CDB is empty.

In such cases, you can bootstrap the CDB data with XML files. There are various uses for this feature. For example, you can implement some default “factory settings” for your module or you might want to pre-load data when creating a new instance for testing.

In particular, some of the provided examples use the CDB init files mechanism to save you from typing out all of the initial configuration commands by hand. They do so by creating a file with the configuration encoded in the XML format.

When starting empty, the CDB will try to initialize the database from all XML files found in the directories specified by the `init-path` and `db-dir` settings in `ncs.conf` (please see [ncs.conf(5)](/guides/resources/man/ncs.conf.5) in Manual Pages for exact details). The loading process scans the files with the `.xml` suffix and adds all the data in a single transaction. In other words, there is no specified order in which the files are processed. This happens early during start-up, during the so-called start phase 1, described in [Starting NSO](https://nso-docs.cisco.com/guides/development/introduction-to-automation/pages/hgWUBFw1TA0R6WyLxOgc#ug.sys_mgmt.starting_ncs).

The content of the init file does not need to be a complete instance document but can specify just a part of the overall data, very much like the contents of the NETCONF `edit-config` operation. However, the end result of applying all the files must still be valid according to the model.

It is a good practice to wrap the data inside a `config` element, as it gives you the option to have multiple top-level data elements in a single file while it remains a valid XML document. Otherwise, you would have to use separate files for each of them. The following example uses the `config` element to fit all the elements into a single file.

{% code title="A Sample CDB init File my-test-data.xml" %}

```xml
<config xmlns="http://tail-f.com/ns/config/1.0">
  <host-name xmlns="http://example.tail-f.com/my-test-model">server-NY-01</host-name>

  <server-admin xmlns="http://example.tail-f.com/my-test-model">
    <name>Ingrid</name>
  </server-admin>
</config>
```

{% endcode %}

There are many ways to generate the XML data. A common approach is to dump existing data with the `ncs_load` utility or the `display xml` filter in the CLI. All of the data in the CDB can be represented (or exported, if you will) in XML. This is no coincidence. XML was the main format for encoding data with NETCONF when YANG was created and you can trace the origin of some YANG features back to XML.

{% code title="Creating init XML File with the ncs\_load Command " %}

```bash
$ ncs_load -F p -p /domains > cdb-init.xml
$ cat cdb-init.xml
<config xmlns="http://tail-f.com/ns/config/1.0">
  <domains xmlns="http://example.tail-f.com/my-test-model">cisco.com</domains>
  <domains xmlns="http://example.tail-f.com/my-test-model">tail-f.com</domains>
</config>
$
```

{% endcode %}


# Basic Automation with Python

Implement basic automation with Python.

You can manipulate data in the CDB with the help of XML files or the UI, however, these approaches are not well suited for programmatic access. NSO includes libraries for multiple programming languages, providing a simpler way for scripts and programs to interact with it. The Python Application Programming Interface (API) is likely the easiest to use.

This section will show you how to read and write data using the Python programming language. With this approach, you will learn how to do basic network automation in just a few lines of code.

## Setup <a href="#d5e305" id="d5e305"></a>

The environment setup that happens during the sourcing of the `ncsrc` file also configures the `PYTHONPATH` environment variable. It allows the Python interpreter to find the NSO modules, which are packaged with the product. This approach also works with Python virtual environments and does not require installing any packages.

Since the `ncsrc` file takes care of setting everything up, you can directly start the Python interactive shell and import the main `ncs` module. This module is a wrapper around a low-level C `_ncs` module that you may also need to reference occasionally. Documentation for both of the modules is available through the built-in `help()` function or separately in the HTML format.

If the `import ncs` statement fails, please verify that you are using a supported Python version and that you have sourced the `ncsrc` beforehand.

Generally, you can run the code from the Python interactive shell but we recommend against it. The code uses nested blocks, which are hard to edit and input interactively. Instead, we recommend you save the code to a file, such as `script.py`, which you can then easily run and rerun with the `python3 script.py` command. If you would still like to interactively inspect or alter the values during the execution, you can use the `import pdb; pdb.set_trace()` statements at the location of interest.

## Transactions <a href="#d5e322" id="d5e322"></a>

With NSO, data reads and writes normally happen inside a transaction. Transactions ensure consistency and avoid race conditions, where simultaneous access by multiple clients could result in data corruption, such as reading half-written data. To avoid this issue, NSO requires you to first start a transaction with a call to `ncs.maapi.single_read_trans()` or `ncs.maapi.single_write_trans()`, depending on whether you want to only read data or read and write data. Both of them require you to provide the following two parameters:

* `user`: The username (string) of the user you wish to connect as
* `context`: Method of access (string), allowing NSO to distinguish between CLI, web UI, and other types of access, such as Python scripts

These parameters specify security-related information that is used for auditing, access authorization, and so on. Please refer to [AAA infrastructure](/guides/administration/management/aaa-infrastructure) for more details.

As transactions use up resources, it is important to clean up after you are done using them. Using a Python `with` code block will ensure that cleanup is automatically performed after a transaction goes out of scope. For example:

```
with ncs.maapi.single_read_trans('admin', 'python') as t:
    ...
```

In this case, the variable `t` stores the reference to a newly started transaction. Before you can actually access the data, you also need a reference to the root element in the data tree for this transaction. That is, the top element, under which all of the data is located. The `ncs.maagic.get_root()` function, with transaction `t` as a parameter, achieves this goal.

{% hint style="info" %}
See [Transactions](/guides/development/core-concepts/transactions) for more details on transactions.
{% endhint %}

## Read and Write Values <a href="#d5e344" id="d5e344"></a>

Once you have the reference to the root element, say in a variable named `root`, navigating the data model becomes straightforward. Accessing a property on `root` selects a child data node with the same name as the property. For example, `root.nacm` gives you access to the `nacm` container, used to define fine-grained access control. Since `nacm` is itself a container node, you can select one of its children using the same approach. So, the code `root.nacm.enable_nacm` refers to another node inside `nacm`, called `enable-nacm`. This node is a leaf, holding a value, which you can print out with the Python `print()` function. Doing so is conceptually the same as using the `show running-config nacm enable-nacm` command in the CLI.

There is a small difference, however. Notice that in the CLI the `enable-nacm` is hyphenated, as this is the actual node name in YANG. But names must not include the hyphen (minus) sign in Python, so the Python code uses an underscore instead.

The following is the full source code that prints the value:

{% code title="Reading a Value in Python" %}

```python
import ncs

with ncs.maapi.single_read_trans('admin', 'python') as t:
    root = ncs.maagic.get_root(t)
    print(root.nacm.enable_nacm)
```

{% endcode %}

As you can see in this example, it is necessary to import only the `ncs` module, which automatically imports all the submodules. Depending on your NSO instance, you might also notice that the value printed is `True`, without any quotation marks. As a convenience, the value gets automatically converted to the best-matching Python type, which in this case is a boolean value (`True` or `False`).

Moreover, if you start a read/write transaction instead of a read-only one, you can also assign a new value to the leaf. Of course, the same validation rules apply as using the CLI and you need to explicitly commit the transaction if you want the changes to persist. A call to the `apply()` method on the transaction object `t` performs this function. Here is an example:

{% code title="Writing a Value in Python" %}

```python
import ncs

with ncs.maapi.single_write_trans('admin', 'python') as t:
    root = ncs.maagic.get_root(t)
    root.nacm.enable_nacm = True
    t.apply()
```

{% endcode %}

## Lists

You can access a YANG list node like how you access a leaf. However, working with a list more resembles working with Python `dict` than a list, even though the name would suggest otherwise. The distinguishing feature is that YANG lists have keys that uniquely identify each list item. So, lists are more naturally represented as a kind of dictionary in Python.

Let's say there is a list of customers defined in NSO, with a YANG schema such as:

```yang
container customers {
  list customer {
    key "id";
    leaf id {
      type string;
    }
  }
}
```

To simplify the code, you might want to assign the value of `root.customers.customer` to a new variable `our_customers`. Then you can easily access individual customers (list items) by their `id`. For example, `our_customers['ACME']` would select the customer with `id` equal to `ACME`. You can check for the existence of an item in a list using the Python `in` operator, for example, `'ACME' in our_customers`. Having selected a specific customer using the square bracket syntax, you can then access the other nodes of this item.

Compared to dictionaries, making changes to YANG lists is quite a bit different. You cannot just add arbitrary items because they must obey the YANG schema rules. Instead, you call the `create()` method on the list object and provide the value for the key. This method creates and returns a new item in the list if it doesn't exist yet. Otherwise, the method returns the existing item. And for item removal, use the Python built-in `del` function with the list object and specify the item to delete. For example, `del our_customers['ACME']` deletes the ACME customer entry.

In some situations, you might want to enumerate all of the list items. Here, the list object can be used with the Python `for` syntax, which iterates through each list item in turn. Note that this differs from standard Python dictionaries, which iterate through the keys. The following example demonstrates this behavior.

{% code title="Using lists with Python" %}

```python
import ncs

with ncs.maapi.single_write_trans('admin', 'python') as t:
    root = ncs.maagic.get_root(t)
    our_customers = root.customers.customer

    new_customer = our_customers.create('ACME')
    new_customer.status = 'active'

    for c in our_customers:
      print(c.id)

    del our_customers['ACME']
    t.apply()
```

{% endcode %}

Now let's see how you can use this knowledge for network automation.

## Showcase - Configuring DNS with Python <a href="#d5e399" id="d5e399"></a>

{% hint style="info" %}
See [examples.ncs/getting-started/basic-automation](https://github.com/NSO-developer/nso-examples/tree/6.7/getting-started/basic-automation) for an example implementation.
{% endhint %}

### **Prerequisites**

* No previous NSO or netsim processes are running. Use the `ncs --stop and ncs-netsim stop` commands to stop them if necessary.

### Step 1 - Start the Routers <a href="#d5e407" id="d5e407"></a>

Leveraging one of the examples included with the NSO installation allows you to quickly gain access to an NSO instance with a few devices already onboarded. The [examples.ncs/device-management](https://github.com/NSO-developer/nso-examples/tree/6.7/device-management) set of examples contains three simulated routers that you can configure.

<div data-with-frame="true"><figure><img src="/files/Tqr4AL2JJZ2R9fNkGhMV" alt="" width="375"><figcaption><p>The Lab Topology</p></figcaption></figure></div>

1. Navigate to the [router-network](https://github.com/NSO-developer/nso-examples/tree/6.7/device-management/router-network) directory with the following command.

   ```bash
   $ cd $NCS_DIR/examples.ncs/device-management/router-network
   ```
2. You can prepare and start the routers by running the `make` and `netsim` commands from this directory.

   ```bash
   $ make clean all && ncs-netsim start
   ```
3. With the routers running, you should also start the NSO instance that will allow you to manage them.

   ```bash
   $ ncs
   ```

In case the `ncs` command reports an error about an address already in use, you have another NSO instance already running that you must stop first (`ncs --stop`).

### Step 2 - Inspect the Device Data Model <a href="#d5e431" id="d5e431"></a>

Before you can use Python to configure the router, you need to know what to configure. The simplest way to find out how to configure the DNS on this type of router is by using the NSO CLI.

```bash
$ ncs_cli -C -u admin
```

1. In the CLI, you can verify that the NSO is managing three routers and check their names with the following command:

   ```cli
   admin@ncs# show devices list
   ```
2. To make sure that the NSO configuration matches the one deployed on routers, also perform a `sync-from` action.

   ```cli
   admin@ncs# devices sync-from
   ```
3. Let's say you would like to configure the DNS server `192.0.2.1` on the `ex1` router. To do this by hand, first enter the configuration mode.

   ```cli
   admin@ncs# config
   ```
4. Then navigate to the NSO copy of the `ex1` configuration, which resides under the `devices device ex1 config` path, and use the `?` and `TAB` keys to explore the available configuration options. You are looking for the DNS configuration.\
   ...

   ```cli
   admin@ncs(config)# devices device ex1 config
   ```
5. Once you have found it, you see the full DNS server configuration path: `devices device ex1 config sys dns server`.

{% hint style="info" %}
As an alternative to using the CLI approach to find this path, you can also consult the data model of the router in the `packages/router/src/yang/` directory.
{% endhint %}

6. As you won't be configuring `ex1` manually at this point, exit the configuration mode.

   ```cli
   admin@ncs(config)# abort
   ```
7. Instead, you will create a Python script to do it, so exit the CLI as well.

   ```cli
   admin@ncs# exit
   ```

### Step 3 - Create the Script <a href="#d5e463" id="d5e463"></a>

You will place the script into the `ex1-dns.py` file.

1. In a text editor, create a new file and add the following text at the start.\\

   ```python
   import ncs
   with ncs.maapi.single_write_trans('admin', 'python') as t:
       root = ncs.maagic.get_root(t)
   ```

   \
   The `root` variable allows you to access configuration in the NSO, much like entering the configuration mode on the CLI does.
2. Next, you will need to navigate to the `ex1` router. It makes sense to assign it to the `ex1_device` variable, which makes it more obvious what it refers to and easier to access in the script.

   ```
       ex1_device = root.devices.device['ex1']
   ```
3. In NSO, each managed device, such as the `ex1` router, is an entry inside the `device` list. The list itself is located in the `devices` container, which is a common practice for lists. The list entry for `ex1` includes another container, `config`where the copy of `ex1` configuration is kept. Assign it to the `ex1_config` variable.

   ```
       ex1_config = ex1_device.config
   ```

   \
   Alternatively, you can assign to `ex1_config` directly, without referring to `ex1_device`, like so:

   ```
       ex1_config = root.devices.device['ex1'].config
   ```

   \
   This is the equivalent of using `devices device ex1 config` on the CLI.
4. For the last part, keep in mind the full configuration path you found earlier. You have to keep navigating to reach the `server` list node. You can do this through the `sys` and `dns` nodes on the `ex1_config` variable.

   ```
       dns_server_list = ex1_config.sys.dns.server
   ```
5. DNS configuration typically allows specifying multiple servers for redundancy and is therefore modeled as a list. You add a new DNS server with the `create()` method on the list object.

   ```
        dns_server_list.create('192.0.2.1')
   ```
6. Having made the changes, do not forget to commit them with a call to `apply()` or they will be lost.

   ```
       t.apply()
   ```

   \
   Alternatively, you can use the `dry-run` parameter with the `apply_params()` to, for example, preview what will be sent to the device.

   ```
       params = t.get_params()
       params.dry_run_native()
       result = t.apply_params(True, params)
       print(result['device']['ex1'])
       t.apply_params(True, t.get_params())
   ```
7. Lastly, add a simple `print` statement to notify you when the script is completed.

   ```
       print('Done!')
   ```

### Step 4 - Run and Verify the Script <a href="#d5e503" id="d5e503"></a>

1. Save the script file as `ex1-dns.py` and run it with the `python3` command.

   ```bash
   $ python3 ex1-dns.py
   ```
2. You should see `Done!` printed out. Then start the NSO CLI to verify the configuration change.

   ```bash
   $ ncs_cli -C -u admin
   ```
3. Finally, you can check the configured DNS servers on `ex1` by using the `show running-config` command.

   ```cli
   admin@ncs# show running-config devices device ex1 config sys dns server
   ```

   \
   If you see the 192.0.2.1 address in the output, you have successfully configured this device using Python!

## A Note on Robustness <a href="#d5e519" id="d5e519"></a>

The code in this chapter is intentionally kept simple to demonstrate the core concepts and lacks robustness in error handling. In particular, it is missing the retry mechanism in case of concurrency conflicts as described in [Handling Conflicts](https://nso-docs.cisco.com/guides/development/introduction-to-automation/pages/8PJW3J1uuhcu5ttTWYkl#ncs.development.concurrency.handling).

## The Magic Behind the API <a href="#d5e523" id="d5e523"></a>

Perhaps you've wondered about the unusual name of Python `ncs.maagic` module? It is not a typo but a portmanteau of the words Management Agent API (MAAPI) and magic. The latter is used in the context of so-called magic methods in Python. The purpose of magic methods is to allow custom code to play nicely with the Python language. An example you might have come across in the past is the `__init__()` method in a class, which gets called whenever you create a new object. This one and similar methods are called magic because they are invoked automatically and behind the scenes (implicitly).

The NSO Python API makes extensive use of such magic methods in the `ncs.maagic` module. Magic methods help this module translate an object-based, user-friendly programming interface into low-level function calls. In turn, the high-level approach to navigating the data hierarchy with `ncs.maagic` objects is called the Python Maagic API.


# Develop a Simple Service

Get started with service development using a simple example.

The device YANG models contained in the Network Element Drivers (NEDs) enable NSO to store device configurations in the CDB and expose a uniform API to the network for automation, such as by Python scripts. The concept of NSO services builds on top of this network API and adds the ability to store service-specific parameters with each service instance.

This section introduces the main service building blocks and shows you how to build one yourself.

## Why Services? <a href="#d5e536" id="d5e536"></a>

Network automation includes provisioning and de-provisioning configuration, even though the de-provisioning part often doesn't get as much attention. It is nevertheless significant since leftover, residual configuration can cause hard-to-diagnose operational problems. Even more importantly, without proper de-provisioning, seemingly trivial changes may prove hard to implement correctly.

Consider the following example. You create a simple script that configures a DNS server on a router, by adding the IP address of the server to the DNS server list. This should work fine for initial provisioning. However, when the IP address of the DNS server changes, the configuration on the router should be updated as well.

Can you still use the same script in this case? Most likely not, since you need to remove the old server from the configuration and add the new one. The original script would just add the new IP address after the old one, resulting in both entries on the device. In turn, the device may experience slow connectivity as the system periodically retries the old DNS IP address and eventually times out.

The following figure illustrates this process, where a simple script first configures the IP address 192.0.2.1 (“.1”) as the DNS server, then later configures 192.0.2.8 (“.8”), resulting in a leftover old entry (“.1”).

<div data-with-frame="true"><figure><img src="/files/zLENyXrUGNMhfiivzTh6" alt="" width="375"><figcaption><p>DNS Configuration with a Simple Script</p></figcaption></figure></div>

In such a situation, the script could perhaps simply replace the existing configuration, by removing all existing DNS server entries before adding the new one. But is this a reliable practice? What if a device requires an additional DNS server that an administrator configured manually? It would be overwritten and lost.

In general, the safest approach is to keep track of the previous changes and only replace the parts that have changed. This, however, is a lot of work and nontrivial to implement yourself. Fortunately, NSO provides such functionality through the FASTMAP algorithm, which is used when deploying services.

The other major benefit of using NSO services for automation is the service interface definition using YANG, which specifies the name and format of the service parameters. Many new NSO users wonder why use a service YANG model when they could just use the Python code or templates directly. While it might be difficult to see the benefits without much prior experience, YANG allows you to write better, more maintainable code, which simplifies the solution in the long run.

Many, if not most, security issues and provisioning bugs stem from unexpected user input. You must always validate user input (service parameter values) and YANG compels you to think about that when writing the service model. It also makes it easy to write the validation rules by using a standardized syntax, specifically designed for this purpose.

Moreover, the separation of concerns into the user interface, validation, and provisioning code allows for better organization, which becomes extremely important as the project grows. It also gives NSO the ability to automatically expose the service functionality through its APIs for integration with other systems.

For these reasons, services are the preferred way of implementing network automation in NSO.

## Service Package <a href="#d5e556" id="d5e556"></a>

As you may already know, services are added to NSO with packages. Therefore, you need to create a package if you want to implement a service of your own. NSO ships with an `ncs-make-package` utility that makes creating packages effortless. Adding the `--service-skeleton python` option creates a service skeleton, that is, an empty service, which you can tailor to your needs. As the last argument, you must specify the package name, which in this case is the service name. The command then creates a new directory with that name and places all the required files in the appropriate subdirectories.

The package contains the two most important parts of the service:

* the service YANG model and
* the service provisioning code also called the mapping logic.

Let's first look at the provisioning part. This is the code that performs the network configuration necessary for your service. The code often includes some parameters, for example, the DNS server IP address or addresses to use if your service is in charge of DNS configuration. So, we say that the code maps the service parameters into the device parameters, which is where the term mapping logic originates from. NSO, with the help of the NED, then translates the device parameters to the actual configuration. This simple tree-to-tree mapping describes how to create the service and NSO automatically infers how to update, remove, or re-deploy the service, hence the name FASTMAP.

<div data-with-frame="true"><figure><img src="/files/r0DCH1OAXURKXnlZdiMu" alt="" width="563"><figcaption><p>Transformation of Service Parameters into Device Configurations</p></figcaption></figure></div>

How do you create the provisioning code and where do you place it? Is it similar to a stand-alone Python script? Indeed, the code is mostly the same. The main difference is that now you don't have to create a session and a transaction yourself because NSO already provides you with one. Through this transaction, the system tracks the changes to the configuration made by your code.

The package skeleton contains a directory called `python`. It holds a Python package named after your service. In the package, the `ServiceCallbacks` class (the `main.py` file) is used for provisioning code. The same file also contains the `Main` class, which is responsible for registering the `ServiceCallbacks` class as a service provisioning code with NSO.

Of the most interest is the `cb_create()` method of the `ServiceCallbacks` class:

```python
def cb_create(self, tctx, root, service, proplist)
```

NSO calls this method for service provisioning. Now, let's see how to evolve a stand-alone automation script into a service. Suppose you have Python code for DNS configuration on a router, similar to the following:

```python
with ncs.maapi.single_write_trans('admin', 'python') as t:
    root = ncs.maagic.get_root(t)

    ex1_device = root.devices.device['ex1']
    ex1_config = ex1_device.config
    dns_server_list = ex1_config.sys.dns.server
    dns_server_list.create('192.0.2.1')

    t.apply()
```

Taking into account the `cb_create()` signature and the fact that the NSO manages the transaction for a service, you won't need the transaction and `root` variable setup. The NSO service framework already takes care of setting up the `root` variable with the right transaction. There is also no need to call `apply()` because NSO does that automatically.

You only have to provide the core of the code (the middle portion in the above stand-alone script) to the `cb_create()`:

```python
def cb_create(self, tctx, root, service, proplist):
    ex1_device = root.devices.device['ex1']
    ex1_config = ex1_device.config
    dns_server_list = ex1_config.sys.dns.server
    dns_server_list.create('192.0.2.1')
```

You can run this code by adding the service package to NSO and provisioning a service instance. It will achieve the same effect as the stand-alone script but with all the benefits of a service, such as tracking changes.

## Service Parameters <a href="#d5e597" id="d5e597"></a>

In practice, all services have some variable parameters. Most often parameter values change from service instance to service instance, as the desired configuration is a little bit different for each of them. They may differ in the actual IP address that they configure or in whether the switch for some feature is on or off. Even the DNS configuration service requires a DNS server IP address, which may be the same across the whole network but could change with time if the DNS server is moved elsewhere. Therefore, it makes sense to expose the variable parts of the service as service parameters. This allows a service operator to set the parameter value without changing the service provisioning code.

With NSO, service parameters are defined in the service model, written in YANG. The YANG module describing your service is part of the service package, located under the `src/yang` path, and customarily named the same as the package. In addition to the module-related statements (description, revision, imports, and so on), a typical service module includes a YANG `list`, named after the service. Having a list allows you to configure multiple service instances with slightly different parameter values. For example, in a DNS configuration service, you might have multiple service instances with different DNS servers. The reason is, that some devices, such as those in the Demilitarized Zone (DMZ), might not have access to the internal DNS servers and would need to use a different set.

The service model skeleton already contains such a list statement. The following is another example, similar to the one in the skeleton:

```yang
list my-svc {
  description "This is an RFS skeleton service";

  key name;
  leaf name {
    tailf:info "Unique service id";
    tailf:cli-allow-range;
    type string;
  }

  uses ncs:service-data;
  ncs:servicepoint my-svc-servicepoint;

  // Devices configured by this service instance
  leaf-list device {
    type leafref {
      path "/ncs:devices/ncs:device/ncs:name";
    }
  }

  // An example generic parameter
  leaf server-ip {
    type inet:ipv4-address;
  }
}
```

Along with the description, the service specifies a key, `name`to uniquely identify each service instance. This can be any free-form text, as denoted by its type (string). The statements starting with `tailf:` are NSO-specific extensions for customizing the user interface NSO presents for this service. After that come two lines, the `uses` and `ncs:servicepoint`, which tells NSO this is a service and not just some ordinary list. At the end, there are two parameters defined, `device` and `server-ip`.

NSO then allows you to add the values for these parameters when configuring a service instance, as shown in the following CLI transcript:

```cli
admin@ncs(config)# my-svc instance1 ?
Possible completions:
  check-sync           Check if device config is according to the service
  commit-queue
  deep-check-sync      Check if device config is according to the service
  device
  < ... output omitted ... >
  server-ip
  < ... output omitted ... >
```

Finally, your Python script can read the supplied values inside the `cb_create()` method via the provided `service` variable. This variable points to the currently-provisioning service instance, allowing you to use code such as `service.server_ip` for the value of the `server-ip` parameter.

## Showcase - A Simple DNS Configuration Service <a href="#d5e620" id="d5e620"></a>

{% hint style="info" %}
See [examples.ncs/getting-started/develop-service](https://github.com/NSO-developer/nso-examples/tree/6.7/getting-started/develop-service) for an example implementation.
{% endhint %}

### Prerequisites

* No previous NSO or netsim processes are running. Use the `ncs --stop` and `ncs-netsim stop` commands to stop them if necessary.
* NSO Local Install with a fresh runtime directory has been created by the `ncs-setup --dest ~/nso-lab-rundir` or a similar command.
* The environment variable `NSO_RUNDIR` points to this runtime directory, such as set by the `export NSO_RUNDIR=~/nso-lab-rundir` command. It enables the below commands to work as-is, without additional substitution needed.

### Step 1 - Prepare Simulated Routers <a href="#d5e635" id="d5e635"></a>

The [examples.ncs/getting-started/develop-service/init](https://github.com/NSO-developer/nso-examples/tree/6.7/getting-started/develop-service/init) holds a package, Makefile, and an XML initialization file you can use for this scenario to start the routers and connect them to your NSO instance.

First, copy the package and files to your `NSO_RUNDIR`:

```bash
$ cp -r $NCS_DIR/examples.ncs/getting-started/develop-service/init/router.in $NSO_RUNDIR/packages/router
$ cp $NCS_DIR/examples.ncs/getting-started/develop-service/init/ncs_init.xml.in $NSO_RUNDIR/ncs-cdb/ncs_init.xml
$ cp $NCS_DIR/examples.ncs/getting-started/develop-service/init/Makefile.in $NSO_RUNDIR/Makefile
```

From the `NSO_RUNDIR` directory, you can start a fresh set of routers by running the following `make` command:

```bash
$ cd $NSO_RUNDIR
$ make showcase-clean-start
< ... output omitted ... >
DEVICE ex0 OK STARTED
DEVICE ex1 OK STARTED
DEVICE ex2 OK STARTED
```

The routers are now running. The required NED package and a CDB initialization file `ncs-cdb/ncs_init.xml`were also added to your NSO instance. The latter contains connection details for the routers and will be automatically loaded on the first NSO start.

In case you're not using a fresh working directory, you may need to use the `ncs_load` command to load the file manually.

### Step 2 - Create a Service Package <a href="#d5e653" id="d5e653"></a>

You create a new service package with the `ncs-make-package` command. Without the `--dest` option, the package is created in the current working directory. Normally you run the command without this option, as it is shorter. For NSO to find and load this package, it has to be placed (or referenced via a symbolic link) in the `packages` subfolder of the NSO running directory.

Change the current working directory before creating the package:

```bash
$ cd $NSO_RUNDIR/packages
```

You need to provide two parameters to `ncs-make-package`. The first is the `--service-skeleton python` option, which selects the Python programming language for scaffolding code. The second parameter is the name of the service. As you are creating a service for DNS configuration, `dns-config` is a fitting name for it. Run the final, full command:

```bash
$ ncs-make-package --service-skeleton python dns-config
```

If you look at the file structure of the newly created package, you will see it contains a number of files.

```
dns-config/
+-- package-meta-data.xml
+-- python
|   '-- dns_config
|       +-- __init__.py
|       '-- main.py
+-- README
+-- src
|   +-- Makefile
|   '-- yang
|       '-- dns-config.yang
+-- templates
'-- test
    +-- < ... output omitted ... >
```

The `package-meta-data.xml` describes the package and tells NSO where to find the code. Inside the `python` folder is a service-specific Python package, where you add your own Python code (to `main.py` file). There is also a `README` file that you can update with the information relevant to your service. The `src` folder holds the source code that you must compile before you can use it with NSO. That's why there is also a `Makefile` that takes care of the compilation process. In the `yang` subfolder is the service YANG module. The `templates` folder can contain additional XML files, discussed later. Lastly, there's the `test` folder where you can put automated testing scripts, which won't be discussed here.

### Step 3 - Add the DNS Server Parameter <a href="#d5e679" id="d5e679"></a>

While you can always hard-code the desired parameters, such as the DNS server IP address, in the Python code, it means you have to change the code every time the parameter value (the IP address) changes. Instead, you can define it as an input parameter in the YANG file. Fortunately, the skeleton already has a leaf called a dummy that you can rename and use for this purpose.

Open the `dns-config.yang`, located inside `dns-config/src/yang/`, in a text or code editor and find the following line:

```yang
    leaf dummy {
```

Replace the word `dummy` with the word `dns-server`, save the file, and return to the shell. Run the `make` command in the `dns-config/src` folder to compile the updated YANG file.

```bash
$ make -C dns-config/src
make: Entering directory 'dns-config/src'
mkdir -p ../load-dir
mkdir -p java/src//
bin/ncsc  `ls dns-config-ann.yang  > /dev/null 2>&1 && echo "-a dns-config-ann.yang"` \
              -c -o ../load-dir/dns-config.fxs yang/dns-config.yang
make: Leaving directory 'dns-config/src'
```

### Step 4 - Add Python Code <a href="#d5e693" id="d5e693"></a>

In a text or code editor, open the `main.py` file, located inside `dns-config/python/dns_config/`. Find the following snippet:

```python
    @Service.create
    def cb_create(self, tctx, root, service, proplist):
        self.log.info('Service create(service=', service._path, ')')
```

Right after the `self.log.info()` call, read the value of the `dns-server` parameter into a `dns_ip` variable:

```
        dns_ip = service.dns_server
```

Mind the 8 spaces in front to make sure that the line is correctly aligned. After that, add the code that configures the `ex1` router:

```
        ex1_device = root.devices.device['ex1']
        ex1_config = ex1_device.config
        dns_server_list = ex1_config.sys.dns.server
        dns_server_list.create(dns_ip)
```

Here, you are using the `dns_ip` variable that contains the operator-provided IP address instead of a hard-coded value. Also, note that there is no need to check if the entry for this DNS server already exists in the list.

In the end, the `cb_create()` method should look like the following:

```python
    @Service.create
    def cb_create(self, tctx, root, service, proplist):
        self.log.info('Service create(service=', service._path, ')')
        dns_ip = service.dns_server
        ex1_device = root.devices.device['ex1']
        ex1_config = ex1_device.config
        dns_server_list = ex1_config.sys.dns.server
        dns_server_list.create(dns_ip)
```

Save the file and let's see the service in action!

### Step 5 - Deploy the Service <a href="#d5e711" id="d5e711"></a>

Start the NSO from the running directory:

```bash
$ cd $NSO_RUNDIR; ncs
```

Then, start the NSO CLI:

```bash
$ ncs_cli -C -u admin
```

If you have started a fresh NSO instance, the packages are loaded automatically. Still, there's no harm in requesting a `package reload` anyway:

```cli
admin@ncs# packages reload
reload-result {
    package dns-config
    result true
}
reload-result {
    package router-nc-1.0
    result true
}
```

As you will be making changes on the simulated routers, make sure NSO has their current configuration with the `devices sync-from` command.

```cli
admin@ncs# devices sync-from
sync-result {
    device ex0
    result true
}
sync-result {
    device ex1
    result true
}
sync-result {
    device ex2
    result true
}
```

Now you can test out your service package by configuring a service instance. First, enter the configuration mode.

```cli
admin@ncs# config
```

Configure a test instance and specify the DNS server IP address:

```cli
admin@ncs(config)# dns-config test dns-server 192.0.2.1
```

The easiest way to see configuration changes from the service code is to use the `commit dry-run` command.

```cli
admin@ncs(config-dns-config-test)# commit dry-run
cli {
    local-node {
        data  devices {
                  device ex1 {
                      config {
                          sys {
                              dns {
             +                    # after server 10.2.3.4
             +                    server 192.0.2.1;
                              }
                          }
                      }
                  }
              }
             +dns-config test {
             +    dns-server 192.0.2.1;
             +}
    }
}
```

The output tells you the new DNS server is being added in addition to an existing one already there. Commit the changes:

```cli
admin@ncs(config-dns-config-test)# commit
```

Finally, change the IP address of the DNS server:

```cli
admin@ncs(config-dns-config-test)# dns-server 192.0.2.8
```

With the help of `commit dry-run` observe how the old IP address gets replaced with the new one, without any special code needed for provisioning.

```cli
admin@ncs(config-dns-config-test)# commit dry-run
cli {
    local-node {
        data  devices {
                  device ex1 {
                      config {
                          sys {
                              dns {
             -                    server 192.0.2.1;
             +                    # after server 10.2.3.4
             +                    server 192.0.2.8;
                              }
                          }
                      }
                  }
              }
              dns-config test {
             -    dns-server 192.0.2.1;
             +    dns-server 192.0.2.8;
              }
    }
}
```

## Service Templates <a href="#d5e746" id="d5e746"></a>

The DNS configuration example intentionally performs very little configuration, a single line really, to focus on the service concepts. In practice, services can become more complex in two different ways. First, the DNS configuration service takes the IP address of the DNS server as an input parameter, supplied by the operator. Instead, the provisioning code could leverage another system, such as an IP Address Management (IPAM), to get the required information. In such cases, you have to add additional logic to your service code to generate the parameters (variables) to be used for configuration.

Second, generating the configuration from the parameters can become more complex when it touches multiple subsystems or spans across multiple devices. An example would be a service that adds a new VLAN, configures an IP address and a DHCP server, and adds the new route to a routing protocol. Or perhaps the service has to be duplicated on two separate devices for redundancy.

An established approach to the second challenge is to use a templating system for configuration generation. Templates separate the process of constructing parameter values from how they are used, adding a degree of flexibility and decoupling. NSO uses XML-based configuration *(config)* templates, which you can invoke from provisioning code or link directly to services. In the latter case, you don't even have to write any Python code.

XML templates are snippets of configuration, similar to the CDB init files, but more powerful. Let's see how you could implement the DNS configuration service using a template instead of navigating the YANG model with Python.

While it is possible to write an XML template from scratch, it has to follow the target YANG model. Fortunately, the NSO CLI can help with generating most parts of the template from changes to the currenly open transaction. First, you'll need a sample instance with the desired configuration. As you are configuring the DNS server on a router and the ex1 device already has one configured, you can reuse that one. Otherwise, you might configure one by hand, using the CLI. You do that by displaying the existing configuration in the format of an XML template and saving it to a file, by piping it through the `display xml-template` and `save` filters, as shown here:

```cli
admin@ncs# show running-config devices device ex1 config sys dns | display xml-template
<config-template xmlns="http://tail-f.com/ns/config/1.0">
  <devices xmlns="http://tail-f.com/ns/ncs">
    <device>
      <name>ex1</name>
      <config>
        <sys xmlns="http://example.com/router">
          <dns>
            <server>
              <address>192.0.2.1</address>
            </server>
          </dns>
        </sys>
      </config>
    </device>
  </devices>
</config-template>
admin@ncs# show running-config devices device ex1 config sys dns | \
    display xml-template | save template.xml
```

The file structure of a package usually contains a `templates` folder and that is where the template belongs. When loading packages, NSO will scan this folder and process any `.xml` files it finds as templates.

Of course, a template with hard-coded values is of limited use, as it would always produce the exact same configuration. It becomes a lot more useful with variable substitution. In its simplest form, you define a variable value in the provisioning (Python) code and reference it from the XML template, by using curly braces and a dollar sign: `{$VARIABLE}`. Also, many users prefer to keep the variable name uppercased to make it stand out more from the other XML elements in the file. For example, in the template XML file for the DNS service, you would likely replace the IP address `192.0.2.1` with the variable `{$DNS_IP}` to control its value from the Python code.

You apply the template by creating a new `ncs.template.Template` object and calling its `apply()` method. This method takes the name of the XML template as the first parameter (no trailing `.xml`), and an object of type `ncs.template.Variables` as the second parameter. Using the `Variables` object, you provide values for the variables in the template.

```
template_vars = ncs.template.Variables()
template_vars.add('VARIABLE', 'some value')

template = ncs.template.Template(service)
template.apply('template', template_vars)
```

Variables in a template can take a more complex form of an XPath expression, where the parameter for the `Template` constructor comes into play. This parameter defines the root node (starting point) when evaluating XPath paths. Use the provided `service` variable, unless you specifically need a different value. It is what the so-called template-based services use as well.

Template-based services are no-code, pure template services that only contain a YANG model and an XML template. Since there is no code to set the variables, they must rely on XPath for the dynamic parts of the template. Such services still have a YANG data model with service parameters, that XPath can access. For example, if you have a parameter leaf defined in the service YANG file by the name `dns-server`, you can refer to its value with the `{/dns-server}` code in the XML template.

Likewise, you can use the same XPath in a template of a Python service. Then you don't have to add this parameter to the variables object but can still access its value in the template, saving you a little bit of Python code.

## Showcase - DNS Configuration Service with Templates <a href="#d5e780" id="d5e780"></a>

{% hint style="info" %}
See [examples.ncs/getting-started/develop-service](https://github.com/NSO-developer/nso-examples/tree/6.7/getting-started/develop-service) for an example implementation.
{% endhint %}

### Prerequisites

* No previous NSO or netsim processes are running. Use the `ncs --stop` and `ncs-netsim stop` commands to stop them if necessary.
* NSO local install with a fresh runtime directory has been created by the `ncs-setup --dest ~/nso-lab-rundir` or similar command.
* The environment variable `NSO_RUNDIR` points to this runtime directory, such as set by the `export NSO_RUNDIR=~/nso-lab-rundir` command. It enables the below commands to work as-is, without additional substitution needed.

### Step 1 - Prepare Simulated Routers <a href="#d5e795" id="d5e795"></a>

The [examples.ncs/getting-started/develop-service/init](https://github.com/NSO-developer/nso-examples/tree/6.7/getting-started/develop-service/init) holds a package, Makefile, and an XML initialization file you can use for this scenario to start the routers and connect them to your NSO instance.

First, copy the package and files to your `NSO_RUNDIR`:

```bash
$ cp -r $NCS_DIR/examples.ncs/getting-started/develop-service/init/router.in $NSO_RUNDIR/packages/router
$ cp $NCS_DIR/examples.ncs/getting-started/develop-service/init/ncs_init.xml.in $NSO_RUNDIR/ncs-cdb/ncs_init.xml
$ cp $NCS_DIR/examples.ncs/getting-started/develop-service/init/Makefile.in $NSO_RUNDIR/Makefile
```

From the `NSO_RUNDIR` directory, you can start a fresh set of routers by running the following `make` command:

```bash
$ cd $NSO_RUNDIR
$ make showcase-clean-start
< ... output omitted ... >
DEVICE ex0 OK STARTED
DEVICE ex1 OK STARTED
DEVICE ex2 OK STARTED
```

The routers are now running. The required NED package and a CDB initialization file, `ncs-cdb/ncs_init.xml`, were also added to your NSO instance. The latter contains connection details for the routers and will be automatically loaded on the first NSO start.

In case you're not using a fresh working directory, you may need to use the `ncs_load` command to load the file manually.

### Step 2 - Create a Service <a href="#d5e813" id="d5e813"></a>

The DNS configuration service that you are implementing will have three parts: the YANG model, the service code, and the XML template. You will put all of these in a package named `dns-config`. First, navigate to the `packages` subdirectory:

```bash
$ cd $NSO_RUNDIR/packages
```

Then, run the following command to set up the service package:

```bash
$ ncs-make-package --build --service-skeleton python dns-config
bin/ncsc  `ls dns-config-ann.yang  > /dev/null 2>&1 && echo "-a dns-config-ann.yang"` \
              -c -o ../load-dir/dns-config.fxs yang/dns-config.yang
```

In case you are building on top of the previous showcase, the package folder may already exist and will be updated.

You can leave the YANG model as is for this scenario but you need to add some Python code that will apply an XML template during provisioning. In a text or code editor open the `main.py` file, located inside `dns-config/python/dns_config/`, and find the definition of the `cb_create()` function:

```python
    @Service.create
    def cb_create(self, tctx, root, service, proplist):
        ...
```

You will define one variable for the template, the IP address of the DNS server. To pass its value to the template, you have to create the `Variables` object and add each variable, along with its value. Replace the body of the `cb_create()` function with the following:

```
        template_vars = ncs.template.Variables()
        template_vars.add('DNS_IP', '192.0.2.1')
```

The `template_vars` object now contains a value for the `DNS_IP` template variable, to be used with the `apply()` method that you are adding next:

```
        template = ncs.template.Template(service)
        template.apply('dns-config-tpl', template_vars)
```

Here, the first argument to `apply()` defines the template to use. In particular, using `dns-config-tpl`, you are requesting the template from the `dns-config-tpl.xml` file, which you will be creating shortly.

This is all the Python code that is required. The final, complete `cb_create` method is as follows:

```python
    @Service.create
    def cb_create(self, tctx, root, service, proplist):
        template_vars = ncs.template.Variables()
        template_vars.add('DNS_IP', '192.0.2.1')
        template = ncs.template.Template(service)
        template.apply('dns-config-tpl', template_vars)
```

### Step 3 - Create a Template

The most straightforward way to create an XML template is by using the NSO CLI. Return to the running directory and start the NSO:

```bash
$ cd $NSO_RUNDIR && ncs --with-package-reload
```

The `--with-package-reload` option will make sure that NSO loads any added packages and save a `packages reload` command on the NSO CLI.

Next, start the NSO CLI:

```bash
$ ncs_cli -C -u admin
```

As you are starting with a new NSO instance, first invoke the `sync-from` action.

```cli
admin@ncs# devices sync-from
sync-result {
    device ex0
    result true
}
sync-result {
    device ex1
    result true
}
sync-result {
    device ex2
    result true
}
```

Next, make sure that the ex1 router already has an existing entry for a DNS server in its configuration.

```cli
admin@ncs# show running-config devices device ex1 config sys dns
devices device ex1
 config
  sys dns server 10.2.3.4
  !
 !
!
```

Pipe the command through the `display xml-template` and `save` CLI filters to save this configuration as an XML template. According to the Python code, you need to create a template file `dns-config-tpl.xml`. Use `packages/dns-config/templates/dns-config-tpl.xml` for the full file path.

```cli
admin@ncs# show running-config devices device ex1 config sys dns \
| display xml-template | save packages/dns-config/templates/dns-config-tpl.xml
```

At this point, you have created a complete template that will provision the 10.2.3.4 as the DNS server on the ex1 device. The only problem is, that the IP address is not the one you have specified in the Python code. To correct that, open the `dns-config-tpl.xml` file in a text editor and replace the line that reads `<address>10.2.3.4</address>` with the following:

```xml
<address>{$DNS_IP}</address>
```

The only static part left in the template now is the target device and it's possible to parameterize that, too. The skeleton, created by the `ncs-make-package` command, already contains a node `device` in the service YANG file. It is there to allow the service operator to choose the target device to be configured.

```
leaf-list device {
  type leafref {
    path "/ncs:devices/ncs:device/ncs:name";
  }
}
```

One way to use the `device` service parameter is to read its value in the Python code and then set up the template parameters accordingly. However, there is a simpler way with XPath. In the template, replace the line that reads `<name>ex1</name>` with the following:

```xml
<name>{/device}</name>
```

The XPath expression inside the curly braces instructs NSO to get the value for the device name from the service instance's data, namely the node called `device`. In other words, when configuring a new service instance, you have to add the device parameter, which selects the router for provisioning. The final XML template is then:

```xml
<config-template xmlns="http://tail-f.com/ns/config/1.0">
  <devices xmlns="http://tail-f.com/ns/ncs">
    <device>
      <name>{/device}</name>
      <config>
        <sys xmlns="http://example.com/router">
          <dns>
            <server>
              <address>{$DNS_IP}</address>
            </server>
          </dns>
        </sys>
      </config>
    </device>
  </devices>
</config-template>
```

### Step 4 - Test the Service <a href="#d5e884" id="d5e884"></a>

Remember to save the template file and return to the NSO CLI. Because you have updated the service code, you have to redeploy it for NSO to pick up the changes:

```cli
admin@ncs# packages package dns-config redeploy
result true]
```

Alternatively, you could call the `packages reload` command, which does a full reload of all the packages.

Next, enter the configuration mode:

```cli
admin@ncs# config
```

As you are using the device node in the service model for target router selection, configure a service instance for the `ex2` router in the following way:

```cli
admin@ncs(config)# dns-config dns-for-ex2 device ex2
```

Finally, using the `commit dry-run` command, observe the `ex2` router being configured with an additional DNS server.

```cli
admin@ncs(config-dns-config-dns-for-ex2)# commit dry-run
```

As a bonus for using an XPath expression to a leaf-list in the service template, you can actually select multiple router devices in a single service instance and they will all be configured.


# Applications in NSO

Build your own applications in NSO.

Services provide the foundation for managing the configuration of a network. But this is not the only aspect of network automation. A holistic solution must also consider various verification procedures, one-time actions, monitoring, and so on. This is quite different from managing configuration. NSO helps you implement such automation use cases through a generic application framework.

This section explores the concept of services as more general NSO applications. It gives an overview of the mechanisms for orchestrating network automation tasks that require more than just configuration provisioning.

## NSO Architecture <a href="#d5e907" id="d5e907"></a>

You have seen two different ways in which you can make a configuration change on a network device. With the first, you make changes directly on the NSO copy of the device configuration. The Device Manager picks up the changes and propagates them to the affected devices.

The purpose of the Device Manager is to manage different devices uniformly. The Device Manager uses the Network Element Drivers (NEDs) to abstract away the different protocols and APIs towards the devices. The NED contains a YANG data model for a supported device. So, each device type requires an appropriate NED package that allows the Device Manager to handle all devices in the same, YANG-model-based way.

The second way to make configuration changes is through services. Here, the Service Manager adds a layer on top of the Device Manager to process the service request and enlists the help of service-aware applications to generate the device changes.

The following figure illustrates the difference between the two approaches.

<div data-with-frame="true"><figure><img src="/files/CG99USbQEuroPZ1lhPKc" alt="" width="375"><figcaption><p>Device and Service Manager</p></figcaption></figure></div>

The Device Manager and the Service Manager are tightly integrated into one transactional engine, using the CDB to store data. Another thing the two managers have in common is packages. Like Device Manager uses NED packages to support specific devices, Service Manager relies on service packages to provide an application-specific mapping for each service type.

However, a network application can consist of more than just a configuration recipe. For example, an integrated service test action can verify the initial provisioning and simplify troubleshooting if issues arise. A simple test might run the `ping` command to verify connectivity. Or an application could only monitor the network and not produce any configuration at all. That is why NSO actually uses an approach where an application chooses what custom code to execute for specific NSO events.

## Callbacks as an Extension Mechanism <a href="#d5e921" id="d5e921"></a>

NSO allows augmenting the base functionality of the system by delegating certain functions to applications. As the communication must happen on demand, NSO implements a system of callbacks. Usually, the application code registers the required callbacks on start-up, and then NSO can invoke each callback as needed. A prime example is a Python service, which registers the `cb_create()` function as a service callback that NSO uses to construct the actual configuration.

<div data-with-frame="true"><figure><img src="/files/3GxUZqo8zA9yx4qEVowe" alt="" width="375"><figcaption><p>Service Callback</p></figcaption></figure></div>

In a Python service skeleton, callback registration happens inside a class `Main`, found in `main.py`:

```python
class Main(ncs.application.Application):
    def setup(self):
        # Service callbacks require a registration for a 'service point',
        # as specified in the corresponding data model.
        #
        self.register_service('my-svc-servicepoint', ServiceCallbacks)
```

In this code, the `register_service()` method registers the `ServiceCallbacks` class to receive callbacks for a service. The first argument defines which service that is. In theory, a single class could even handle service callbacks for multiple services but that is not a common practice.

On the other hand, it is also possible that no code registered a callback for a given service. This is quite often a result of a misspelling or a bug in the code that causes the application code to crash. In these situations, NSO presents an error if you try to use the service:

```
Error: no registration found for callpoint my-svc-servicepoint/service_create of type=external
```

This error refers to the concept of a service point. Service points are declared in the service YANG model and allow NSO to distinguish ordinary data from services. They instruct NSO to invoke FASTMAP and the service callbacks when a service instance is being provisioned. That means the service skeleton YANG file also contains a service point definition, such as the following:

```yang
list my-svc {
  description "This is an RFS skeleton service";

  uses ncs:service-data;
  ncs:servicepoint my-svc-servicepoint;
}
```

Service point therefore links the definition in the model with custom code. Some methods in the code will have names starting with `cb_`, for instance, the `cb_create()` method, letting you know quickly that they are an implementation of a callback.

NSO implements additional callbacks for each service point, that may be required in some specific circumstances. Most of these callbacks perform work outside of the automatic change tracking, so you need to consider that before using them. The section [Service Callbacks](https://nso-docs.cisco.com/guides/development/introduction-to-automation/pages/Bw8TviXCSEsM9XjXBC1d#ch_svcref.cbs) offers more details.

As well as services, other extensibility options in NSO also rely on callbacks and `callpoints`, a generalized version of a service point. Two notable examples are validation callbacks, to implement additional validation logic to that supported by YANG, and custom actions. The section [Overview of Extension Points](#overview-of-extension-points) provides a comprehensive list and an overview of when to use each.

In summary, you implement custom behavior in NSO by providing the following three parts:

* A YANG model directing NSO to use callbacks, such as a service point for services.
* Registration of callbacks, telling NSO to call into your code at a given point.
* The implementation of each callback with your custom logic.

This way, an application in NSO can implement all the required functionality for a given use case (configuration management and otherwise) by registering the right callbacks.

## Overview of Extension Points

NSO supports a number of extension points for custom callbacks:

<table data-full-width="false"><thead><tr><th>Type</th><th>Supported In</th><th>YANG Extension</th><th>Description</th></tr></thead><tbody><tr><td>Service</td><td>Python, Java, Erlang</td><td><code>ncs:servicepoint</code></td><td>Transforms a list or container into a model for service instances. When the configuration of a service instance changes, NSO invokes Service Manager and FASTMAP, which may call service create and similar callbacks. See <a href="/pages/8ue8V7101TddJ17baqbw">Developing a Simple Service</a> for an introduction.</td></tr><tr><td>Action</td><td>Python, Java, Erlang</td><td><code>tailf:actionpoint</code></td><td>Defines callbacks when an action or RPC is invoked. See <a href="/pages/tRcyIzz683UB2NxeMnOj">Actions</a> for an introduction.</td></tr><tr><td>Validation</td><td>Python, Java, Erlang</td><td><code>tailf:validate</code></td><td>Defines callbacks for additional validation of data when the provided YANG functionality, such as <code>must</code> and <code>unique</code> statements are insufficient. See the respective API documentation for examples; the section <a href="/pages/BSsWUCJFFK53rRYvXVDw#validation-point-handler">ValidationPoint Handler</a> (Python), the section <a href="/pages/Uzy6qvKpLQF47FSwk0S2#d5e3761">Validation Callbacks</a> (Java), and <a href="/pages/fnCkPPtLHhIFVZ7ERbuq">Embedded Erlang applications</a> (Erlang).</td></tr><tr><td>Data Provider</td><td>Java, Python (low-level API with experimental high-level API), Erlang</td><td><code>tailf:callpoint</code></td><td>Defines callbacks for transparently accessing external data (data not stored in the CDB) or callbacks for special processing of data nodes (transforms, set, and transaction hooks). Requires careful implementation and understanding of transaction intricacies. Rarely used in NSO.</td></tr></tbody></table>

Each extension point in the list has a corresponding YANG extension that defines to which part of the data model the callbacks apply, as well as the individual name of the call point. The name is required during callback registration and helps distinguish between multiple uses of the extension. Each extension generally specifies multiple callbacks, however, you often need to implement only the main one, e.g. create for services or action for actions.

## Monitoring for Change <a href="#ch_apps.kickers" id="ch_apps.kickers"></a>

Services and actions are examples of something that happens directly as a result of a user (or other northbound agent) request. That is, a user takes an active role in starting service instantiation or invoking an action. Contrast this to a change that happens in the network and requires the orchestration system to take some action. In this latter case, the system monitors the notifications that the network generates, such as losing a link, and responds to the new data.

NSO provides out-of-the-box support for the automation of not only notifications but also changes to the operational and configuration data, using the concept of kickers. With kickers, you can watch for a particular change to occur in the system and invoke a custom action that handles the change.

The kicker system is further described in [Kicker](/guides/development/advanced-development/kicker).

## Running Application Code <a href="#ncs.development.applications.running" id="ncs.development.applications.running"></a>

Services, actions, and other features all rely on callback registration. In Python code, the class responsible for registration derives from the `ncs.application.Application`. This allows NSO to manage the application code as appropriate, such as starting and stopping in response to NSO events. These events include package load or unload and NSO start or stop events.

While the Python package skeleton names the derived class `Main`, you can choose a different name if you also update the `package-meta-data.xml` file accordingly. This file defines a component with the name of the Python class to use:

```xml
<ncs-package xmlns="http://tail-f.com/ns/ncs-packages">
  < ... output omitted ... >

  <component>
    <name>main</name>
    <application>
      <python-class-name>dns_config.main.Main</python-class-name>
    </application>
  </component>
</ncs-package>
```

When starting the package, NSO reads the class name from `package-meta-data.xml`, starts the Python interpreter, and instantiates a class instance. The base `Application` class takes care of establishing communication with the NSO process and calling the `setup` and `teardown` methods. The two methods are a good place to do application-specific initialization and cleanup, along with any callback registrations you require.

The communication between the application process and NSO happens through a dedicated control socket, as described in the section called [IPC Ports](/guides/administration/advanced-topics/ipc-connection) in Administration. This setup prevents a faulty application from bringing down the whole system along with it and enables NSO to support different application environments.

In fact, NSO can manage applications written in Java or Erlang in addition to those in Python. If you replace the `python-class-name` element of a component with `java-class-name` in the `package-meta-data.xml` file, NSO will instead try to run the specified Java class in the managed Java VM. If you wanted to, you could implement all of the same services and actions in Java, too. For example, see [Service Actions](https://nso-docs.cisco.com/guides/development/introduction-to-automation/pages/1ErWKzeM15TwAei4yXWm#ch_services.actions) to compare Python and Java code.

Regardless of the programming language you use, the high-level approach to automation with NSO does not change, registering and implementing callbacks as part of your network application. Of course, the actual function calls (the API) and other specifics differ for each language. The [NSO Python VM](/guides/development/core-concepts/nso-virtual-machines/nso-python-vm), [NSO Java VM](/guides/development/core-concepts/nso-virtual-machines/nso-java-vm), and [Embedded Erlang Applications](/guides/development/core-concepts/nso-virtual-machines/embedded-erlang-applications) cover the details. Even so, the concepts of actions, services, and YANG modeling remain the same.

As you have seen, everything in NSO is ultimately tied to the YANG model, making YANG knowledge such a valuable skill for any NSO developer.

## Application Updates

As your NSO application evolves, you will create newer versions of your application package, which will replace the existing one. If the application becomes sufficiently complex, you might even split it across multiple packages.

When you replace a package, NSO must redeploy the application code and potentially replace the package-provided part of the YANG schema. For the latter, NSO can perform the data migration for you, as long as the schema is backward compatible. This process is documented in [Automatic Schema Upgrades and Downgrades](https://nso-docs.cisco.com/guides/development/introduction-to-automation/pages/FxpCNgv5QKnfWrJw4nXf#ug.cdb.upgrade) and is automatic when you request a reload of the package with `packages reload` or a similar command.

If your schema changes are not backward compatible, you can implement a data migration procedure, which NSO invokes when upgrading the schema. Among other things, this allows you to reuse and migrate the data that is no longer present in the new schema. You can specify the migration procedure as part of the `package-meta-data.xml` file, using a component of the `upgrade` type. See [The Upgrade Component](https://nso-docs.cisco.com/guides/development/introduction-to-automation/pages/jwFVej1l7kjXuaM9axfN#ncs.development.pythonvm.upgrade) (Python) and [examples.ncs/service-management/upgrade-service](https://github.com/NSO-developer/nso-examples/tree/6.7/service-management/upgrade-service) example (Java) for details.

Note that changing the schema in any way requires you to recompile the `.fxs` files in the package, which is typically done by running `make` in the package's `src` folder.

However, if the schema does not change, you can request that only the application code and templates be redeployed by using the `packages package <my-pkg> redeploy` command.


# Core Concepts

Key concepts in NSO development.


# Services

Implement network automation in your NSO deployment using services.

Services are the cornerstone of network automation with NSO. A service is not just a reusable recipe for provisioning network configurations; it allows you to manage the full configuration life cycle with minimal effort.

This section examines in greater detail how services work, how to design them, and the different ways to implement them.

{% hint style="success" %}
For a quicker introduction and a simple showcase of services, see [Develop a Simple Service](/guides/development/introduction-to-automation/develop-a-simple-service).
{% endhint %}

In NSO, the term service has a special meaning and represents an automation construct that orchestrates the 'create', 'modify', and 'delete' of a service instance into the resulting native commands to devices in the network. In its simplest form, a service takes some input parameters and maps them to device-specific configurations. It is a recipe or a set of instructions.

Much like you can bake many cakes using a single cake recipe, you can create many service instances using the same service. But unlike cakes, having the recipe produce exactly the same output, is not very useful. That is why service instances define a set of input parameters, which the service uses to customize the produced configuration.

A network engineer on the CLI, or an API call from a northbound system, provides the values for input parameters when requesting a new service instance, and NSO uses the service recipe, called a 'service mapping', to configure the network.

<div data-with-frame="true"><figure><img src="/files/lzFmh71XENpFeq62LZ99" alt="" width="375"><figcaption><p>A High-level View of Services in NSO</p></figcaption></figure></div>

A similar process takes place when deleting the service instance or modifying the input parameters. The main task of a service is therefore: from a given set of input parameters, calculate the minimal set of device operations to achieve the desired service change. Here, it is very important that the service supports any change; create, delete, and update of any service parameter.

Device configuration is usually the primary goal of a service. However, there may be other supporting functions that are expected from the service, such as service-specific actions. The complete service application, implementing all the service functionality, is packaged in an NSO service package.

The following definitions are used throughout this section:

* **Service type**: Often referred to simply as a service; denotes a specific type of service, such as "L2 VPN", "L3 VPN", "Firewall", or "DNS".
* **Service instance**: A specific instance of a service type, such as "L3 VPN for ACME" or "Firewall for user X".
* **Service model**: The schema definition for a service type, defined in YANG. It specifies the names and format of input parameters for the service.
* **Service mapping**: The instructions that implement a service by mapping the input parameters for a service instance to device configuration.
* **Device configuration**: Network devices are configured to perform network functions. A service instance results in corresponding device configuration changes.
* **Service application**: The code and models implementing the complete service functionality, including service mapping, actions, models for auxiliary data, and so on.

## Service Mapping <a href="#d5e1405" id="d5e1405"></a>

Developing a service that transforms a service instance request to the relevant device configurations is done differently in NSO than in most other tools on the market. As a service developer, you create a mapping from a YANG service model to the corresponding device YANG model.

This is a declarative, model-to-model mapping. Irrespective of the underlying device type and its native device interface, the mapping is towards a YANG device model and not the native CLI (or any other protocol/API). As you write the service mapping, you do not have to worry about the syntax of different CLI commands or in which order these commands are sent to the device. It is all taken care of by the NSO device manager and device NEDs. Implementing a service in NSO is reduced to transforming the input data structure, described in YANG, to device data structures, also described in YANG.

Who writes the models?

* Developing the service model is part of developing the service application and is covered later in this section.
* Every device NED comes with a corresponding device YANG model. This model has been designed by the NED developer to capture the configuration data that is supported by the device.

A service application then has two primary artifacts: a YANG service model and a mapping definition to the device YANG, as illustrated in the following figure.

<div data-with-frame="true"><figure><img src="/files/f1zrpAFiFsZ3wzDP9Ygd" alt="" width="188"><figcaption><p>Service Model and Mapping</p></figcaption></figure></div>

To reiterate:

* The mapping is not defined using workflows, or sequences of device commands.
* The mapping is not defined in the native device interface language.

This approach may seem somewhat unorthodox at first but allows NSO to streamline and greatly simplify how you implement services.

A common problem for traditional automation systems is that a set of instructions needs to be defined for every possible service instance change. Take, for example, a VPN service. During a service life cycle, you want to:

1. Create the initial VPN.
2. Add a new site or leg to the VPN.
3. Remove a site or leg from the VPN.
4. Modify the parameters of a VPN leg, such as the IP addresses used.
5. Change the interface used for the VPN on a device.
6. ...
7. Delete the VPN.

The possible run-time changes for an existing service instance are numerous. If a developer must define instructions for every possible change, such as a script or a workflow, the task is daunting, error-prone, and never-ending.

NSO reduces this problem to a single data-mapping definition for the "create" scenario. At run-time, NSO renders the minimum resulting change for any possible change in the service instance. It achieves this with the FASTMAP algorithm.

Another challenge in traditional systems is that a lot of code goes into managing error scenarios. The NSO built-in transaction manager takes that burden away from the developer of the service application by providing automatic rollback of incomplete changes.

Another benefit of this approach is that NSO can automatically generate the northbound APIs and database schema from the YANG models, enabling a true DevOps way of working with service models. A new service model can be defined as part of a package and loaded into NSO. An existing service model can be modified, and the package upgraded, and all northbound APIs and user interfaces are automatically regenerated to reflect the new or updated models.


# Implementing Services

Explore service development in detail.

## A Template is All You Need <a href="#ch_services.just_template" id="ch_services.just_template"></a>

To demonstrate the simplicity a pure model-to-model service mapping affords, let us consider the most basic approach to providing the mapping: the service XML template. The XML template is an XML-encoded file that tells NSO what configuration to generate when someone requests a new service instance.

The first thing you need is the relevant device configuration (or configurations if multiple devices are involved). Suppose you must configure `192.0.2.1` as a DNS server on the target device. Using the NSO CLI, you first enter the device configuration, then add the DNS server. For a Cisco IOS-based device:

```bash
admin@ncs# config
Entering configuration mode terminal
admin@ncs(config)# devices device c1 config
admin@ncs(config-config)# ip name-server 192.0.2.1
admin@ncs(config-config)# top
admin@ncs(config)#
```

Note here that the configuration is not yet committed. You can use the `show configuration` command and pipe it through the `display xml-template` filter to produce the configuration in the format of an XML template.

```xml
admin@ncs(config)# show configuration | display xml-template

<config-template xmlns="http://tail-f.com/ns/config/1.0">
  <devices xmlns="http://tail-f.com/ns/ncs">
    <device>
      <name>c1</name>
      <config>
        <ip xmlns="urn:ios">
          <name-server>192.0.2.1</name-server>
        </ip>
      </config>
    </device>
  </devices>
</config-template>
```

The interesting portion is the part between `<devices>` and `</devices>` tags.

Another way to get the XML template output is to list the existing device configuration in NSO by piping it through the `display xml-template` filter:

```xml
admin@ncs# show running-config devices device c1 config ip name-server | display xml-template

<config-template xmlns="http://tail-f.com/ns/config/1.0">
  <devices xmlns="http://tail-f.com/ns/ncs">
    <device>
      <name>c1</name>
      <config>
        <ip xmlns="urn:ios">
          <name-server>192.0.2.1</name-server>
        </ip>
      </config>
    </device>
  </devices>
</config-template>
```

If there is a lot of data, it is easy to save the output to a file using the `save` pipe in the CLI, instead of copying and pasting it by hand:

```bash
admin@ncs# show running-config devices device c1 config ip name-server | display xml-template\
 | save dns-template.xml
```

The last command saves the configuration for a device in the `dns-template.xml` file using XML template format. To use it in a service, you need a service package.

You create an empty, skeleton service with the `ncs-make-package` command, such as:

```bash
ncs-make-package --build --no-test --service-skeleton template dns
```

The command generates the minimal files necessary for a service package, here named `dns`. One of the files is `dns/templates/dns-template.xml`, which is where the configuration in the format of an XML template goes.

```xml
<config-template xmlns="http://tail-f.com/ns/config/1.0"
                 servicepoint="dns">
  <devices xmlns="http://tail-f.com/ns/ncs">
    <!-- ... more statements here ... -->
  </devices>
</config-template>
```

If you look closely, there is one difference from the `show running-config` output: the `config-template` XML root tag in the template file has the `servicepoint` attribute. Other than that, you can use the XML template formatted configuration from the CLI as-is.

Bringing the two XML documents together gives the final `dns/templates/dns-template.xml` XML template:

#### **Static DNS Configuration Template Example:**

{% code title="Example: Static DNS Configuration Template" %}

```xml
<config-template xmlns="http://tail-f.com/ns/config/1.0"
                 servicepoint="dns">
  <devices xmlns="http://tail-f.com/ns/ncs">
    <device>
      <name>c1</name>
      <config>
        <ip xmlns="urn:ios">
          <name-server>192.0.2.1</name-server>
        </ip>
      </config>
    </device>
  </devices>
</config-template>
```

{% endcode %}

The service is now ready to use in NSO. Start the [examples.ncs/service-management/implement-a-service/dns-v1](https://github.com/NSO-developer/nso-examples/tree/6.7/service-management/implement-a-service/dns-v1) example to set up a live NSO system with such a service and inspect how it works. Try configuring two different instances of the `dns` service.

```bash
$ cd $NCS_DIR/examples.ncs/service-management/implement-a-service/dns-v1
$ make demo
```

The problem with this service is that it always does the same thing because it always generates exactly the same configuration. It would be much better if the service could configure different devices. The updated version, v1.1, uses a slightly modified template:

```xml
<config-template xmlns="http://tail-f.com/ns/config/1.0"
                 servicepoint="dns">
  <devices xmlns="http://tail-f.com/ns/ncs">
    <device>
      <name>{/name}</name>
      <config>
        <ip xmlns="urn:ios">
          <name-server>192.0.2.1</name-server>
        </ip>
      </config>
    </device>
  </devices>
</config-template>
```

The changed part is `<name>{/name}</name>`, which now uses the `{/name}` code instead of a hard-coded `c1` value. The curly braces indicate that NSO should evaluate the enclosed expression and use the resulting value in its place. The `/name` expression is an XPath expression, referencing the service YANG model. In the model, `name` is the name you give each service instance. In this case, the instance name doubles for identifying the target device.

```cli
admin@ncs# config
Entering configuration mode terminal
admin@ncs(config)# dns c2
admin@ncs(config-dns-c2)# commit dry-run

cli {
    local-node {
        data  devices {
                  device c2 {
                      config {
                          ip {
             +                name-server 192.0.2.1;
                          }
                      }
                  }
              }
             +dns c2 {
             +}
    }
}
```

In the output, the instance name used was `c2` and that is why the service performs DNS configuration for the c2 device.

The template actually allows a decent amount of programmability through XPath and special XML processing instructions. For example:

```xml
<config-template xmlns="http://tail-f.com/ns/config/1.0"
                 servicepoint="dns">
  <devices xmlns="http://tail-f.com/ns/ncs">
    <device>
      <name>{/name}</name>
      <config>
        <ip xmlns="urn:ios">
          <?if {starts-with(/name, 'c1')}?>
            <name-server>192.0.2.1</name-server>
          <?else?>
            <name-server>192.0.2.2</name-server>
          <?end?>
        </ip>
      </config>
    </device>
  </devices>
</config-template>
```

In the preceding printout, the XPath `starts-with()` function is used to check if the device name starts with a specific prefix. Then one set of configuration items is used, and a different one otherwise. For additional available instructions and the complete set of template features, see [Templates](/guides/development/core-concepts/templates).

However, most provisioning tasks require some kind of input to be useful. Fortunately, you can define any number of input parameters in the service model that you can then reference from the template; either to use directly in the configuration or as something to base provisioning decisions on.

## Service Model Captures Inputs <a href="#ch_services.input" id="ch_services.input"></a>

The YANG service model specifies the input parameters a service in NSO takes. For a specific service model think of the parameters that a northbound system sends to NSO or the parameters that a network engineer needs to enter in the NSO CLI.

Even a service as simple as the DNS configuration service usually needs some parameters, such as the target device. The service model gives each parameter a name and defines validation rules, ensuring the client-provided values fit what the service expects.

Suppose you want to add a parameter for the target device to the simple DNS configuration service. You need to construct an appropriate service model, adding a YANG leaf to capture this input.

{% hint style="info" %}
This task requires some basic YANG knowledge. Review the section [Data Modeling Basics](/guides/development/introduction-to-automation/cdb-and-yang#d5e154) for a primer on the main building blocks of the YANG language.
{% endhint %}

The service model is located in the `src/yang/servicename.yang` file in the package. It typically resembles the following structure:

```yang
  list servicename {
    key name;

    uses ncs:service-data;
    ncs:servicepoint "servicename";

    leaf name {
      type string;
    }

    // ... other statements ...
  }
```

The list named after the package (`servicename` in the example) is the interesting part.

The `uses ncs:service-data` and `ncs:servicepoint` statements differentiate this list from any standard YANG list and make it a service. Each list item in NSO represents a service instance of this type.

The `uses ncs:service-data` part allows the system to store internal state and provide common service actions, such as `re-deploy` and `get-modifications` for each service instance.

The `ncs:servicepoint` identifies which part of the system is responsible for the service mapping. For a template-only service, it is the XML template that uses the same service point value in the `config-template` element.

The `name` leaf serves as the key of the list and is primarily used to distinguish service instances from each other.

The remaining statements describe the functionality and input parameters that are specific to this service. This is where you add the new leaf for the target device parameter of the DNS service:

```yang
  list dns {
    key name;

    uses ncs:service-data;
    ncs:servicepoint "dns";

    leaf name {
      type string;
    }

    leaf target-device {
      type string;
    }
  }
```

Use the [examples.ncs/service-management/implement-a-service/dns-v2](https://github.com/NSO-developer/nso-examples/tree/6.7/service-management/implement-a-service/dns-v2) example to explore how this model works and try to discover what deficiencies it may have.

```bash
$ cd $NCS_DIR/examples.ncs/service-management/implement-a-service/dns-v2
$ make demo
```

In its current form, the model allows you to specify any value for `target-device`, including none at all! Obviously, this is not good as it breaks the provisioning of the service. But even more importantly, not validating the input may allow someone to use the service in the way you have not intended and perhaps bring down the network.

You can guard against invalid input with the help of additional YANG statements. For example:

```yang
    leaf target-device {
      mandatory true;
      type string {
        length "2";
        pattern "c[0-2]";
      }
    }
```

Now this parameter is mandatory for every service instance and must be one of the string literals: `c0`, `c1`, or `c2`. This format is defined by the regular expression in the `pattern` statement. In this particular case, the `length` restriction is redundant but demonstrates how you can combine multiple restrictions. You can even add multiple `pattern` statements to handle more complex cases.

What if you wanted to make the DNS server address configurable too? You can add another leaf to the service model:

```yang
    leaf dns-server-ip {
      type inet:ipv4-address {
        pattern "192\\.0\\.2\\..*";
      }
    }
```

There are three notable things about this leaf:

* There is no mandatory statement, meaning the value for this leaf is optional. The XML template will be designed to provide some default value if none is given.
* The type of the leaf is `inet:ipv4-address`, which restricts the value for this leaf to an IP address.
* The `inet:ipv4-address` type is further restricted using a regular expression to only allow IP addresses from the 192.0.2.0/24 range.

YANG is very powerful and allows you to model all kinds of values and restrictions on the data. In addition to the ones defined in the YANG language ([RFC 7950, section 9](https://datatracker.ietf.org/doc/html/rfc7950#section-9)), predefined types describing common networking concepts, such as those from the `inet` namespace ([RFC 6991](https://datatracker.ietf.org/doc/html/rfc6991#section-2)), are available to you out of the box. It is much easier to validate the inputs when so many options are supported.

The one missing piece for the service is the XML template. You can take the Example [Static DNS Configuration Template](#static-dns-configuration-template-example) as a base and tweak it to reference the defined inputs.

Using the code `{`*`XYZ`*`}` or `{/`*`XYZ`*`}` in the template, instructs NSO to look for the value in the service instance data, in the node with the name *`XYZ`*. So, you can refer to the target-device input parameter as defined in YANG with the `{/target-device}` code in the XML template.

{% hint style="info" %}
The code inside the curly brackets actually contains an XPath 1.0 expression with the service instance data as its root, so an absolute path (with a slash) and a relative one (without it) refer to the same node in this case, and you can use either.
{% endhint %}

The final, improved version of the DNS service template that takes into account the new model, is:

```xml
<config-template xmlns="http://tail-f.com/ns/config/1.0"
                 servicepoint="dns">
  <devices xmlns="http://tail-f.com/ns/ncs">
    <device>
      <name>{/target-device}</name>
      <config>
        <ip xmlns="urn:ios">
          <?if {/dns-server-ip}?>
            <!-- If dns-server-ip is set, use that. -->
            <name-server>{/dns-server-ip}</name-server>
          <?else?>
            <!-- Otherwise, use the default one. -->
            <name-server>192.0.2.1</name-server>
          <?end?>
        </ip>
      </config>
    </device>
  </devices>
</config-template>
```

The following figure captures the relationship between the YANG model and the XML template that ultimately produces the desired device configuration.

<div data-with-frame="true"><figure><img src="/files/KbprnFn7MSHuVSljfRmz" alt="" width="563"><figcaption><p>XML Template and Model Relationship</p></figcaption></figure></div>

The complete service is available in the [examples.ncs/service-management/implement-a-service/dns-v2.1](https://github.com/NSO-developer/nso-examples/tree/6.7/service-management/implement-a-service/dns-v2.1) example. Feel free to investigate on your own how it differs from the initial, no-validation service.

```bash
$ cd $NCS_DIR/examples.ncs/service-management/implement-a-service/dns-v2.1
$ make demo
```

## Extracting the Service Parameters <a href="#ch_services.model" id="ch_services.model"></a>

When the service is simple, constructing the YANG model and creating the service mapping (the XML template) is straightforward. Since the two components are mostly independent, you can start your service design with either one.

If you write the YANG model first, you can load it as a service package into NSO (without having any mapping defined) and iterate on it. This way, you can try the model, which is the interface to the service, with network engineers or northbound systems before investing the time to create the mapping. This model-first approach is also sometimes called top-down.

The alternative is to create the mapping first. Especially for developers new to NSO, the template-first, or bottom-up, approach is often easier to implement. With this approach, you templatize the configuration and extract the required service parameters from the template.

Experienced NSO developers naturally combine the two approaches, without much thinking. However, if you have trouble modeling your service at first, consider following the template-first approach demonstrated here.

For the following example, suppose you want the service to configure IP addressing on an ethernet interface. You know what configuration is required to do this manually for a particular ethernet interface. For a Cisco IOS-based device you would use the commands, such as:

```bash
admin@ncs# config
Entering configuration mode terminal
admin@ncs(config)# devices device c1 config
admin@ncs(config-config)# interface GigabitEthernet 0/0
admin@ncs(config-if)# ip address 192.168.5.1 255.255.255.0
```

To transform this configuration into a reusable service, complete the following steps:

* Create an XML template with hard-coded values.
* Replace each value specific to this instance with a parameter reference.
* Add each parameter to the YANG model.
* Add parameter validation.
* Consolidate and clean up the YANG model as necessary.

Start by generating the configuration in the format of an XML template, making use of the `display xml-template` filter. Note that the XML template will not necessarily be a one-to-one mapping of the CLI commands; the XML reflects the device YANG model which can be more complex but the commands on the CLI can hide some of this complexity.

The transformation to a template also requires you to add the `servicepoint` attribute to the `config-template` XML root tag, which produces the resulting XML template:

```xml
<config-template xmlns="http://tail-f.com/ns/config/1.0"
                 servicepoint="iface-servicepoint">
  <devices xmlns="http://tail-f.com/ns/ncs">
    <device>
      <name>c1</name>
      <config>
        <interface xmlns="urn:ios">
          <GigabitEthernet>
            <name>0/0</name>
            <ip>
              <address>
                <primary>
                  <address>192.168.5.1</address>
                  <mask>255.255.255.0</mask>
                </primary>
              </address>
            </ip>
          </GigabitEthernet>
        </interface>
      </config>
    </device>
  </devices>
</config-template>
```

However, this template has all the values hard-coded and only configures one specific interface on one specific device.

Now you must replace all the dynamic parts that vary from service instance to service instance with references to the relevant parameters. In this case, it is data specific to each device: which interface and which IP address to use.

Suppose you pick the following names for the variable parameters:

1. `device`: The network device to configure.
2. `interface`: The network interface on the selected device.
3. `ip-address`: The IP address to use on the selected interface.

Generally, you can make up any name for a parameter but it is best to follow the same rules that apply for naming variables in programming languages, such as making the name descriptive but not excessively verbose. It is customary to use a hyphen (minus sign) to concatenate words and use all-lowercase (“kebab-case”), which is the convention used in the YANG language standards.

<div data-with-frame="true"><figure><img src="/files/y7JDMPzqi1l26H0xZvjn" alt="" width="375"><figcaption><p>Making a Configuration Template</p></figcaption></figure></div>

The corresponding template then becomes:

```xml
<config-template xmlns="http://tail-f.com/ns/config/1.0"
                 servicepoint="iface-servicepoint">
  <devices xmlns="http://tail-f.com/ns/ncs">
    <device>
      <name>{/device}</name>
      <config>
        <interface xmlns="urn:ios">
          <GigabitEthernet>
            <name>{/interface}</name>
            <ip>
              <address>
                <primary>
                  <address>{/ip-address}</address>
                  <mask>255.255.255.0</mask>
                </primary>
              </address>
            </ip>
          </GigabitEthernet>
        </interface>
      </config>
    </device>
  </devices>
</config-template>
```

Having completed the template, you can add all the parameters, three in this case, to the service model.

<div data-with-frame="true"><figure><img src="/files/sKAzL7dFJjFFvBhkBwxR" alt="" width="375"><figcaption><p>Extracting Service Model from Template in a Bottom-up Approach</p></figcaption></figure></div>

The partially completed model is now:

```yang
  list iface {
    key name;

    uses ncs:service-data;
    ncs:servicepoint "iface-servicepoint";

    leaf name {
      type string;
    }

    leaf device { ... }

    leaf interface { ... }

    leaf ip-address { ... }
  }
```

Missing are the data type and other validation statements. At this point, you could fill out the model with generic `type string` statements, akin to the `name` leaf. This is a useful technique to test out the service in early development. But here you can complete the model directly, as it contains only three parameters.

You can use a `leafref` type leaf to refer to a device by its name in the NSO. This type uses dynamic lookup at the specified path to enumerate the available values. For the `device` leaf, it lists every value for a device name that NSO knows about. If there are two devices managed by NSO, named `rtr-sjc-01` and `rtr-sto-01`, either “`rtr-sjc-01`” or “`rtr-sto-01`” are valid values for such a leaf. This is a common way to refer to devices in NSO services.

```yang
    leaf device {
      mandatory true;
      type leafref {
        path "/ncs:devices/ncs:device/ncs:name";
      }
    }
```

In a similar fashion, restrict the valid values of the other two parameters.

```yang
    leaf interface {
      mandatory true;
      type string {
        pattern "[0-9]/[0-9]+";
      }
    }

    leaf ip-address {
      mandatory true;
      type inet:ipv4-address;
    }
  }
```

You would typically create the service package skeleton with the `ncs-make-package` command and update the model in the `.yang` file. The model in the skeleton might have some additional example leafs that you do not need and should remove to finalize the model. That gives you the final, full-service model:

```yang
  list iface {
    key name;

    uses ncs:service-data;
    ncs:servicepoint "iface-servicepoint";

    leaf name {
      type string;
    }

    leaf device {
      mandatory true;
      type leafref {
        path "/ncs:devices/ncs:device/ncs:name";
      }
    }

    leaf interface {
      mandatory true;
      type string {
        pattern "[0-9]/[0-9]+";
      }
    }

    leaf ip-address {
      mandatory true;
      type inet:ipv4-address;
    }
  }
```

The [examples.ncs/service-management/implement-a-service/iface-v1](https://github.com/NSO-developer/nso-examples/tree/6.7/service-management/implement-a-service/iface-v1) example contains the complete YANG module with this service model in the `packages/iface-v1/src/yang/iface.yang` file, as well as the corresponding service template in `packages/iface-v1/templates/iface-template.xml`.

## FASTMAP and Service Life Cycle <a href="#ch_services.fastmap" id="ch_services.fastmap"></a>

The YANG model and the mapping (the XML template) are the two main components required to implement a service in NSO. The hidden part of the system that makes such an approach feasible is called FASTMAP.

FASTMAP covers the complete service life cycle: creating, changing, and deleting the service. It requires a minimal amount of code for mapping from a service model to a device model.

FASTMAP is based on generating changes from an initial create operation. When the service instance is created the reverse of the resulting device configuration is stored together with the service instance. If an NSO user later changes the service instance, NSO first applies (in an isolated transaction) the reverse diff of the service, effectively undoing the previous create operation. Then it runs the logic to create the service again and finally performs a diff against the current configuration. Only the result of the diff is then sent to the affected devices.

{% hint style="warning" %}
It is therefore very important that the service create code produces the same device changes for a given set of input parameters every time it is executed. See [Persistent Opaque Data](https://nso-docs.cisco.com/guides/development/core-concepts/pages/Bw8TviXCSEsM9XjXBC1d#ch_svcref.opaque) for techniques to achieve this.
{% endhint %}

If the service instance is deleted, NSO applies the reverse diff of the service, effectively removing all configuration changes the service did on the devices.

<div data-with-frame="true"><figure><img src="/files/71iMpqctpCb0x5Cbeai0" alt="" width="563"><figcaption><p>FASTMAP Create a Service</p></figcaption></figure></div>

Assume we have a service model that defines a service with attributes X, Y, and Z. The mapping logic calculates that attributes A, B, and C must be set on the devices. When the service is instantiated, the previous values of the corresponding device attributes A, B, and C are stored with the service instance in the CDB. This allows NSO to bring the network back to the state before the service was instantiated.

Now let us see what happens if one service attribute is changed. Perhaps the service attribute Z is changed. NSO will execute the mapping as if the service was created from scratch. The resulting device configurations are then compared with the actual configuration and the minimal diff is sent to the devices. Note that this is managed automatically, there is no code to handle the specific "change Z" operation.

<div data-with-frame="true"><figure><img src="/files/ONUFJvpzJMuoHIVfeYJk" alt="" width="563"><figcaption><p>FASTMAP Change a Service</p></figcaption></figure></div>

When a user deletes a service instance, NSO retrieves the stored device configuration from the moment before the service was created and reverts to it.

<div data-with-frame="true"><figure><img src="/files/0h1MaubvKYJQ8DpIqpEH" alt="" width="563"><figcaption><p>FASTMAP Delete a Service</p></figcaption></figure></div>

## Templates and Code

For a complex service, you may realize that the input parameters for a service are not sufficient to render the device configuration. Perhaps the northbound system only provides a subset of the required parameters. For example, the other system wants NSO to pick an IP address and does not pass it as an input parameter. Then, additional logic or API calls may be necessary but XML templates provide no such functionality on their own.

The solution is to augment XML templates with custom code. Or, more accurately, create custom provisioning code that leverages XML templates. Alternatively, you can also implement the mapping logic completely in the code and not use templates at all. The latter, forgoing the templates altogether, is less common, since templates have a number of beneficial properties.

Templates separate the way parameters are applied, which depends on the type of target device, from calculating the parameter values. For example, you would use the same code to find the IP address to apply on a device, but the actual configuration might differ whether it is a Cisco IOS (XE) device, an IOS XR, or another vendor entirely.

Moreover, if you use templates, NSO can automatically validate the templates being compatible with the used NEDs, which allows you to sidestep whole groups of bugs.

NSO offers multiple programming languages to implement the code. The `--service-skeleton` option of the `ncs-make-package` command influences the selection of the programming language and if the generated code should contain sample calls for applying an XML template.

Suppose you want to extend the template-based ethernet interface addressing service to also allow specifying the netmask. You would like to do this in the more modern, CIDR-based single number format, such as is used in the 192.168.5.1/24 format (the /24 after the address). However, the generated device configuration takes the netmask in the dot-decimal format, such as 255.255.255.0, so the service needs to perform some translation. And that requires a custom service code.

Such a service will ultimately contain three parts: the service YANG model, the translation code, and the XML template. The model and the template serve the same purpose as before, while custom code provides fine-grained control over how templates are applied and the data available to them.

<div data-with-frame="true"><figure><img src="/files/SMDFdPCODwdaT8BVRPU6" alt="" width="375"><figcaption><p>Code and Template Service Compared to Template-only Service</p></figcaption></figure></div>

Since the service is based on the previous interface addressing service, you can save yourself a lot of work by starting with the existing YANG model and XML template.

The service YANG model needs an additional `cidr-netmask` leaf to hold the user-provided netmask value:

```yang
  list iface {
    key name;

    uses ncs:service-data;
    ncs:servicepoint "iface-servicepoint";

    leaf name {
      type string;
    }

    leaf device {
      mandatory true;
      type leafref {
        path "/ncs:devices/ncs:device/ncs:name";
      }
    }

    leaf interface {
      mandatory true;
      type string {
        pattern "[0-9]/[0-9]+";
      }
    }

    leaf ip-address {
      mandatory true;
      type inet:ipv4-address;
    }

    leaf cidr-netmask {
      default 24;
      type uint8 {
        range "0..32";
      }
    }
  }
```

This leaf stores a small number (of `uint8` type), with values between 0 and 32. It also specifies a default of 24, which is used when the client does not supply a value for this parameter.

The previous XML template also requires only minor tweaks. A small but important change is the removal of the `servicepoint` attribute on the top element. Since it is gone, NSO does not apply the template directly for each service instance. Instead, your custom code registers itself on this servicepoint and is responsible for applying the template.

The reason for it being this way is that the code will supply the value for the additional variable, here called `NETMASK`. This is the other change that is necessary in the template: referencing the `NETMASK` variable for the netmask value:

```xml
<config-template xmlns="http://tail-f.com/ns/config/1.0">
  <devices xmlns="http://tail-f.com/ns/ncs">
    <device>
      <name>{/device}</name>
      <config>
        <interface xmlns="urn:ios">
          <GigabitEthernet>
            <name>{/interface}</name>
            <ip>
              <address>
                <primary>
                  <address>{/ip-address}</address>
                  <mask>{$NETMASK}</mask>
                </primary>
              </address>
            </ip>
          </GigabitEthernet>
        </interface>
      </config>
    </device>
  </devices>
</config-template>
```

Unlike references to other parameters, `NETMASK` does not represent a data path but a variable. It must start with a dollar character (`$`) to distinguish it from a path. As shown here, variables are often written in all-uppercase, making it easier to quickly tell whether something is a variable or a data path.

Variables get their values from different sources but the most common one is the service code. You implement the service code using a programming language, such as Java or Python.

The following two procedures create an equivalent service that acts identically from a user's perspective. They only differ in the language used; they use the same logic and the same concepts. Still, the final code differs quite a bit due to the nature of each programming language. Generally, you should pick one language and stick with it. If you are unsure which one to pick, you may find Python slightly easier to understand because it is less verbose.

### Templates and Python Code <a href="#d5e1723" id="d5e1723"></a>

The usual way to start working on a new service is to first create a service skeleton with the `ncs-make-package` command. To use Python code for service logic and XML templates for applying configuration, select the `python-and-template` option. For example:

```bash
ncs-make-package --no-test --service-skeleton python-and-template iface
```

To use the prepared YANG model and XML template, save them into the `iface/src/yang/iface.yang` and `iface/templates/iface-template.xml` files. This is exactly the same as for the template-only service.

What is different, is the presence of the `python/` directory in the package file structure. It contains one or more Python packages (not to be confused with NSO packages) that provide the service code.

The function of interest is the `cb_create()` function, located in the `main.py` file that the package skeleton created. Its purpose is the same as that of the XML template in the template-only service: generate configuration based on the service instance parameters. This code is also called 'the create code'.

The create code usually performs the following tasks:

* Read service instance parameters.
* Prepare configuration variables.
* Apply one or more XML templates.

Reading instance parameters is straightforward with the help of the `service` function parameter, using the Maagic API. For example:

```python
    def cb_create(self, tctx, root, service, proplist):
        cidr_mask = service.cidr_netmask
```

Note that the hyphen in `cidr-netmask` is replaced with the underscore in `service.cidr_netmask` as documented in [Python API Overview](/guides/development/core-concepts/api-overview/python-api-overview).

The way configuration variables are prepared depends on the type of the service. For the interface addressing service with netmask, the netmask must be converted into dot-decimal format:

```
        quad_mask = ipaddress.IPv4Network((0, cidr_mask)).netmask
```

The code makes use of the built-in Python `ipaddress` package for conversion.

Finally, the create code applies a template, with only minimal changes to the skeleton-generated sample; the names and values for the `vars.add()` function, which are specific to this service.

```
        vars = ncs.template.Variables()
        vars.add('NETMASK', quad_mask)
        template = ncs.template.Template(service)
        template.apply('iface-template', vars)
```

If required, your service code can call `vars.add()` multiple times, to add as many variables as the template expects.

The first argument to the `template.apply()` call is the name of the XML template. Template name is the file path relative to the `templates` subdirectory, without the .xml suffix. It allows you to apply multiple, different templates for a single service instance. Separating the configuration into multiple templates based on functionality, called feature templates, is a great practice with bigger, complex configurations.

The complete create code for the service is:

```python
    def cb_create(self, tctx, root, service, proplist):
        cidr_mask = service.cidr_netmask

        quad_mask = ipaddress.IPv4Network((0, cidr_mask)).netmask

        vars = ncs.template.Variables()
        vars.add('NETMASK', quad_mask)
        template = ncs.template.Template(service)
        template.apply('iface-template', vars)
```

You can test it out in the [examples.ncs/service-management/implement-a-service/iface-v2-py](https://github.com/NSO-developer/nso-examples/tree/6.7/service-management/implement-a-service/iface-v2-py) example.

### Templates and Java Code <a href="#d5e1768" id="d5e1768"></a>

The usual way to start working on a new service is to first create a service skeleton with the `ncs-make-package` command. To use Java code for service logic and XML templates for applying the configuration, select the `java-and-template` option. For example:

```bash
ncs-make-package --no-test --service-skeleton java-and-template iface
```

To use the prepared YANG model and XML template, save them into the `iface/src/yang/iface.yang` and `iface/templates/iface-template.xml` files. This is exactly the same as for the template-only service.

What is different, is the presence of the `src/java` directory in the package file structure. It contains a Java package (not to be confused with NSO packages) that provides the service code and build instructions for the `ant` tool to compile the Java code.

The function of interest is the `create()` function, located in the `ifaceRFS.java` file that the package skeleton created. Its purpose is the same as that of the XML template in the template-only service: generate configuration based on the service instance parameters. This code is also called 'the create code'.

The create code usually performs the following tasks:

* Read service instance parameters.
* Prepare configuration variables.
* Apply one or more XML templates.

Reading instance parameters is done with the help of the `service` function parameter, using [NAVU API](https://nso-docs.cisco.com/guides/development/core-concepts/pages/Uzy6qvKpLQF47FSwk0S2#ug.java_api_overview.navu). For example:

```java
    public Properties create(ServiceContext context,
                             NavuNode service,
                             NavuNode ncsRoot,
                             Properties opaque)
                             throws ConfException {

        String cidr_mask_str = service.leaf("cidr-netmask").valueAsString();
        int cidr_mask = Integer.parseInt(cidr_mask_str);
```

The way configuration variables are prepared depends on the type of the service. For the interface addressing service with netmask, the netmask must be converted into dot-decimal format:

```java
        long tmp_mask = 0xffffffffL << (32 - cidr_mask);
        String quad_mask =
            ((tmp_mask >> 24) & 0xff) + "." +
            ((tmp_mask >> 16) & 0xff) + "." +
            ((tmp_mask >> 8) & 0xff) + "." +
            ((tmp_mask >> 0) & 0xff);
```

The create code applies a template, with only minimal changes to the skeleton-generated sample; the names and values for the `myVars.putQuoted()` function are different since they are specific to this service.

```java
        Template myTemplate = new Template(context, "iface-template");
        TemplateVariables myVars = new TemplateVariables();
        myVars.putQuoted("NETMASK", quad_mask);
        myTemplate.apply(service, myVars);
```

If required, your service code can call `myVars.putQuoted()` multiple times, to add as many variables as the template expects.

The second argument to the `Template` constructor is the name of the XML template. Template name is the file path relative to the `templates` subdirectory, without the .xml suffix. It allows you to instantiate and apply multiple, different templates for a single service instance. Separating the configuration into multiple templates based on functionality, called feature templates, is a great practice with bigger, complex configurations.

Finally, you must also return the `opaque` object and handle various exceptions for the function. If exceptions are propagated out of the create code, you should transform them into NSO specific ones first, so the UI can present the user with a meaningful error message.

The complete create code for the service is then:

```java
    public Properties create(ServiceContext context,
                             NavuNode service,
                             NavuNode ncsRoot,
                             Properties opaque)
                             throws ConfException {

        try {
            String cidr_mask_str = service.leaf("cidr-netmask").valueAsString();
            int cidr_mask = Integer.parseInt(cidr_mask_str);

            long tmp_mask = 0xffffffffL << (32 - cidr_mask);
            String quad_mask = ((tmp_mask >> 24) & 0xff) +
                "." + ((tmp_mask >> 16) & 0xff) +
                "." + ((tmp_mask >> 8) & 0xff) +
                "." + ((tmp_mask) & 0xff);

            Template myTemplate = new Template(context, "iface-template");
            TemplateVariables myVars = new TemplateVariables();
            myVars.putQuoted("NETMASK", quad_mask);
            myTemplate.apply(service, myVars);
        } catch (Exception e) {
            throw new DpCallbackException(e.getMessage(), e);
        }
        return opaque;
    }
```

You can test it out in the [examples.ncs/service-management/implement-a-service/iface-v2-java](https://github.com/NSO-developer/nso-examples/tree/6.7/service-management/implement-a-service/iface-v2-java) example.

## Configuring Multiple Devices <a href="#ch_services.devs" id="ch_services.devs"></a>

A service instance may require configuration on more than just a single device. In fact, it is quite common for a service to configure multiple devices.

<div data-with-frame="true"><figure><img src="/files/jYOHfVmZ4buv2s7eO4jN" alt="" width="375"><figcaption><p>Service Provisioning Multiple Devices</p></figcaption></figure></div>

There are a few ways in which you can achieve this for your services:

* **In code**: Using API, such as Python Maagic or Java NAVU, navigate the data model to individual device configurations under each `devices device DEVNAME config` and set the required values.
* **In code with templates**: Apply the template multiple times with different values, such as the device name.
* **With templates only**: use `foreach` or automatic (implicit) loops.

The generally recommended approach is to use either code with templates or templates with `foreach` loops. They are explicit and also work well when you configure devices of different types. Using only code extends less well to the latter case, as it requires additional logic and checks for each device type.

Automatic, implicit loops in templates are harder to understand since the syntax looks like the one for normal leafs. A common example is a device definition as a leaf-list in the service YANG model, such as:

```yang
    leaf-list device {
      type leafref {
        path "/ncs:devices/ncs:device/ncs:name";
      }
    }
```

Because it is a leaf-list, the following template applies to all the selected devices, using an implicit loop:

```xml
<config-template xmlns="http://tail-f.com/ns/config/1.0"
                 servicepoint="servicename">
  <devices xmlns="http://tail-f.com/ns/ncs">
    <device>
      <name>{/device}</name>
      <config>
        <!-- ... -->
     </config>
    </device>
  </devices>
</config-template>
```

It performs the same as the one, which loops through the devices explicitly:

```xml
<config-template xmlns="http://tail-f.com/ns/config/1.0"
                 servicepoint="servicename">
  <devices xmlns="http://tail-f.com/ns/ncs">
    <?foreach {/device}?>
      <device>
        <name>{.}</name>
        <config>
          <!-- ... -->
      </config>
      </device>
    <?end?>
  </devices>
</config-template>
```

Being explicit, the latter is usually much easier to understand and maintain for most developers. The [examples.ncs/service-management/implement-a-service/dns-v3](https://github.com/NSO-developer/nso-examples/tree/6.7/service-management/implement-a-service/dns-v3) demonstrates this syntax in the XML template.

### Supporting Different Device Types <a href="#ch_services.devs_types" id="ch_services.devs_types"></a>

Applying the same template works fine as long as you have a uniform network with similar devices. What if two different devices can provide the same service but require different configuration? Should you create two different services in NSO? No. Services allow you to abstract and hide the device specifics through a device-independent service model, while still allowing customization of device configuration per device type.

<div data-with-frame="true"><figure><img src="/files/DgoZCERbSu4DJ0GdcNIR" alt="" width="375"><figcaption><p>Service Provisioning Multiple Device Types</p></figcaption></figure></div>

One way to do this is to apply a different XML template from the service code, depending on the device type. However, the same is also possible through XML templates alone.

When NSO applies configuration elements in the template, it checks the XML namespaces that are used. If the target device does not support a particular namespace, NSO simply skips that part of the template. Consequently, you can put configuration for different device types in the same XML template and only the relevant parts will be applied.

Consider the following example:

```xml
<config-template xmlns="http://tail-f.com/ns/config/1.0">
  <devices xmlns="http://tail-f.com/ns/ncs">
    <device>
      <name>{/device}</name>
      <config>
        <!-- Part for device with the cisco-ios NED -->
        <interface xmlns="urn:ios">
          <GigabitEthernet>
            <name>{/interface}</name>
            <!-- ... -->
          </GigabitEthernet>
        </interface>

        <!-- Part for device with the router-nc NED -->
        <sys xmlns="http://example.com/router">
          <interfaces>
            <interface>
              <name>{/interface}</name>
              <!-- ... -->
            </interface>
          </interfaces>
        </sys>
     </config>
    </device>
  </devices>
</config-template>
```

Due to the `xmlns="urn:ios"` attribute, the first part of the template (the `interface GigabitEthernet`) will only apply to Cisco IOS-based device. While the second part (the `sys interfaces interface`) will only apply to the netsim-based router-nc-type devices, as defined by the `xmlns` attribute on the `sys` element.

In case you need to further limit what configuration applies where and namespace-based filtering is too broad, you can also use the `if-ned-id` XML processing instruction. Each NED package in NSO defines a unique NED-ID, which distinguishes between different device types (and possibly firmware versions). Based on the configured ned-id of the device, you can apply different parts of the XML template. For example:

```xml
<config-template xmlns="http://tail-f.com/ns/config/1.0">
  <devices xmlns="http://tail-f.com/ns/ncs">
    <device>
      <name>{/device}</name>
      <config>
        <?if-ned-id cisco-ios-cli-3.0:cisco-ios-cli-3.0?>
        <interface xmlns="urn:ios">
          <GigabitEthernet>
            <name>{/interface}</name>
            <!-- ... -->
          </GigabitEthernet>
        </interface>
        <?end?>
      </config>
    </device>
  </devices>
</config-template>
```

The preceding template applies configuration for the interface only if the selected device uses the `cisco-ios-cli-3.0` NED-ID. You can find the full code as part of the [examples.ncs/service-management/implement-a-service/iface-v3](https://github.com/NSO-developer/nso-examples/tree/6.7/service-management/implement-a-service/iface-v3) example.

## Shared Service Settings and Auxiliary Data <a href="#ch_services.data" id="ch_services.data"></a>

In the previous sections, we have looked at service mapping when the input parameters are enough to generate the corresponding device configurations. In many situations, this is not the case. The service mapping logic may need to reach out to other data in order to generate the device configuration. This is common in the following scenarios:

* Policies: Often a set of policies is defined that is shared between service instances. The policies, such as QoS, have data models of their own (not service models) and the mapping code reads data from those.
* Topology information: the service mapping might need to know how devices are connected, such as which network switches lie between two routers.
* Resources such as VLAN IDs or IP addresses, which might not be given as input parameters. They may be modeled separately in NSO or fetched from an external system.

It is important to design the service model considering the above requirements: what is input and what is available from other sources. In the latter case, in terms of implementation, an important distinction is made between accessing the existing data and allocating new resources. You must take special care for resource allocation, such as VLAN or IP address assignment, as discussed later on. For now, let us focus on using pre-existing shared data.

One example of such use is to define QoS policies "on the side." Only a reference to an existing QoS policy is supplied as input. This is a much better approach than giving all QoS parameters to every service instance. But note that, if you modify the QoS definitions the services are referring to, this will not immediately change the existing deployed service instances. In order to have the service implement the changed policies, you need to perform a **re-deploy** of the service.

A simpler example is a modified DNS configuration service that allows selecting from a predefined set of DNS servers, instead of supplying the DNS server directly as a service parameter. The main benefit in this case is that clients have no need to be aware of the actual DNS servers (and their IPs). In addition, this approach simplifies the management for the network operator, as all the servers are kept in a single place.

What is required to implement such as service? There are two parts. The first is the model and data that defines the available DNS server options, which are shared (used) across all the DNS service instances. The second is a modification to the service inputs and mapping logic to use this data.

For the first part, you must create a data model. If the shared data is specific to one service type, such as the DNS configuration, you can define it alongside the service instance model, in the service package. But sometimes this data may be shared between multiple types of service. Then it makes more sense to create a separate package for the shared data models.

In this case, define a new top-level container in the service's YANG file as:

```yang
  container dns-options {
    list dns-option {
      key name;

      leaf name {
        type string;
      }

      leaf-list servers {
        type inet:ipv4-address;
      }
    }
  }
```

Note that the container is defined outside the service list because this data is not specific to individual service instances:

```yang
  container dns-options {
    // ...
  }

  list dns {
    key name;

    uses ncs:service-data;
    ncs:servicepoint "dns";

    // ...
  }
```

The `dns-options` container includes a list of `dns-option` items. Each item defines a set of DNS servers (`leaf-list`) and a name for this set.

Once the shared data model is compiled and loaded into NSO, you can define the available DNS server sets:

```cli
admin@ncs(config)# dns-options dns-option lon servers 192.0.2.3
admin@ncs(config-dns-option-lon)# top
admin@ncs(config)# dns-options dns-option sto servers 192.0.2.3
admin@ncs(config-dns-option-sto)# top
admin@ncs(config)# dns-options dns-option sjc servers [ 192.0.2.5 192.0.2.6 ]
admin@ncs(config-dns-option-sjc)# commit
```

You must also update the service instance model to allow clients to pick one of these DNS servers:

```yang
  list dns {
    key name;

    uses ncs:service-data;
    ncs:servicepoint "dns";

    leaf name {
      type string;
    }

    leaf target-device {
      type string;
    }

    // Replace the old, explicit IP with a reference to shared data
    // leaf dns-server-ip {
    //   type inet:ip-address {
    //     pattern "192\.0.\.2\..*";
    //   }
    // }
    leaf dns-servers {
      mandatory true;
      type leafref {
        path "/dns-options/dns-option/name";
      }
    }
  }
```

Different ways exist to model the service input for `dns-servers`. The first option you might think about might be using a string type and a pattern to limit the inputs to one of `lon`, `sto`, or `sjc`. Another option would be to use a YANG `enum` type. But both of these have the drawback that you need to change the YANG model if you add or remove available `dns-option` items.

Using a `leafref` allows NSO to validate inputs for this leaf by comparing them to the values, returned by the `path` XPath expression. So, whenever you update the `/dns-options/dns-option` items, the change is automatically reflected in the valid `dns-server` values.

At the same time, you must also update the mapping to take advantage of this service input parameter. The service XML template is very similar to the previous one. The main difference is the way in which the DNS addresses are read from the CDB, using the special `deref()` XPath function:

```xml
<config-template xmlns="http://tail-f.com/ns/config/1.0"
                 servicepoint="dns">
  <devices xmlns="http://tail-f.com/ns/ncs">
    <device>
      <name>{/target-device}</name>
      <config>
        <ip xmlns="urn:ios">
          <name-server>{deref(/dns-servers)/../servers}</name-server>
        </ip>
      </config>
    </device>
  </devices>
</config-template>
```

The `deref()` function “jumps” to the item selected by the leafref. Here, leafref's path points to `/dns-options/dns-option/name`, so this is where `deref(/dns-servers)` ends: at the name leaf of the selected dns-option item.

The following code, which performs the same thing but in a more verbose way, further illustrates how the DNS server value is obtained:

```xml
        <ip xmlns="urn:ios">
          <?set dns_option = {/dns-servers}?>   <!-- Set $dns_option to e.g. 'lon' -->
          <?set-root-node {/}?>                 <!-- Make '/' point to datastore root,
                                                     instead of service instance   -->
          <name-server>{/dns-options/dns-option[name=$dns_option]/servers}</name-server>
        </ip>
```

The complete service is available in the [examples.ncs/service-management/implement-a-service/dns-v3](https://github.com/NSO-developer/nso-examples/tree/6.7/service-management/implement-a-service/dns-v3) example.

## Service Actions <a href="#ch_services.actions" id="ch_services.actions"></a>

NSO provides some service actions out of the box, such as **re-deploy** or **check-sync**. You can also add others. A typical use case is to implement some kind of a self-test action that tries to verify the service is operational. The latter could use **ping** or similar network commands, as well as verify device operational data, such as routing table entries.

This action supplements the built-in `check-sync` or `deep-check-sync` action, which checks for the required device configuration.

For example, a DNS configuration service might perform a domain lookup to verify the Domain Name System is working correctly. Likewise, an interface configuration service could ping an IP address or check the interface status.

The action consists of the YANG model for action inputs and outputs, as well as the action code that is executed when a client invokes the action.

Typically, such actions are defined per service instance, so you model them under the service list:

```yang
  list iface {
    key name;

    uses ncs:service-data;
    ncs:servicepoint "iface-servicepoint";

    leaf name { /* ... */ }
    leaf device { /* ... */ }
    leaf interface { /* ... */ }
    // ... other statements omitted ...

    action test-enabled {
      tailf:actionpoint iface-test-enabled;
      output {
        leaf status {
          type enumeration {
            enum up;
            enum down;
            enum unknown;
          }
        }
      }
    }
  }
```

The action needs no special inputs; because it is defined on the service instance, it can find the relevant interface to query. The output has a single leaf, called `status`, which uses an `enumeration` type for explicitly defining all the possible values it can take (`up`, `down`, or `unknown`).

Note that using the `action` statement requires you to also use the `yang-version 1.1` statement in the YANG module header (see [Actions](/guides/development/core-concepts/actions)).

### Action Code in Python <a href="#d5e1944" id="d5e1944"></a>

NSO Python API contains a special-purpose base class, `ncs.dp.Action`, for implementing actions. In the `main.py` file, add a new class that inherits from it, and implements an action callback:

```python
class IfaceActions(Action):
    @Action.action
    def cb_action(self, uinfo, name, kp, input, output, trans):
        ...
```

The callback receives a number of arguments, one of them being `kp`. It contains a keypath value, identifying the data model path, to the service instance in this case, it was invoked on.

The keypath value uniquely identifies each node in the data model and is similar to an XPath path, but encoded a bit differently. You can use it with the `ncs.maagic.cd()` function to navigate to the target node.

```
        root = ncs.maagic.get_root(trans)
        service = ncs.maagic.cd(root, kp)
```

The newly defined `service` variable allows you to access all of the service data, such as `device` and `interface` parameters. This allows you to navigate to the configured device and verify the status of the interface. The method likely depends on the device type and is not shown in this example.

The action class implementation then resembles the following:

```python
class IfaceActions(Action):
    @Action.action
    def cb_action(self, uinfo, name, kp, input, output, trans):
        root = ncs.maagic.get_root(trans)
        service = ncs.maagic.cd(root, kp)

        device = root.devices.device[service.device]

        status = 'unknown'    # Replace with your own code that checks
                              # e.g. operational status of the interface

        output.status = status
```

Finally, do not forget to register this class on the action point in the `Main` application.

```python
class Main(ncs.application.Application):
    def setup(self):
        ...
        self.register_action('iface-test-enabled', IfaceActions)
```

You can test the action in the [examples.ncs/service-management/implement-a-service/iface-v4-py](https://github.com/NSO-developer/nso-examples/tree/6.7/service-management/implement-a-service/iface-v4-py) example.

### Action Code in Java <a href="#d5e1966" id="d5e1966"></a>

Using the Java programming language, all callbacks, including service and action callback code, are defined using annotations on a callback class. The class NSO looks for is specified in the `package-meta-data.xml` file. This class should contain an `@ActionCallback()` annotated method that ties it back to the action point in the YANG model:

```java
    @ActionCallback(callPoint="iface-test-enabled",
                    callType=ActionCBType.ACTION)
    public ConfXMLParam[] test_enabled(DpActionTrans trans, ConfTag name,
                                       ConfObject[] kp, ConfXMLParam[] params)
    throws DpCallbackException {
        // ...
    }
```

The callback receives a number of arguments, one of them being `kp`. It contains a keypath value, identifying the data model path, to the service instance in this case, it was invoked on.

The keypath value uniquely identifies each node in the data model and is similar to an XPath path, but encoded a bit differently. You can use it with the `com.tailf.navu.KeyPath2NavuNode` class to navigate to the target node.

```
            NavuContext context = new NavuContext(maapi);
            NavuContainer service =
                (NavuContainer)KeyPath2NavuNode.getNode(kp, context);
```

The newly defined `service` variable allows you to access all of the service data, such as `device` and `interface` parameters. This allows you to navigate to the configured device and verify the status of the interface. The method likely depends on the device type and is not shown in this example.

The complete implementation requires you to supply your own Maapi read transaction and resembles the following:

```java
    @ActionCallback(callPoint="iface-test-enabled",
                    callType=ActionCBType.ACTION)
    public ConfXMLParam[] test_enabled(DpActionTrans trans, ConfTag name,
                                       ConfObject[] kp, ConfXMLParam[] params)
    throws DpCallbackException {
        try (Maapi maapi = new Maapi(UnixDomainSocketAddress.of(Conf.NCS_PATH))) {
            maapi.startUserSession("admin", "system");

            NavuContext context = new NavuContext(maapi);
            context.startRunningTrans(Conf.MODE_READ);

            NavuContainer root = new NavuContainer(context);
            NavuContainer service =
                (NavuContainer)KeyPath2NavuNode.getNode(kp, context);

            String status = "unknown";    // Replace with your own code that
                                          // checks e.g. operational status of
                                          // the interface

            String nsPrefix = name.getPrefix();
            return new ConfXMLParam[] {
                new ConfXMLParamValue(nsPrefix, "status", new ConfBuf(status)),
            };
        } catch (Exception e) {
            throw new DpCallbackException(name.toString() + " action failed",
                e);
        }
    }
```

You can test the action in the [examples.ncs/service-management/implement-a-service/iface-v4-java](https://github.com/NSO-developer/nso-examples/tree/6.7/service-management/implement-a-service/iface-v4-java) example.

## Operational Data <a href="#ch_services.oper" id="ch_services.oper"></a>

In addition to device configuration, services may also provide operational status or statistics. This is operational data, modeled with `config false` statements in YANG, and cannot be directly set by clients. Instead, clients can only read this data, for example to check service health.

What kind of data a service exposes depends heavily on what the service does. Perhaps the interface configuration service needs to provide information on whether a network interface was enabled and operational at the time of the last check (because such a check could be expensive).

Taking `iface` service as a base, consider how you can extend the instance model with another operational leaf to hold the interface status data as of the last check.

```yang
  list iface {
    key name;

    uses ncs:service-data;
    ncs:servicepoint "iface-servicepoint";

    // ... other statements omitted ...

    action test-enabled {
      tailf:actionpoint iface-test-enabled;
      output {
        leaf status {
          type enumeration {
            enum up;
            enum down;
            enum unknown;
          }
        }
      }
    }

    leaf last-test-result {
      config false;
      type enumeration {
            enum up;
            enum down;
            enum unknown;
      }
    }
  }
```

The new leaf `last-test-result` is designed to store the same data as the `test-enabled` action returns. Importantly, it also contains a `config false` substatement, making it operational data.

When faced with duplication of type definitions, as seen in the preceding code, the best practice is to consolidate the definition in a single place and avoid potential discrepancies in the future. You can use a `typedef` statement to define a custom YANG data type.

{% hint style="info" %}
The `typedef` statements should come before data statements, such as containers and lists in the model.
{% endhint %}

```
  typedef iface-status-type {
    type enumeration {
          enum up;
          enum down;
          enum unknown;
    }
  }
```

Once defined, you can use the new type as you would any other YANG type. For example:

```yang
    leaf last-test-status {
      config false;
      type iface-status-type;
    }

    action test-enabled {
      tailf:actionpoint iface-test-enabled;
      output {
        leaf status {
          type iface-status-type;
      }
    }
```

Users can then view operational data with the help of the `show` command. The data is also available through other NB interfaces, such as NETCONF and RESTCONF.

```cli
admin@ncs# show iface test-instance1 last-test-status
iface test-instance1 last-test-status up
```

But where does the operational data come from? The service application code provides this data. In this example, the `last-test-status` leaf captures the result of the enabled check, which is implemented as a custom action. So, here it is the action code that sets the leaf's value.

This approach works well when operational data is updated based on some event, such as a received notification or a user action, and NSO is used to cache its value.

For cases, where this is insufficient, NSO also allows producing operational data on demand, each time a client requests it, through the Data Provider API. See [DP API](https://nso-docs.cisco.com/guides/development/core-concepts/pages/Uzy6qvKpLQF47FSwk0S2#ug.java_api_overview.dp) for this alternative approach.

### Writing Operational Data in Python <a href="#d5e2012" id="d5e2012"></a>

Unlike configuration data, which always requires a transaction, you can write operational data to NSO with or without a transaction. Using a transaction allows you to easily compose multiple writes into a single atomic operation but has some small performance penalty due to transaction overhead.

If you avoid transactions and write data directly, you must use the low-level CDB API, which requires manual connection management and does not support Maagic API for data model navigation.

```python
with contextlib.closing(socket.socket(family=socket.AF_UNIX)) as s:
    _ncs.cdb.connect(s, _ncs.cdb.DATA_SOCKET, path=_ncs.PATH)
    _ncs.cdb.start_session(s, _ncs.cdb.OPERATIONAL)
    _ncs.cdb.set_elem(s, 'up', '/iface{test-instance1}/last-test-status')
```

The alternative, transaction-based approach uses high-level MAAPI and Maagic objects:

```python
with ncs.maapi.single_write_trans('admin', 'python', db=ncs.OPERATIONAL) as t:
    root = ncs.maagic.get_root(t)
    root.iface['test-instance1'].last_test_status = 'up'
    t.apply()
```

When used as part of the action, the action code might be as follows:

```python
    def cb_action(self, uinfo, name, kp, input, output, trans):
        with ncs.maapi.single_write_trans('admin', 'python',
                                          db=ncs.OPERATIONAL) as t:
            root = ncs.maagic.get_root(t)
            service = ncs.maagic.cd(root, kp)

            # ...
            service.last_test_status = status
            t.apply()

        output.status = status
```

Note that you have to start a new transaction in the action code, even though `trans` is already supplied, since `trans` is read-only and cannot be used for writes.

Another thing to keep in mind with operational data is that NSO by default does not persist it to storage, only keeps it in RAM. One way for the data to survive NSO restarts is to use the `tailf:persistent` statement, such as:

```yang
    leaf last-test-status {
      config false;
      type iface-status-type;
      tailf:cdb-oper {
        tailf:persistent true;
      }
    }
```

You can also register a function with the service application class to populate the data on package load, if you are not using `tailf:persistent`.

```python
class ServiceApp(Application):
    def setup(self):
        ...
        self.register_fun(init_oper_data, lambda _: None)


def init_oper_data(state):
    state.log.info('Populating operational data')
    with ncs.maapi.single_write_trans('admin', 'python',
                                      db=ncs.OPERATIONAL) as t:
        root = ncs.maagic.get_root(t)
        # ...
        t.apply()

    return state
```

The [examples.ncs/service-management/implement-a-service/iface-v5-py](https://github.com/NSO-developer/nso-examples/tree/6.7/service-management/implement-a-service/iface-v5-py) example implements such code.

### Writing Operational Data in Java <a href="#d5e2032" id="d5e2032"></a>

Unlike configuration data, which always requires a transaction, you can write operational data to NSO with or without a transaction. Using a transaction allows you to easily compose multiple writes into a single atomic operation but has some small performance penalty due to transaction overhead.

If you avoid transactions and write data directly, you must use the low-level CDB API, which does not support NAVU for data model navigation.

```java
SocketAddress address = UnixDomainSocketAddress.of(Conf.NCS_PATH);
try (Cdb cdb = new Cdb("IfaceServiceOperWrite", address)) {
    CdbSession session = cdb.startSession(CdbDBType.CDB_OPERATIONAL);

    String status = "up";
    ConfPath path = new ConfPath("/iface{%s}/last-test-status",
        "test-instance1");
    session.setElem(ConfEnumeration.getEnumByLabel(path, status), path);

    session.endSession();
}
```

The alternative, transaction-based approach uses high-level MAAPI and NAVU objects:

```java
try (Maapi maapi = new Maapi(UnixDomainSocketAddress.of(Conf.NCS_PATH))) {
    maapi.startUserSession("admin", "system");

    NavuContext context = new NavuContext(maapi);
    context.startOperationalTrans(Conf.MODE_READ_WRITE);

    NavuContainer root = new NavuContainer(context);
    NavuContainer service =
        (NavuContainer)KeyPath2NavuNode.getNode(kp, context);

    // ...
    service.leaf("last-test-status").set(status);
    context.applyClearTrans();
}
```

Note the use of the `context.startOperationalTrans()` function to start a new transaction against the operational data store. In other respects, the code is the same as for writing configuration data.

Another thing to keep in mind with operational data is that NSO by default does not persist it to storage, only keeps it in RAM. One way for the data to survive NSO restarts is to model the data with the `tailf:persistent` statement, such as:

```yang
    leaf last-check-status {
      config false;
      type iface-status-type;
      tailf:cdb-oper {
        tailf:persistent true;
      }
    }
```

You can also register a custom `com.tailf.ncs.ApplicationComponent` class with the service application to populate the data on package load, if you are not using `tailf:persistent`. Please refer to [The Application Component Type](/guides/development/core-concepts/nso-virtual-machines/nso-java-vm#d5e1255) for details.

The [examples.ncs/service-management/implement-a-service/iface-v5-java](https://github.com/NSO-developer/nso-examples/tree/6.7/service-management/implement-a-service/iface-v5-java) example implements such code.

## Nano Services for Provisioning with Side Effects <a href="#ncs.development.reactive_fastmap" id="ncs.development.reactive_fastmap"></a>

A FASTMAP service cannot perform explicit function calls with side effects. The only action a service is allowed to take is to modify the configuration of the current transaction. For example, a service may not invoke an action to generate authentication key files or start a virtual machine. All such actions must occur before the service is created and provided as input parameters. This restriction is because the FASTMAP code may be executed as part of a `commit dry-run`, or the commit may fail, in which case the side effects would have to be undone.

Nano services use a technique called reactive FASTMAP (RFM) and provide a framework to safely execute actions with side effects by implementing the service as several smaller (nano) steps or stages. Reactive FASTMAP can also be implemented directly using the CDB subscribers, but nano services offer a more streamlined and robust approach for staged provisioning.

The services discussed previously in this section were modeled to give all required parameters to the service instance. The mapping logic code could immediately do its work. Sometimes this is not possible. Two examples that require staged provisioning where a nano service step executing an action is the best practice solution:

* Allocating a resource from an external system, such as an IP address, or generating an authentication key file using an external command. It is impossible to do this allocation from within the normal FASTMAP `create()` code since there is no way to deallocate the resource on commit, abort, or failure and when deleting the service. Furthermore, the `create()` code runs within the transaction lock. The time spent in services `create()` code should be as short as possible.
* The service requires the start of one or more Virtual Machines, Virtual Network Functions. The VMs do not yet exist, and the `create()` code needs to trigger something that starts the VMs, and then later, when the VMs are operational, configure them.

The basic concepts of nano services are covered in detail by [Nano Services for Staged Provisioning](/guides/development/core-concepts/nano-services). The example in [examples.ncs/getting-started/netsim-sshkey](https://github.com/NSO-developer/nso-examples/tree/6.7/getting-started/netsim-sshkey) implements SSH public key authentication setup using a nano service. The nano service uses the following steps in a plan that produces the `generated`, `distributed`, and `configured` states:

1. Generates the NSO SSH client authentication key files using the OpenSSH `ssh-keygen` utility from a nano service side-effect action implemented in Python.
2. Distributes the public key to the netsim (ConfD) network elements to be stored as an authorized key using a Python service `create()` callback.
3. Configures NSO to use the public key for authentication with the netsim network elements using a Python service `create()` callback and service template.
4. Test the connection using the public key through a nano service side-effect executed by the NSO built-in **connect** action.

Upon deletion of the service instance, NSO restores the configuration. The only delete step in the plan is the `generated` state side-effect action that deletes the key files. The example is described in more detail in [Developing and Deploying a Nano Service](/guides/administration/installation-and-deployment/development-to-production-deployment/develop-and-deploy-a-nano-service).

The `basic-vrouter`, `netsim-vrouter`, and `mpls-vpn-vrouter` examples in the [examples.ncs/nano-services](https://github.com/NSO-developer/nso-examples/tree/6.7/nano-services) directory start, configure, and stop virtual devices. In addition, the `mpls-vpn-vrouter` example manages Layer3 VPNs in a service provider MPLS network consisting of physical and virtual devices. Using a Network Function Virtualization (NFV) setup, the L3VPN nano service instructs a VM manager nano service to start a virtual device in a multi-step process consisting of the following:

1. When the L3VPN nano service `pe-create` state step create or delete a `/vm-manager/start` service configuration instance, the VM manager nano service instructs a VNF-M, called ESC, to start or stop the virtual device.
2. Wait for the ESC to start or stop the virtual device by monitoring and handling events. Here NETCONF notifications.
3. Mount the device in the NSO device tree.
4. Fetch the ssh-keys and perform a `sync-from` on the newly created device.

See the `mpls-vpn-vrouter` example for details on how the `l3vpn.yang` YANG model `l3vpn-plan` `pe-created` state and `vm-manager.yang` `vm-plan` for more information. `vm-manager` plan states with a nano-callback have their callbacks implemented by the `escstart.java` `escstart` class. Nano services are documented in [Nano Services for Staged Provisioning](/guides/development/core-concepts/nano-services).

## Service Troubleshooting <a href="#ncs.development.services.tshoot" id="ncs.development.services.tshoot"></a>

Service troubleshooting is an inevitable part of any NSO development process and eventually a part of their operational tasks as well. By their nature, NSO services are composed primarily out of user-defined code, models, and templates. This gives you plenty of opportunities to make unintended mistakes in mapping code, use incorrect indentations, create invalid configuration templates, and much more. Not only that, they also rely on southbound communication with devices of many different versions and vendors, which presents you with yet another domain that can cause issues in your NSO services.

This is why it is important to have a systematic approach when debugging and troubleshooting your services:

* **Understand the problem** - First, you need to make sure that you fully understand the issue you are trying to troubleshoot. Why is this issue happening? When did it first occur? Does it happen only on specific deployments or devices? What is the error message like? Is it consistent and can it be replicated? What do the logs say?
* **Identify the root cause** - When you understand the issues, their triggers, conditions, and any additional insights that NSO allows you to inspect, you can start breaking down the problem to identify its root cause.
* **Form and implement the solution** - Once the root cause (or several of them) is found, you can focus on producing a suitable solution. This might be a simple NSO operation, modification of service package codebase, a change in southbound connectivity of managed devices, and any other action or combination required to achieve a working service.

### Common Troubleshooting Steps <a href="#d5e2130" id="d5e2130"></a>

You can use these general steps to give you a high-level idea of how to approach troubleshooting your NSO services:

1. Ensure that your NSO instance is installed and running properly. You can verify the overall status with `ncs --status` shell command. To find out more about installation problems and potential runtime issues, check [Troubleshooting](https://nso-docs.cisco.com/guides/development/core-concepts/pages/hgWUBFw1TA0R6WyLxOgc#ug.sys_mgmt.tshoot) in Administration.\
   \
   If you encounter a blank CLI when you connect to NSO you must also make sure that your user is added to the correct NACM group (for example `ncsadmin`) and that the rules for this group allow the user to view and edit your service through CLI. You can find out more about groups and authorization rules in [AAA Infrastructure](/guides/administration/management/aaa-infrastructure) in Administration.
2. Verify that you are using the latest version of your packages. This means copying the latest packages into load path, recompiling the package YANG models and code with the `make` command, and reloading the packages. In the end, you must expect the NSO packages to be successfully reloaded to proceed with troubleshooting. You can read more about loading packages in [Loading Packages](/guides/development/advanced-development/developing-packages#loading-packages). If nothing else, successfully reloading packages will at least make sure that you can use and try to create service instances through NSO.\
   \
   Compiling packages uses the `ncsc` compiler internally, which means that this part of the process reveals any syntax errors that might exist in YANG models or Java code. You do not need to rely on `ncsc` for compile-level errors though and should use specialized tools such as `yanger` for YANG, and one of the many IDEs and syntax validation tools for Java.

   ```
   yang/demo.yang:32: error: expected keyword 'type' as substatement to 'leaf'
   make: *** [Makefile:41: ../load-dir/demo.fxs] Error 1
   ```

   ```
       [javac] /nso-run/packages/demo/src/java/src/com/example/demo/demoRFS.java:52: error: ';' expected
       [javac]         Template myTemplate = new Template(context, "demo-template")
       [javac]                                                                          ^
       [javac] 1 error
       [javac] 1 warning

   BUILD FAILED
   ```

   \
   Additionally, reloading packages can also supply you with some valuable information. For example, it can tell you that the package requires a higher version of NSO which is specified in the `package-meta-data.xml` file, or about any Python-related syntax errors.

   ```bash
   admin@ncs# packages reload
   Error: Failed to load NCS package: demo; requires NCS version 6.3
   ```

   ```cli
   admin@ncs# packages reload
   reload-result {
       package demo
       result false
       info SyntaxError: invalid syntax
   }
   ```

   \
   Last but not least, package reloading also provides some information on the validity of your XML configuration templates based on the NED namespace you are using for a specific part of the configuration, or just general syntactic errors in your template.

   ```bash
   admin@ncs# packages reload
   reload-result {
       package demo1
       result false
       info demo-template.xml:87 missing tag: name
   }
   reload-result {
       package demo2
       result false
       info demo-template.xml:11 Unknown namespace: 'ios-xr'
   }
   reload-result {
       package demo3
       result false
       info demo-template.xml:12: The XML stream is broken. Run-away < character found.
   }
   ```
3. Examine what the template and XPath expressions evaluate to. If some service instance parameters are missing or are mapped incorrectly, there might be an error in the service template parameter mapping or in their XPath expressions. Use the CLI pipe command `debug template` to show all the XPath expression results from your service configuration templates or `debug xpath` to output all XPath expression results for the current transaction (e.g., as a part of the YANG model as well).

   \
   In addition, you can use the `xpath eval` command in CLI configuration mode to test and evaluate arbitrary XPath expressions. The same can be done with `ncs_cmd` from the command shell. To see all the XPath expression evaluations in your system, you can also enable and inspect the `xpath.trace` log. You can read more about debugging templates and XPath in [Debugging Templates](/guides/development/core-concepts/templates#debugging-templates). If you are using multiple versions of the same NED, make sure that you are using the correct processing instructions as described in [Namespaces and Multi-NED Support](https://nso-docs.cisco.com/guides/development/core-concepts/pages/Whkl286OHBVy52qnpzLh#ch_templates.multined) when applying different bits of configuration to different versions of devices.

   ```bash
   admin@ncs# devtools true
   admin@ncs# config
   Entering configuration mode terminal
   admin@ncs(config)# xpath eval /devices/device
   admin@ncs(config)# xpath eval /devices/device[name='r0']
   ```
4. Validate that your custom service code is performing as intended. Depending on your programming language of choice, there might be different options to do that. If you are using Java, you can find out more on how to configure logging for the internal Java VM Log4j in [Logging](/guides/development/core-concepts/nso-virtual-machines/nso-java-vm#logging). You can use a debugger as well, to see the service code execution line by line. To learn how to use Eclipse IDE to debug Java package code, read [Using Eclipse to Debug the Package Java Code](https://nso-docs.cisco.com/guides/development/core-concepts/pages/Yb0PdnXkYaOCtAjNU3ch#ug.package_dev.java_debugger). The same is true for Python. NSO uses the standard `logging` module for logging, which can be configured as per instructions in [Debugging of Python Packages](/guides/development/core-concepts/nso-virtual-machines/nso-python-vm#debugging-of-python-packages). Python debugger can be set up as well with `debugpy` or `pydevd-pycharm` modules.
5. Inspect NSO logs for hints. NSO features extensive logging functionality for different components, where you can see everything from user interactions with the system to low-level communications with managed devices. For best results, set the logging level to DEBUG or lower. To learn what types of logs there are and how to enable them, consult [Logging](https://nso-docs.cisco.com/guides/development/core-concepts/pages/hgWUBFw1TA0R6WyLxOgc#ug.ncs_sys_mgmt.logging) in Administration.

   \
   Another useful option is to append `| details` to your service commits, which shows the relevant [Progress Trace](/guides/development/advanced-development/progress-trace).

   ```cli
   admin@ncs(config)# commit | details
   ... lots of logs ...
   Commit complete.
   ```

   \
   Alternatively, you can append a custom label to your service commits. The label can be used to correlate events from a commit in logs and notifications during focused troubleshooting sessions. It is provided as a [commit parameter](/guides/operation-and-usage/operations/lifecycle-operations#d5e5048), e.g. via CLI:

   ```cli
   admin@ncs(config)# commit label myServiceCommit
   Commit complete.
   ```
6. Measuring the time it takes for specific commands to complete can also give you some hints about what is going on. You can do this by using the `timecmd`, which requires the dev tools to be enabled.

   ```bash
   admin@ncs# devtools true
   admin@ncs(config)# timecmd commit
   Commit complete.
   Command executed in 5.31 seconds.
   ```

   \
   Another useful tool to examine how long a specific event or command takes is the progress trace. See how it is used in [Progress Trace](/guides/development/advanced-development/progress-trace).
7. Double-check your service points in the model, templates, and in code. Since configuration templates don't get applied if the servicepoint attribute doesn't match the one defined in the service model or are not applied from the callbacks registered to specific service points, make sure they match and that they are not missing. Otherwise, you might notice errors such as the following ones.

   ```bash
   admin@ncs# packages reload
   reload-result {
       package demo
       result false
       info demo-template.xml:2 Unknown servicepoint: notdemo
   }
   ```

   ```cli
   admin@ncs(config-demo-s1)# commit dry-run
   Aborted: no registration found for callpoint demo/service_create of type=external
   ```
8. Verify YANG imports and namespaces. If your service depends on NED or other YANG files, make sure their path is added to where the compiler can find them. If you are using the standard service package skeleton, you can add to that path by editing your service package `Makefile` and adding the following line.

   ```
   YANGPATH += ../../my-dependency/src/yang \
   ```

   \
   Likewise, when you use data types from other YANG namespaces in either your service model definition or by referencing them in XPath expressions.

   ```
   // Following XPath might trigger an error if there is collision for the 'interfaces' node with other modules
   path "/ncs:devices/ncs:device['r0']/config/interfaces/interface";
   yang/demo.yang:25: error: the node 'interfaces' from module 'demo' (in node 'config' from 'tailf-ncs') is not found

   // And the following XPath will not, since it uses namespace prefixes
   path "/ncs:devices/ncs:device['r0']/config/iosxr:interfaces/iosxr:interface";
   ```
9. Trace the southbound communication. If the service instance creation results in a different configuration than would be expected from the NSO point of view, especially with custom NED packages, you can try enabling the southbound tracing (either per device or globally).

   ```bash
   admin@ncs(config)# devices global-settings trace pretty
   admin@ncs(config)# devices global-settings trace-dir ./my-trace
   admin@ncs(config)# commit
   ```

***

**Next Steps**

{% content-ref url="/pages/Bw8TviXCSEsM9XjXBC1d" %}
[Services Deep Dive](/guides/development/advanced-development/developing-services/services-deep-dive)
{% endcontent-ref %}


# Actions

Automate non-configuration tasks with NSO.

The most common way to implement non-configuration automation in NSO is using actions. An action represents a task or an operation that a user of the system can invoke on demand, such as downloading a file, resetting a device, or performing a test.

Like configuration elements, actions must also be defined in the YANG model. Each action is described by the `action` YANG statement that specifies what are its inputs and outputs, if any. Inputs allow a user of the action to provide additional information to the action invocation, while outputs provide information to the caller. Actions are a form of a Remote Procedure Call (RPC) and have historically evolved from NETCONF RPCs. It's therefore unsurprising that with NSO you implement both in a similar manner.

Let's look at an example action definition:

```
action my-test {
  tailf:actionpoint my-test-action;
  input {
    leaf test-string {
      type string;
    }
  }
  output {
    leaf has-nso {
      type boolean;
    }
  }
}
```

The first thing to notice in the code is that, just like services use a service point, actions use an `actionpoint`. It is denoted by the `tailf:actionpoint` statement and tells NSO to execute a callback registered to this name. The callback mechanism allows you to provide custom action implementation.

Correspondingly, your code needs to register a callback to this action point, by calling the `register_action()`, as demonstrated here:

```python
def setup(self):
    self.register_action('my-test-action', MyTestAction)
```

The `MyTestAction` class, referenced in the call, is responsible for implementing the actual action logic and should inherit from the `ncs.dp.Action` base class. The base class will take care of calling the `cb_action()` class method when users initiate the action. The `cb_action()` is where you put your own code. The following code shows a trivial implementation of an action, that checks whether its input contains the string “`NSO`”:

```python
class MyTestAction(Action):
    @Action.action
    def cb_action(self, uinfo, name, kp, input, output, trans):
        self.log.info('Action invoked: ', name)
        output.has_nso = 'NSO' in input.test_string
```

The `input` and `output` arguments contain input and output data, respectively, which matches the definition in the action YANG model. The example shows the value of a simple Python `in` string check that is assigned to an output value.

The `name` argument has the name of the called action (such as `my-test`), to help you distinguish which action was called in the case where you would register the same class for multiple actions. Similarly, an action may be defined on a list item and the `kp` argument contains the full keypath (a tuple) to an instance where it was called.

Finally, the `uinfo` contains information on the user invoking the action and the `trans` argument represents a transaction, that you can use to access data other than input. This transaction is read-only, as configuration changes should normally be done through services instead. Still, the action may need some data from NSO, such as an IP address of a device, which you can access by using `trans` with the `ncs.maagic.get_root()` function and navigate to the relevant information.

{% hint style="info" %}
If, for any reason, your action requires a new, read-write transaction, please also read through [NSO Concurrency Model](/guides/development/core-concepts/nso-concurrency-model) to learn about the possible pitfalls.
{% endhint %}

Further details and the format of the arguments can be found in the NSO Python API reference.

The last thing to note in the above action code definition is the use of the decorator `@Action.action`. Its purpose is to set up the function arguments correctly, so variables such as `input` and `output` behave like other Python Maagic objects. This is no different from services, where decorators are required for the same reason.

## Showcase - Implementing Device Count Action <a href="#d5e1000" id="d5e1000"></a>

{% hint style="info" %}
See [examples.ncs/getting-started/applications-nso](https://github.com/NSO-developer/nso-examples/tree/6.7/getting-started/applications-nso) for an example implementation.
{% endhint %}

### Prerequisites

* No previous NSO or netsim processes are running. Use the `ncs --stop` and `ncs-netsim stop` commands to stop them if necessary.
* NSO local install with a fresh runtime directory has been created by the `ncs-setup --dest ~/nso-lab-rundir` or similar command.
* The environment variable `NSO_RUNDIR` points to this runtime directory, such as set by the `export NSO_RUNDIR=~/nso-lab-rundir` command. It enables the below commands to work as-is, without additional substitution needed.

### Step 1 - Create a New Python Package <a href="#d5e1015" id="d5e1015"></a>

One of the most common uses of NSO actions is automating network and service tests but they are also a good choice for any other non-configuration task. Being able to quickly answer questions, such as how many network ports are available (unused) or how many devices currently reside in a given subnet, can greatly simplify the network planning process. Coding these computations as actions in NSO makes them accessible on-demand to a wider audience.

For this scenario, you will create a new package for the action, however actions can also be placed into existing packages. A common example is adding a self-test action to a service package.

First, navigate to the `packages` subdirectory:

```bash
$ cd $NSO_RUNDIR/packages
```

Create a package skeleton with the `ncs-make-package` command and the `--action-example` option. Name the package `count-devices`, like so:

```bash
$ ncs-make-package --service-skeleton python --action-example count-devices
```

This command creates a YANG module file, where you will place a custom action definition. In a text or code editor open the `count-devices.yang` file, located inside `count-devices/src/yang/`. This file already contains an example action which you will remove. Find the following line (after module imports):

```
  description
```

Delete this line and all the lines following it, to the very end of the file. The file should now resemble the following:

```yang
module count-devices {

  namespace "http://example.com/count-devices";
  prefix count-devices;

  import ietf-inet-types {
    prefix inet;
  }
  import tailf-common {
    prefix tailf;
  }
  import tailf-ncs {
    prefix ncs;
  }
```

### Step 2 - Define a New Action in YANG <a href="#d5e1035" id="d5e1035"></a>

To model an action, you can use the `action` YANG statement. It is part of the YANG standard from version 1.1 onward, requiring you to also define `yang-version 1.1` in the YANG model. So, add the following line at the start of the module, right before `namespace` statement:

```
  yang-version 1.1;
```

Note that in YANG version 1.0, actions used the NSO-specific `tailf:action` extension, which you may still find in some YANG models.

Now, go to the end of the file and add a `custom-actions` container with the `count-devices` action, using the `count-devices-action` action point. The input is an IP subnet and the output is the number of devices managed by NSO in this subnet.

```yang
  container custom-actions {
    action count-devices {
      tailf:actionpoint count-devices-action;
      input {
        leaf in-subnet {
          type inet:ipv4-prefix;
        }
      }
      output {
        leaf result {
          type uint16;
        }
      }
    }
  }
```

Also, add the closing bracket for the module at the end:

```
}
```

Remember to finally save the file, which should now be similar to the following:

```yang
module count-devices {

  yang-version 1.1;
  namespace "http://example.com/count-devices";
  prefix count-devices;

  import ietf-inet-types {
    prefix inet;
  }
  import tailf-common {
    prefix tailf;
  }
  import tailf-ncs {
    prefix ncs;
  }

  container custom-actions {
    action count-devices {
      tailf:actionpoint count-devices-action;
      input {
        leaf in-subnet {
          type inet:ipv4-prefix;
        }
      }
      output {
        leaf result {
          type uint16;
        }
      }
    }
  }
}
```

### Step 3 - Implement the Action Logic <a href="#d5e1053" id="d5e1053"></a>

The action code is implemented in a dedicated class, that you will put in a separate file. Using an editor, create a new, empty file `count_devices_action.py` in the `count-devices/python/count_devices/` subdirectory.

At the start of the file, import the packages that you will need later on and define the action class with the `cb_action()` method:

```python
from ipaddress import IPv4Address, IPv4Network
import socket
import ncs
from ncs.dp import Action

class CountDevicesAction(Action):
    @Action.action
    def cb_action(self, uinfo, name, kp, input, output, trans):
```

Then initialize the `count` variable to `0` and construct a reference to the NSO data root, since it is not part of the method arguments:

```
        count = 0
        root = ncs.maagic.get_root(trans)
```

Using the `root` variable, you can iterate through the devices managed by NSO and find their (IPv4) address:

```
        for device in root.devices.device:
            address = socket.gethostbyname(device.address)
```

If the IP address comes from the specified subnet, increment the count:

```
            if IPv4Address(address) in IPv4Network(input.in_subnet):
                count = count + 1
```

Lastly, assign the count to the result:

```
        output.result = count
```

### Step 4 - Register Callback <a href="#d5e1071" id="d5e1071"></a>

Your custom Python code is ready; however, you still need to link it to the `count-devices` action. Open the `main.py` from the same directory in a text or code editor and delete all the content already in there.

Next, create a class called `Main` that inherits from the `ncs.application.Application` base class. Add a single class method `setup()` that takes no additional arguments.

```python
import ncs

class Main(ncs.application.Application):
    def setup(self):
```

Inside the `setup()` method call the `register_action()` as follows:

```python
        self.register_action('count-devices-action', CountDevicesAction)
```

This line instructs NSO to use the `CountDevicesAction` class to handle invocations of the `count-devices-action` action point. Also, import the `CountDevicesAction` class from the `count_devices_action` module.

The complete `main.py` file should then be similar to the following:

```python
import ncs
from count_devices_action import CountDevicesAction

class Main(ncs.application.Application):
    def setup(self):
        self.register_action('count-devices-action', CountDevicesAction)
```

### Step 5 - And... Action! <a href="#d5e1093" id="d5e1093"></a>

With all of the code ready, you are one step away from testing the new action, but to do that, you will need to add some devices to NSO. So, first, add a couple of simulated routers to the NSO instance:

```bash
$ cd $NCS_DIR/examples.ncs/device-management/router-network
```

```bash
$ make all
$ cp ncs-cdb/ncs_init.xml $NSO_RUNDIR/ncs-cdb/
```

```bash
$ cp -a packages/router $NSO_RUNDIR/packages/
```

Before the packages can be loaded, you must compile them:

```bash
$ cd $NSO_RUNDIR
```

```bash
$ make -C packages/router/src && make -C packages/count-devices/src
make: Entering directory 'packages/router/src'
< ... output omitted ... >
make: Leaving directory 'packages/router/src'
make: Entering directory 'packages/count-devices/src'
mkdir -p ../load-dir
mkdir -p java/src//
bin/ncsc  `ls count-devices-ann.yang  > /dev/null 2>&1 && echo "-a count-devices-ann.yang"` \
              -c -o ../load-dir/count-devices.fxs yang/count-devices.yang
make: Leaving directory 'packages/count-devices/src'
```

You can start the NSO now and connect to the CLI:

```bash
$ ncs --with-package-reload && ncs_cli -C -u admin
```

Finally, invoke the action:

```bash
$ admin@ncs# custom-actions count-devices in-subnet 127.0.0.0/16
result 3
```

You can use the `show devices list` command to verify that the result is correct. You can alter the address of any device and see how it affects the result. You can even use a hostname, such as `localhost`.

{% hint style="info" %}
Other examples of action implementations can be found under [examples.ncs/sdk-api](https://github.com/NSO-developer/nso-examples/tree/6.7/sdk-api).
{% endhint %}


# Templates

Simplify change management in your network using templates.

NSO comes with a flexible and powerful built-in templating engine, which is based on XML. The templating system simplifies how you apply configuration changes across devices of different types and provides additional validation against the target data model. Templates are a convenient, declarative way of updating structured configuration data and allow you to avoid lots of boilerplate code.

You will most often find this type of configuration templates used in services, which is why they are sometimes also called service templates. However, we mostly refer to them simply as XML templates, since they are defined in XML files.

NSO loads templates as part of a package, looking for XML files in the `templates` directory and its subdirectories. You then apply an XML template through API or by connecting it with a service through a service point, allowing NSO to use it whenever a service instance needs updating.

{% hint style="info" %}
XML templates are distinct from so-called “device templates”, which are dynamically created and applied as needed by the operator, for example in the CLI. There are also other types of templates in NSO, unrelated to XML templates described here.
{% endhint %}

## Structure of a Template <a href="#ch_templates.structure" id="ch_templates.structure"></a>

Template is an XML file with the `config-template` root element, residing in the `http://tail-f.com/ns/config/1.0` namespace. The root contains configuration elements according to NSO YANG schema and XML processing instructions.

Configuration element structure is very much like the one you would find in a NETCONF message since it uses the same encoding rules defined by YANG. Additionally, each element can specify a `tags` attribute that refines how the configuration is applied.

A typical template for configuring an NSO-managed device is:

```xml
<config-template xmlns="http://tail-f.com/ns/config/1.0">
  <devices xmlns="http://tail-f.com/ns/ncs">
    <device tags="nocreate">
      <name>{/name}</name>
      <config tags="merge">
        <!-- ... -->
      </config>
    </device>
  </devices>
</config-template>
```

The first line defines the root node. It contains elements that follow the same structure as that used by the CDB, in particular, the `devices device <name> config` path in the CLI. In the printout, two elements, `device` and `config`, also have a `tags` attribute.

You can write this structure by studying the YANG schema if you wish. However, a more typical approach is to start with manipulating NSO configuration by hand, such as through the NSO CLI or web UI. Then, generate the XML structure with the help of NSO output filters, using the `show ... | display xml-template` and similar commands. You can also reuse the existing configuration, such as the one loaded with the `ncs_load` utility. For a worked, step-by-step example, refer to the section [A Template is All You Need](https://nso-docs.cisco.com/guides/development/core-concepts/pages/1ErWKzeM15TwAei4yXWm#ch_services.just_template).

```bash
admin@ncs(config)# devices device rtr01 config ...
admin@ncs(config-device-rtr01)# show configuration | display xml-template
<config-template xmlns="http://tail-f.com/ns/config/1.0">
  <devices xmlns="http://tail-f.com/ns/ncs">
    <device>
      <name>rtr01</name>
      <config>
        <!-- ... -->
      </config>
    </device>
  </devices>
</config-template>
admin@ncs(config-device-rtr01)# commit
admin@ncs# show running-config devices device rtr01 config ... | display xml-template
<config-template xmlns="http://tail-f.com/ns/config/1.0">
  <devices xmlns="http://tail-f.com/ns/ncs">
    <device>
      <name>rtr01</name>
      <config>
        <!-- ... -->
      </config>
    </device>
  </devices>
</config-template>
```

Having the basic structure in place, you can then fine-tune the template by adding different processing instructions and tags, as well as replacing static values with variable references using the XPath syntax.

Note that a single template can configure multiple devices of different type, services, or any other configurable data in NSO; basically the same as you can do in a CLI commit. But a single, gigantic template can become a burden to maintain. That is why many developers prefer to split up bigger configurations into multiple feature templates, either by functionality or by device type.

Finally, every XML template has a name. The name of the template is the file path relative to the `templates` directory of the package, without the `.xml` extension. The name allows you to reference the template from the code later on. In case multiple packages define a template with the same path, you disambiguate between them by prepending *`<package name>`*`:` to the name. (Note that any colon or backslash characters in the package name or the file path must be backslash escaped.)

## Generating a Template From Configuration <a href="#ch_templates.templatize" id="ch_templates.templatize"></a>

To simplify template creation, NSO features the `/services/create-template` action that can find common structural patterns in a set of device configurations and create a configuration template and the corresponding service YANG model based on it.

In addition to extracting patterns from configuration already present in NSO, the action can also consume configuration snippets directly. Snippets can be supplied either from a file on the NSO server filesystem or as inline payload data. Supported formats are NETCONF-style XML wrapped in a `<config>` element, Cisco XR style CLI (`cli-c`), Juniper curly-brace CLI (`cli-j`), and Juniper set commands (`cli-j-cmd`). Delete operations in the input, such as Cisco-style `no` commands or XML `operation="remove"` attributes, are translated into `delete` tags in the generated service template.

The algorithm works by traversing the data depth-first, keeping track of the rate of occurrence of configuration nodes, and any values that compare equal. Values that do not compare equal are parameterized and service input parameters are created for these paths in the YANG model. For example:

{% code overflow="wrap" %}

```bash
admin@ncs# services create-template name policy-map-srv path [ /devices/device[device-type/cli/ned-id='cisco-ios-cli-3.0:cisco-ios-cli-3.0']/config/policy-map ] include-doc
template <config-template xmlns="http://tail-f.com/ns/config/1.0"
                           servicepoint="policy-map-srv">
            <devices xmlns="http://tail-f.com/ns/ncs">
              <device tags="nocreate">
                <name>{/device}</name>
                <config>
                  <policy-map xmlns="urn:ios"
                              tags="merge"
                              foreach="{/policy-map}">
                    <name>{name}</name>
                    <class foreach="{class}">
                      <name>{name}</name>
                      <drop/>
                      <estimate>
                        <bandwidth>
                          <delay-one-in>
                            <doi>500</doi>
                            <milliseconds>100</milliseconds>
                          </delay-one-in>
                        </bandwidth>
                      </estimate>
                      <priority>
                        <percent>33</percent>
                      </priority>
                    </class>
                  </policy-map>
                </config>
              </device>
            </devices>
          </config-template>

yang-module module policy-map-srv {
  yang-version 1.1;
  namespace "http://com/example/policy-map-srv";
  prefix policy-map-srv;

  import tailf-ncs {
    prefix ncs;
  }
  import tailf-common {
    prefix tailf;
  }

  list policy-map-srv {
    key name;

    uses ncs:service-data;
    ncs:servicepoint policy-map-srv;

    leaf name {
      type string;
    }

    leaf-list device {
      type leafref {
        path "/ncs:devices/ncs:device/ncs:name";
      }
    }

    list policy-map {
      key "name";
      description
        "Configure QoS Policy Map";
      leaf name {
        type string;
      }
      list class {
        key "name";
        description
          "policy criteria";
        leaf name {
          type union {
            type string;
            type enumeration {
              enum class-default {
                description
                  "System default class matching otherwise unclassified packet";
              }
            }
          }
        }
      }
    }
  }
}
```

{% endcode %}

The action takes a number of arguments to control how the resulting template looks:

* `name` - The name of the new service.
* `path` - A list of XPath 1.0 expressions pointing into `/devices/device/config` to create the template from. The template is only created from the paths that are common in the node-set.
* `match-rate` - Device configuration is included in the resulting template based on the rate of occurrence given by this setting. By giving different rates the user can decide how often configuration needs to occur for it to be included in the template.
* `exclude-service-config` - Exclude configuration that is already under service management. This is useful when the intention is to detect common configuration that can be turned into a service.
* `make-package` - Create a service package including the generated template and YANG module. The package is created in the parent directory specified by `in-directory`, but is not built. The package needs to be built separately by running `make` in its `src/` subdirectory. The user has the freedom of making modifications to the generated files.
* `augment` - An XPath 1.0 location path to be included as an augment statement in the generated YANG module.
* `include-doc` - Include descriptions derived from device schema in the generated YANG module.
* `import-user-modules` - Import device YANG modules and their defined types in the generated YANG module.
* `collapse-list-keys` - Decides what lists to parameterize, either `all`, `automatic` (default), or those specified by the `list-path` parameter. The default is to find lists that differ among the device configurations.

The [examples.ncs/service-management/implement-a-service/dns-v3](https://github.com/NSO-developer/nso-examples/tree/6.7/service-management/implement-a-service/dns-v3) environment can be used to try the command.

{% code overflow="wrap" %}

```bash
$ cd $NCS_DIR/examples.ncs/service-management/implement-a-service/dns-v3
$ make demo
admin@ncs# services create-template name policy-map-srv path [ /devices/device[device-type/cli/ned-id='cisco-ios-cli-3.0:cisco-ios-cli-3.0']/config ]
```

{% endcode %}

## Generating the XML Template Structure <a href="#ch_templates.templatize" id="ch_templates.templatize"></a>

`/services/create-template` requires you to reference existing configurations in NSO. If such configuration is not readily available to you and you want to avoid manually creating sample configuration in NSO first, you can use the `sample-xml-skeleton` functionality of the **yanger** utility to generate sample XML data directly:

```bash
$ cd $NCS_DIR/packages/neds/cisco-ios-cli-3.8/
$ yanger -f sample-xml-skeleton \
    --sample-xml-skeleton-doctype=config \
    --sample-xml-skeleton-path='/ip/name-server' \
    --sample-xml-skeleton-defaults \
    src/yang/tailf-ned-cisco-ios.yang
<?xml version='1.0' encoding='UTF-8'?>
<config xmlns="urn:ietf:params:xml:ns:netconf:base:1.0">
  <ip xmlns="urn:ios">
    <name-server>
      <name-server-list>
        <address/>
      </name-server-list>
      <vrf>
        <name/>
        <name-server-list>
          <address/>
        </name-server-list>
      </vrf>
    </name-server>
  </ip>
</config>
```

You can replace the value of *`--sample-xml-skeleton-path`* with the path to the part of the configuration you want to generate.

In case the target data model contains submodules, or references other non-built-in modules, you must also tell `yanger` where to find additional modules with the *`-p`* parameter, such as adding `-p src/yang/` to the invocation.

## Values in a Template <a href="#ch_templates.values" id="ch_templates.values"></a>

Some XML elements, notably those that represent leafs or leaf-lists, specify element text content as values that you wish to configure, such as:

```xml
      <name>rtr01</name>
```

NSO converts the string value to the actual value type of the YANG model automatically when the template is applied.

Along with hard-coded, static content (`rtr01`), the value may also contain curly brackets (`{...}`), which the templating engine treats as XPath 1.0 expressions.

The simplest form of an XPath expression is a plain XPath variable:

```xml
      <name>{$CE}</name>
```

A value can contain any number of `{...}` expressions and strings. The end result is the concatenation of all the strings and XPath expressions. For example, `<description>Link to PE: {$PE} - {$PE_INT_NAME}</description>` might evaluate to `<description>Link to PE: pe0 - GigabitEthernet0/0/0/3</description>` if you set `PE` to `pe0` and `PE_INT_NAME` to `GigabitEthernet0/0/0/3` when applying the template.

You set the values for variables in the code where you apply the template. However, if you set the value to an empty string, the corresponding statement is ignored (in this case you may use the XPath function `string()` to set a node to the actual empty string).

NSO also sets some predefined variables, which you can reference:

* `$DEVICE`: The name of the current device. Cannot be overridden.
* `$TEMPLATE_NAME`: The name of the current template. Cannot be overridden.
* `$SCHEMA_OPAQUE`: Defined if the template is registered for a servicepoint (the top node in the template has `servicepoint` attribute) and the corresponding `ncs:servicepoint` statement in the YANG model has `tailf:opaque` substatement. Set to the value of the `tailf:opaque` statement.
* `$OPERATION`: Defined if the template is registered for a servicepoint with the `cbtype` attribute set to `pre-/post-modification` (see [Service Callpoints and Templates](#ch_templates.servicepoint)). Contains the requested service operation; create, update, or delete.

The `{...}` expression can also be any other valid XPath 1.0 expression. To address a reachable node, you might for example use:

```
/endpoint/ce/device
```

Or to select a leaf node, `device`:

```
../ce/device
```

NSO then uses the value of this leaf, say `ce5`, when constructing the value of the expression.

However, there are some special cases. If the result of the expression is a node-set (e.g. multiple leafs), and the target is a leaf list or a list's key leaf, the template configures multiple destination nodes. This handling allows you to set multiple values for a leaf list or set multiple list items.

Similarly, if the result is an empty node set, nothing is set (the set operation is ignored).

Finally, what nodes are reachable in the XPath expression, and how, depends on the root node and context used in the template. See [XPath Context in Templates](#ch_templates.contexts).

## Conditional Statements <a href="#ch_templates.conditionals" id="ch_templates.conditionals"></a>

The `if`, and the accompanying `elif`, `else`, processing instructions make it possible to apply parts of the template, based on a condition. For example:

```xml
<policy-map xmlns="urn:ios" tags="merge">
  <name>{$POLICY_NAME}</name>
  <class>
    <name>{$CLASS_NAME}</name>
    <?if {qos-class/priority = 'realtime'}?>
      <priority-realtime>
        <percent>{$CLASS_BW}</percent>
      </priority-realtime>
    <?elif {qos-class/priority = 'critical'}?>
      <priority-critical>
        <percent>{$CLASS_BW}</percent>
      </priority-critical>
    <?else?>
      <bandwidth>
        <percent>{$CLASS_BW}</percent>
      </bandwidth>
    <?end?>
    <set>
      <ip>
        <dscp>{$CLASS_DSCP}</dscp>
      </ip>
    </set>
  </class>
</policy-map>
```

The preceding template shows how to produce different configuration, for network bandwidth management in this case, when different `qos-class/priority` values are specified.

In particular, the sub-tree containing the `priority-realtime` tag will only be evaluated if `qos-class/priority` in the `if` processing instruction evaluates to the string `'realtime'`.

The subtree under the `elif` processing instruction will be executed if the preceding `if` expression evaluated to `false`, i.e. `qos-class/priority` is not equal to the string `'realtime'`, but '`critical'` instead.

The subtree under the `else` processing instruction will be executed when both the preceding `if` and `elif` expressions evaluated to `false`, i.e. `qos-class/priority` is not `'realtime'` nor `'critical'`.

In your own code you can of course use just a subset of these instructions, such as a simple `if` - `end` conditional evaluation. But note that every conditional evaluation must end with the `end` processing instruction, to allow nesting multiple conditionals.

The evaluation of the XPath statements used in the `if` and `elif` processing instructions follow the XPath standard for computing boolean values. In summary, the conditional expression will evaluate to false when:

* The argument evaluates to an empty node-set.
* The value of the argument is either an empty string or numeric zero.
* The argument is of boolean type and evaluates to false, such as using the `not(true())` function.

## Loop Statements <a href="#ch_templates.loops" id="ch_templates.loops"></a>

The `foreach` and `for` processing instructions allow you to avoid needless repetition: they iterate over a set of values and apply statements in a sub-tree several times. For example:

```xml
<ip xmlns="urn:ios">
  <route>
  <?foreach {/tunnel}?>
    <ip-route-forwarding-list>
      <prefix>{network}</prefix>
      <mask>{netmask}</mask>
      <forwarding-address>{tunnel-endpoint}</forwarding-address>
    </ip-route-forwarding-list>
  <?end?>
  </route>
</ip>
```

The printout shows the use of `foreach` to configure a set of IP routes (the list `ip-route-forwarding-list`) for a Cisco network router. If there is a `tunnel` list in the service model, the `{/tunnel}` expression selects all the items from the list. If this is a non-empty set, then the sub-tree containing `ip-route-forwarding-list` is evaluated once for every item in that node set.

For each iteration, the initial context is set to one node, that is, the node being processed in that iteration. The XPath function `current()` retrieves this initial context if needed. Using the context, you can access the node data with relative XPath paths, e.g. the `{network}` code in the example refers to `/tunnel[...]/network` for the current item.

`foreach` only supports a single XPath expression as its argument and the result needs to be a node-set, not a simple value. However, you may use XPath union operator to join multiple node sets in a single expression when required: `{some-list-1 | some-leaf-list-2}`.

Similarly, `for` is a processing instruction that uses a variable to control the iteration, in line with traditional programming languages. For example, the following template disables the first four (0-3) interfaces on a Cisco router:

```xml
<interface xmlns="urn:ios">
  <?for i=0; {$i < 4}; i={$i + 1}?>
    <FastEthernet>
      <name>0/{$i}</name>
      <shutdown/>
    </FastEthernet>
  <?end?>
</interface>
```

In this example, three semicolon-separated clauses follow the `for` keyword:

* The first clause is the initial step executed before the loop is entered the first time. The format of the clause is that of a variable name followed by an equals sign and an expression. The latter may combine literal strings and XPath expressions surrounded by `{}`. The expression is evaluated in the same way as the XML tag contents in templates. This clause is optional.
* The second clause is the progress condition. The loop will execute as long as this condition evaluates to true, using the same rules as the `if` processing instruction. The format of this clause is an XPath expression surrounded by `{}`. This clause is mandatory.
* The third clause is executed after each iteration. It has the same format as the first clause (variable assignment) and is optional.

The `foreach` and `for` expressions make the loop explicit, which is why they are the first choice for most programmers. Alternatively, under certain circumstances, the template invokes an implicit loop, as described in [XPath Context in Templates](#ch_templates.contexts).

## Template Operations <a href="#ch_templates.operations" id="ch_templates.operations"></a>

The most common use-case for templates is to produce new configuration but other behavior is possible too. This is accomplished by setting the `tags` attribute on XML elements.

NSO supports the following `tags` values, colloquially referred to as “tags”:

* `merge`: Merge with a node if it exists, otherwise create the node. This is the default operation if no operation is explicitly set.

  ```xml
  <config tags="merge">
    <interface xmlns="urn:ios">
    ...
  ```
* `replace`: Replace a node if it exists, otherwise create the node.

  ```xml
      <GigabitEthernet tags="replace">
        <name>{link/interface-number}</name>
        <description tags="merge">Link to PE</description>
        ...
  ```
* `create`: Creates a node. The node must not already exist. An error is raised if the node exists.

  ```xml
      <GigabitEthernet tags="create">
        <name>{link/interface-number}</name>
        <description tags="merge">Link to PE</description>
        ...
  ```
* `nocreate`: Merge with a node if it exists. If it does not exist, it will *not* be created.

  ```xml
      <GigabitEthernet tags="nocreate">
        <name>{link/interface-number}</name>
        <description tags="merge">Link to PE</description>
        ...
  ```
* `delete`: Delete the node.

  ```xml
      <GigabitEthernet tags="delete">
        <name>{link/interface-number}</name>
        <description tags="merge">Link to PE</description>
        ...
  ```

Tags `merge` and `nocreate` are inherited to their sub-nodes until a new tag is introduced.

Tags `create` and `replace` are not inherited and only apply to the node they are specified on. Children of the nodes with `create` or `replace` tags have `merge` behavior.

Tag `delete` applies only to the current node; any children (except keys specifying the list/leaf-list entry to delete) are ignored.

Optionally, you can use the `child-tags` or the `inherit` attribute together with the `tags` attribute on XML elements to specify operation on the children nodes separately from the current node.

NSO supports the same type of values for `child-tags` as for `tags`, i.e., `merge`, `replace`, `create`, `nocreate`, `delete`. The `child-tags` attribute specifies which operation should be applied to all sub-nodes (until a new tag is introduced) regardless of what operation is being set to the current node.

The `inherit` attribute value can be either `true` or `false`. It specifies whether the operation (i.e., the `tags` value) on the current node should be inherited by its sub-nodes.

If both `child-tags` and `inherit` attributes are set, `child-tags` would take precedence over `inherit`.

Here are some examples of different combinations of `tags` with `child-tags` and/or `inherit`:

* `tags="nocreate" child-tags="merge"`: The parent node `<GigabitEthernet>` will have `nocreate` behavior while the children nodes `<name>` and `<description>` will have `merge` behavior.

  ```xml
      <GigabitEthernet tags="nocreate" child-tags="merge">
        <name>{link/interface-number}</name>
        <description>Link to PE</description>
        ...
  ```
* `tags="nocreate" inherit="false"`: The parent node `<GigabitEthernet>` will have `nocreate` behavior which is not inherited to its children nodes due to `inherit="false"`. The children nodes `<name>` and `<description>` will have the default operation `merge` since no operation is explicitly set.

  ```xml
      <GigabitEthernet tags="nocreate" inherit="false">
        <name>{link/interface-number}</name>
        <description>Link to PE</description>
        ...
  ```
* `tags="create" child-tags="nocreate"`: The parent node `<GigabitEthernet>` will have `create` behavior while its children nodes on all the sub-levels `<name>`, `<description>`, `<manually-set>`, `<presence-container>` and `<dummy>` (except `<state>`) will have `nocreate` behavior due to `child-tags="nocreate"` which affects subtree of the current node (until a new `tags="merge"` is introduced on the sub-node `state` of which it will have a new operation `merge`).

  ```xml
      <GigabitEthernet tags="create" child-tags="nocreate">
        <name>{link/interface-number}</name>
        <description>Link to PE</description>
        <manually-set>
          <presence-container>
            <dummy>123</dummy>
            <state tags="merge">enabled</state>
          </presence-container>
        </manually-set>
        ...
  ```
* `tags="replace" inherit="true"`: The parent node `<GigabitEthernet>` will have `replace` behavior which is inherited to its children nodes due to `inherit="true"`. The children nodes `<name>` and `<description>` will have the inherited operation `replace` since no operation is explicitly set. The inheritance is cascaded to children nodes on all the sub-levels (until a new tag is introduced).

  ```xml
      <GigabitEthernet tags="replace" inherit="true">
        <name>{link/interface-number}</name>
        <description>Link to PE</description>
        ...
  ```
* `tags="replace" inherit="true" child-tags="nocreate"`: The parent node `<GigabitEthernet>` will have `replace` behavior which is not inherited to its children nodes even though `inherit="true"`. This is because `child-tags="nocreate"` takes precedence over `inherit="true"`. So, the children nodes `<name>` and `<description>` will have `nocreate` behavior.

  ```xml
      <GigabitEthernet tags="replace" inherit="true" child-tags="nocreate">
        <name>{link/interface-number}</name>
        <description>Link to PE</description>
        ...
  ```

## Operations on Ordered Lists and Leaf-lists <a href="#ch_templates.order_ops" id="ch_templates.order_ops"></a>

For ordered-by-user lists and leaf lists, where item order is significant, you can use the `insert` attribute to specify where in the list, or leaf-list, the node should be inserted. You specify whether the node should be inserted first or last in the node-set, or before or after a specific instance.

For example, if you have a list of rules, such as ACLs, you may need to ensure a particular order:

```xml
<rule insert="first">
  <name>{$FIRSTRULE}</name>
</rule>
<rule insert="last">
  <name>{$LASTRULE}</name>
</rule>
<rule insert="after" value={$FIRSTRULE}>
  <name>{$SECONDRULE}</name>
</rule>
<rule insert="before" value={$LASTRULE}>
  <name>{$SECONDTOLASTRULE}</name>
</rule>
```

However, it is not uncommon that there are multiple services managing the same ordered-by user list or leaf-list. The relative order of elements inserted by these services might not matter, but there are some constraints on element positions that need to be fulfilled.

Following the ACL rules example, suppose that initially the list contains only the "deny-all" rule:

```xml
<rule>
  <name>deny-all</name>
  <ip>0.0.0.0</ip>
  <mask>0.0.0.0</mask>
  <action>deny</action>
</rule>
```

There are services that prepend permit rules to the beginning of the list using the `insert="first"` operation. If there are two services creating one entry each, say 10.0.0.0/8 and 192.168.0.0/24 respectively, then the resulting configuration looks like this:

```xml
<rule>
  <name>service-2</name>
  <ip>192.168.0.0</ip>
  <mask>255.255.255.0</mask>
  <action>permit</action>
</rule>
<rule>
  <name>service-1</name>
  <ip>10.0.0.0</ip>
  <mask>255.0.0.0</mask>
  <action>permit</action>
</rule>
<rule>
  <ip>0.0.0.0</ip>
  <mask>0.0.0.0</mask>
  <action>deny</action>
</rule>
```

Note that the rule for the second service comes first because it was configured last and inserted as the first item in the list.

If you now try to check-sync the first service (10.0.0.0/8), it will report as out-of-sync, and re-deploying it would move the 10.0.0.0/8 rule first. But what you really want is to ensure the deny-all rule comes last. This is when the `guard` attribute comes in handy.

If both the `insert` and `guard` attributes are specified on a list entry in a template, then the template engine first checks whether the list entry already exists in the resulting configuration between the target position (as indicated by the `insert` attribute) and the position of an element indicated by the `guard` attribute:

* If the element exists and fulfills this constraint, then its position is preserved. If a template list entry results in multiple configuration list entries, then all of them need to exist in the configuration in the same order as calculated by the template, and all of them need to fulfill the guard constraint in order for their position to be preserved.
* If the list entry/entries do not exist, are not in the same order, or do not fulfill the constraint, then the list is reordered as instructed by the insert statement.

So, in the ACL example, the template can specify the guard as follows:

```xml
<rule insert="first" guard="deny-all">
  <name>{$NAME}</name>
  <ip>{$IP}</ip>
  <mask>{$MASK}</mask>
  <action>permit</action>
</rule>
```

A guard can be specified literally (e.g. `guard="deny-all"` if "name" is the key of the list) or using an XPath expression (e.g. `guard="{$LASTRULE}"`). If the guard evaluates to a node-set consisting of multiple elements, then only the first element in this node-set is considered as the guard. The constraint defined by the `guard` is evaluated as follows:

* If the guard evaluates to an empty node-set (i.e. the node indicated by the guard does not exist in the target configuration), then the constraint is not fulfilled.
* If `insert="first"`, then the constraint is fulfilled if the element exists in the configuration *before* the element indicated by the guard.
* If `insert="last"`, then the constraint is fulfilled if the element exists in the configuration after the element indicated by the guard.
* If `insert="after"`, then the constraint is fulfilled if the element exists in the configuration before the element indicated by the `guard`, but after the element indicated by the `value` attribute.
* If `insert="before"`, then the constraint is fulfilled if the element exists in the configuration after the element indicated by the `guard`, but before the element indicated by the or `value` attribute.

## Macros in Templates <a href="#ch_templates.macros" id="ch_templates.macros"></a>

Templates support macros - named XML snippets that facilitate reuse and simplify complex templates. When you call a previously defined macro, the templating engine inserts the macro data, expanded with the values of the supplied arguments. The following example demonstrates the use of a macro.

{% code title="Example: Template with Macros" %}

```xml
  1 <config-template xmlns="http://tail-f.com/ns/config/1.0">
      <?macro GbEth name='{/name}' ip mask='255.255.255.0'?>
        <GigabitEthernet>
          <name>$name</name>
  5       <ip>
            <address>
              <primary>
                <address>$ip</address>
                <mask>$mask</mask>
 10           </primary>
            </address>
          </ip>
        </GigabitEthernet>
      <?endmacro?>
 15 
      <?macro GbEthDesc name='{/name}' ip mask='255.255.255.0' desc?>
        <?expand GbEth name='$name' ip='$ip' mask='$mask'?>
        <GigabitEthernet>
          <name>$name</name>
 20       <description>$desc</description>
        </GigabitEthernet>
      <?endmacro?>
    
      <devices xmlns="http://tail-f.com/ns/ncs">
 25     <device tags="nocreate">
          <name>{/device}</name>
          <config tags="merge">
            <interface xmlns="urn:ios">
              <?expand GbEthDesc name='0/0/0/0' ip='10.250.1.1'
 30                              desc='Link to core'?>
            </interface>
          </config>
        </device>
      </devices>
 35 </config-template>}
```

{% endcode %}

When using macros, be mindful of the following:

* A macro must be a valid chunk of XML, or a simple string without any XML markup. So, a macro cannot contain only start-tags or only end-tags, for example.
* Each macro is defined between the `<?macro?>` and `<?endmacro?>` processing instructions, immediately following the `<config-template>` tag in the template.
* A macro definition takes a name and an optional list of parameters. Each parameter may define a default value.

  In the preceding example, a macro is defined as:

  ```xml
    <?macro GbEth name='{/name}' ip mask='255.255.255.0'?>
  ```

  Here, `GbEth` is the name of the macro. This macro takes three parameters, `name`, `ip`, and `mask`. The parameters `name` and `mask` have default values, and `ip` does not.

  The default value for `mask` is a fixed string, while the one for `name` by default gets its value through an XPath expression.
* A macro can be expanded in another location in the template using the `<?expand?>` processing instruction. As shown in the example (line 29), the `<?expand?>` instruction takes the name of the macro to expand, and an optional list of parameters and their values.

  The parameters in the macro definition are replaced with the values given during expansion. If a parameter is not given any value during expansion, the default value is used. If there is no default value in the definition, not supplying a value causes an error.
* Macro definitions cannot be nested - that is, a macro definition cannot contain another macro definition. But a macro definition can have `<?expand?>` instructions to expand another macro within this macro (line 17 in the example).

  The macro expansion and the parameter replacement work on just strings - there is no schema validation or XPath evaluation at this stage. A macro expansion just inserts the macro definition at the expansion site.
* Macros can be defined in multiple files, and macros defined in the same package are visible to all templates in that package. This means that a template file could have just the definitions of macros, and another file in the same package could use those macros.

When reporting errors in a template using macros, the line numbers for the macro invocations are also included, so that the actual location of the error can be traced. For example, an error message might resemble `service.xml:19:8 Invalid parameters for processing instruction set.` - meaning that there was a macro expansion on line 19 in `service.xml` and an error occurred at line 8 in the file defining that macro.

## XPath Context in Templates <a href="#ch_templates.contexts" id="ch_templates.contexts"></a>

When the evaluation of a template starts, the XPath context node and root node are both set to either the service instance data node (with a template-only service) or the node specified with the API call to apply the template (usually the service instance data node as well).

The root node is used as the starting point for evaluating absolute paths starting with `/` and puts a limit on where you can navigate with `../`.

You can access data outside the current root node subtree by dereferencing a leafref type leaf or by changing the root node from within the template.

To change the root node within the template, use the `set-root-node` XML processing instruction. The instruction takes an XPath expression as a parameter and this expression is evaluated in a special context, where the root node is the root of the datastore. This makes it possible to change to a node outside the current evaluation context.

For example: `<?set-root-node {/}?>` changes the accessible tree to the whole data store. Note that, as all processing instructions, the effect of `set-root-node` only applies until the closing parent tag.

The context node refers to the node that is used as the starting point for navigation with relative paths, such as `../device` or `device`.

You can change the current context node using the `set-context-node` or other context-related processing instructions. For example: `<?set-context-node {..}?>` changes the context node to the parent of the current context node.

There is a special case where NSO automatically changes the evaluation context as it progresses through and applies the template, which makes it easier to work with lists. There are two conditions required to trigger this special case:

1. The value being set in the template is the key of a list.
2. The XPath expression used for this key evaluates to a node set, not a value.

To illustrate, consider the following example.

Suppose you are using the template to configure interfaces on a device. Target device YANG model defines the list of interfaces as:

```yang
  list interface {
    key "name";
    leaf name {
      type string;
    }
    leaf address {
      type inet:ip-address;
    }
  }
```

You also use a service model that allows configuring multiple links:

```
  // ...
  container links {
    list link {
      key "intf-name";
      leaf intf-name {
        type string;
      }
      leaf intf-addr {
        type inet:ip-address;
      }
    }
  }
```

The context-changing mechanism allows you to configure the device interface with the specified address using the template:

```xml
  <interface>
    <name>{/links/link[0]/intf-name}</name>
    <address>{intf-addr}</address>
  </interface>
```

The `/links/link[0]/intf-name` evaluates to a node and the evaluation context node is changed to the parent of this node, `/links/link[0]`, because `name` is a key leaf. Now you can refer to `/links/link[0]/intf-addr` with a simple relative path `{intf-addr}`.

The true power and usefulness of context changing becomes evident when used together with XPath expressions that produce node sets with multiple nodes. You can create a template that configures multiple interfaces with their corresponding addresses (note the use of `link` instead of `link[0]`):

```xml
  <interface>
    <name>{/links/link/intf-name}</name>
    <address>{intf-addr}</address>
  </interface>
```

The first expression returns a node set possibly including multiple leafs. NSO then configures multiple list items (interfaces), based on their name. The context change mechanism triggers as well, making `{intf-addr}` refer to the corresponding leaf in the same link definition. Alternatively, you can achieve the same outcome with a loop (see [Loop Statements](#ch_templates.loops)).

However, in some situations, you may not desire to change the context. You can avoid it by making the XPath expression return a value instead of a node/node-set. The simplest way is to use the XPath `string()` function, for example:

```xml
  <interface>
    <name>{string(/links-list/intf-name)}</name>
  </interface>
```

## Namespaces and Multi-NED Support <a href="#ch_templates.multined" id="ch_templates.multined"></a>

When a device makes itself known to NSO, it presents a list of capabilities (see [Capabilities, Modules, and Revision Management](https://nso-docs.cisco.com/guides/development/core-concepts/pages/auKQMOAF2p1jiGYJBweP#user_guide.devicemanager.capas)), which includes what YANG modules that particular device supports. Since each YANG module defines a unique XML namespace, this information can be used in a template.

Hence, a template may include configuration for many diverse devices. The templating system streamlines this by applying only those pieces of the template that have a namespace matching the one advertised by the device (see [Supporting Different Device Types](https://nso-docs.cisco.com/guides/development/core-concepts/pages/1ErWKzeM15TwAei4yXWm#ch_services.devs_types)).

Additionally, the system performs validation of the template against the specified namespace when loading the template as part of the package load sequence, allowing you to detect a lot of the errors at load time instead of at run time.

In case the namespace matching is insufficient, such as when you want to check for a particular version of a NED, you can use special processing instructions `if-ned-id` or `if-ned-id-match`. See [Processing Instructions Reference](#ch_templates.xml_instructions) for details and [Supporting Different Device Types](https://nso-docs.cisco.com/guides/development/core-concepts/pages/1ErWKzeM15TwAei4yXWm#ch_services.devs_types) for an example.

However, strict validation against the currently loaded schema may become a problem for developing generic, reusable templates that should run in different environments with different sets of NEDs and NED versions loaded. For example, an NSO instance having fewer NED versions than the template is designed for may result in some elements not being recognized, while having more NED versions may introduce ambiguities.

In order to allow templates to be reusable while at the same time keeping as many errors as possible detectable at load time, NSO has a concept of `supported-ned-ids`. This is a set of NED IDs the package developer declares in the `package-meta-data.xml` file, indicating all NEDs the XML templates contained in this package are designed to support. This gives NSO a hint on how to interpret the template.

{% code title="Example: Package Declaring supported-ned-id" %}

```xml
<ncs-package xmlns="http://tail-f.com/ns/ncs-packages">
  <name>mypackage</name>
  <!-- ... -->

  <!-- Exact NED id match, requires namespace -->
  <supported-ned-id xmlns:id="http://tail-f.com/ns/ned-id/cisco-ios-cli-3.0">
    id:cisco-ios-cli-3.0
  </supported-ned-id>

  <!-- Regex-based NED id match -->
  <supported-ned-id-match>router-nc-1\..*</supported-ned-id-match>
</ncs-package>
```

{% endcode %}

Namely, if a package declares a list of supported-ned-ids, then the templates in this package are interpreted as if no other ned-ids are loaded in the system. If such a template is attempted to be applied to a device with ned-id outside the supported list, then a run-time error is generated because this ned-id was not considered when the template was loaded. This allows us to ignore ambiguities in the data model introduced by additional NEDs that were not considered during template development.

If a package declares a list of supported-ned-ids and the runtime system does not have one or more declared NEDs loaded, then the template engine uses the so-called relaxed loading mode, which means it ignores any unknown namespaces and `<?if-ned-id?>` clauses containing exclusively unknown ned-ids, assuming that these parts of the template are not applicable in the current running system. Note, however, that `<supported-ned-id-match>` in the current implementation only filters the list of currently loaded NEDs and does not result in relaxed loading mode.

Because relaxed loading mode performs less strict validation and potentially prevents some errors from being detected, the package developer should always make sure to test in the system with all the supported ned-ids loaded, i.e. when the loading mode is `strict`. The loading mode can be verified by looking at the value of `template-loading-mode` leaf for the corresponding package under `/packages/package` list.

If the package does not declare any `supported-ned-ids`, then the templates are loaded in `strict` mode, using the full set of currently loaded NED IDs. This may make the package less reusable between different systems, but is usually fine in environments where the package is intended to be used in runtime systems fully under the control of the package developer.

## Passing Deep Structures from API <a href="#d5e2638" id="d5e2638"></a>

When applying the template via API, you typically pass parameters to a template through variables, as described in [Templates and Code](/guides/development/core-concepts/implementing-services#templates-and-code) and [Values in a Template](#ch_templates.values). One limitation of this mechanism is that a variable can only hold one string value. Yet, sometimes there is a need to pass not just a single value, but a list, map, or even more complex data structures from API to the template.

One way to achieve this is to use smaller templates, such as invoking the template repeatedly, one by one for each list item (or perhaps pair-by-pair in the case of a map). However, there are certain disadvantages to this approach. One of them is the performance: every invocation of the template from the API requires a context switch between the user application process and the NSO core process, which can be costly. Another disadvantage is that the logic is split between Java or Python code and the template, which makes it harder to understand and implement.

An alternative approach described in this section involves modeling the required auxiliary data as operational data and populating it in the code, before applying the template. For a service, the service callback code in Java or Python first populates the auxiliary data and then passes control to the template, which handles the main service configuration logic. The auxiliary data is accessible in the template, by means of XPath, just like any other service input data.

There are different approaches to modeling the auxiliary data. It can reside in the service tree as it is private to the service instance; either integrated in the existing data tree or as a separate subtree under the service instance. It can also be located outside of the service instance, however, it is important to keep in mind that operational data cannot be shared by multiple services because there are no refcounters or backpointers stored on operational data.

After the service is deployed, the auxiliary leafs remain in the database which facilitates debugging because they can be seen via all northbound interfaces. If this is not the intention, they can be hidden with the help of `tailf:hidden` statement. Because operational data is also a part of FASTMAP diff, these values will be deleted when the service is deleted and need to be recomputed when the service is re-deployed. This also means that in most cases there should be no need to write any additional code to clean up this data.

One example of a task that is hard to solve in the template by native XPath functions is converting a network prefix into a network mask or vice versa. Below is a snippet of a data model that is part of a service input data and contains a list of interfaces along with IP addresses to be configured on those interfaces. If the input IP address contains a prefix, but the target device accepts an IP address with a network mask instead, then you can use an auxiliary operational leaf to pass the mask (calculated from the prefix) to the template.

```yang
list interface {
  key name;
  leaf name {
    type string;
  }
  leaf address {
    type tailf:ipv4-address-and-prefix-length;
    description
      "IP address with prefix in the following format, e.g.: 10.2.3.4/24";
  }
  leaf mask {
    config false;
    type inet:ipv4-address;
    description
      "Auxiliary data populated by service code, represents network mask
       corresponding to the prefix in the address field, e.g.: 255.255.255.0";
  }
}
```

The code that calls the template needs to populate the mask. For example, using the Python Maagic API in a service:

```python
    def cb_create(self, tctx, root, service, proplist):
        interface_list = service.interface
        for intf in interface_list:
            prefix = intf.address.split('/')[1]
            intf.mask = ipaddress.IPv4Network(0, int(prefix)).netmask

        # Template variables don't need to contain mask
        # as it is passed via (operational) database
        template = ncs.template.Template(service)
        template.apply('iface-template')
```

The corresponding `iface-template` might then be as simple as:

```xml
      <interface>
        <name>{/interface/name}</name>
        <ip-address>{substring-before(address, '/')}</ip-address>
        <ip-mask>{mask}</ip-mask>
      </interface>
```

### Service Callpoints and Templates <a href="#ch_templates.servicepoint" id="ch_templates.servicepoint"></a>

The archetypical use case for XML templates is service provisioning and NSO allows you to directly invoke a template for a service, without writing boilerplate code in Python or Java. You can take advantage of this feature by configuring the `servicepoint` attribute on the root `config-template` element. For example:

```xml
<config-template xmlns="http://tail-f.com/ns/config/1.0"
                 servicepoint="some-service">
  <!-- ... -->
</config-template>
```

Adding the attribute registers this template for the given servicepoint, defined in the YANG service model. Without any additional attributes, the registration corresponds to the standard *create* service callback.

{% hint style="info" %}
While the template (file) name is not referred to in this case, it must still be unique in an NSO node.
{% endhint %}

In a similar manner, you can register templates for each state of a nano service, using `componenttype` and `state` attributes. The section [Nano Service Callbacks](https://nso-docs.cisco.com/guides/development/core-concepts/pages/vJ0F1vOS4MeZ3S4VAOGS#ug.nano_services.callbacks) contains examples.

Services also have pre- and post-modification callbacks, further described in [Service Callbacks](https://nso-docs.cisco.com/guides/development/core-concepts/pages/Bw8TviXCSEsM9XjXBC1d#ch_svcref.cbs), which you can also implement with templates. Simply put, pre- and post-modification templates are applied before and after applying the main service template.

These pre- and post-modification templates can only be used in classic (non-nano) services when the create callback is implemented as a template. That is, they cannot be used together with create callbacks implemented in Java or Python. If you want to mix the two approaches for the same service, consider using nano services.

To define a template as pre- or post-modification, appropriately configure the `cbtype` attribute, along with `servicepoint`. The `cbtype` attribute supports these three values:

* `pre-modification`
* `create`
* `post-modification`

{% hint style="info" %}
NSO supports only a single registration for each servicepoint and callback type. Therefore, you cannot register multiple templates for the same `servicepoint/cbtype` combination.
{% endhint %}

The `$OPERATION` variable is set internally by NSO in pre- and post-modification templates to contain the service operation, i.e., create, update, or delete, that triggered the callback. The `$OPERATION` variable can be used together with template conditional statements (see [Conditional Statements](#ch_templates.conditionals)) to apply different parts of the template depending on the triggering operation. Note that the service data is not available in the pre- or post-modification callbacks when `$OPERATION = 'delete'` since the service has been deleted already in the transaction context where the template is applied.

{% code title="Example: Post-modification Template" %}

```xml
<config-template xmlns="http://tail-f.com/ns/config/1.0"
                 servicepoint="some-service"
                 cbtype="post-modification">
  <?if {$OPERATION = 'create'}?>
    <devices xmlns="http://tail-f.com/ns/ncs">
      <device>
        <name>{/device}</name>
        <config>
          <!-- ... -->
        </config>
      </device>
    </devices>
  <?elif {$OPERATION = 'update'}?>
    <!-- ... -->
  <?else?>
    <!-- $OPERATION = 'delete' -->
    <!-- ... -->
  <?end?>
</config-template>
```

{% endcode %}

## Debugging Templates

You can request additional information when applying templates in order to understand what is going on. When applying or committing a template in the CLI, the `debug` pipe command enables debug information:

```bash
admin@ncs(config)# commit dry-run | debug template
```

```bash
admin@ncs(config)# commit dry-run | debug xpath
```

The `debug xpath` option outputs *all* XPath evaluations for the transaction, and is not limited to the XPath expressions inside templates.

The `debug template` option outputs XPath expression results from the template, under which context expressions are evaluated, what operation is used, and how it affects the configuration, for all templates that are invoked. You can narrow it down to only show debugging information for a template of interest:

```bash
admin@ncs(config)# commit dry-run | debug template l3vpn
```

Additionally, the template and xpath debugging can be combined:

<pre class="language-bash"><code class="lang-bash"><strong>admin@ncs(config)# commit dry-run | debug template | debug xpath
</strong></code></pre>

For XPath evaluation, you can also inspect the XPath trace log if it is enabled (e.g. with `tail -f logs/xpath.trace`). XPath trace is enabled in the `ncs.conf` configuration file and is enabled by default for the examples.

Another option to help you get the XPath selections right is to use the NSO CLI `show` command with the `xpath` display flag to find out the correct path to an instance node. This shows the name of the key elements and also the namespace changes.

```bash
admin@ncs# show running-config devices device c0 config ios:interface | display xpath
/devices/device[name='c0']/config/ios:interface/FastEthernet[name='1/0']
/devices/device[name='c0']/config/ios:interface/FastEthernet[name='1/1']
/devices/device[name='c0']/config/ios:interface/FastEthernet[name='1/2']
/devices/device[name='c0']/config/ios:interface/FastEthernet[name='2/1']
/devices/device[name='c0']/config/ios:interface/FastEthernet[name='2/2']
```

When using more complex expressions, the **ncs\_cmd** utility can be used to experiment with and debug expressions. **ncs\_cmd** is used in a command shell. The command does not print the result as XPath selections but is still of great use when debugging XPath expressions. The following example selects FastEthernet interface names on the device `c0`:

```bash
$ ncs_cmd -c "x /devices/device[name='c0']/config/ios:interface/FastEthernet/name"
/devices/device{c0}/config/interface/FastEthernet{1/0}/name [1/0]
/devices/device{c0}/config/interface/FastEthernet{1/1}/name [1/1]
/devices/device{c0}/config/interface/FastEthernet{1/2}/name [1/2]
/devices/device{c0}/config/interface/FastEthernet{2/1}/name [2/1]
/devices/device{c0}/config/interface/FastEthernet{2/2}/name [2/2]
```

### Example Debug Template Output <a href="#d5e2727" id="d5e2727"></a>

The following text walks through the output of the `debug template` command for a dns-v3 example service, found in [examples.ncs/service-management/implement-a-service/dns-v3](https://github.com/NSO-developer/nso-examples/tree/6.7/service-management/implement-a-service/dns-v3). To try it out for yourself, start the example with `make demo` and configure a service instance:

```bash
admin@ncs# config
admin@ncs(config)# load merge example.cfg
admin@ncs(config)# commit dry-run | debug template
```

The XML template used in the service is simple but non-trivial:

```xml
  1 <config-template xmlns="http://tail-f.com/ns/config/1.0"
                     servicepoint="dns">
      <devices xmlns="http://tail-f.com/ns/ncs">
        <?foreach {/target-device}?>
  5     <device>
          <name>{.}</name>
          <config>
            <ip xmlns="urn:ios">
              <?if {/dns-server-ip}?>
 10             <!-- If dns-server-ip is set, use that. -->
                <name-server>{/dns-server-ip}</name-server>
              <?else?>
                <!-- Otherwise, use the default one. -->
                <name-server>192.0.2.1</name-server>
 15           <?end?>
            </ip>
          </config>
        </device>
        <?end?>
 20   </devices>
    </config-template>
```

Applying the template produces a substantial amount of output. Let's interpret it piece by piece. The output starts with:

```
Processing instruction 'foreach': evaluating the node-set \
    (from file "dns-template.xml", line 4)
Evaluating "/target-device" (from file "dns-template.xml", line 4)
Context node: /dns[name='instance1']
Result:
For /dns[name='instance1']/target-device[.='c1'], it evaluates to []
For /dns[name='instance1']/target-device[.='c2'], it evaluates to []
```

The templating engine found the `foreach` in the `dns-template.xml` file at line 4. In this case, it is the only `foreach` block in the file but in general, there might be more. The `{/target-device}` expression is evaluated using the `/dns[name='instance1']` context, resulting in the complete `/dns[name='instance1']/target-device` path. Note that the latter is based on the root node (not shown in the output), not the context node (which happens to be the same as the root node at the start of template evaluation).

NSO found two nodes in the leaf-list for this expression, which you can verify in the CLI:

```bash
admin@ncs(config)# show full-configuration dns instance1 target-device | display xpath
/dns[name='instance1']/target-device [ c1 c2 ]
```

Next comes:

```
Processing instruction 'foreach': next iteration: \
    context /dns[name='instance1']/target-device[.='c1'] \
    (from file "dns-template.xml", line 4)
Evaluating "." (from file "dns-template.xml", line 6)
Context node: /dns[name='instance1']/target-device[.='c1']
Result:
For /dns[name='instance1']/target-device[.='c1'], it evaluates to "c1"
```

The template starts with the first iteration of the loop with the `c1` value. Since the node was an item in a leaf-list, the context refers to the actual value. If instead, it was a list, the context would refer to a single item in the list.

```
Operation 'merge' on existing node: /devices/device[name='c1'] \
    (from file "dns-template.xml", line 6)
```

This line signifies the system “applied” line 6 in the template, selecting the `c1` device for further configuration. The line also informs you the device (the item in the /devices/device list with this name) exists.

```
Processing instruction 'if': evaluating the condition \
    (from file "dns-template.xml", line 9)
Evaluating conditional expression "boolean(/dns-server-ip)" \
    (from file "dns-template.xml", line 9)
Context node: /dns[name='instance1']/target-device[.='c1']
Result: true - continuing
```

The template then evaluates the `if` condition, resulting in processing of the lines 10 and 11 in the template:

```
Processing instruction 'if': recursing (from file "dns-template.xml", line 9)
Evaluating "/dns-server-ip" (from file "dns-template.xml", line 11)
Context node: /dns[name='instance1']/target-device[.='c1']
Result:
For /dns[name='instance1'], it evaluates to "192.0.2.110"
Operation 'merge' on non-existing node: \
    /devices/device[name='c1']/config/ios:ip/name-server[.='192.0.2.110'] \
    (from file "dns-template.xml", line 11)
```

The last line shows how a new value is added to the target leaf-list, that was not there (non-existing) before.

```
Processing instruction 'else': skipping (from file "dns-template.xml", line 12)
Processing instruction 'foreach': next iteration: \
    context /dns[name='instance1']/target-device[.='c2'] \
    (from file "dns-template.xml", line 4)
```

As the `if` statement matched, the `else` part does not apply and a new iteration of the loop starts, this time with the `c2` value.

Now the same steps take place for the other, `c2`, device:

```
Evaluating "." (from file "dns-template.xml", line 6)
Context node: /dns[name='instance1']/target-device[.='c2']
Result:
For /dns[name='instance1']/target-device[.='c2'], it evaluates to "c2"
Operation 'merge' on existing node: /devices/device[name='c2'] \
    (from file "dns-template.xml", line 6)
Processing instruction 'if': evaluating the condition \
    (from file "dns-template.xml", line 9)
Evaluating conditional expression "boolean(/dns-server-ip)" \
    (from file "dns-template.xml", line 9)
Context node: /dns[name='instance1']/target-device[.='c2']
Result: true - continuing
Processing instruction 'if': recursing (from file "dns-template.xml", line 9)
Evaluating "/dns-server-ip" (from file "dns-template.xml", line 11)
Context node: /dns[name='instance1']/target-device[.='c2']
Result:
For /dns[name='instance1'], it evaluates to "192.0.2.110"
Operation 'merge' on non-existing node: \
    /devices/device[name='c2']/config/ios:ip/name-server[.='192.0.2.110'] \
    (from file "dns-template.xml", line 11)
Processing instruction 'else': skipping (from file "dns-template.xml", line 12)
```

Finally, the template processing completes as there are no more nodes in the loop, and NSO outputs the new dry-run configuration:

```
cli {
    local-node {
        data  devices {
                  device c1 {
                      config {
                          ip {
             -                name-server 192.0.2.1;
             +                name-server 192.0.2.1 192.0.2.110;
                          }
                      }
                  }
                  device c2 {
                      config {
                          ip {
             +                name-server 192.0.2.110;
                          }
                      }
                  }
              }
             +dns instance1 {
             +    target-device [ c1 c2 ];
             +    dns-server-ip 192.0.2.110;
             +}
    }
}
```

## Processing Instructions Reference <a href="#ch_templates.xml_instructions" id="ch_templates.xml_instructions"></a>

NSO template engine supports a number of XML processing instructions to allow more dynamic templates:

<table data-full-width="false"><thead><tr><th valign="top">Syntax</th><th valign="top">Description</th></tr></thead><tbody><tr><td valign="top"><pre><code>    &#x3C;?set v = value?>
</code></pre></td><td valign="top">Allows you to assign a new variable or manipulate the existing value of a variable <code>v</code>. If used to create a new variable, the scope of visibility of this variable is limited to the parent tag of the processing instruction or the current processing instruction block. Specifically, if a new variable is defined inside a loop, then it is discarded at the end of each iteration.</td></tr><tr><td valign="top"><pre><code>    &#x3C;?if {expression}?>
        ...
    &#x3C;?elif {expression}?>
        ...
    &#x3C;?else?>
        ...
    &#x3C;?end?>
</code></pre></td><td valign="top">Processing instruction block that allows conditional execution based on the boolean result of the expression. For a detailed description, see <a href="#ch_templates.conditionals">Conditional Statements</a>.</td></tr><tr><td valign="top"><pre><code>    &#x3C;?foreach {expression}?>
        ...
    &#x3C;?end?>
</code></pre></td><td valign="top">The expression must evaluate to a (possibly empty) XPath node-set. The template engine will then iterate over each node in the node set by changing the XPath current context node to this node and evaluating all children tags within this context. For the detailed description, see <a href="#ch_templates.loops">Loop Statements</a>.</td></tr><tr><td valign="top"><pre data-overflow="wrap"><code>    &#x3C;?for v = start_value; {progress condition}; v = next_value?>
        ...
    &#x3C;?end?>
</code></pre></td><td valign="top"><p>This processing instruction allows you to iterate over the same set of template tags by changing a variable value. The variable visibility scope obeys the same rules as the <code>set</code> processing instruction, except the variable value, is carried over to the next iteration instead of being discarded at the end of each iteration.</p><p>Only the condition expression is mandatory; either or both of the initial and next value assignment can be omitted, e.g.,</p><pre><code>    &#x3C;?for ; {condition}; ?>
</code></pre><p>For a detailed description, see <a href="#ch_templates.loops">Loop Statements</a>.</p></td></tr><tr><td valign="top"><pre><code>   &#x3C;?copy-tree {source}?>
</code></pre></td><td valign="top">This instruction is analogous to <code>copy_tree()</code> function available in the MAAPI API. The parameter is an XPath expression that must evaluate to exactly one node in the data tree and indicate the source path to copy from. The target path is defined by the position of the <code>copy-tree</code> instruction in the template within the current context.</td></tr><tr><td valign="top"><pre><code>    &#x3C;?set-root-node {expression}?>
</code></pre></td><td valign="top">Allows you to manipulate the root node of the XPath-accessible tree. This expression is evaluated in an XPath context where the accessible tree is the entire datastore, which means that it is possible to select a root node outside the currently accessible tree. The current context node remains unchanged. The expression must evaluate to exactly one node in the data tree.</td></tr><tr><td valign="top"><pre><code>    &#x3C;?set-context-node {expression}?>
</code></pre></td><td valign="top">Allows you to manipulate the current context node used to evaluate XPath expressions in the template. The expression is evaluated within the current XPath context and must evaluate to exactly one node in the data tree.</td></tr><tr><td valign="top"><pre><code>    &#x3C;?save-context name?>
</code></pre></td><td valign="top">Store both the current context node and the root node of the XPath accessible tree with <em><code>name</code></em> being the key to access it later. It is possible to switch to this context later using <code>switch-context</code> with the name. Multiple contexts can be stored simultaneously under different names. Using save-context with the same name multiple times will result in the stored context being overwritten.</td></tr><tr><td valign="top"><pre><code>    &#x3C;?switch-context name?>
</code></pre></td><td valign="top">Used to switch to a context stored using <code>save-context</code> with the specified name. This means that both the current context node and the root node of the XPath accessible tree will be changed to the stored values. <code>switch-context</code> does not remove the context from the storage and can be used as many times as needed; however, using it with a name that does not exist in the storage causes an error.</td></tr><tr><td valign="top"><pre><code>    &#x3C;?if-ned-id ned-ids?>
        ...
    &#x3C;?elif-ned-id ned-ids?>
        ...
    &#x3C;?else?>
        ...
    &#x3C;?end?>
</code></pre></td><td valign="top"><p>If there are multiple versions of the same NED expected to be loaded in the system, which define different versions of the same namespace, this processing instruction helps to resolve ambiguities in the schema between different versions of the NED. The part of the template following this processing instruction, up to matching <code>elif-ned-id</code>, <code>else</code> or <code>end</code> processing instruction, is only applied to devices with the <code>ned-id</code> matching one of the <code>ned-ids</code> specified as a parameter to this processing instruction. If there are no ambiguities to resolve, then this processing instruction is not required. The <code>ned-ids</code> must contain one or more qualified NED ID identities separated by spaces.</p><p><br>The <code>elif-ned-id</code> is optional and used to define a part of the template that applies to devices with another set of <code>ned-ids</code> than previously specified. Multiple <code>elif-ned-id</code> instructions are allowed in a single block of <code>if-ned-id</code> instructions. The set of ned-ids specified as a parameter to <code>elif-ned-id</code> instruction must be non-intersecting with the previously specified ned-ids in this block.</p><p>The <code>else</code> processing instruction should be used with care in this context, as the set of the <code>ned-ids</code> it handles depends on the set of <code>ned-ids</code> loaded in the system, which can be hard to predict at the time of developing the template. To mitigate this problem, it is recommended that the package containing this template defines a set of <code>supported-ned-ids</code> as described in <a href="#ch_templates.multined">Namespaces and Multi-NED Support</a>.</p></td></tr><tr><td valign="top"><pre><code>    &#x3C;?if-ned-id-match regex?>
        ...
    &#x3C;?elif-ned-id-match regex?>
        ...
    &#x3C;?else?>
        ...
    &#x3C;?end?>
</code></pre></td><td valign="top">The <code>if-ned-id-match</code> and <code>elif-ned-id-match</code> processing instructions work similarly to <code>if-ned-id</code> and <code>elif-ned-id</code> but they accept a regular expression as an argument instead of a list of ned-ids. The regular expression is matched against all of the <code>ned-ids</code> supported by the package. If the <code>if-ned-id-match</code> processing instruction is nested inside of another <code>if-ned-id-match</code> or <code>if-ned-id</code> processing instruction, then the regular expression will only be matched against the subset of ned-ids matched by the encompassing processing instruction. The <code>if-ned-id-match</code> and <code>elif-ned-id-match</code> processing instructions are only allowed inside a device's mounted configuration subtree rooted at /devices/device/config.</td></tr><tr><td valign="top"><pre><code>    &#x3C;?macro name params...?>
        ...
    &#x3C;?endmacro?>
</code></pre></td><td valign="top">Define a new macro with the specified name and optional parameters. Macro definitions must come at the top of the template, right after the <code>config-template</code> tag. For a detailed description see <a href="#ch_templates.macros">Macros in Templates</a>.</td></tr><tr><td valign="top"><pre><code>    &#x3C;?expand name params...?>
</code></pre></td><td valign="top">Insert and expand the named macro, using the specified values for parameters. For a detailed description, see <a href="#ch_templates.macros">Macros in Templates</a>.</td></tr></tbody></table>

The variable value in both `set` and `for` processing instructions are evaluated in the same way as the values within XML tags in a template (see [Values in a Template](#ch_templates.values)). So, it can be a mix of literal values and XPath expressions surrounded by `{...}`.

The variable value is always stored as a string, so any XPath expression will be converted to literal using the XPath `string()` function. Namely, if the expression results in an integer or a boolean, then the resulting value would be a string representation of the integer or boolean. If the expression results in a node set, then the value of the variable is a concatenated string of values of nodes in this node set.

It is important to keep in mind that while in some cases XPath converts the literal to another type implicitly (for example, in an expression `{$x < 3}` a value x='1' is converted to integer 1 implicitly), in other cases an explicit conversion is needed. For example, using the expression `{$x > $y}`, if x='9' and y='11', the result of the expression is true due to alphabetic order as both variables are strings. In order to compare the values as numbers, an explicit conversion of at least one argument is required: `{number($x) > $y}`.

## XPath Functions <a href="#d5e2911" id="d5e2911"></a>

This section lists a few useful functions, available in XPath expressions. The list is not exhaustive; please refer to the [XPath standard](https://www.w3.org/TR/1999/REC-xpath-19991116/#corelib), [YANG standard](https://datatracker.ietf.org/doc/html/rfc7950#section-10), and NSO-specific extensions in [XPATH FUNCTIONS](/guides/resources/man/tailf_yang_extensions.5#xpath-functions) in Manual Pages for a full list.

<details>

<summary>Type Conversion</summary>

* [bit-is-set()](https://datatracker.ietf.org/doc/html/rfc7950#section-10.6.1)
* [boolean()](https://www.w3.org/TR/1999/REC-xpath-19991116/#function-boolean)
* [enum-value()](https://datatracker.ietf.org/doc/html/rfc7950#section-10.5.1)
* [number()](https://www.w3.org/TR/1999/REC-xpath-19991116/#function-number)
* [string()](https://www.w3.org/TR/1999/REC-xpath-19991116/#function-string)

</details>

<details>

<summary>String Handling</summary>

* [concat()](https://www.w3.org/TR/1999/REC-xpath-19991116/#function-concat)
* [contains()](https://www.w3.org/TR/1999/REC-xpath-19991116/#function-contains)
* [normalize-space()](https://www.w3.org/TR/1999/REC-xpath-19991116/#function-normalize-space)
* [re-match()](https://datatracker.ietf.org/doc/html/rfc7950#section-10.2.1)
* [starts-with()](https://www.w3.org/TR/1999/REC-xpath-19991116/#function-starts-with)
* [substring()](https://www.w3.org/TR/1999/REC-xpath-19991116/#function-substring)
* [substring-after()](https://www.w3.org/TR/1999/REC-xpath-19991116/#function-substring-after)
* [substring-before()](https://www.w3.org/TR/1999/REC-xpath-19991116/#function-substring-before)
* [translate()](https://www.w3.org/TR/1999/REC-xpath-19991116/#function-translate)

</details>

<details>

<summary>Model Navigation</summary>

* [current()](https://datatracker.ietf.org/doc/html/rfc7950#section-10.1.1)
* [deref()](https://datatracker.ietf.org/doc/html/rfc7950#section-10.3.1)
* [last()](https://www.w3.org/TR/1999/REC-xpath-19991116/#function-last)
* [sort-by()](/guides/resources/man/tailf_yang_extensions.5) in Manual Pages

</details>

<details>

<summary>Other</summary>

* [compare()](/guides/resources/man/tailf_yang_extensions.5) in Manual Pages
* [count()](https://www.w3.org/TR/1999/REC-xpath-19991116/#function-count)
* [max()](/guides/resources/man/tailf_yang_extensions.5) in Manual Pages
* [min()](/guides/resources/man/tailf_yang_extensions.5) in Manual Pages
* [not()](https://www.w3.org/TR/1999/REC-xpath-19991116/#function-not)
* [sum()](https://www.w3.org/TR/1999/REC-xpath-19991116/#function-sum)

</details>


# Nano Services

Implement staged provisioning in your network using nano services.

Typical NSO services perform the necessary configuration by using the `create()` callback, within a transaction tracking the changes. This approach greatly simplifies service implementation, but it also introduces some limitations. For example, all provisioning is done at once, which may not be possible or desired in all cases. In particular, network functions implemented by containers or virtual machines often require provisioning in multiple steps.

Another limitation is that the service mapping code must not produce any side effects. Side effects are not tracked by the transaction and therefore cannot be automatically reverted. For example, imagine that there is an API call to allocate an IP address from an external system as part of the `create()` code. The same code runs for every service change or a service re-deploy, even during a `commit dry-run`, unless you take special precautions. So, a new IP address would be allocated every time, resulting in a lot of waste, or worse, provisioning failures.

Nano services help you overcome these limitations. They implement a service as several smaller (nano) steps or stages, by using a technique called reactive FASTMAP (RFM), and provide a framework to safely execute actions with side effects. Reactive FASTMAP can also be implemented directly, using the CDB subscribers, but nano services offer a more streamlined and robust approach for staged provisioning.

The section starts by gradually introducing the nano service concepts in a typical use case. To aid readers working with nano services for the first time, some of the finer points are omitted in this part and discussed later on, in [Implementation Reference](#ug.nano_services.impl). The latter is designed as a reference to aid you during implementation, so it focuses on recapitulating the workings of nano services at the expense of examples. The rest of the chapter covers individual features with associated use cases and the complete working examples, which you may find in the [examples.ncs/nano-services](https://github.com/NSO-developer/nso-examples/tree/6.7/nano-services) folder.

## Basic Concepts <a href="#d5e9577" id="d5e9577"></a>

Services ideally perform the configuration all at once, with all the benefits of a transaction, such as automatic rollback and cleanup on errors. For nano services, this is not possible in the general case. Instead, a nano service performs as much configuration as possible at the moment and leaves the rest for later. When an event occurs that allows more work to be done, the nano service instance restarts provisioning, by using a re-deploy action called `reactive-re-deploy`. It allows the service to perform additional configuration that was not possible before. The process of automatic re-deploy, called reactive FASTMAP, is repeated until the service is fully provisioned.

This is most evident with, for example, virtual machine (VM) provisioning, during virtual network function (VNF) orchestration. Consider a service that deploys and configures a router in a VM. When the service is first instantiated, it starts provisioning a router VM. However, it will likely take some time before the router has booted up and is ready to accept a new configuration. In turn, the service cannot configure the router just yet. The service must wait for the router to become ready. That is the event that triggers a re-deploy and the service can finish configuring the router, as the following figure illustrates:

<div data-with-frame="true"><figure><img src="/files/XgR9eu8lYPwqMIzXxL1w" alt="" width="563"><figcaption><p>Virtual Router Provisioning Steps</p></figcaption></figure></div>

While each step of provisioning happens inside a transaction and is still atomic, the whole service is not. Instead of a simple fully-provisioned or not-provisioned-at-all status, a nano service can be in a number of other *states*, depending on how far in the provisioning process it is.

The figure shows that the router VM goes through multiple states internally, however, only two states are important for the service. These two are shown as arrows, in the lower part of the figure. When a new service is configured, it requests a new VM deployment. Having completed this first step, it enters the “VM is requested but still provisioning” state. In the following step, the VM is configured and so enters the second state, where the router VM is deployed and fully configured. The states obviously follow individual provisioning steps and are used to report progress. What is more, each state tracks if an error occurred during provisioning.

For these reasons, service states are central to the design of a nano service. A list of different states, their order, and transitions between them is called a plan outline and governs the service behavior.

### Plan Outline <a href="#d5e9593" id="d5e9593"></a>

By default, the plan outline consists of a single component, the `self` component, with the two states `init` and `ready`. It can be used to track the progress of the service as a whole. You can add any number of additional components and states to form the nano service.

The following YANG snippet, also part of the [examples.ncs/nano-services/basic-vrouter](https://github.com/NSO-developer/nso-examples/tree/6.7/nano-services/basic-vrouter) example, shows a plan outline with the two VM-provisioning states presented above:

```yang
module vrouter {
  prefix vr;

  identity vm-requested {
    base ncs:plan-state;
  }

  identity vm-configured {
    base ncs:plan-state;
  }

  identity vrouter {
    base ncs:plan-component-type;
  }

  ncs:plan-outline vrouter-plan {
    description "Plan for configuring a VM-based router";

    ncs:component-type "vr:vrouter" {
      ncs:state "vr:vm-requested";
      ncs:state "vr:vm-configured";
    }
  }
}
```

The first part contains a definition of states as identities, deriving from the `ncs:plan-state` base. These identities are then used with the `ncs:plan-outline`, inside an `ncs:component-type` statement. Also, note that it is customary to use past tense for state names, for example, `configured-vm` or `vm-configured` instead of `configure-vm` and `configuring-vm`.

At present, the plan contains one component and two states but no logic. If you wish to do any provisioning for a state, the state must declare a special nano create callback, otherwise, it just acts as a checkpoint. The nano create callback is similar to an ordinary create service callback, allowing service code or templates to perform configuration. To add a callback for a state, extend the definition in the plan outline:

```
ncs:state "vr:vm-requested" {
  ncs:create {
    ncs:nano-callback;
  }
}
```

The service automatically enters each state one by one when a new service instance is configured. However, for the `vm-configured` state, the service should wait until the router VM has had the time to boot and is ready to accept a new configuration. An `ncs:pre-condition` statement in YANG provides this functionality. Until the condition becomes fulfilled, the service will not advance to that state.

The following YANG code instructs the nano service to check the value of the `vm-up-and-running` leaf, before entering and performing the configuration for a state.

```
ncs:state "vr:vm-configured" {
  ncs:create {
    ncs:nano-callback;
    ncs:pre-condition {
      ncs:monitor "$SERVICE" {
        ncs:trigger-expr "vm-up-and-running = 'true'";
      }
    }
  }
}
```

### Per-State Configuration <a href="#d5e9614" id="d5e9614"></a>

The main reason for defining multiple nano service states is to specify what part of the overall configuration belongs in each state. For the VM-router example, that entails splitting the configuration into a part for deploying a VM on a virtual infrastructure and a part for configuring it. In this case, a router VM is requested simply by adding an entry to a list of VM requests, while making the API calls is left to an external component, such as the VNF Manager.

If a state defines a nano callback, you can register a configuration template to it. The XML template file is very similar to an ordinary service template but requires additional `componenttype` and `state` attributes in the `config-template` root element. These attributes identify which component and state in the plan outline the template belongs to, for example:

```xml
<config-template xmlns="http://tail-f.com/ns/config/1.0"
                 servicepoint="vrouter-servicepoint"
                 componenttype="vr:vrouter"
                 state="vr:vm-configured">
  <devices xmlns="http://tail-f.com/ns/ncs">
    <!-- ... -->
  </devices>
</config-template>
```

Likewise, you can implement a callback in the service code. The registration requires you to specify the component and state, as the following Python example demonstrates:

```python
class NanoApp(ncs.application.Application):
    def setup(self):
        self.register_nano_service('vrouter-servicepoint',  # Service point
                                   'vr:vrouter',            # Component
                                   'vr:vm-requested',       # State
                                   NanoServiceCallbacks)
```

The selected `NanoServiceCallbacks` class then receives callbacks in the `cb_nano_create()` function:

```python
class NanoServiceCallbacks(ncs.application.NanoService):
    @ncs.application.NanoService.create
    def cb_nano_create(self, tctx, root, service, plan, component, state,
                       proplist, component_proplist):
        ...
```

The `component` and `state` parameters allow the function to distinguish calls for different callbacks when registered for more than one.

For most flexibility, each state defines a separate callback, allowing you to implement some with a template and others with code, all as part of the same service. You may even use Java instead of Python, as explained in [Nano Service Callbacks](#ug.nano_services.callbacks).

### Link Plan Outline to Service <a href="#d5e9633" id="d5e9633"></a>

The set of states used in the plan outline describes the stages that a service instance goes through during provisioning. Naturally, these are service-specific, which presents a problem if you just want to tell whether a service instance is still provisioning or has already finished. It requires the knowledge of which state is the last, final one, making it hard to check in a generic way.

That is why each service component must have the built-in `ncs:init` state as the first state and `ncs:ready` as the last state. Using the two built-in states allows for interoperability with other services and tools. The following is a complete four-state plan outline for the VM-based router service, with the two states added:

```
ncs:plan-outline vrouter-plan {
  description "Plan for configuring a VM-based router";

  ncs:component-type "vr:vrouter" {
    ncs:state "ncs:init";
    ncs:state "vr:vm-requested" {
      ncs:create {
        ncs:nano-callback;
      }
    }
    ncs:state "vr:vm-configured" {
      ncs:create {
        ncs:nano-callback;
        ncs:pre-condition {
          ncs:monitor "$SERVICE" {
            ncs:trigger-expr "vm-up-and-running = 'true'";
          }
        }
      }
    }
    ncs:state "ncs:ready";
  }
}
```

For the service to use it, the plan outline must be linked to a service point with the help of a `behavior tree`. The main purpose of a behavior tree is to allow a service to dynamically instantiate components, based on service parameters. Dynamic instantiation is not always required and the behavior tree for a basic, static, single-component scenario boils down to the following:

```
ncs:service-behavior-tree vrouter-servicepoint {
  description "A static, single component behavior tree";
  ncs:plan-outline-ref "vr:vrouter-plan";
  ncs:selector {
    ncs:create-component "'vrouter'" {
      ncs:component-type-ref "vr:vrouter";
    }
  }
}
```

This behavior tree always creates a single `“vrouter”` component for the service. The service point is provided as an argument to the `ncs:service-behavior-tree` statement, while the `ncs:plan-outline-ref` statement provides the name for the plan outline to use.

The following figure visualizes the resulting service plan and its states.

<div data-with-frame="true"><figure><img src="/files/hLtCwqCvUoazxFXVmf6L" alt="" width="375"><figcaption><p>Virtual Router Provisioning Plan</p></figcaption></figure></div>

Along with the behavior tree, a nano service also relies on the `ncs:nano-plan-data` grouping in its service model. It is responsible for storing state and other provisioning details for each service instance. Other than that, the nano service model follows the standard YANG definition of a service:

```yang
list vrouter {
  description "Trivial VM-based router nano service";

  uses ncs:nano-plan-data;
  uses ncs:service-data;
  ncs:servicepoint vrouter-servicepoint;

  key name;
  leaf name {
    type string;
  }

  leaf vm-up-and-running {
    type boolean;
    config false;
  }
}
```

This model includes the operational `vm-up-and-running` leaf, that the example plan outline depends on. In practice, however, a plan outline is more likely to reference values provided by another part of the system, such as the actual, externally provided, state of the provisioned VM.

### Service Instantiation <a href="#d5e9658" id="d5e9658"></a>

A nano service does not directly use its service point for configuration. Instead, the service point invokes a behavior tree to generate a plan, and the service starts executing according to this plan. As it reaches a certain state, it performs the relevant configuration for that state.

For example, when you create a new instance of the VM-router service, the `vm-up-and-running` leaf is not set, so only the first part of the service runs. Inspecting the service instance plan reveals the following:

```cli
admin@ncs# show vrouter vr-01 plan
                                                                                     POST
                  BACK                                                               ACTION
TYPE     NAME     TRACK  GOAL  STATE          STATUS       WHEN                 ref  STATUS
---------------------------------------------------------------------------------------------
self     self     false  -     init           reached      2023-08-11T07:45:20  -    -
                               ready          not-reached  -                    -    -
vrouter  vrouter  false  -     init           reached      2023-08-11T07:45:20  -    -
                               vm-requested   reached      2023-08-11T07:45:20  -    -
                               vm-configured  not-reached  -                    -    -
                               ready          not-reached  -                    -    -
```

Since neither the `init` nor the `vm-requested` states have any pre-conditions, they are reached right away. In fact, NSO can optimize it into a single transaction (this behavior can be disabled if you use forced commits, discussed later on).

But the process has stopped at the `vm-configured` state, denoted by the `not-reached` status in the output. It is waiting for the pre-condition to become fulfilled with the help of a kicker. The job of the kicker is to watch the value and perform an action, the reactive re-deploy, when the conditions are satisfied. The kickers are managed by the nano service subsystem: when an unsatisfied precondition is encountered, a kicker is configured, and when the precondition becomes satisfied, the kicker is removed.

You may also verify, through the `get-modifications` action, that only the first part, the creation of the VM, was performed:

```cli
admin@ncs# vrouter vr-01 get-modifications
cli {
    local-node {
        data +vm-instance vr-01 {
              +    type csr-small;
              +}

    }
}
```

At the same time, a kicker was installed under the `kickers` container but you may need to use the `unhide debug` command to inspect it. More information on kickers in general is available in [Kicker](/guides/development/advanced-development/kicker).

At a later point in time, the router VM becomes ready, and the `vm-up-and-running` leaf is set to a `true` value. The installed kicker notices the change and automatically calls the `reactive-re-deploy` action on the service instance. In turn, the service gets fully deployed.

```cli
admin@ncs# show vrouter vr-01 plan
                                                                                 POST
                  BACK                                                           ACTION
TYPE     NAME     TRACK  GOAL  STATE          STATUS   WHEN                 ref  STATUS
-----------------------------------------------------------------------------------------
self     self     false  -     init           reached  2023-08-11T07:45:20  -    -
                               ready          reached  2023-08-11T07:47:36  -    -
vrouter  vrouter  false  -     init           reached  2023-08-11T07:45:20  -    -
                               vm-requested   reached  2023-08-11T07:45:20  -    -
                               vm-configured  reached  2023-08-11T07:47:36  -    -
                               ready          reached  2023-08-11T07:47:36  -    -
```

The `get-modifications` output confirms this fact. It contains the additional IP address configuration, performed as part of the `vm-configured` step:

```cli
admin@ncs# vrouter vr-01 get-modifications
cli {
    local-node {
        data +vm-instance vr-01 {
             +    type    csr-small;
             +    address 198.51.100.1;
             +}
    }
}
```

The `ready` state has no additional pre-conditions, allowing NSO to reach it along with the `vm-configured` state. This effectively breaks the provisioning process into two steps. To break it down further, simply add more states with corresponding pre-conditions and create logic.

Other than staged provisioning, nano services act the same as other services, allowing you to use the service check-sync and similar actions, for example. But please note the un-deploy and re-deploy actions may behave differently than expected, as they deal with provisioning. Chiefly, a re-deploy reevaluates the pre-conditions, possibly generating a different configuration if a pre-condition depends on operational values that have changed. The un-deploy action, on the other hand, removes all of the recorded modifications, along with the generated plan.

## Benefits and Use Cases <a href="#d5e9694" id="d5e9694"></a>

Every service in NSO has a YANG definition of the service parameters, a service point name, and an implementation of the service point `create()` callback. Normally, when a service is committed, the FASTMAP algorithm removes all previous data changes internally and presents the service data to the `create()` callback as if this was the initial create. When the `create()` callback returns, the FASTMAP algorithm compares the result and calculates a reverse diff-set from the data changes. This reverse diff-set contains the operations that are needed to restore the configuration data to the state as it was before the service was created. The reverse diff-set is required, for instance, if the service is deleted or modified.

This fundamental principle is what makes the implementation of services and the `create()` callback simple. In turn, a lot of the NSO functionality relies on this mechanism.

However, in the reactive FASTMAP pattern, the `create()` callback is re-entered several times by using the subsequent `reactive-re-deploy` calls. Storing all changes in a single reverse diff-set then becomes an impediment. For instance, if a staged delete is necessary, there is no way to single out which changes each RFM step performed.

A nano service abandons the single reverse diff-set by introducing `nano-plan-data` and a new `NanoCreate()` callback. The `nano-plan-data` YANG grouping represents an executable plan that the system can follow to provision the service. It has additional storage for reverse diff-set and pre-conditions per state, for each component of the plan.

This is illustrated in the following figure:

<div data-with-frame="true"><figure><img src="/files/L8yqx6RkxDMMHXcFxK9Q" alt="" width="563"><figcaption><p>Per-state FASTMAP with nano services</p></figcaption></figure></div>

You can still use the service `get-modifications` action to visualize all data changes performed by the service as an aggregate. In addition, each state also has its own `get-modifications` action that visualizes the data changes for that particular state. It allows you to more easily identify the state and, by extension, the code that produced those changes.

Before nano services became available, RFM services could only be implemented by creating a CDB subscriber. With the subscriber approach, the service can still leverage the plan-data grouping, which `nano-plan-data` is based on, to report the progress of the service under the resulting `plan` container. But the `create()` callback becomes responsible for creating the plan components, their states, and setting the status of the individual states as the service creation progresses.

Moreover, implementing a staged delete with a subscriber often requires keeping the configuration data outside of the service. The code is then distributed between the service `create()` callback and the correlated CDB subscriber. This all results in several sources that potentially contain errors that are complicated to track down. Nano services, on the other hand, do not require any use of CDB subscribers or other mechanisms outside of the service code itself to support the full-service life cycle.

## Backtracking and Staged Delete <a href="#d5e9725" id="d5e9725"></a>

Resource de-provisioning is an important part of the service life cycle. The FASTMAP algorithm ensures that no longer needed configuration changes in NSO are removed automatically but that may be insufficient by itself. For example, consider the case of a VM-based router, such as the one described earlier. Perhaps provisioning of the router also involves assigning a license from a central system to the VM and that license must be returned when the VM is decommissioned. If releasing the license must be done by the VM itself, simply destroying it will not work.

Another example is the management of a web server VM for a web application. Here, each VM is part of a larger pool of servers behind a load balancer that routes client requests to these servers. During de-provisioning, simply stopping the VM interrupts the currently processing requests and results in client timeouts. This can be avoided with a graceful shutdown, which stops the load balancer from sending new connections to the server and waits for the current ones to finish, before removing the VM.

Both examples require two distinct steps for de-provisioning. Can nano services be of help in this case? Certainly. In addition to the state-by-state provisioning of the defined components, the nano service system in NSO is responsible for back-tracking during their removal. This process traverses all reached states in the reverse order, removing the changes previously done for each state one by one.

<div data-with-frame="true"><figure><img src="/files/YwBjyxuY8ZUyGDZdLuhY" alt="" width="563"><figcaption><p>Staged Delete with Backtracking</p></figcaption></figure></div>

In doing so, the back-tracking process checks for a 'delete pre-condition' of a state. A delete pre-condition is similar to the create pre-condition, but only relevant when back-tracking. If the condition is not fulfilled, the back-tracking process stops and waits until it becomes satisfied. Behind the scenes, a kicker is configured to restart the process when that happens.

If the state's delete pre-condition is fulfilled, back-tracking first removes the state's 'create' changes recorded by FASTMAP and then invokes the nano `delete()` callback, if defined. The main use of the callback is to override or veto the default status calculation for a back-tracking state. That is why you can't implement the `delete()` callback with a template, for example. Very importantly, `delete()` changes are not kept in a service's reverse diff-set and may stay even after the service is completely removed. In general, you are advised to avoid writing any configuration data because this callback is called under a removal phase of a plan component where new configuration is seldom expected.

Since the 'create' configuration is automatically removed, without the need for a separate `delete()` callback, these callbacks are used only in specific cases and are not very common. Regardless, the `delete()` callback may run as part of the `commit dry-run` command, so it must not invoke further actions or cause side effects.

Backtracking is invoked when a component of a nano service is removed, such as when deleting a service. It is also invoked when evaluating a plan and a reached state's 'create' pre-condition is no longer satisfied. In this case, the affected component is temporarily set to a back-tracking mode for as long as it contains such nonconforming states. It allows the service to recover and return to a well-defined state.

<div data-with-frame="true"><figure><img src="/files/qeF5tCTh7qYSWtefx2EY" alt="" width="375"><figcaption><p>Backtracking on no longer satisfied pre-condition</p></figcaption></figure></div>

To implement the delete pre-condition or the `delete()` callback, you must add the `ncs:delete` statement to the relevant state in the plan outline. Applying it to the web server example above, you might have:

```
    ncs:state "vr:vm-requested" {
      ncs:create { ... }
      ncs:delete {
        ncs:pre-condition {
          ncs:monitor "$SERVICE" {
            ncs:trigger-expr "requests-in-processing = '0'";
          }
        }
      }
    }
    ncs:state "vr:vm-configured" {
      ncs:create { ... }
      ncs:delete {
        ncs:nano-callback;
      }
    }
```

While, in general, the `delete()` callback should not produce any configuration, the graceful shutdown scenario is one of the few exceptional cases where this may be required. Here, the `delete()` callback allows you to re-configure the load balancer to remove the server from actively accepting new connections, such as marking it 'under maintenance'. The 'delete' pre-condition allows you to further delay the VM removal until the ongoing requests are completed.

Similar to the `create()` callback, the `ncs:nano-callback` statement instructs NSO to also process a `delete()` callback. A Python class that you have registered for the nano service must then implement the following method:

```python
    @NanoService.delete
    def cb_nano_delete(self, tctx, root, service, plan, component, state,
                       proplist, component_proplist):
        ...
```

As explained, there are some uncommon cases where additional configuration with the `delete()` callback is required. However, a more frequent use of the `ncs:delete` statement is in combination with side-effect actions.

## Managing Side Effects <a href="#d5e9769" id="d5e9769"></a>

In some scenarios, side effects are an integral part of the provisioning process and cannot be avoided. The aforementioned example on license management may require calling a specific device action. Even so, the `create()` or `delete()` callbacks, nano service or otherwise, are a bad fit for such work. Since these callbacks are invoked during the transaction commit, no RPCs or other access outside of the NSO datastore are allowed. If allowed, they would break the core NSO functionality, such as a dry run, where side effects are not expected.

A common solution is to perform these actions outside of the configuration transaction. Nano services provide this functionality through the post-actions mechanism, using a `post-action-node` statement for a state. It is a definition of an action that should be invoked after the state has been reached and the commit performed. To ensure the latter, NSO will commit the current transaction before executing the post-action and advancing to the next state.

The service's plan state data also carries a post-action status leaf, which reflects whether the action was executed and if it was successful. The leaf will be set to `not-reached`, `create-reached`, `delete-reached`, or `failed`, depending on the case and result. If the action is still executing, then the leaf will show either a `create-init` or `delete-init` status instead.

Moreover, post actions can be run either asynchronously (default) or synchronously. To run them synchronously, add a `sync` statement to the post-action statement. When a post action is run asynchronously, further states will not wait for the action to finish, unless you define an explicit `post-action-status` precondition. While for a synchronous post action, later states in the same component will be invoked only after the post action is run successfully.

The exception to this setting is when a component switches to a backtracking mode. In that case, the system will not wait for any create post action to complete (synchronous or not) but will start executing backtracking right away. It means a delete callback or a delete post action for a state may run before its synchronous create post action has finished executing.

The side-effect-queue and a corresponding kicker are responsible for invoking the actions on behalf of the nano service and reporting the result in the respective state's post-action-status leaf. The following figure shows an entry is made in the side-effect-queue (2) after the state is reached (1) and its post-action status is updated (3) once the action finishes executing.

<div data-with-frame="true"><figure><img src="/files/FDIaLJikwqr7Fn3EXAWm" alt="" width="375"><figcaption><p>Post-action Execution Through side-effect-queue</p></figcaption></figure></div>

You can use the `show side-effect-queue` command to inspect the queue. The queue will run multiple actions in parallel and keep the failed ones for you to inspect. Please note that High Availability (HA) setups require special consideration: the side effect queue is disabled when High Availability is enabled and the High Availability mode is `NONE`. See [Mode of Operation](https://nso-docs.cisco.com/guides/development/core-concepts/pages/qG3CMifhI63daJ1BZfmB#ha.moo) for more details.

In case of a failure, a post action sets the post-action-status accordingly and, if the action is synchronous, the nano service stops progressing. To retry the failed action, you can perform the action `reschedule`.

```bash
$ ncs_cli -u admin
admin@ncs> show side-effect-queue side-effect status
ID  STATUS
------------
2   failed

[ok][2023-08-15 11:01:10]
admin@ncs> request side-effect-queue side-effect 2 reschedule
side-effect-id 2
[ok][2023-08-15 11:01:18]
```

Or, execute a (reactive) re-deploy, which will also restart the nano service if it was stopped.

Using the post-action mechanism, it is possible to define side effects for a nano service in a safe way. A post-action is only executed one time. That is if the post-action-status is already at the `create-reached` in the create case or `delete-reached` in the delete case, then new calls of the post-actions are suppressed. In dry-run operations, post-actions are never called.

These properties make post actions useful in a number of scenarios. A widely applicable use case is invoking a service self-test as part of initial service provisioning.

Another example, requiring the use of post-actions, is the IP address allocation scenario from the chapter introduction. By its nature, the allocation or assignment call produces a side effect in an external system: it marks the assigned IP address in use. The same is true for releasing the address. Since NSO doesn't know how to reverse these effects on its own, they can't be part of any `create()` callback. Instead, the API calls can be implemented as post-actions.

The following snippet of a plan outline defines a `create` and `delete` post-action to handle IP management:

```
      ncs:state "ncs:init" {
        ncs:create {
          ncs:post-action-node "$SERVICE" {
            ncs:action-name "allocate-ip";
            ncs:sync;
          }
        }
      }
      ncs:state "vr:ip-allocated" {
        ncs:delete {
          ncs:post-action-node "$SERVICE" {
            ncs:action-name "release-ip";
          }
        }
      }
```

Let's see how this plan manifests during provisioning. After the first (`init`) state is reached and committed, it fires off an allocation action on the service instance, called `allocate-ip`. The job of the `allocate-ip` action is to communicate with the external system, the IP Address Management (IPAM), and allocate an address for the service instance. This process may take a while, however, it does not tie up NSO, since it runs outside of the configuration transaction and other configuration sessions can proceed in the meantime.

The `$SERVICE` XPath variable is automatically populated by the system and allows you to easily reference the service instance. There are other automatic variables defined. You can find the complete list inside the `tailf-ncs-plan.yang` submodule, in the `$NCS_DIR/src/ncs/yang/` folder.

Due to the `ncs:sync` statement, service provisioning can continue only after the allocation process (the action) completes. Once that happens, the service resumes processing in the `ip-allocated` state, with the IP value now available for configuration.

On service deprovisioning, the back-tracking mechanism works backwards through the states. When it is the ip-allocated state's turn to deprovision, NSO reverts any configuration done as part of this state, and then runs the `release-ip` action, defined inside the `ncs:delete` block. Of course, this only happens if the state previously had a reached status. Implemented as a post-action, `release-ip` can safely use the external IPAM API to deallocate the IP address, without impacting other sessions.

The actions, as defined in the example, do not take any parameters. When needed, you may pass additional parameters from the service's `opaque` and `component_proplist` object. These parameters must be set in advance, for example in some previous create callback. For details, please refer to the YANG definition of `post-action-input-params` in the `tailf-ncs-plan.yang` file.

### Multiple and Dynamic Plan Components <a href="#d5e9827" id="d5e9827"></a>

The discussion on basic concepts briefly mentions the role of a nano behavior tree but it does not fully explore its potential. Let's now consider in which situations you may find a non-trivial behavior tree beneficial.

Suppose that you are implementing a service that requires not one but two VMs. While you can always add more states to the component, these states are processed sequentially. However, you might want to provision the two VMs in parallel, since they take a comparatively long time, and it makes little sense having to wait until the first one is finished before starting with the second one. Nano services provide an elegant solution to this challenge in the form of multiple plan components: provisioning of each VM can be tracked by a separate plan component, allowing the two to advance independently, in parallel.

If the two VMs go through the same states, you can use a single component type in the plan outline for both. It is the job of the behavior tree to create or synthesize actual components for each service instance. Therefore, you could use a behavior tree similar to the following example:

```
ncs:service-behavior-tree multirouter-servicepoint {
  description "A 2-VM behavior tree";
  ncs:plan-outline-ref "vr:multirouter-plan";
  ncs:selector {
    ncs:create-component "'vm1'" {
      ncs:component-type-ref "vr:router-vm";
    }
    ncs:create-component "'vm2'" {
      ncs:component-type-ref "vr:router-vm";
    }
  }
}
```

The two `ncs:create-component` statements instruct NSO to create two components, named `vm1` and `vm2`, of the same `vr:router-vm` type. Note the required use of single quotes around component names, because the value is actually an XPath expression. The quotes ensure the name is used verbatim when the expression is evaluated.

With multiple components in place, the implicit `self` component reflects the cumulative status of the service. The `ready` state of the `self` component will never have its status set to `reached` until all other components have the `ready` state status set to `reached` and all post-actions have been run, too. Likewise, during backtracking, the `init` state will never be set to `not-reached` until all other components have been fully backtracked and all delete post actions have been run. Additionally, the `self` `ready` or `init` state status will be set to `failed` if any other state has a `failed` status or a failed post-action, thus signaling that something has failed while executing the service instance.

As you can see, all the `ncs:create-component` statements are placed inside an `ncs:selector` block. A selector is a so-called control flow node. It selects a group of components and allows you to decide whether they are created or not, based on a pre-condition. The pre-condition can reference a service parameter, which in turn controls if the relevant components are provisioned for this service instance. The mechanism enables you to dynamically produce just the necessary plan components.

The pre-condition is not very useful on the top selector node, but selectors can also be nested. For example, having a `use-virtual-devices` configuration leaf in the service YANG model, you could modify the behavior tree to the following:

```
ncs:service-behavior-tree multirouter-servicepoint {
  description "A conditional 2-VM behavior tree";
  ncs:plan-outline-ref "vr:multirouter-plan";
  ncs:selector {
    ncs:create-component "'router'" { ... }
    ncs:selector {
      ncs:pre-condition {
        ncs:monitor "$SERVICE" {
          ncs:trigger-expr "use-virtual-devices = 'true'";
        }
      }
      ncs:create-component "'vm1'" { ... }
      ncs:create-component "'vm2'" { ... }
    }
  }
}
```

The described behavior tree always synthesizes the `router` component and evaluates the child selector. However, the child selector only synthesizes the two VM components if the service configuration requested so by setting the `use-virtual-devices` to `true`.

What is more, if the pre-condition value changes, the system re-evaluates the behavior tree and starts the backtracking operation for any removed components.

For even more complex cases, where a variable number of components needs to be synthesized, the `ncs:multiplier` control flow node becomes useful. Its `ncs:foreach` statement selects a set of elements and each element is processed in the following way:

* If the optional `when` statement is not satisfied, the element is skipped.
* All `variable` statements are evaluated as XPath expressions for this element, to produce a unique name for the component and any other element-specific values.
* All `ncs:create-component` and other control flow nodes are processed, creating the necessary components for this element.

The multiplier node is often used to create a component for each item in a list. For example, if the service model contains a list of VMs, with a key `name`, then the following code creates a component for each of the items:

```
ncs:multiplier {
  ncs:foreach "vms" {
    ncs:variable "NAME" {
      ncs:value-expr "concat('vm-', name)";
    }
    ncs:create-component "$NAME" { ... }
  }
}
```

In this particular case, it might be possible to avoid the variable altogether, by using the expression for the `create-component` statement directly. However, defining a variable also makes it available to service `create()` callbacks.

This is extremely useful, since you can access these values, as well as the ones from the service opaque object, directly in the nano service XML templates. The opaque, especially, allows you to separate the logic in code from applying the XML templates.

## Netsim Router Provisioning Example <a href="#d5e9876" id="d5e9876"></a>

The [examples.ncs/nano-services/netsim-vrouter](https://github.com/NSO-developer/nso-examples/tree/6.7/nano-services/netsim-vrouter) folder contains a complete implementation of a service that provisions a netsim device instance, onboards it to NSO, and pushes a sample interface configuration to the device. Netsim device creation is neither instantaneous nor side-effect-free and thus requires the use of a nano service. It more closely resembles a real-world use case for nano services.

To see how the service is used through a prearranged scenario, execute the `make demo` command from the example folder. The scenario provisions and de-provisions multiple netsim devices to show different states and behaviors, characteristic of nano services.

The service, called `vrouter`, defines three component types in the `src/yang/vrouter.yang` file:

* `vr:vrouter`: A “day-0” component that creates and initializes a netsim process as a virtual router device.
* `vr:vrouter-day1`: A “day-1” component for configuring the created device and tracking NETCONF notifications.

As the name implies, the day-0 component must be provisioned before the day-1 component. Since the two provision in sequence, in general, a single component would suffice. However the components are kept separate to illustrate component dependencies.

The behavior tree synthesizes each of the components for a service instance using some service-specific names. To do so, the example defines three variables to hold different names:

```
      // vrouter name
      ncs:variable "NAME" {
        ncs:value-expr "current()/name";
      }
      // vrouter component name
      ncs:variable "D0NAME" {
        ncs:value-expr "concat(current()/name, '-day0')";
      }
      // vrouter day1 component name
      ncs:variable "D1NAME" {
        ncs:value-expr "concat(current()/name, '-day1')";
      }
```

The `vr:vrouter` (day-0) component has a number of plan states that it goes through during provisioning:

* ncs:init
* vr:requested
* vr:onboarded
* ncs:ready

The init and ready states are required as the first and last state in all components for correct overall state tracking in `ncs:self`. They have no additional logic tied to them.

The `vr:requested` state represents the first step in virtual router provisioning. While it does not perform any configuration itself (no nano-callback statement), it calls a post-action that does all the work. The following is a snippet of the plan outline for this state:

```
      ncs:state "vr:requested" {
        ncs:create {
          // Call a Python action to create and start a netsim vrouter
          ncs:post-action-node "$SERVICE" {
            ncs:action-name "create-vrouter";
            ncs:result-expr "result = 'true'";
            ncs:sync;
          }
        }
      }
```

The `create-router` action calls the Python code inside the `python/vrouter/main.py` file, which runs a couple of system commands, such as the `ncs-netsim create-device` and the `ncs-netsim start` commands. These commands do the same thing as you would if you performed the task manually from the shell.

The `vr:requested` state also has a `delete` post-action, analogous to `create`, which stops and removes the netsim device during service de-provisioning or backtracking.

Inspecting the Python code for these post actions will reveal that a semaphore is used to control access to the common netsim resource. It is needed because multiple `vrouter` instances may run the create and delete action callbacks in parallel. The Python semaphore is shared between the delete and create action processes using a Python multiprocessing manager, as the example configures the NSO Python VM to start the actions in multiprocessing mode. See [The Application Component](https://nso-docs.cisco.com/guides/development/core-concepts/pages/jwFVej1l7kjXuaM9axfN#ncs.development.pythonvm.cthread) for details.

In `vr:onboarded`, the nano Python callback function from the `main.py` file adds the relevant NSO device entry for a newly created netsim device. It also configures NSO to receive notifications from this device through a NETCONF subscription. When the NSO configuration is complete, the state transitions into the `reached` status, denoting the onboarding has completed successfully.

The `vr:vrouter` component handles so-called day-0 provisioning. Alongside this component, the `vr:vrouter-day1` component starts provisioning in parallel. During provisioning, it transitions through the following states:

* `ncs:init`
* `vr:configured`
* `vr:deployed`
* `ncs:ready`

The component reaches the `init` state right away. However, the `vr:configured` state has a precondition:

```
      ncs:state "vr:configured" {
        ncs:create {
          // Wait for the onboarding to complete
          ncs:pre-condition {
            ncs:monitor  "$SERVICE/plan/component[type='vr:vrouter']" +
                         "[name=$D0NAME]/state[name='vr:onboarded']" {
              ncs:trigger-expr "post-action-status = 'create-reached'";
            }
          }
          // Invoke a service template to configure the vrouter
          ncs:nano-callback;
        }
      }
```

Provisioning can continue only after the first component, `vr:vrouter`, has executed its `vr:onboarded` post-action. The precondition demonstrates how one component can depend on another component reaching some particular state or successfully executing a post-action.

The `vr:onboarded` post-action performs a `sync-from` command for the new device. After that happens, the `vr:configured` state can push the device configuration according to the service parameters, by using an XML template, `templates/vrouter-configured.xml`. The service simply configures an interface with a VLAN ID and a description.

Similarly, the `vr:deployed` state has its own precondition, which makes use of the `ncs:any` statement. It specifies either (any) of the two monitor statements will satisfy the precondition.

One of them checks the last received NETCONF notification contains a `link-status` value of `up` for the configured interface. In other words, it will wait for the interface to become operational.

However, relying solely on notifications in the precondition can be problematic, as the received notifications list in NSO can be cleared and would result in unintentional backtracking on a service re-deploy. For this reason, there is the other monitor statement, checking the device live-status.

Once either of the conditions is satisfied, it marks the end of provisioning. Perhaps the use of notifications in this case feels a little superficial but it illustrates a possible approach to waiting for the steady state, such as routing adjacencies to form and alike.

Altogether, the example shows how to use different nano service mechanisms in a single, complex, multistage service that combines configuration and side effects. The example also includes a Python script that uses the RESTCONF protocol to configure a service instance and monitor its provisioning status. You are encouraged to configure a service instance yourself and explore the provisioning process in detail, including service removal. Regarding removal, have you noticed how nano services can de-provision in stages, but the service instance is gone from the configuration right away?

## Zombie Services <a href="#d5e9953" id="d5e9953"></a>

By removing the service instance configuration from NSO, you start a service de-provisioning process. For an ordinary service, a stored reverse diff-set is applied, ensuring that all of the service-induced configuration is removed in the same transaction. For nano services, having a staged, multistep service delete operation, is not possible. The provisioned states must be backtracked one by one, often across multiple transactions. With the service instance deleted, NSO must track the de-provisioning progress elsewhere.

For this reason, NSO mutates a nano service instance when it is removed. The instance is transformed into a zombie service, which represents the original service that still requires de-provisioning. Once the de-provisioning is complete, with all the states backtracked, the zombie is automatically removed.

Zombie service instances are stored with their service data, their plan states, and diff-sets in a `/ncs:zombies/services` list. When a service mutates to a zombie, all plan components are set to back-tracking mode and all service pre-condition kickers are rewritten to reference the zombie service instead. Also, the nano service subsystem now updates the zombie plan states as de-provisioning progresses. You can use the `show zombies service` command to inspect the plan.

Under normal conditions, you should not see any zombies, except for the service instances that are actively de-provisioning. However, if an error occurs, the de-provisioning process will stop with an error status and a zombie will remain. With a zombie present, NSO will not allow creating the same service instance in the configuration tree. The zombie must be removed first.

After addressing the underlying problem, you can restart the de-provisioning process with the `re-deploy` or the `reactive-re-deploy` actions. The difference between the two is which user the action uses. The `re-deploy` uses the current user that initiated the action whilst the `reactive-re-deploy` action keeps using the same user that last modified the zombie service.

These zombie actions behave a bit differently than their normal service counterparts. In particular, the zombie variants perform the following steps to better serve the de-provisioning process:

1. Start a temporary transaction in which the service is reinstated (created). The service plan will have the same status as it had when it mutated.
2. Back-track plan components in a normal fashion, that is, removing device changes for states with delete pre-conditions satisfied.
3. If all components are completely back-tracked, the zombie is removed from the zombie list. Otherwise, the service and the current plan states are stored back into the zombie list, with new kickers waiting to activate the zombie when some delete pre-condition is satisfied.

In addition, zombie services support the `resurrect` action. The action reinstates the zombie back in the configuration tree as a real service, with the current plan status, and reverts plan components back from back-tracking to normal mode. It is an “undo” for a nano service delete.

In some situations, especially during nano service development, a zombie may get stuck because of a misconfigured precondition or similar issues. A re-deploy is unlikely to help in that case and you may need to forcefully remove the problematic plan component. The `force-back-track` action performs this job and allows you to backtrack to a specific state if specified. But beware that using the action avoids calling any post-actions or delete callbacks for the forcefully backtracked states, even though the recorded configuration modifications are reverted. It can and will leave your systems in an inconsistent or broken state if you are not careful.

## Using Notifications to Track the Plan and its Status <a href="#d5e9981" id="d5e9981"></a>

When a service is provisioned in stages, as nano services are, the success of the initial commit no longer indicates the service is provisioned. Provisioning may take a while and may fail later, requiring you to consult the service plan to observe the service status. This makes it harder to tell when a service finishes provisioning, for example. Fortunately, services provide a set of notifications that indicate important events in the service's life-cycle, including a successful completion. These events enable NETCONF and RESTCONF clients to subscribe to events instead of polling the plan and commit queue status.

The built-in service-state-changes NETCONF/RESTCONF stream is used by NSO to generate northbound notifications for services, including nano services. The event stream is enabled by default in `ncs.conf`, however, individual notification events must be explicitly configured to be sent.

### The `plan-state-change` Notification <a href="#d5e9986" id="d5e9986"></a>

When a service's plan component changes state, the `plan-state-change` notification is generated with the new state of the plan. It includes the status, which indicates one of not-reached, reached, or failed. The notification is sent when the state is `created`, `modified`, or `deleted`, depending on the configuration. For reference on the structure and all the fields present in the notification, please see the YANG model in the `tailf-ncs-plan.yang` file.

As a common use case, an event with status `reached` for the `self` component `ready` state signifies that all nano service components have reached their `ready` state and provisioning is complete. A simple example of this scenario is included in the [examples.ncs/nano-services/netsim-vrouter](https://github.com/NSO-developer/nso-examples/tree/6.7/nano-services/netsim-vrouter) `demo_rc.py` Python script, using RESTCONF.

To enable the plan-state-change notifications to be sent, you must enable them for a specific service in NSO. For example, can load the following configuration into the CDB as an XML initialization file:

```xml
<services xmlns="http://tail-f.com/ns/ncs">
  <plan-notifications>
    <subscription>
      <name>nano1</name>
      <service-type>/vr:vrouter</service-type>
      <component-type>self</component-type>
      <state>ready</state>
      <operation>modified</operation>
    </subscription>
    <subscription>
      <name>nano2</name>
      <service-type>/vr:vrouter</service-type>
      <component-type>self</component-type>
      <state>ready</state>
      <operation>created</operation>
    </subscription>
  </plan-notifications>
</services>
```

This configuration enables notifications for the self component's ready state when created or modified.

### The `service-commit-queue-event` Notification <a href="#d5e10003" id="d5e10003"></a>

When a service is committed through the commit queue, this notification acts as a reference regarding the state of the service. Notifications are sent when the service commit queue item is waiting to run, executing, waiting to be unlocked, completed, failed, or deleted. More details on the `service-commit-queue-event` notification content can be found in the YANG model inside `tailf-ncs-services.yang` .

For example, the `failed` event can be used to detect that a nano service instance deployment failed because a configuration change committed through the commit queue has failed. Measures to resolve the issue can then be taken and the nano service instance can be re-deployed. A simple example of this scenario is included in the [examples.ncs/nano-services/netsim-vrouter](https://github.com/NSO-developer/nso-examples/tree/6.7/nano-services/netsim-vrouter) `demo_rc.py` Python script where the service is committed through the commit queue, using RESTCONF. By design, the configuration commit to a device fails, resulting in a `commit-queue-notification` with the `failed` event status for the commit queue item.

To enable the service-commit-queue-event notifications to be sent, you can load the following example configuration into NSO, as an XML initialization file or some other way:

```xml
<services xmlns="http://tail-f.com/ns/ncs">
  <commit-queue-notifications>
    <subscription>
      <name>nano1</name>
      <service-type>/vr:vrouter</service-type>
    </subscription>
  </commit-queue-notifications>
</services>
```

### Examples of `service-state-changes` Stream Subscriptions <a href="#d5e10016" id="d5e10016"></a>

The following examples demonstrate the usage and sample events for the notification functionality, described in this section, using RESTCONF, NETCONF, and CLI northbound interfaces.

RESTCONF subscription request using `curl`:

```bash
$ curl -isu admin:admin -X GET -H "Accept: text/event-stream"
    http://localhost:8080/restconf/streams/service-state-changes/json

data: {
data:   "ietf-restconf:notification": {
data:     "eventTime": "2021-11-16T20:36:06.324322+00:00",
data:     "tailf-ncs:service-commit-queue-event": {
data:       "service": "/vrouter:vrouter[name='vr7']",
data:       "id": 1637135519125,
data:       "label": "vr7",
data:       "status": "completed",
data:     }
data:   }
data: }

data: {
data:   "ietf-restconf:notification": {
data:     "eventTime": "2021-11-16T20:36:06.728911+00:00",
data:     "tailf-ncs:plan-state-change": {
data:       "service": "/vrouter:vrouter[name='vr7']",
data:       "component": "self",
data:       "state": "tailf-ncs:ready",
data:       "operation": "modified",
data:       "status": "reached",
data:     }
data:   }
data: }
```

See [Streams](https://nso-docs.cisco.com/guides/development/core-concepts/pages/TItWBhukD9D6FJkD3eWB#ncs.northbound.restconf.streams) in Northbound APIs for further reference.

NETCONF creates subscription using `netconf-console`:

```
$ netconf-console create-subscription=service-state-changes

<?xml version="1.0" encoding="UTF-8"?>
<notification xmlns="urn:ietf:params:xml:ns:netconf:notification:1.0">
  <eventTime>2021-11-16T20:36:06.324322+00:00</eventTime>
  <service-commit-queue-event xmlns="http://tail-f.com/ns/ncs">
    <service xmlns:vr="http://com/example/vrouter">/vr:vrouter[vr:name='vr7']</service>
    <id>1637135519125</id>
    <label>vr7</label>
    <status>completed</status>
  </service-commit-queue-event>
</notification>
<?xml version="1.0" encoding="UTF-8"?>
<notification xmlns="urn:ietf:params:xml:ns:netconf:notification:1.0">
  <eventTime>2021-11-16T20:36:06.728911+00:00</eventTime>
  <plan-state-change xmlns="http://tail-f.com/ns/ncs">
    <service xmlns:vr="http://com/example/vrouter">/vr:vrouter[vr:name='vr7']</service>
    <component>self</component>
    <state>ready</state>
    <operation>modified</operation>
    <status>reached</status>
  </plan-state-change>
</notification>
```

See [Notification Capability](https://nso-docs.cisco.com/guides/development/core-concepts/pages/TItWBhukD9D6FJkD3eWB#ug.netconf_agent.notif) in Northbound APIs for further reference.

CLI shows received notifications using `ncs_cli`:

```bash
$ ncs_cli -u admin -C <<<'show notification stream service-state-changes'

notification
 eventTime 2021-11-16T20:36:06.324322+00:00
 service-commit-queue-event
  service /vrouter[name='vr7']
  id 1637135519125
  label vr7
  status completed
 !
!
notification
 eventTime 2021-11-16T20:36:06.728911+00:00
 plan-state-change
  service /vrouter[name='vr7']
  component self
  state ready
  operation modified
  status reached
 !
!
```

### The `label` in the Notification <a href="#d5e10037" id="d5e10037"></a>

You have likely noticed the `label` field in the example notifications above. The `label` is an optional but very useful parameter when committing the service configuration; it helps you correlate events from the commit in the `service-state-changes` stream notifications. The above notifications, taken from the [examples.ncs/nano-services/netsim-vrouter](https://github.com/NSO-developer/nso-examples/tree/6.7/nano-services/netsim-vrouter) example, are emitted after applying a RESTCONF plain patch:

```
$ curl -isu admin:admin -X PATCH
  -H "Content-type: application/yang-data+json"
  'http://localhost:8080/restconf/data?commit-queue=sync&label=vr7'
  -d '{ "vrouter:vrouter": [ { "name": "vr7" } ] }'
```

Note that the `label` is specified as part of the URL.

## Developing and Updating a Nano Service <a href="#d5e10046" id="d5e10046"></a>

At times, especially when you use an iterative development approach or simply due to changing requirements, you might need to update (change) an existing nano service and its implementation. In addition to other service update best practices, such as model upgrades, you must carefully consider the nano-service-specific aspects. The following discussion mostly focuses on migrating an already provisioned service instance to a newer version; however, the same concepts also apply while you are initially developing the service.

In the simple case, updating the model of a nano service and getting the changes to show up in an already created instance is a matter of executing a normal re-deploy. This will synthesize any new components and provision them, along with the new configuration, just like you would expect from a non-nano service.

A major difference occurs if a service instance is deleted and is in a zombie state when the nano service is updated. You should be aware that no synthetization is done for that service instance. The only goal of a deleted service is to revert any changes made by the service instance. Therefore, in that case, the synthetization is not needed. It means that, if you've made changes to callbacks, post-actions, or pre-conditions, those changes will not be applied to zombies of the nano service. If a service instance requires the new changes to be applied, you must re-deploy it before it is deleted.

When updating nano services, you also need to be aware that any old callbacks, post actions and any other models that the service depends on, need to be available in the new nano service package until all service instances created before the update have either been updated (through a re-deploy) or fully deleted. Therefore, you must take great care with any updates to a service if there are still zombies left in the system.

### Adding Components <a href="#d5e10052" id="d5e10052"></a>

Adding new components to the behavior tree will create the new components during the next re-deploy (synthetization) and execute the states in the new components as is normally done.

### Removing Components <a href="#d5e10055" id="d5e10055"></a>

When removing components from the behavior tree, the components that are removed are set to backtracking and are backtracked fully before they are removed from the plan.

When you remove a component, do so carefully so that any callbacks, post actions or any other model data that the component depends on are not removed until all instances of the old component are removed.

If the identity for a component type is removed, then NSO removes the component from the database when upgrading the package. If this happens, the component is not backtracked and the reverse diffsets are not applied.

### Replacing Components <a href="#d5e10060" id="d5e10060"></a>

Replacing components in the behavior tree is the same as having unrelated components that are deleted and added in the same update. The deleted components are backtracked as far as possible, and then the added components are created and their states executed in order.

In some cases, this is not the desired behavior when replacing a component. For example, if you only want to rename a component, backtracking and then adding the component again might make NSO push unnecessary changes to the network or run delete callbacks and post actions that should not be run. To remedy this, you might add the `ncs:deprecates-component` statements to the new component, detailing which components it replaces. NSO then skips the backtracking of the old component and just applies all reverse diffsets of the deprecated component. In the same re-deploy, it then executes the new component as usual. Therefore, if the new component produces the same configuration as the old component, nothing is pushed to the network.

If any of the deprecated components are backtracking, the backtracking will be handled before the component is removed. When there are multiple components that are deprecated in the same update, the components will not be removed, as detailed above, until all of them are done backtracking (if any one of them are backtracking).

### Adding and Removing States <a href="#d5e10066" id="d5e10066"></a>

When adding or removing states in a component, the component is backtracked before a new component with the new states is added and executed. If the updated component produces the same configuration as the old one (and no preconditions halt the execution), this should lead to no configuration being pushed to the network. So, if changes to the states are done, you need to take care when writing the preconditions and post actions for a component if no new changes should be pushed to the network.

Any changes to the already present states that are kept in the updated component will not have their configuration updated until the new component is created, which happens after the old one has been fully backtracked.

### Modifying States <a href="#d5e10070" id="d5e10070"></a>

For a component where only the configuration for one or more states have changed, the synthetization process will update the component with the new configuration and make sure that any new callbacks or similar are called during future execution of the component.

## Implementation Reference <a href="#ug.nano_services.impl" id="ug.nano_services.impl"></a>

The text in this section sums up as well as adds additional detail on the way nano services operate, which you will hopefully find beneficial during implementation.

To reiterate, the purpose of a nano service is to break down an RFM service into its isolated steps. It extends the normal `ncs:servicepoint` YANG mechanism and requires the following:

* A YANG definition of the service input parameters, with a service point name and the additional nano-plan-data grouping.
* A YANG definition of the plan component types and their states in a plan outline.
* A YANG definition of a behavior tree for the service. The behavior tree defines how and when to instantiate components in the plan.
* Code or templates for individual state transfers in the plan.

When a nano service is committed, the system evaluates its behavior tree. The result of this evaluation is a set of components that form the current plan for the service. This set of components is compared with the previous plan (before the commit). If there are new components, they are processed one by one.

For each component in the plan, it is executed state by state in the defined order. Before entering a new state, the create pre-condition for the state is evaluated if it exists. If a create pre-condition exists and if it is not satisfied, the system stops progressing this component and jumps to the next one. A kicker is then defined for the pre-condition that was not satisfied. Later, when this kicker triggers and the pre-condition is satisfied, it performs a `reactive-re-deploy` and the kicker is removed. This kicker mechanism becomes a self-sustained RFM loop.

If a state's pre-conditions are met, the callback function or template associated with the state is invoked, if it exists. If the callback is successful, the state is marked as `reached`, and the next state is executed.

A component, that is no longer present but was in the previous plan, goes into back-tracking mode, during which the goal is to remove all reached states and eventually remove the component from the plan. Removing state data changes is performed in a strict reverse order, beginning with the last reached state and taking into account a delete pre-condition if defined.

A nano service is expected to have a component. All components are expected to have `ncs:init` as its first state and `ncs:ready` as its last state. A component-type can have any number of specific states in between `ncs:init` and `ncs:ready`.

### Back-Tracking <a href="#d5e10107" id="d5e10107"></a>

Back-tracking is completely automatic and occurs in the following scenarios:

* **State pre-condition not satisfied**: A `reached` state's pre-condition is no longer satisfied, and there are subsequent states that are reached and contain reverse diff-sets.
* **Plan component is removed**: When a plan component is removed and has reached states that contain reverse diff-sets.
* **Service is deleted**: When a service is deleted, NSO will set all plan components to back-tracking mode before deleting the service.

For each RFM loop, NSO traverses each component and state in order. For each non-satisfied create pre-condition, a kicker is started that monitors and triggers when the pre-condition becomes satisfied.

<div data-with-frame="true"><figure><img src="/files/xGvTj5CACAMqKd4owvrm" alt="" width="375"><figcaption></figcaption></figure></div>

While traversing the states, a `create` pre-condition that was previously satisfied may become unsatisfied. If there are subsequent reached states that contain reverse diff-sets, then the component must be set to back-tracking mode. The back-tracking mode has as its goal to revert all changes up to the state that originally failed to satisfy its `create` pre-condition. While back-tracking, the delete pre-condition for each state is evaluated, if it exists. If the delete pre-condition is satisfied, the state's reverse diff-set is applied, and the next state is considered. If the delete pre-condition is not satisfied, a kicker is created to monitor this delete pre-condition. When the kicker triggers, a `reactive-re-deploy` is called and the back-tracking will continue until the goal is reached.

<div data-with-frame="true"><figure><img src="/files/yaaj3WteK2YKhjvaABNp" alt="" width="375"><figcaption></figcaption></figure></div>

When the back-tracking plan component has reached its goal state, the component is set to normal mode again. The state's create pre-condition is evaluated and if it is satisfied the state is entered or otherwise a kicker is created as described above.

<div data-with-frame="true"><figure><img src="/files/dt4QM6he2QoC4STBTv0x" alt="" width="375"><figcaption></figcaption></figure></div>

In some circumstances, a complete plan component is removed (for example, if the service input parameters are changed). If this happens, the plan component is checked if it contains reached states that contain reverse diff-sets.

<div data-with-frame="true"><figure><img src="/files/UqwQsAhv42BLHjuoF3b3" alt="" width="375"><figcaption></figcaption></figure></div>

If the removed component contains reached states with reverse diff-sets, the deletion of the component is deferred and the component is set to back-tracking mode.

<div data-with-frame="true"><figure><img src="/files/AoapX9EtWDayFLpqJ2wL" alt="" width="375"><figcaption></figcaption></figure></div>

In this case, there is no specified goal state for the back-tracking. This means that when all the states have been reverted, the component is automatically deleted.

<div data-with-frame="true"><figure><img src="/files/m6pjbGf86oEA8MWi2r2A" alt="" width="375"><figcaption></figcaption></figure></div>

If a service is deleted, all components are set to back-tracking mode. The service becomes a zombie, storing away its plan states so that the service configuration can be removed.

All components of a deleted service are set in backtracking mode.

<div data-with-frame="true"><figure><img src="/files/JDJIFr8fiA05o7TEuo0b" alt="" width="375"><figcaption></figcaption></figure></div>

When a component becomes completely back-tracked, it is removed.

<div data-with-frame="true"><figure><img src="/files/WEOd3VQdZdFfhiNFH6ec" alt="" width="375"><figcaption></figcaption></figure></div>

When all components in the plan are deleted, the service is removed.

<div data-with-frame="true"><figure><img src="/files/OgHiLOEFUBUDHoKvTcoi" alt="" width="375"><figcaption></figcaption></figure></div>

### Behavior Tree <a href="#d5e10173" id="d5e10173"></a>

A nano service behavior tree is a data structure defined for each service type. Without a behavior tree defined for the service point, the nano service cannot execute. It is the behavior tree that defines the currently executing nano-plan with its components.

{% hint style="info" %}
This is in stark contrast to plan-data used for logging purposes where the programmer needs to write the plan and its components in the `create()` callback. For nano services, it is not allowed to define the nano plan in any other way than by a behavior tree.
{% endhint %}

The purpose of a behavior tree is to have a declarative way to specify how the service's input parameters are mapped to a set of component instances.

A behavior tree is a directed tree in which the nodes are classified as control flow nodes and execution nodes. For each pair of connected nodes, the outgoing node is called parent and the incoming node is called child. A control flow node has zero or one parent and at least one child and the execution nodes have one parent and no children.

There is exactly one special control flow node called the root, which is the only control flow node without a parent.

This definition implies that all interior nodes are control flow nodes, and all leaves are execution nodes. When creating, modifying, or deleting a nano service, NSO evaluates the behavior tree to render the current nano plan for the service. This process is called synthesizing the plan.

The control flow nodes have a different behavior, but in the end, they all synthesize its children in zero or more instances. When the a control flow node is synthesized, the system executes its rules for synthesizing the node's children. Synthesizing an execution node adds the corresponding plan component instance to the nano service's plan.

All control flow and execution nodes may define pre-conditions, which must be satisfied to synthesize the node. If a pre-condition is not satisfied, a kicker is started to monitor the pre-condition.

All control flow and execution nodes may define an observe monitor which results in a kicker being started for the monitor when the node is synthesized.

If an invocation of an RFM loop (for example, a re-deploy) synthesizes the behavior tree and a pre-condition for a child is no longer satisfied, the sub-tree with its plan-components is removed (that is, the plan-components are set to back-tracking mode).

The following control flow nodes are defined:

* **Selector**: A selector node has a set of children which are synthesized as described above.
* **Multiplier**: A multiplier has a 'foreach\_'\_ mechanism that produces a list of elements. For each resulting element, the children are synthesized as described above. This can be used, for example, to create several plan-components of the same type.

There is just one type of execution node:

* **Create component**: The create-component execution node creates an instance of the component type that it refers to in the plan.

It is recommended to keep the behavior tree as flat as possible. The most trivial case is when the behavior tree creates a static nano-plan, that is, all the plan-components are defined and never removed. The following is an example of such a behavior tree:

<div data-with-frame="true"><figure><img src="/files/BsK9cqoXDnyV29Eht4TR" alt="" width="563"><figcaption><p>Behavior Tree with a Static nano-plan</p></figcaption></figure></div>

Having a selector on root implies that all plan-components are created if they don't have any pre-conditions, or for which the pre-conditions are satisfied.

An example of a more elaborated behavior tree is the following:

<div data-with-frame="true"><figure><img src="/files/2qZkocTCL8lqxvfeHT1m" alt="" width="563"><figcaption><p>Elaborated Behavior Tree</p></figcaption></figure></div>

This behavior tree has a selector node as the root. It will always synthesize the "base-config" plan component and then evaluate then pre-condition for the selector child. If that pre-condition is satisfied, it then creates four other plan-components.

The multiplier control flow node is used when a plan component of a certain type should be cloned into several copies depending on some service input parameters. For this reason, the multiplier node defines a `foreach`, a `when`, and a `variable`. The `foreach` is evaluated and for each node in the nodeset that satisfies the `when`, the `variable` is evaluated as the outcome. The value is used for parameter substitution to a unique name for a duplicated plan component.

<div data-with-frame="true"><figure><img src="/files/seyI80Wd1wp09PaVtUBn" alt="" width="563"><figcaption></figcaption></figure></div>

The value is also added to the nano service opaque which enables the individual state nano service `create()` callbacks to retrieve the value.

Variables might also have “when” expressions, which are used to decide if the variable should be added to the list of variables or not.

### Nano Service Pre-Condition <a href="#d5e10234" id="d5e10234"></a>

Pre-conditions are what drive the execution of a nano service. A pre-condition is a prerequisite for a state to be executed or a component to be synthesized. If the pre-condition is not satisfied, it is then turned into a kicker which in turn re-deploys the nano service once the condition is fulfilled.

When working with pre-conditions, you need to be aware that they work a bit differently when used as a kicker to redeploy the service and when they are used in the execution of the service. When the pre-condition is used in the re-deploy kicker, it then works as explained in the kicker documentation (that is, the trigger expression is evaluated before and after the change-set of the commit when the monitored nodeset is changed). When used during the execution of a nano service, you can only evaluate it on the current state of the database, which means that it only checks that the monitor returns a nodeset of one or more nodes and that trigger expression (if there is one) is fulfilled for any of the nodes in the nodeset.

Support for pre-conditions checking, if a node has been deleted, is handled a bit differently due to the difference in how the pre-condition is evaluated. Kickers always trigger for changed nodes (add, deleted, or modified) and can check that the node was deleted in the commit that triggered the kicker. While in the nano service evaluation, you only have the current state of the database and the monitor expression will not return any nodes for evaluation of the trigger expression, consequently evaluating the pre-condition to false. To support deletes in both cases, you can create a pre-condition with a monitor expression and a child node `ncs:trigger-on-delete` which then both create a kicker that checks for deletion of the monitored node and also does the right thing in the nano service evaluation of the pre-condition. For example, you could have the following component:

```
            ncs:component "base-config" {
              ncs:state "init" {
                ncs:delete {
                  ncs:pre-condition {
                    ncs:monitor "/devices/device[name='test']" {
                      ncs:trigger-on-delete;
                    }
                  }
                }
              }
              ncs:state "ready";
            }
```

The component would only trigger the init states delete pre-condition when the device named test is deleted.

It is possible to add multiple monitors to a pre-condition by using the `ncs:all` or `ncs:any` extensions. Both extensions take one or multiple monitors as argument. A pre-condition using the `ncs:all` extension is satisfied if all monitors given as arguments evaluate to true. A pre-condition using the `ncs:any` extension is satisfied if at least one of the monitors given as argument evaluates to true. The following component uses the `ncs:all` and `ncs:any` extensions for its self state's create and delete pre-condition, respectively:

```
          ncs:component "base-config" {
            ncs:state "init" {
              ncs:create {
                ncs:pre-condition {
                  ncs:all {
                    ncs:monitor $SERVICE/syslog {
                      ncs:trigger-expr: "current() = true"
                    }
                    ncs:monitor $SERVICE/dns {
                      ncs:trigger-expr: "current() = true"
                      }
                    }
                  }
                }
              }
              ncs:delete {
                ncs:pre-condition {
                  ncs:any {
                    ncs:monitor $SERVICE/syslog {
                      ncs:trigger-expr: "current() = false"
                    }
                    ncs:monitor $SERVICE/dns {
                      ncs:trigger-expr: "current() = false"
                      }
                    }
                  }
                }
              }
            }
            ncs:state "ready";
          }
```

### Nano Service Opaque and Component Properties <a href="#d5e10252" id="d5e10252"></a>

The service opaque is a name-value list that can optionally be created/modified in some of the service callbacks, and then travels the chain of callbacks (pre-modification, create, post-modification). It is returned by the callbacks and stored persistently in the service private data. Hence, the next service invocation has access to the current opaque and can make subsequent read/write operations to the same object. The object is usually called `opaque` in Java and `proplist` in Python callbacks.

The nano services handle the opaque in a similar fashion, where a callback for every state has access to and can modify the opaque. However, the behavior tree can also define variables, which you can use in preconditions or to set component names. These variables are also available in the callbacks, as component properties. The mechanism is similar but separate from the opaque. While the opaque is a single service-instance-wide object set only from the service code, component variables are set in and scoped according to the behavior tree. That is, component properties contain only the behavior tree variables which are in scope when a component is synthesized.

For example, take the following behavior tree snippet:

```
     ncs:selector {
      ncs:variable "VAR1" {
        ncs:value-expr "'value1'";
      }
      ncs:create-component "'base-config'" {
        ncs:component-type-ref "t:base-config";
      }
      ncs:selector {
        ncs:variable "VAR2" {
          ncs:value-expr "'value2'";
        }
        ncs:create-component "'component1'" {
          ncs:component-type-ref "t:my-component";
        }
      }
    }
```

The callbacks for states in the `“base-config”` component only see the `VAR1` variable, while those in “component1” see both `VAR1` and `VAR2` as component properties.

Additionally, both the service opaque and component variables (properties) are used to look up substitutions in nano service XML templates and in the behavior tree. If used in the behavior tree, the same rules apply for the opaque as for component variables. So, a value needs to contain single quotes if you wish to use it verbatim in preconditions and similar constructs, for example:

```
proplist.append(('VARX', "'some value'"))
```

Using this scheme at an early state, such as the `“base-config”` component's `“ncs:init”`, you can have a callback that sets name-value pairs for all other states that are then implemented solely with templates and preconditions.

### Nano Service Callbacks <a href="#ug.nano_services.callbacks" id="ug.nano_services.callbacks"></a>

The nano service can have several callback registrations, one for each plan component state. But note that some states may have no callbacks at all. The state may simply act as a checkpoint, that some condition is satisfied, using pre-condition statements. A component's `ncs:ready` state is a good example of this.

The drawback with this flexible callback registration is that there must be a way for the NSO Service Manager to know if all expected nano service callbacks have been registered. For this reason, all nano service plan component states that require callbacks are marked with this information. When the plan is executed and the callback markings in the plan mismatch with the actual registrations, this results in an error.

All callback registrations in NSO require a daemon to be instantiated, such as a Python or Java process. For nano services, it is allowed to have many daemons where each daemon is responsible for a subset of the plan state callback registrations. The neat thing here is that it becomes possible to mix different callback types (Template/Python/Java) for different plan states.

<div data-with-frame="true"><figure><img src="/files/hHfAziKCHfaJwIV7vL4W" alt="" width="563"><figcaption></figcaption></figure></div>

The mixed callback feature caters to the case where most of the callbacks are templates and only some are Java or Python. This works well because nano services try to resolve the template parameters using the nano service opaque when applying a template. This is a unique functionality for nano services that makes Java or Python apply-template callbacks unnecessary.

You can implement nano service callbacks as Templates as well as Python, Java, Erlang, and C code. The following examples cover the implementation of Template, Python and Java.

A plan state template, if defined, replaces the need of a `create()` callback. In this case, there are no `delete()` callbacks and the status definitions must in this case be handled by the states delete pre-condition. The template must in addition to the `servicepoint` attribute, have a `componenttype` and a `state` attribute to be registered on the plan state:

```xml
<config-template xmlns="http://tail-f.com/ns/config/1.0"
                 servicepoint="my-servicepoint"
                 componenttype="my:some-component"
                 state="my:some-state">
  <devices xmlns="http://tail-f.com/ns/ncs">
    <!-- ... -->
  </devices>
</config-template>
```

Specific to nano services, you can use parameters, such as `$SOMEPARAM` in the template. The system searches for the parameter value in the service opaque and in the component properties. If it is not defined, applying the template will fail.

A Python `create()` callback is very similar to its ordinary service counterpart. The difference is that it has additional arguments. `plan` refers to the synthesized plan, while `component` and `state` specify the component and state for which it is invoked. The `proplist` argument is the nano service opaque (same naming as for ordinary services) and `component_proplist` contains component variables, along with their values.

```python
class NanoServiceCallbacks(ncs.application.NanoService):

    @ncs.application.NanoService.create
    def cb_nano_create(self, tctx, root, service, plan, component, state,
                       proplist, component_proplist):
        ...

    @ncs.application.NanoService.delete
    def cb_nano_delete(self, tctx, root, service, plan, component, state,
                       proplist, component_proplist):
        ...
```

In the majority of cases, you should not need to manage the status of nano states yourself. However, should you need to override the default behavior, you can set the status explicitly, in the callback, using code similar to the following :

```
plan.component[component].state[state].status = 'failed'
```

The Python nano service callback needs a registration call for the specific service point, `componentType`, and state that it should be invoked for.

```python
class Main(ncs.application.Application):

    def setup(self):
        ...
        self.register_nano_service('my-servicepoint',
                                   'my:some-component',
                                   'my:some-state',
                                   NanoServiceCallbacks)
```

For Java, annotations are used to define the callbacks for the component states. The registration of these callbacks is performed by the ncs-java-vm. The `NanoServiceContext` argument contains methods for retrieving the component and state for the invoked callback as well as methods for setting the resulting plan state status.

```java
public class myRFS {

    @NanoServiceCallback(servicePoint="my-servicepoint",
                         componentType="my:some-component",
                         state="my:some-state",
                         callType=NanoServiceCBType.CREATE)
    public Properties createSomeComponentSomeState(
                                    NanoServiceContext context,
                                    NavuNode service,
                                    NavuNode ncsRoot,
                                    Properties opaque,
                                    Properties componentProperties)
                                    throws DpCallbackException {
        // ...
    }

    @NanoServiceCallback(servicePoint="my-servicepoint",
                         componentType="my:some-component",
                         state="my:some-state",
                         callType=NanoServiceCBType.DELETE)
    public Properties deleteSomeComponentSomeState(
                                    NanoServiceContext context,
                                    NavuNode service,
                                    NavuNode ncsRoot,
                                    Properties opaque,
                                    Properties componentProperties)
                                    throws DpCallbackException {
        // ...
    }
```

Several `componentType` and state callbacks can be defined in the same Java class and are then registered by the same `daemon`.

#### Generic Service Callbacks <a href="#d5e10312" id="d5e10312"></a>

In some scenarios, there is a need to be able to register a callback for a certain state in several components with different component types. For this reason, it is possible to register a callback with a wildcard, using “\*” as the component type. The invoked state sends the actual component name to the callback, allowing the callback to still distinguish component types if required.

In Python, the component type is provided as an argument to the callback (`component`) and a generic callback is registered with an asterisk for a component, such as:

```python
self.register_nano_service('my-servicepoint', '*', state, ServiceCallbacks)
```

In Java, you can perform the registration in the method annotation, as before. To retrieve the calling component type, use the `NanoServiceContext.getComponent()` method. For example:

```java
    @NanoServiceCallback(servicePoint="my-servicepoint",
                         componentType="*", state="my:some-state",
                         callType=NanoServiceCBType.CREATE)
    public Properties genericNanoCreate(NanoServiceContext context,
                                        NavuNode service,
                                        NavuNode ncsRoot,
                                        Properties opaque,
                                        Properties componentProperties)
                                        throws DpCallbackException {

        String currentComponent = context.getComponent();
        // ...
    }
```

The generic callback can then act for the registered state in any component type.

#### Nano Service Pre/Post Modifications <a href="#d5e10323" id="d5e10323"></a>

The ordinary service pre/post modification callbacks still exist for nano services. They are registered as for an ordinary service and are invoked before the behavior tree synthetization and after the last component/state invocation.

Registration of the ordinary `create()` will not fail for a nano service. But they will never be invoked.

### Forced Commits <a href="#d5e10328" id="d5e10328"></a>

When implementing a nano service, you might end up in a situation where a commit is needed between states in a component to make sure that something has happened before the service can continue executing. One example of such behavior is if the service is dependent on the notifications from a device. In such a case, you can set up a notification kicker in the first state and then trigger a forced commit before any later states can proceed, therefore making sure that all future notifications are seen by the later states of the component.

To force a commit in between two states of a component, add the `ncs:force-commit` tag in a `ncs:create` or `ncs:delete` tag. See the following example:

```
              ncs:component "base-config" {
                ncs:state "init" {
                  ncs:create {
                    ncs:force-commit;
                  }
                }
                ncs:state "ready" {
                  ncs:delete {
                    ncs:force-commit;
                  }
                }
              }
```

### Plan Location <a href="#d5e10337" id="d5e10337"></a>

When defining a nano service, it is assumed that the plan is stored under the service path, as `ncs:plan-data` is added to the service definition. When the service instance is deleted, the plan is moved to the zombie instead, since the instance has been removed and the plan cannot be stored under it anymore. When writing other services or when working with a nano service in general, you need to be aware that the plan for a service might be in one of these two places depending on if the service instance has been deleted or not.

To make it easier to work with a service, you can define a custom location for the plan and its history. In the `ncs:service-behaviour-tree`, you can specify that the plan should be stored outside of the service by setting the `ncs:plan-location` tag to a custom location. The location where the plan should be stored must be either a list or a container and include the `ncs:plan-data` tag. The plan data is then created in this location, no matter if the service instance has been deleted (turned into a zombie) or not, making it easy to base decisions on the state of the service as all plan queries can query the same plan.

You can use XPath with the `ncs:plan-location` statement. The XPath is evaluated based on the nano service context. When the list or container, which contains the plan, is nested under another list, the outer list instance must exist before creating the nano service. At the same time, the outer list instance of the plan location must also remain intact for further service's life-cycle management, such as redeployment, deletion, etc. Otherwise, an error will be returned and logged, and any service interaction (create, re-deploy, delete, etc.) won't succeed.

{% code title="Nano services custom plan location example" %}

```
    identity base-config {
    base ncs:plan-component-type;
  }
  
  list custom {
    description "Custom plan location example service.";

    key name;
    leaf name {
      tailf:info "Unique service id";
      tailf:cli-allow-range;
      type string;
    }

    uses ncs:service-data;
    ncs:servicepoint custom-plan-servicepoint;
  }

  list custom-plan {
    description "Custom plan location example plan.";

    key name;
    leaf name {
      tailf:info "Unique service id";
      tailf:cli-allow-range;
      type string;
    }

    uses ncs:nano-plan-data;
  }

  ncs:plan-outline custom-plan {
    description
      "Custom plan location example outline";

    ncs:component-type "p:base-config" {
      ncs:state "ncs:init";
      ncs:state "ncs:ready";
    }
  }

  ncs:service-behavior-tree custom-plan-location-servicepoint {
    description
      "Custom plan location example service behaviour three.";

    ncs:plan-outline-ref custom:custom-plan;
    ncs:plan-location "/custom-plan";

    ncs:selector {
      ncs:create-component "'base-config'" {
        ncs:component-type-ref "p:base-config";
      }
    }
  }
```

{% endcode %}

### Nano Services and Commit Queue

The commit queue feature, described in [Commit Queue](https://nso-docs.cisco.com/guides/development/core-concepts/pages/auKQMOAF2p1jiGYJBweP#user_guide.devicemanager.commit-queue), allows for increased overall throughput of NSO by committing configuration changes into an outbound queue item instead of directly to affected devices. Nano services are aware of the commit queue and will make use of it, however, this interaction requires additional consideration.

When the commit queue is enabled and there are outstanding commit queue items, the network is lagging behind the CDB. The CDB is forward-looking and shows the desired state of the network. Hence, the nano plan shows the desired state as well, since changes to reach this state may not have been pushed to the devices yet.

To keep the convergence of the nano service in sync with the commit queue, nano services behave more asynchronously:

* A nano service state does not make any progression while the service has an outstanding commit queue item. The outstanding item is listed under `plan/commit-queue` for the service, in normal or in zombie mode.
* On completion of the commit queue item, the nano plan comes in sync with the network. The outstanding commit queue item is removed from the list above and the system issues a `reactive-re-deploy` action to resume the progression of the nano service.
* Post-actions are delayed, while there is an outstanding commit queue item.
* Deleting a nano service always (even without a commit queue) creates a zombie and schedules its re-deploy to perform backtracking. Again, the re-deploy and, consequently, removal will not take place while there is an outstanding commit queue item.

The reason for such behavior is that commit queue items can fail. In case of a failure, the CDB and the network have diverged. In turn, the nano plan may have diverged and not reflect the actual network state if the failed commit queue item contained changes related to the nano service.

What is worse, the network may be left in an inconsistent state. To counter that, NSO supports multiple recovery options for the commit queue. Since NSO release 5.7, using the `rollback-on-error` is the recommended option, as it undoes all the changes that are part of the same transaction. If the transaction includes the initial service instance creation, the instance is removed as well. That is usually not desired for nano services. A nano service will avoid such removal by only committing the service intent (the instance configuration) in the initial transaction. In this case, the service avoids potential rollback, as it does not perform any device configuration in the same transaction but progresses solely through (reactive) re-deploy.

While error recovery helps keeping the network consistent, the end result remains that the requested change was not deployed. If a commit queue item with nano service-related changes fails, that signifies a failure for the nano service and NSO does the following:

* Service progression stops.
* The nano plan is marked as failed by creating the `failed` leaf under the plan.
* The scheduled post-actions are canceled. Canceled post actions stay in the `side-effect-queue` with status `canceled` and are not going to be executed.

After such an event, manual intervention is required. If not using the `rollback-on-error` option or the rollback transaction fails, consult [Commit Queue](https://nso-docs.cisco.com/guides/development/core-concepts/pages/auKQMOAF2p1jiGYJBweP#user_guide.devicemanager.commit-queue) for the correct procedure to follow. Once the cause of the commit queue failure is resolved, you can manually resume the service progression by invoking the `reactive-re-deploy` action on a nano service or a zombie.

The `service-commit-queue-event` helps detect that a nano service instance deployment failed because a configuration change committed through the commit queue has failed. See [The service-commit-queue-event Notification](#d5e10003) section for details.

## Graceful Link Migration Example <a href="#d5e10385" id="d5e10385"></a>

You can find another nano service example under [examples.ncs/nano-services/link-migration](https://github.com/NSO-developer/nso-examples/tree/6.7/nano-services/link-migration). The example illustrates a situation with a simple VPN link that should be set up between two devices. The link is considered established only after it is tested and a `test-passed` leaf is set to `true`. If the VPN link changes, the new endpoints must be set up before removing the old endpoints, to avoid disturbing customer traffic during the operation.

The package named `link` contains the nano service definition. The service has a list containing at most one element, which constitutes the VPN link and is keyed on a-device a-interface b-device b-interface. The list element corresponds to a component type `link:vlan-link` in the nano service plan.

{% code title="Example: Link Migration Example Plan" %}

```
  identity vlan-link {
    base ncs:plan-component-type;
  }

  identity dev-setup {
    base ncs:plan-state;
  }

  ncs:plan-outline link:link-plan {
    description
      "Make before brake vlan plan";

    ncs:component-type "link:vlan-link" {
      ncs:state "ncs:init";
      ncs:state "link:dev-setup" {
        ncs:create {
          ncs:nano-callback;
        }
      }
      ncs:state "ncs:ready" {
        ncs:create {
          ncs:pre-condition {
            ncs:monitor "$SERVICE/endpoints" {
              ncs:trigger-expr "test-passed = 'true'";
            }
          }
        }
        ncs:delete {
          ncs:pre-condition {
            ncs:monitor "$SERVICE/plan" {
              ncs:trigger-expr
                "component[type = 'vlan-link'][back-track = 'false']"
              + "/state[name = 'ncs:ready'][status = 'reached']"
              + " or not(component[back-track = 'false'])";
            }
          }
        }
      }
    }
  }
```

{% endcode %}

In the plan definition, note that there is only one nano service callback registered for the service. This callback is defined for the `link:dev-setup` state in the `link:vlan-link` component type. In the plan, it is represented as follows:

```
        ncs:state "link:dev-setup" {
          ncs:create {
            ncs:nano-callback;
          }
        }
```

The callback is a template. You can find it under packages/link/templates as `link-template.xml`.

For the state `ncs:ready` in the `link:vlan-link` component type there are both a `create` and a `delete` pre-condition. The `create` pre-condition for this state is as follows:

```
        ncs:create {
          ncs:pre-condition {
            ncs:monitor "$SERVICE/endpoints" {
              ncs:trigger-expr "test-passed = 'true'";
            }
          }
        }
```

This pre-condition implies that the components based on this component type are not considered finished until the `test-passed` leaf is set to a `true` value. The pre-condition implements the requirement that after the initial setup of a link configured by the `link:dev-setup` state, a manual test and setting of the `test-passed` leaf is performed before the link is considered finished.

The `delete` pre-condition for the same state is as follows:

```
        ncs:delete {
          ncs:pre-condition {
            ncs:monitor "$SERVICE/plan" {
              ncs:trigger-expr
                "component[type = 'vlan-link'][back-track = 'false']"
              + "/state[name = 'ncs:ready'][status = 'reached']"
              + " or not(component[back-track = 'false'])";
            }
          }
        }
```

This pre-condition implies that before you start deleting (back-tracking) an old component, the new component must have reached the `ncs:ready` state, that is, after being successfully tested. The first part of the pre-condition checks the status of the `vlan-link` components. Since there can be at most one link configured in the service instance, the only non-backtracking component, other than self, is the new link component. However, that condition on its own prevents the component to be deleted when deleting the service. So, the second part, after the `or` statement, checks if all components are back-tracking, which signifies service deletion. This approach illustrates a "create-before-break" scenario where the new link is created first, and only when it is set up, the old one is removed.

{% code title="Example: Link Migration Example Behavior Tree" %}

```
  ncs:service-behavior-tree link-servicepoint {
    description
      "Make before brake vlan example";

    ncs:plan-outline-ref "link:link-plan";

    ncs:selector {
      ncs:multiplier {
        ncs:foreach "endpoints" {
          ncs:variable "VALUE" {
            ncs:value-expr "concat(a-device, '-', a-interface,
                                   '-', b-device, '-', b-interface)";
          }
        }
        ncs:create-component "$VALUE" {
          ncs:component-type-ref "link:vlan-link";
        }
      }
    }
```

{% endcode %}

The `ncs:service-behavior-tree` is registered on the servicepoint `link-servicepoint` that is defined by the nano service. It refers to the plan definition named `link:link-plan`. The behavior tree has a selector on top, which chooses to synthesize its children depending on their pre-conditions. In this tree, there are no pre-conditions, so all children will be synthesized.

The `multiplier` control node chooses a node set. A variable named `VALUE` is created with a unique value for each node in that node-set and creates a component of the `link:vlan-link` type for each node in the chosen node-set. The name for each individual component is the value of the variable `VALUE`.

Since the chosen node-set is the "endpoints" list that can contain at most one element, it produces only one component. However, if the link in the service is changed, that is, the old list entry is deleted and a new one is created, then the multiplier creates a component with a new name.

This forces the old component (which is no longer synthesized) to be back-tracked and the plan definition above handles the "create-before-break" behavior of the back-tracking.

To run the example, do the following:

Build the example:

```bash
$ cd examples.ncs/nano-services/link-migration
$ make all
```

Start the example:

```bash
$ cd ncs-netsim restart
$ ncs
```

Run the example:

```bash
$ ncs_cli -C -u admin
admin@ncs(config)# devices sync-from
sync-result {
    device ex0
    result true
}
sync-result {
    device ex1
    result true
}
sync-result {
    device ex2
    result true
}
admin@ncs(config)# config
Entering configuration mode terminal
```

Now you create a service that sets up a VPN link between devices `ex1` and `ex2`, and is completed immediately since the `test-passed` leaf is set to `true`.

```bash
admin@ncs(config)# link t2 unit 17 vlan-id 1
admin@ncs(config-link-t2)# link t2 endpoints ex1 eth0 ex2 eth0 test-passed true
admin@ncs(config-endpoints-ex1/eth0/ex2/eth0)# commit
admin@ncs(config-endpoints-ex1/eth0/ex2/eth0)# top
```

You can inspect the result of the commit:

```cli
admin@ncs(config)# exit
admin@ncs# link t2 get-modifications
cli  devices {
          device ex1 {
              config {
                  r:sys {
                      interfaces {
                          interface eth0 {
     +                        unit 17 {
     +                            vlan-id 1;
     +                        }
                          }
                      }
                  }
              }
          }
          device ex2 {
              config {
                  r:sys {
                      interfaces {
                          interface eth0 {
     +                        unit 17 {
     +                            vlan-id 1;
     +                        }
                          }
                      }
                  }
              }
          }
      }
```

The service sets up the link between the devices. Inspect the plan:

```cli
admin@ncs# show link t2 plan component * state * status
NAME               STATE      STATUS
---------------------------------------
self               init       reached
                   ready      reached
ex1-eth0-ex2-eth0  init       reached
                   dev-setup  reached
                   ready      reached
```

All components in the plan have reached their `ready` state.

Now, change the link by changing the interface on one of the devices. To do this, you must remove the old list entry in "endpoints" and create a new one.

```bash
admin@ncs# config
Entering configuration mode terminal
admin@ncs(config)# no link t2 endpoints ex1 eth0 ex2 eth0
admin@ncs(config)# link t2 endpoints ex1 eth0 ex2 eth1
```

Commit a dry-run to inspect what happens:

```cli
admin@ncs(config-endpoints-ex1/eth0/ex2/eth1)# commit dry-run
cli  devices {
         device ex1 {
             config {
                 r:sys {
                     interfaces {
                         interface eth0 {
                         }
                     }
                 }
             }
         }
         device ex2 {
             config {
                 r:sys {
                     interfaces {
    +                    interface eth1 {
    +                        unit 17 {
    +                            vlan-id 1;
    +                        }
    +                    }
                     }
                 }
             }
         }
     }
     link t2 {
    -    endpoints ex1 eth0 ex2 eth0 {
    -        test-passed true;
    -    }
    +    endpoints ex1 eth0 ex2 eth1 {
    +    }
     }
```

Upon committing, the service just adds the new interface and does not remove anything at this point. The reason is that the `test-passed` leaf is not set to `true` for the new component. Commit this change and inspect the plan:

```bash
admin@ncs(config-endpoints-ex1/eth0/ex2/eth1)# commit
admin@ncs(config-endpoints-ex1/eth0/ex2/eth1)# top
admin@ncs(config)# exit
admin@ncs# show link t2 plan
                                                                   ...
                              BACK                                 ...
NAME               TYPE       TRACK  GOAL  STATE      STATUS       ...
-------------------------------------------------------------------...
self               self       false  -     init       reached      ...
                                           ready      reached      ...
ex1-eth0-ex2-eth1  vlan-link  false  -     init       reached      ...
                                           dev-setup  reached      ...
                                           ready      not-reached  ...
ex1-eth0-ex2-eth0  vlan-link  true   -     init       reached      ...
                                           dev-setup  reached      ...
                                           ready      reached      ...
```

Notice that the new component `ex1-eth0-ex2-eth1` has not reached its `ready` state yet. Therefore, the old component `ex1-eth0-ex2-eth0` still exists in back-track mode but is still waiting for the new component to finish.

If you check what the service has configured at this point, you get the following:

```cli
admin@ncs# link t2 get-modifications
cli  devices {
          device ex1 {
              config {
                  r:sys {
                      interfaces {
                          interface eth0 {
     +                        unit 17 {
     +                            vlan-id 1;
     +                        }
                          }
                      }
                  }
              }
          }
          device ex2 {
              config {
                  r:sys {
                      interfaces {
                          interface eth0 {
     +                        unit 17 {
     +                            vlan-id 1;
     +                        }
                          }
     +                    interface eth1 {
     +                        unit 17 {
     +                            vlan-id 1;
     +                        }
     +                    }
                      }
                  }
              }
          }
      }
```

Both the old and the new link exist at this point. Now, set the `test-passed` leaf to `true` to force the new component to reach its ready state.

```bash
admin@ncs(config)# link t2 endpoints ex1 eth0 ex2 eth1 test-passed true
admin@ncs(config-endpoints-ex1/eth0/ex2/eth1)# commit
```

If you now check the service plan, you see the following:

```bash
admin@ncs(config-endpoints-ex1/eth0/ex2/eth1)# top
admin@ncs(config)# exit
admin@ncs# show link t2 plan
                                                               ...
                              BACK                             ...
NAME               TYPE       TRACK  GOAL  STATE      STATUS   ...
---------------------------------------------------------------...
self               self       false  -     init       reached  ...
                                           ready      reached  ...
ex1-eth0-ex2-eth1  vlan-link  false  -     init       reached  ...
                                           dev-setup  reached  ...
                                           ready      reached  ...
```

The old component has been completely backtracked and is removed because the new component is finished. You should also check the service modifications. You should see that the old link endpoint is removed:

```cli
admin@ncs# link t2 get-modifications
cli  devices {
          device ex1 {
              config {
                  r:sys {
                      interfaces {
                          interface eth0 {
     +                        unit 17 {
     +                            vlan-id 1;
     +                        }
                          }
                      }
                  }
              }
          }
          device ex2 {
              config {
                  r:sys {
                      interfaces {
     +                    interface eth1 {
     +                        unit 17 {
     +                            vlan-id 1;
     +                        }
     +                    }
                      }
                  }
              }
          }
      }
```


# Packages

Run user code in NSO using packages.

All user code that needs to run in NSO must be part of a package. A package is basically a directory of files with a fixed file structure. A package consists of code, YANG modules, custom Web UI widgets, etc., that are needed to add an application or function to NSO. Packages are a controlled way to manage the loading and versions of custom applications.

A package is a directory where the package name is the same as the directory name. At the top level of this directory, a file called `package-meta-data.xml` must exist. The structure of that file is defined by the YANG model `$NCS_DIR/src/ncs/yang/tailf-ncs-packages.yang`. A package may also be a tar archive with the same directory layout. The tar archive can be either uncompressed with the suffix `.tar`, or gzip-compressed with the suffix `.tar.gz` or `.tgz`. The archive file should also follow some naming conventions. There are two acceptable naming conventions for archive files, one is that after the introduction of CDM in the NSO 5.1, it can be named by `ncs-<ncs-version>-<package-name>-<package-version>.<suffix>`, e.g. `ncs-5.3-my-package-1.0.tar.gz` and the other is `<package-name>-<package-version>.<suffix>`, e.g. `my-package-1.0.tar.gz`.

* `package-name`: should use letters, and digits and may include underscores (`_`) or dashes (`-`), but no additional punctuation, and digits can not follow underscores or dashes immediately.
* `package-version`: should use numbers and dot (`.`).

<div data-with-frame="true"><figure><img src="/files/ZOtGZfuCsXuALxpcqo9p" alt="" width="563"><figcaption><p>Package Model</p></figcaption></figure></div>

Packages are composed of components. The following types of components are defined: NED, Application, and Callback.

The file layout of a package is:

```xml
           <package-name>/package-meta-data.xml
                    load-dir/
                    shared-jar/
                    private-jar/
                    webui/
                    templates/
                    src/
                    doc/
                    netsim/
```

The `package-meta-data.xml` defines several important aspects of the package, such as the name, dependencies on other packages, the package's components, etc. This will be thoroughly described later in this section.

When NSO starts, it needs to search for packages to load. The `ncs.conf` parameter `/ncs-config/load-path` defines a list of directories. At initial startup, NSO searches these directories for packages and copies the packages to a private directory tree in the directory defined by the `/ncs-config/state-dir` parameter in `ncs.conf`, and loads and starts all the packages found. All .fxs (compiled YANG files) and .ccl (compiled CLI spec files) files found in the directory `load-dir` in a package are loaded. On subsequent startups, NSO will by default only load and start the copied packages - see [Loading Packages](/guides/development/advanced-development/developing-packages#loading-packages) for different ways to get NSO to search the load path for changed or added packages.

A package usually contains Java code. This Java code is loaded by a class loader in the NSO Java VM. A package that contains Java code must compile the Java code so that the compilation results are divided into .jar files where code, that is supposed to be shared among multiple packages, is compiled into one set of .jar files, and code that is private to the package itself is compiled into another set of .jar files. The shared and the common jar files shall go into the `shared-jar` directory and the `private-jar` directory, respectively. By putting for example the code for a specific service in a private jar, NSO can dynamically upgrade the service without affecting any other service.

The optional `webui` directory contains the WEB UI customization files.

## An Example Package <a href="#d5e4949" id="d5e4949"></a>

The NSO example collection for contains a number of small self-contained examples. The collection resides at `$NCS_DIR/examples.ncs` Each of these examples defines a package. Let's take a look at some of these packages. The example [examples.ncs/device-management/aggregated-stats](https://github.com/NSO-developer/nso-examples/tree/6.7/device-management/aggregated-stats) has a package `./packages/stats`. The `package-meta-data.xml` file for that package looks like this:

{% code title="An Example Package" %}

```xml
<ncs-package xmlns="http://tail-f.com/ns/ncs-packages">
  <name>stats</name>
  <package-version>1.0</package-version>
  <description>Aggregating statistics from the network</description>
  <ncs-min-version>3.0</ncs-min-version>
  <required-package>
    <name>router-nc-1.0</name>
  </required-package>
  <component>
    <name>stats</name>
    <callback>
      <java-class-name>com.example.stats.Stats</java-class-name>
    </callback>
  </component>
</ncs-package>
```

{% endcode %}

The file structure in the package looks like this:

```
|----package-meta-data.xml
|----private-jar
|----shared-jar
|----src
|    |----Makefile
|    |----yang
|    |    |----aggregate.yang
|    |----java
|         |----build.xml
|         |----src
|              |----com
|                   |----example
|                        |----stats
|                             |----namespaces
|                             |----Stats.java
|----doc
|----load-dir
```

## The `package-meta-data.xml` File <a href="#d5e4962" id="d5e4962"></a>

The `package-meta-data.xml` file defines the name of the package, additional settings, and one component. Its settings are defined by the `$NCS_DIR/src/ncs/yang/tailf-ncs-packages.yang` YANG model, where the *package* list name gets renamed to `ncs-package`. See the `tailf-ncs-packages.yang` module where all options are described in more detail. To get an overview, use the IETF RFC 8340-based YANG tree diagram.

```bash
$ yanger -f tree tailf-ncs-packages.yang
```

```
submodule: tailf-ncs-packages (belongs-to tailf-ncs)
  +--ro packages
     +--ro package* [name] <-- renamed to "ncs-package" in package-meta-data.xml
        +--ro name                      string
        +--ro package-version           version
        +--ro display-name?             string
        +--ro description?              string
        +--ro ncs-min-version*          version
        +--ro ncs-max-version*          version
        +--ro single-sign-on-url?       string
        +--ro python-package!
        |  +--ro vm-name?           string
        |  +--ro callpoint-model?   enumeration
        +--ro directory?                string
        +--ro templates*                string
        +--ro template-loading-mode?    enumeration
        +--ro supported-ned-id*         union
        +--ro supported-ned-id-match*   string
        +--ro required-package* [name]
        |  +--ro name           string
        |  +--ro min-version?   version
        |  +--ro max-version?   version
        +--ro component* [name]
           +--ro name                 string
           +--ro description?         string
           +--ro entitlement-tag?     string
           +--ro (type)
              +--:(ned)
              |  +--ro ned
              |     +--ro (ned-type)
              |     |  +--:(netconf)
              |     |  |  +--ro netconf
              |     |  |     +--ro ned-id?   identityref
              |     |  +--:(snmp)
              |     |  |  +--ro snmp
              |     |  |     +--ro ned-id?   identityref
              |     |  +--:(cli)
              |     |  |  +--ro cli
              |     |  |     +--ro ned-id             identityref
              |     |  |     +--ro java-class-name    string
              |     |  +--:(generic)
              |     |     +--ro generic
              |     |        +--ro ned-id                 identityref
              |     |        +--ro java-class-name        string
              |     |        +--ro management-protocol?   string
              |     +--ro device
              |     |  +--ro vendor              string
              |     |  +--ro product-family*     string
              |     |  +--ro operating-system*   string
              |     +--ro option* [name]
              |        +--ro name     string
              |        +--ro value?   string
              +--:(upgrade)
              |  +--ro upgrade
              |     +--ro (type)
              |        +--:(java)
              |        |  +--ro java-class-name?     string
              |        +--:(python)
              |           +--ro python-class-name?   string
              +--:(callback)
              |  +--ro callback
              |     +--ro java-class-name*   string
              +--:(application)
                 +--ro application
                    +--ro (type)
                    |  +--:(java)
                    |  |  +--ro java-class-name      string
                    |  +--:(python)
                    |     +--ro python-class-name    string
                    +--ro start-phase?               enumeration
```

{% hint style="info" %}
The order of the XML entries in a `package-meta-data.xml` must be in the same order as the model shown above.
{% endhint %}

A sample package configuration is taken from the [examples.ncs/nano-services/netsim-vrouter](https://github.com/NSO-developer/nso-examples/tree/6.7/nano-services/netsim-vrouter) example:

```bash
$ ncs_load -o -Fp -p /packages
```

```xml
<config xmlns="http://tail-f.com/ns/config/1.0">
  <packages xmlns="http://tail-f.com/ns/ncs">
    <package>
      <name>router-nc-1.1</name>
      <package-version>1.1</package-version>
      <description>Generated netconf package</description>
      <ncs-min-version>5.7</ncs-min-version>
      <directory>./state/packages-in-use/1/router</directory>
      <component>
        <name>router</name>
        <ned>
          <netconf>
            <ned-id xmlns:router-nc-1.1="http://tail-f.com/ns/ned-id/router-nc-1.1">
            router-nc-1.1:router-nc-1.1</ned-id>
          </netconf>
          <device>
            <vendor>Acme</vendor>
          </device>
        </ned>
      </component>
      <oper-status>
        <up/>
      </oper-status>
    </package>
    <package>
      <name>vrouter</name>
      <package-version>1.0</package-version>
      <description>Nano services netsim virtual router example</description>
      <ncs-min-version>5.7</ncs-min-version>
      <python-package>
        <vm-name>vrouter</vm-name>
        <callpoint-model>threading</callpoint-model>
      </python-package>
      <directory>./state/packages-in-use/1/vrouter</directory>
      <templates>vrouter-configured</templates>
      <template-loading-mode>strict</template-loading-mode>
      <supported-ned-id xmlns:router-nc-1.1="http://tail-f.com/ns/ned-id/router-nc-1.1">
      router-nc-1.1:router-nc-1.1</supported-ned-id>
      <required-package>
        <name>router-nc-1.1</name>
        <min-version>1.1</min-version>
      </required-package>
      <component>
        <name>nano-app</name>
        <description>Nano service callback and post-actions example</description>
        <application>
          <python-class-name>vrouter.nano_app.NanoApp</python-class-name>
          <start-phase>phase2</start-phase>
        </application>
      </component>
      <oper-status>
        <up/>
      </oper-status>
    </package>
  </packages>
</config>
```

Below is a brief list of the configurables in the `tailf-ncs-packages.yang` YANG model that applies to the metadata file. A more detailed description can be found in the YANG model:

* `name` - the name of the package. All packages in the system must have unique names.
* `package-version` - the version of the package. This is for administrative purposes only, NSO cannot simultaneously handle two versions of the same package.
* `ncs-min-version` - the oldest known NSO version where the package works.
* `ncs-max-version` - the latest known NSO version where the package works.
* `python-package` - Python-specific package data.
  * `vm-name` - the Python VM name for the package. Default is the package `vm-name`. Packages with the same `vm-name` run in the same Python VM. Applicable only when `callpoint-model = threading`.
  * `callpoint-model` - A Python package runs Services, Nano Services, and Actions in the same OS process. If the `callpoint-model` is set to `multiprocessing` each will get a separate worker process. Running Services, Nano Services, and Actions in parallel can, depending on the application, improve the performance at the cost of complexity. See [The Application Component](https://nso-docs.cisco.com/guides/development/core-concepts/pages/jwFVej1l7kjXuaM9axfN#ncs.development.pythonvm.cthread) for details.
* `directory` - the path to the directory of the package.
* `templates` - the templates defined by the package.
* `template-loading-mode` - control if the templates are interpreted in strict or relaxed mode.
* `supported-ned-id` - the list of ned-ids supported by this package. An example of the expected format taken from the [examples.ncs/nano-services/netsim-vrouter](https://github.com/NSO-developer/nso-examples/tree/6.7/nano-services/netsim-vrouter) example:

  ```xml
  <supported-ned-id xmlns:router-nc-1.1="http://tail-f.com/ns/ned-id/router-nc-1.1">
  router-nc-1.1:router-nc-1.1</supported-ned-id>
  ```
* `supported-ned-id-match` - the list of regular expressions for ned-ids supported by this package. Ned-ids in the system that matches at least one of the regular expressions in this list are added to the `supported-ned-id` list. The following example demonstrates how all minor versions with a major number of 1 of the `router-nc` NED can be added to a package's list of supported ned-ids:

  ```xml
  <supported-ned-id-match>router-nc-1.\d+:router-nc-1.\d+</supported-ned-id-match>
  ```
* `required-package` - a list of names of other packages that are required for this package to work.
* `component` - Each package defines zero or more components.

## Components <a href="#d5e5042" id="d5e5042"></a>

Each component in a package has a name. The names of all the components must be unique within the package. The YANG model for packages contains:

```
....
list component {
  key name;
  leaf name {
    type string;
  }
  ...
  choice type {
    mandatory true;
    case ned {
      ...
    }
    case callback {
      ...
    }
    case application {
      ...
    }
    case upgrade {
      ...
    }
    ....
  }
  ....
```

Lots of additional information can be found in the YANG module itself. The mandatory choice that defines a component must be one of `ned`, `callback`, `application`, or `upgrade`.

### Component Types

#### **NED**

A Network Element Driver component is used southbound of NSO to communicate with managed devices (described in [Network Element Drivers (NEDs](/guides/development/advanced-development/developing-neds)). The easiest NED to understand is the NETCONF NED which is built into NSO.

There are four different types of NEDs:

* **NETCONF**: used for NETCONF-enabled devices such as Juniper routers, ConfD-powered devices, or any device that speaks proper NETCONF and also has YANG models. Plenty of packages in the NSO example collection have NETCONF NED components, for example, [examples.ncs/device-management/router-network](https://github.com/NSO-developer/nso-examples/tree/6.7/device-management/router-network) under `packages/router`.
* **SNMP**: Used for SNMP devices.

  The example [examples.ncs/device-management/snmp-ned](https://github.com/NSO-developer/nso-examples/tree/6.7/device-management/snmp-ned) has a package that has an SNMP NED component.
* **CLI**: used for CLI devices. The [examples.ncs/device-management/cli-ned](https://github.com/NSO-developer/nso-examples/tree/6.7/device-management/cli-ned) example has a package called `router-cli-1.0` that defines a NED component of type CLI.
* **Generic**: used for generic NED devices. The example [examples.ncs/device-management/generic-xmlrpc-ned](https://github.com/NSO-developer/nso-examples/tree/6.7/device-management/generic-xmlrpc-ned) has a package called `xml-rpc` which defines a NED component of type generic.

A CLI NED and a generic NED component must also come with additional user-written Java code, whereas a NETCONF NED and an SNMP NED have no Java code.

#### Callback

This defines a component with one or many Java classes that implement callbacks using the Java callback annotations.

If we look at the components in the `stats` package above, we have:

```xml
  <component>
    <name>stats</name>
    <callback>
      <java-class-name>
        com.example.stats.Stats
      </java-class-name>
    </callback>
  </component>
```

The `Stats` class here implements a read-only data provider. See [DP API](https://nso-docs.cisco.com/guides/development/core-concepts/pages/Uzy6qvKpLQF47FSwk0S2#ug.java_api_overview.dp).

The `callback` type of component is used for a wide range of callback-type Java applications, where one of the most important are the Service Callbacks. The following list of Java callback annotations applies to callback components.

* `ServiceCallback` to implement service-to-device mappings. See the example: [examples.ncs/service-management/rfs-service](https://github.com/NSO-developer/nso-examples/tree/6.7/service-management/rfs-service) See [Developing NSO Services](/guides/development/advanced-development/developing-services) for a thorough introduction to services.
* `ActionCallback` to implement user-defined `tailf:actions` or YANG RPC and actions. See the examples: [examples.ncs/sdk-api/actions-python](https://github.com/NSO-developer/nso-examples/tree/6.7/sdk-api/actions-py) and [examples.ncs/sdk-api/actions-java](https://github.com/NSO-developer/nso-examples/tree/6.7/sdk-api/actions-java).
* `DataCallback` to implement the data getters and setters for a data provider. See the example [examples.ncs/device-management/aggregated-stats](https://github.com/NSO-developer/nso-examples/tree/6.7/device-management/aggregated-stats).
* `TransCallback` to implement the transaction portions of a data provider callback. See the example [examples.ncs/device-management/aggregated-stats](https://github.com/NSO-developer/nso-examples/tree/6.7/device-management/aggregated-stats).
* `DBCallback` to implement an external database. See the example: [examples.ncs/sdk-api/external-db](https://github.com/NSO-developer/nso-examples/tree/6.7/sdk-api/external-db).
* `SnmpInformResponseCallback` to implement an SNMP listener - See the example [examples.ncs/device-management/snmp-notification-receiver](https://github.com/NSO-developer/nso-examples/tree/6.7/device-management/snmp-notification-receiver).
* `TransValidateCallback`*,* `ValidateCallback` to implement a user-defined validation hook that gets invoked on every commit.
* `AuthCallback` to implement a user hook that gets called whenever a user is authenticated by the system.
* `AuthorizationCallback` to implement an authorization hook that allows/disallows users to do operations and/or access data. Note, that this callback should normally be avoided since, by nature, invoking a callback for any operation and/or data element is a performance impairment.

A package that has a `callback` component usually has some YANG code and then also some Java code that relates to that YANG code. By convention, the YANG and the Java code reside in a src directory in the component. When the source of the package is built, any resulting `fxs` files (compiled YANG files) must reside in the `load-dir` of package and any resulting Java compilation results must reside in the `shared-jar` and `private-jar` directories. Study the [examples.ncs/device-management/aggregated-stats](https://github.com/NSO-developer/nso-examples/tree/6.7/device-management/aggregated-stats) example to see how this is achieved.

#### Application

Used to cover Java applications that do not fit into the callback type. Typically this is functionality that should be running in separate threads and work autonomously.

The example [examples.ncs/sdk-api/cdb-java](https://github.com/NSO-developer/nso-examples/tree/6.7/sdk-api/cdb-java) contains three components that are of type `application`. These components must also contain a `java-class-name` element. For application components, that Java class must implement the `ApplicationComponent` Java interface.

#### Upgrade

Used to migrate data for packages where the yang model has changed and the automatic CDB upgrade is not sufficient. The upgrade component consists of a Java class with a main method that is expected to run one time only.

The example [examples.ncs/service-management/upgrade-service](https://github.com/NSO-developer/nso-examples/tree/6.7/service-management/upgrade-service) illustrates user CDB upgrades using `upgrade` components.

## Creating Packages <a href="#ug.packages.creating" id="ug.packages.creating"></a>

NSO ships with a tool `ncs-make-package` that can be used to create packages. [Package Development](/guides/development/advanced-development/developing-packages) discusses in depth how to develop a package.

### Creating a NETCONF NED Package <a href="#d5e5156" id="d5e5156"></a>

This use case applies if we have a set of YANG files that define a managed device. If we wish to develop an EMS solution for an existing device *and* that device has YANG files and also speaks NETCONF, we need to create a package for that device to be able to manage it. Assuming all YANG files for the device are stored in `./acme-router-yang-files`, we can create a package for the router as:

```bash
  $ ncs-make-package --netconf-ned ./acme-router-yang-files acme
  $ cd acme/src; make
```

The above command will create a package called `acme` in `./acme`. The `acme` package can be used for two things; managing real `acme` routers, and as input to the `ncs-netsim` tool to simulate a network of `acme` routers.

In the first case, managing real acme routers, all we need to do is to put the newly generated package in the load-path of NSO, start NSO with package reload (see [Loading Packages](/guides/development/advanced-development/developing-packages#loading-packages)), and then add one or more acme routers as managed devices to NSO. The `ncs-setup` tool can be used to do this:

```bash
 $ ncs-setup --ned-package ./acme --dest ./ncs-project
```

The above command generates a directory `./ncs-project` which is suitable for running NSO. Assume we have an existing router at the IP address `10.2.3.4` and that we can log into that router over the NETCONF interface using the username `bob`, and password `secret`. The following session shows how to set up NSO to manage this router:

```bash
 $ cd ./ncs-project
 $ ncs
 $ ncs_cli -u admin
 > configure
 > set devices authgroups group southbound-bob umap admin \
        remote-name bob remote-password secret
 > set devices device acme1 authgroup southbound-bob address 10.2.3.4
 > set devices device acme1 device-type netconf
 > commit
```

We can also use the newly generated `acme` package to simulate a network of `acme` routers. During development, this is especially useful. The `ncs-netsim` tool can create a simulated network of `acme` routers as:

```bash
 $ ncs-netsim create-network ./acme 5 a --dir ./netsim
 $ ncs-netsim start
DEVICE a0 OK STARTED
DEVICE a1 OK STARTED
DEVICE a2 OK STARTED
DEVICE a3 OK STARTED
DEVICE a4 OK STARTED
 $
```

Finally, `ncs-setup` can be used to initialize an environment where NSO is used to manage all devices in an `ncs-netsim` network:

```bash
 $ ncs-setup --netsim-dir ./netsim --dest ncs-project
```

### Creating an SNMP NED Package <a href="#d5e5190" id="d5e5190"></a>

Similarly, if we have a device that has a set of MIB files, we can use `ncs-make-package` to generate a package for that device. An SNMP NED package can, similarly to a NETCONF NED package, be used to both manage real devices and also be fed to `ncs-netsim` to generate a simulated network of SNMP devices.

Assuming we have a set of MIB files in `./mibs`, we can generate a package for a device with those mibs as:

```bash
 $ ncs-make-package --snmp-ned ./mibs acme
 $ cd acme/src; make
```

### Creating a CLI NED Package or a Generic NED Package <a href="#d5e5199" id="d5e5199"></a>

For CLI NEDs and Generic NEDs, we cannot (yet) generate the package. Probably the best option for such packages is to start with one of the examples. A good starting point for a CLI NED is the [examples.ncs/device-management/cli-ned](https://github.com/NSO-developer/nso-examples/tree/6.7/device-management/cli-ned) and a good starting point for a Generic NED is the example [examples.ncs/device-management/generic-xmlrpc-ned](https://github.com/NSO-developer/nso-examples/tree/6.7/device-management/generic-xmlrpc-ned).

### Creating a Service Package or a Data Provider Package <a href="#d5e5204" id="d5e5204"></a>

The `ncs-make-package` can be used to generate empty skeleton packages for a data provider and a simple service. The flags `--service-skeleton` and `--data-provider-skeleton`.

Alternatively, one of the examples can be modified to provide a good starting point. For example [examples.ncs/service-management/rfs-service](https://github.com/NSO-developer/nso-examples/tree/6.7/service-management/rfs-service).


# Transactions

Describe and understand transactions.

Similar to relational database systems, transactions are at the core of NSO. Transactions ensure consistency and avoid race conditions where simultaneous access by multiple clients could result in data corruption. NSO even extends the concept of transactions to the network, enabling better error recovery during provisioning and simplifying network automation.

Transactions are used for both reads and writes: for reads, they provide a consistent view that does not change half-way through the transaction; for writes, they guarantee all-or-nothing behavior, so partial updates are never persisted. NSO transactions provide ([ACID](https://en.wikipedia.org/wiki/ACID)) properties.

When dealing with operational data in NSO, you can choose to use transactions or not. However, all configuration changes for managed devices and the NSO system itself must go through transactions.

A normal NSO transaction has multiple phases:

* **Creation**: A new transaction is created (opened). This can be a result of an API call, a northbound interface (NBI) request, or some internal process.
* **Reading and writing data**: Also called the "work" phase, this is where the client can safely access the data. If the transaction was created as a read-write one (and not a read-only one), the client can make the desired changes. These changes are temporary and local to the transaction; if the transaction is aborted or client disconnects, the changes are lost. Some of the changes are validated as they are made and client can request a full validation. In the end, all changes must satisfy all model constraints.
* **Commit/Apply**: To persist the changes, the client commits the transaction. NSO ensures the changes are valid and enacts the changes on the affected systems.
* **Closing**: The transaction is closed and cannot be used any longer. Transaction-specific resources are freed.

A transaction generally only advances forward through these phases, not back; for example, once it is committed, you must start a new one to make further changes.

## Transaction Lifetime

Transactions can start explicitly or implicitly. They are meant to be relatively short-lived, since they take up resources.

You can start a transaction explictly with a MAAPI call, such as `ncs.maapi.single_write_trans()` / `com.tailf.maapi.Maapi.startTrans()`, or with JSON-RPC `new_trans`. Each transaction gets a unique identifier, called `th` (for transaction handle). You must close the transaction when you are done, e.g. by using the `with` statement (Python) or calling `Maapi.finishTrans()` (Java) / `delete_trans` (JSON-RPC).

NSO also starts and closes transactions implicitly in certain situations, for example:

* For configure mode in the CLI.
* Advancing a nano service with redeploy.
* For each NBI RESTCONF operation, such as a `PATCH` request.
* For each NBI NETCONF operation, such as `edit-config`.

{% hint style="info" %}
As an exception, NSO recognizes the proprietary `:transactions` NETCONF capability, used in LSA setups and netsim devices, which defines a separate `start-transaction` NETCONF operation and allows explicit management of transaction lifecycle.
{% endhint %}

A CLI configure session can live for a relatively long time and so can the transaction (though a new transaction is started after every successful commit). If a transaction runs into a conflict with other concurrent transactions, NSO tries to restart it from a checkpoint. In case that is not possible, the transaction must restart with the latest datastore state and reapply all the changes. The CLI keeps track of the entered commands and tries to reapply them automatically, but other clients (e.g. a Python script) must [retry](https://nso-docs.cisco.com/guides/development/core-concepts/pages/8PJW3J1uuhcu5ttTWYkl#ncs.development.concurrency.handling) with a new transaction.

## Transaction Commit

When a client commits (applies) a transaction, it starts a complex process in NSO that needs to ensure valid and consistent data across concurrent transactions and multiple systems (managed devices). This process goes through multiple stages, as shown by the progress trace (e.g. using `commit | details` in the CLI). The detailed output breaks up the transaction into four distinct phases:

1. validate phase
2. write-start phase
3. prepare phase
4. commit phase

These phases deal with how the network-wide transactions work:

The validate phase prepares and validates the new configuration (including NSO copy of device configurations), then the CDB processes the changes and prepares them for local storage in the write-start phase.

The prepare phase sends out the changes to the network through the Device Manager and the HA system. The changes are staged and validated. This process uses the candidate data store if the device supports it. Otherwise, the changes are activated immediately.

If all systems acknowledge that they have received the new configuration successfully, enter the commit phase, marking the new NSO configuration as active and activating or committing the staged configuration on remote devices. Otherwise, enter the abort phase, discarding changes, and ask NEDs to revert activated changes on devices that do not support transactions (e.g. without candidate data store).

<figure><img src="/files/e8qirIaYO2jPQcaMwKsn" alt="" width="375"><figcaption><p>Typical Transaction Phases</p></figcaption></figure>

There are also two types of locks involved with the transaction that are of interest to the NSO developer; the service write lock and the transaction lock. Transaction lock is an exclusive (global) lock that is required to serialize all the transactions in NSO (e.g. to check for conflicts, persist data to disk etc.). [NSO Concurrency Model](/guides/development/core-concepts/nso-concurrency-model) provides details on conflict checking and [Scaling and Performance Optimization](/guides/development/advanced-development/scaling-and-performance-optimization) a discussion of the lock's impact on performance.

Service lock is a per-service-type lock for serializing services to minimize conflicts. See [Services Deep Dive](/guides/development/advanced-development/developing-services/services-deep-dive) to learn more.

{% hint style="info" %}
Note that transaction lock is distinct from [global running-datastore locks](/guides/administration/advanced-topics/locks), which are managed by the client agents and not tied to the transaction lifetime. For example, a NETCONF client can acquire a global lock via `lock` operation without starting a transaction.
{% endhint %}

The first phase, historically called validation, does more than just validate data. When the transaction starts applying, NSO captures the initial intent and creates a rollback file, which allows one to reverse (or roll back) the intent. For example, the rollback file might contain the information that you changed a service instance parameter but it would not contain the service-produced device changes.

Then the first, partial validation takes place. It ensures the service input parameters are valid according to the service YANG model, so the service code can safely use provided parameter values.

Next, NSO runs transaction hooks and performs the necessary transforms, which alter the data before it is saved, for example encrypting passwords. This is also where the Service Manager invokes FASTMAP and service mapping callbacks, recording the resulting changes. NSO takes service write locks in this stage, too.

After transforms, there are no more changes to the configuration data, and the full validation starts, including YANG model constraints over the complete configuration, custom validation through validation points, and configuration policies (see [Policies](/guides/operation-and-usage/operations/basic-operations#d5e319) in Operation and Usage).

<figure><img src="/files/Q8Wz220dIgquQHty2WJH" alt="" width="375"><figcaption><p>Stages of Transaction Validation Phase</p></figcaption></figure>

Throughout the phase, the transaction engine makes checkpoints, so it can restart the transaction faster in case of concurrency conflicts. The check for conflicts happens at the end of this first phase when NSO also takes the global transaction lock. Concurrency is further discussed in [NSO Concurrency Model](/guides/development/core-concepts/nso-concurrency-model).

Similarly, we can further break down the other phases to get the main transaction commit steps:

1. **Generate rollback**: NSO first generates a rollback file, capturing and reversing the intent of the work phase, allowing the system to later undo the transaction if required.
2. **Transaction hooks**: Define transaction extension points for code to prepare the final complete data. This includes:

* Pre-transform validation for service input data.
* Custom hooks and transforms to update existing and fill in additional data.
* Service Manager callbacks for service mapping logic, calling service `create()` code and similar.

3. **Validation**: Complete validation of the final change against the YANG models and custom validation logic.
4. **Lock**: Takes an exclusive transaction lock, required for resolving concurrent transaction conflicts. Since only one transaction can acquire such a lock at a time, it affects transactional throughput. The parts that run inside the transaction lock are collectively called the critical section.
5. **Conflict check**: Checks if other transactions that run in parallel carry conflicting operations. The transaction may abort if conflicts cannot be automatically resolved.
6. **Write-start** and **Prepare**: Notifies all participating systems of the changes. In particular:

* Send all changes to CDB and any custom data provider.
* Evaluate relevant kickers.
* Perform Device Manager prepare, sending device-related changes to the NEDs or commit queue. If a live device rejects the new configuration, transaction is aborted and no changes at all are persisted.

7. **Commit**: First commits (persists) the CDB and data provider changes, then invokes Device Manager commit to commit (persist) the new configuration on devices.
8. **Notify**: Invokes kickers and notifies CDB subscribers.
9. **Unlock**: First service, then transaction lock is released and critical section ends.

<figure><img src="/files/dsbW59A7h4gVHTBJ6fvB" alt="" width="563"><figcaption><p>Stages of a Transaction Commit</p></figcaption></figure>

The following is a partial, annotated progress trace, demonstrating these steps for a sample service instantiation.

```
# --- Transaction applies due to `commit` in CLI ---
applying transaction for running datastore usid=64 tid=249 trace-id=2b3...
 2026-06-09T10:49:29.266 waiting to apply... ok (0.000 s)
entering validate phase
 # --- Generate rollback ---
 2026-06-09T10:49:29.267 creating rollback file... ok (0.004 s)
 ...
 # --- Pre-validation for service input ---
 2026-06-09T10:49:29.273 run pre-transform validation: ok (0.001 s)
 ...
 # --- Custom hooks and transforms ---
 2026-06-09T10:49:29.273 run transforms and transaction hooks...
 ...
 # --- Service mapping logic ---
 2026-06-09T10:49:29.276 service /test: creating service... ok (0.010 s)
 2026-06-09T10:49:29.287 run transforms and transaction hooks: ok (0.013 s)
 ...
 # --- Validation ---
 2026-06-09T10:49:29.289 run validation over the changeset... ok (0.001 s)
 2026-06-09T10:49:29.291 run dependency-triggered validation... ok (0.000 s)
 ...
 # --- Lock ---
 2026-06-09T10:49:29.293 taking transaction lock... ok (0.000 s)
 2026-06-09T10:49:29.293 holding transaction lock...
 # --- Conflict check ---
 2026-06-09T10:49:29.293 check for read-write conflicts... ok (0.000 s)
 ...
leaving validate phase (0.028 s)
entering write-start phase
 # --- Send changes to CDB ---
 2026-06-09T10:49:29.297 cdb: write changeset... ok (0.000 s)
 # --- Evaluate relevant kickers ---
 2026-06-09T10:49:29.297 check data kickers... ok (0.000 s)
leaving write-start phase (0.004 s)
entering prepare phase
 ...
 # --- Send device-related changes ---
 2026-06-09T10:49:29.299 device-manager: prepare
 2026-06-09T10:49:29.417 device c1: push configuration...
 ...
leaving prepare phase (0.153 s)
entering commit phase
 # --- Commit ---
 2026-06-09T10:49:29.452 cdb: commit
 2026-06-09T10:49:29.452 cdb: switch to new running... ok (0.000 s)
 2026-06-09T10:49:29.453 device-manager: commit
 2026-06-09T10:49:29.456 device c1: push configuration: ok (0.038 s)
 ...
 # --- Notify ---
 2026-06-09T10:49:29.461 invoking data kickers... ok (0.000 s)
 # --- Unlock ---
 2026-06-09T10:49:29.462 holding transaction lock: ok (0.169 s)
 ...
applying transaction for running datastore usid=64 tid=249 trace-id=2b3... (0.196 s)
Commit complete.
```

You can additionally observe from the start of the output that a transaction might have to wait before being allowed to proceed due to `transaction-limits` settings in the `ncs.conf`.

## Extension Points

NSO transactions support three main extension points for implementing custom functionality using callbacks, as listed in [Overview of Extension Points](/guides/development/introduction-to-automation/applications-in-nso#overview-of-extension-points):

* Service: for implementation of classic and nano service provisioning logic.
* Validation: for constraining data beyond the capabilities of YANG.
* Data provider: for components that supply or process data and integrate tightly with NSO transactions. Enables storage of data outside NSO and similar use cases.

Of the three, only data providers are aware of the transaction lifecycle and support callbacks for transaction state transitions (phases). NSO uses the two-phase commit protocol to ensure that all participants perform the requested operations, which is required to ensure transactional properties.

In addition, NSO supports some specific callbacks from internal systems, such as the transaction or the authorization engine. These may be called during processing of transactions, but have very narrow use (and require careful consideration as they can easily negatively affect performance).

## Nested Transactions

NSO supports creating a transaction on top of another transaction, often called **trans-in-trans**. It uses an already open parent transaction as its backend instead of the `running` or `operational` datastore.

This is useful when you want to stage a group of related edits and then either merge all staged edits into the parent transaction, or discard the whole staged group, without affecting other edits already present in the parent transaction.

Another use case is to work around restrictions in data callbacks, such as using `maapi_load_config()` in a nested transaction where it cannot be used directly on the original attached transaction.

Important semantics:

* Applying the nested transaction does not commit to datastore or devices but only updates the parent transaction.
* Finishing the nested transaction without apply discards nested edits.
* Final validation and network commit happen when the parent transaction commits.

Nested transactions can be validated explicitly but can always be applied, even if the validation failed due to invalid configuration. Applying a nested transaction does not itself run a full datastore commit flow.

### Python Example

```python
import ncs

with ncs.maapi.Maapi() as m:
    m.start_user_session("admin", "system")

    with m.start_write_trans() as parent:
        # Parent edits can happen here.

        with m.start_trans_in_trans(parent.th, ncs.READ_WRITE) as nested:
            # Stage a related set of edits in isolation.
            root = ncs.maagic.get_root(nested)
            root.test__items.item.create("example")

            # Optional explicit validation of nested changes.
            nested.validate(True)

            # Merge nested edits into the parent transaction.
            nested.apply()

        # Parent can continue with other edits, then commits once.
        parent.apply()
```

In the example, if `nested.apply()` is skipped, nested edits are discarded when the nested transaction closes, while the parent transaction remains open and can still be committed or aborted independently.


# Using CDB

Concepts in usage of the Configuration Database (CDB).

When using CDB to store the configuration data, the applications need to be able to:

1. Read configuration data from the database.
2. React to changes to the database. There are several possible writers to the database, such as the CLI, NETCONF sessions, the Web UI, either of the NSO sync commands, alarms that get written into the alarm table, NETCONF notifications that arrive at NSO or the NETCONF agent.

The figure below illustrates the architecture of when the CDB is used. The Application components read configuration data and subscribe to changes to the database using a simple RPC-based API. The API is part of the Java library and is fully documented in the Javadoc for CDB.

<div data-with-frame="true"><figure><img src="/files/Ow24olbh0sFQaFXV1WZg" alt="" width="563"><figcaption><p>NSO CDB Architecture Scenario</p></figcaption></figure></div>

While CDB is the default data store for configuration data in NSO, it is possible to use an external database, if needed. See the example [examples.ncs/sdk-api/external-db](https://github.com/NSO-developer/nso-examples/tree/6.7/sdk-api/external-db) for details.

In the following, we will use the files in [examples.ncs/service-management/mpls-vpn-java](https://github.com/NSO-developer/nso-examples/tree/6.7/service-management/mpls-vpn-java) as a source for our examples. Refer to `README` in that directory for additional details.

## The NSO Data Model <a href="#ncs.ug.architecture.im" id="ncs.ug.architecture.im"></a>

NSO is designed to manage devices and services. NSO uses YANG as the overall modeling language. YANG models describe the NSO configuration, the device configurations, and the configuration of services. Therefore it is vital to understand the data model for NSO including these aspects. The YANG models are available in `$NCS_DIR/src/ncs/yang` and are structured as follows.

`tailf-ncs.yang` is the top module that includes the following sub-modules:

* `tailf-ncs-common.yang`: common definitions.
* `tailf-ncs-packages.yang`: this sub-module defines the management of packages that are run by NSO. A package contains custom code, models, and documentation for any function added to the NSO platform. It can for example be a service application or a southbound integration to a device.
* `tailf-ncs-devices.yang`: This is a core model of NSO. The device model defines everything a user can do with a device that NSO speaks to via a Network Element Driver, NED.
* `tailf-ncs-services.yang`: Services represent anything that spans across devices. This can for example be MPLS VPN, MEF e-line, BGP peer, or website. NSO provides several mechanisms to handle services in general which are specified by this model. Also, it defines placeholder containers under which developers, as an option, can augment their specific services.
* `tailf-ncs-snmp-notification-receiver.yang`: NSO can subscribe to SNMP notifications from the devices. The subscription is specified by this model.
* `tailf-ncs-java-vm.yang`: Custom code that is part of a package is loaded and executed by the NSO Java VM. This is managed by this model. Further, when browsing `$NCS_DIR/src/ncs/yang` you will find models for all aspects of NSO functionality, for example
* `tailf-ncs-alarms.yang`: This model defines how NSO manages alarms. The source of an alarm can be anything like an NSO state change, SNMP, or NETCONF notification.
* `tailf-ncs-snmp.yang`: This model defines how to configure the NSO northbound SNMP agent.
* `tailf-ncs-config.yang`: This model describes the layout of the NSO config file, usually called `ncs.conf`
* `tailf-ncs-packages.yang:` This model describes the layout of the file `package-meta-data.xml`. All user code, data models MIBS, and Java code are always contained in an NSO package. The `package-meta-data.xml` file must always exist in a package and describe the package.

These models will be illustrated and briefly explained below. Note that the figures only contain some relevant aspects of the model and are far from complete. The details of the model are explained in the respective sections.

A good way to learn the model is to start the NSO CLI and use tab completion to navigate the model. Note that depending if you are in operation mode or configuration mode different parts of the model will show up. Also try using TAB to get a list of actions at the level you want, for example, `devices TAB`.

Another way to learn and explore the NSO model is to use the Yanger tool to render a tree output from the NSO model: `yanger -f tree --tree-depth=3 tailf-ncs.yang`. This will show a tree for the complete model. Below is a truncated example:

{% code title="Example: Using yanger" %}

```bash
$ yanger -f tree --tree-depth=3 tailf-ncs.yang
module: tailf-ncs
   +--rw ssh
   |  +--rw host-key-verification?   ssh-host-key-verification-level
   |  +--rw private-key* [name]
   |     +--rw name          string
   |     +--rw key-data      ssh-private-key
   |     +--rw passphrase?   tailf:aes-256-cfb-128-encrypted-string
   +--rw cluster
   |  +--rw remote-node* [name]
   |  |  +--rw name             node-name
   |  |  +--rw address?         inet:host
   |  |  +--rw port?            inet:port-number
   |  |  +--rw ssh
   |  |  +--rw authgroup        -> /cluster/authgroup/name
   |  |  +--rw trace?           trace-flag
   |  |  +--rw username?        string
   |  |  +--rw notifications
   |  |  +--ro device* [name]
   |  +--rw authgroup* [name]
   |  |  +--rw name           string
   |  |  +--rw default-map!
   |  |  +--rw umap* [local-user]
   |  +--rw commit-queue
   |  |  +--rw enabled?   boolean
   |  +--ro enabled?        boolean
   |  +--ro connection*
   |     +--ro remote-node?   -> /cluster/remote-node/name
   |     +--ro address?       inet:ip-address
   |     +--ro port?          inet:port-number
   |     +--ro channels?      uint32
   |     +--ro local-user?    string
   |     +--ro remote-user?   string
   |     +--ro status?        enumeration
   |     +--ro trace?         enumeration
...
```

{% endcode %}

## Addressing Data Using Keypaths

As CDB stores hierarchical data as specified by a YANG model, data is addressed by a path to the key. We call this a keypath. A keypath provides a path through the configuration data tree. A keypath can be either absolute or relative. An absolute keypath starts from the root of the tree, while a relative path starts from the "current position" in the tree. They are differentiated by the presence or absence of a leading `/`. Navigating the configuration data tree is thus done in the same way as a directory structure. It is possible to change the current position with for example the `CdbSession.cd()` method. Several of the API methods take a keypath as a parameter.

YANG elements that are lists of other YANG elements can be traversed using two different path notations. Consider the following YANG model fragment:

{% code title="Example: L3 VPN YANG Extract" %}

```yang
module l3vpn {

  namespace "http://com/example/l3vpn";
  prefix l3vpn;


        ...

  container topology {
    list role {
      key "role";
      tailf:cli-compact-syntax;
      leaf role {
        type enumeration {
          enum ce;
          enum pe;
          enum p;
        }
      }

      leaf-list device {
        type leafref {
          path "/ncs:devices/ncs:device/ncs:name";
        }
      }
    }

    list connection {
      key "name";
      leaf name {
        type string;
      }
      container endpoint-1 {
        tailf:cli-compact-syntax;
        uses connection-grouping;
      }
      container endpoint-2 {
        tailf:cli-compact-syntax;
        uses connection-grouping;
      }
      leaf link-vlan {
        type uint32;
      }
    }
  }
```

{% endcode %}

We can use the method `CdbSession.getNumberOfInstances()` to find the number of elements in a list has, and then traverse them using a standard index notation, i.e., `<path to list>[integer]`. The children of a list are numbered starting from 0. Looking at the example above (L3 VPN YANG Extract) the path `/l3vpn:topology/connection[2]/endpoint-1` refers to the `endpoint-1` leaf of the third `connection`. This numbering is only valid during the current CDB session. CDB is always locked for writing during a read session.

We can also refer to list instances using the values of the keys of the list. In a YANG model, you specify which leafs (there can be several) are to be used for keys by using the `key <name>` statement at the beginning of the list. In our case a `connection` has the `name` leaf as the key. So the path `/l3vpn:topology/connection{c1}/endpoint-2` refers to the `endpoint-2` leaf of the `connection` whose name is “c1”.

A YANG list may have more than one key. The syntax for the keys is a space-separated list of key values enclosed within curly brackets: `{Key1 Key2 ...}`

Which version of the list element referencing to use depends on the situation. Indexing with an integer is convenient when looping through all elements. As a convenience all methods expecting keypaths accept formatting characters and accompanying data items. For example, you can use `CdbSession.getElem("server[%d]/ifc{%s}/mtu", 2, "eth0")` to fetch the MTU of the third server instance's interface named "eth0". Using relative paths and `CdbSession.pushd()` it is possible to write code that can be re-used for common sub-trees.

The current position also includes the namespace. To read elements from a different namespace use the prefix qualified tag for that element like in `l3vpn:topology`.

## Subscriptions <a href="#ug.cdb.subscriptions" id="ug.cdb.subscriptions"></a>

The CDB subscription mechanism allows an external program to be notified when some part of the configuration changes. When receiving a notification it is also possible to iterate through the changes written to CDB. Subscriptions are always towards the running data store (it is not possible to subscribe to changes to the startup data store). Subscriptions towards operational data (see [Operational Data in CDB](#ug.cdb.opdata)) kept in CDB are also possible, but the mechanism is slightly different.

The first thing to do is to inform CDB which paths we want to subscribe to. Registering a path returns a subscription point identifier. This is done by acquiring a subscriber instance by calling `CdbSubscription Cdb.newSubscription()` method. For the subscriber (or `CdbSubscription` instance) the paths are registered with the `CdbSubscription.subscribe()` that that returns the actual subscription point identifier. A subscriber can have multiple subscription points, and there can be many different subscribers. Every point is defined through a path - similar to the paths we use for read operations, with the exception that instead of fully instantiated paths to list instances we can selectively use tagpaths.

When a client is done defining subscriptions it should inform NSO that it is ready to receive notifications by calling `CdbSubscription.subscribeDone()`, after which the subscription socket is ready to be polled.

We can subscribe either to specific leaves, or entire subtrees. Explaining this by example we get:

* `/ncs:devices/global-settings/trace`: Subscription to a leaf. Only changes to this leaf will generate a notification.
* `/ncs:devices`: Subscription to the subtree rooted at `/ncs:devices`. Any changes to this subtree will generate a notification. This includes additions or removals of `device` instances, as well as changes to already existing `device` instances.
* `/ncs:devices/device{"ex0"}/address`: Subscription to a specific element in a list. A notification will be generated when the device `ex0` changes its IP address.
* `/ncs:devices/device/address`: Subscription to a leaf in a list. A notification will be generated leaf `address` is changed in any device instance.

When adding a subscription point the client must also provide a priority, which is an integer (a smaller number means a higher priority). When data in CDB is changed, this change is part of a transaction. A transaction can be initiated by a `commit` operation from the CLI or an `edit-config` operation in NETCONF resulting in the running database being modified. As the last part of the transaction CDB will generate notifications in lock-step priority order. First, all subscribers at the lowest numbered priority are handled, once they all have replied and synchronized by calling `CdbSubscription.sync()` the next set - at the next priority level - is handled by CDB. Not until all subscription points have been acknowledged is the transaction complete. This implies that if the initiator of the transaction was for example a `commit` command in the CLI, the command will hang until notifications have been acknowledged.

{% hint style="info" %}
Do not start a new write transaction or call commit from inside a CDB subscriber callback. A CDB subscriber is triggered by a transaction that already holds the NSO transaction lock. NSO sends a change notification to the subscriber and waits for the subscriber to acknowledge it before releasing the lock. If the subscriber code tries to commit another change before returning, that new commit must wait for the same lock, while the original transaction is waiting for the subscriber acknowledgment. This causes a deadlock.

\
If subscriber logic needs to make additional configuration changes, defer that work until after the subscriber callback has returned. For example, signal a separate worker, action, or external process to open a new transaction after the notification has been acknowledged.
{% endhint %}

Note that even though the notifications are delivered within the transaction, a subscriber can't reject the changes (since this would break the two-phase commit protocol used by the NSO backplane towards all data providers).

As a subscriber has read its subscription notifications using `CdbSubscription.read()`, it can iterate through the changes that caused the particular subscription notification using the `CdbSubscription.diffIterate()` method. It is also possible to start a new read-session to the `CdbDBType.CDB_PRE_COMMIT_RUNNING` database to read the running database as it was before the pending transaction.

To view registered subscribers use the `ncs --status` command.

## Sessions <a href="#d5e2708" id="d5e2708"></a>

It is important to note that CDB is locked for writing during a read session using the Java API. A session starts with `CdbSession Cdb.startSession()` and the lock is not released until the `CdbSession.endSession()` (or the `Cdb.close()`) call. CDB will also automatically release the lock if the socket is closed for some other reason, such as program termination.

## Loading Initial Data into CDB <a href="#ug.cdb.init" id="ug.cdb.init"></a>

When NSO starts for the first time, the CDB database is empty. The location of the database files used by CDB is given in `ncs.conf`. At first startup, when CDB is empty, i.e., no database files are found in the directory specified by `<db-dir>` (`./ncs-cdb` as given by the example below (CDB Init)), CDB will try to initialize the database from all XML documents found in the same directory.

{% code title="Example: CDB Init" %}

```xml
<!-- Where the database (and init XML) files are kept -->
<cdb>
    <db-dir>./ncs-cdb</db-dir>
</cdb>
```

{% endcode %}

This feature can be used to reset the configuration to factory settings.

Given the YANG model in the example above (L3 VPN YANG Extract), the initial data for `topology` can be found in `topology.xml` as seen in the example below (Initial Data for Topology).

{% code title="Example: Initial Data for Topology" %}

```xml
<config xmlns="http://tail-f.com/ns/config/1.0">
  <topology xmlns="http://com/example/l3vpn">
    <role>
      <role>ce</role>
      <device>ce0</device>
      <device>ce1</device>
      <device>ce2</device>
    ...
    </role>
    <role>
      <role>pe</role>
      <device>pe0</device>
      <device>pe1</device>
      <device>pe2</device>
      <device>pe3</device>
    </role>
    ...
    <connection>
      <name>c0</name>
      <endpoint-1>
        <device>ce0</device>
        <interface>GigabitEthernet0/8</interface>
        <ip-address>192.168.1.1/30</ip-address>
      </endpoint-1>
      <endpoint-2>
        <device>pe0</device>
        <interface>GigabitEthernet0/0/0/3</interface>
        <ip-address>192.168.1.2/30</ip-address>
      </endpoint-2>
      <link-vlan>88</link-vlan>
    </connection>
    <connection>
      <name>c1</name>
    ...
```

{% endcode %}

Another example of using these features is when initializing the AAA database. This is described in [AAA infrastructure](/guides/administration/management/aaa-infrastructure).

All files ending in `.xml` will be loaded (in an undefined order) and committed in a single transaction when CDB enters start phase 1 (see [Starting NSO](https://nso-docs.cisco.com/guides/development/core-concepts/pages/hgWUBFw1TA0R6WyLxOgc#ug.sys_mgmt.starting_ncs) for more details on start phases). The format of the init files is rather lax in that it is not required that a complete instance document following the data model is present, much like the NETCONF `edit-config` operation. It is also possible to wrap multiple top-level tags in the file with a surrounding config tag, as shown in the example below (Wrapper for Multiple Top-Level Tags) like this:

{% code title="Example: Wrapper for Multiple Top-Level Tags" %}

```xml
<config xmlns="http://tail-f.com/ns/config/1.0">
  ...
</config>
```

{% endcode %}

{% hint style="info" %}
The actual names of the XML files do not matter, i.e., they do not need to correspond to the part of the YANG model being initialized.
{% endhint %}

## Operational Data in CDB <a href="#ug.cdb.opdata" id="ug.cdb.opdata"></a>

In addition to handling configuration data, CDB can also take care of operational data such as alarms and traffic statistics. By default, operational data is not persistent and thus not kept between restarts. In the YANG model annotating a node with `config false` will mark the subtree rooted at that node as operational data. Reading and writing operational data is done similarly to ordinary configuration data, with the main difference being that you have to specify that you are working against operational data. Also, the subscription model is different.

### Subscriptions <a href="#d5e2749" id="d5e2749"></a>

Subscriptions towards the operational data in CDB are similar to the above, but because the operational data store is designed for light-weight access, does not have transactions, and normally avoids the use of any locks, there are several differences - in particular:

* Subscription notifications are only generated if the writer obtains the “subscription lock”, by using the `Cdb.startSession()` method with the CdbLockType.LOCK\_REQUEST flag.
* Subscriptions are registered with the `CdbSubscription.subscribe()` method with the flag `CdbSubscriptionType.SUB_OPERATIONAL` rather than `CdbSubscriptionType.SUB_RUNNING`.
* No priorities are used.
* Neither the writer that generated the subscription notifications nor other writes to the same data are blocked while notifications are being delivered. However, the subscription lock remains in effect until notification delivery is complete.
* The previous value for the modified leaf is not available when using the `CdbSubscriber.diffIterate()` method.

Essentially a write operation towards the operational data store, combined with the subscription lock, takes on the role of a transaction for configuration data as far as subscription notifications are concerned. This means that if operational data updates are done with many single-element write operations, this can potentially result in a lot of subscription notifications. Thus it is a good idea to use the multi-element `CdbSession.setObject()` etc methods for updating operational data that applications subscribe to.

Since write operations that do not attempt to obtain the subscription lock are allowed to proceed even during notification delivery, it is the responsibility of the applications using the operational data store to obtain the lock as needed when writing. If subscribers should be able to reliably read the exact data that resulted from the write that triggered their subscription, the subscription lock must always be obtained when writing that particular set of data elements. One possibility is of course to obtain the lock for all writes to operational data, but this may have an unacceptable performance impact.

## Example <a href="#d5e2773" id="d5e2773"></a>

We will take a first look at the [examples.ncs/sdk-api/cdb-java](https://github.com/NSO-developer/nso-examples/tree/6.7/sdk-api/cdb-java) example. This example is an NSO project with two packages: `cdb` and `router`.

### Example packages

* `router`: A NED package with a simple but still realistic model of a network device. The only component in this package is the NED component that uses NETCONF to communicate with the device. This package is used in many NSO examples including [examples.ncs/device-management/router-network](https://github.com/NSO-developer/nso-examples/tree/6.7/device-management/router-network) which is an introduction to NSO device manager, NSO netsim, and this router package.
* `cdb`: This package has an even simpler YANG model to illustrate some aspects of CDB data retrieval. The package consists of five application components:
  * Plain CDB Subscriber: This CDB subscriber subscribes to changes under the path `/devices/device{ex0}/config`. Whenever a change occurs there, the code iterates through the change and prints the values.
  * CdbCfgSubscriber: A more advanced CDB subscriber that subscribes to changes under the path `/devices/device/config/sys/interfaces/interface`.
  * OperSubscriber: An operational data subscriber that subscribes to changes under the path `/t:test/stats-item`.

The [examples.ncs/sdk-api/cdb-py](https://github.com/NSO-developer/nso-examples/tree/6.7/sdk-api/cdb-py) and [examples.ncs/sdk-api/cdb-java](https://github.com/NSO-developer/nso-examples/tree/6.7/sdk-api/cdb-java) examples `packages/cdb` package includes the YANG model in the in the example below:.

{% code title="Example: Simple Config Data" %}

```yang
module test {
  namespace "http://example.com/test";
  prefix t;

  import tailf-common {
    prefix tailf;
  }

  description "This model is used as a simple example model
               illustrating some aspects of CDB subscriptions
               and CDB operational data";

  revision 2012-06-26 {
    description "Initial revision.";
  }

  container test {
    list config-item {
      key ckey;
      leaf ckey {
        type string;
      }
      leaf i {
        type int32;
      }
    }
    list stats-item {
      config false;
      tailf:cdb-oper;
      key skey;
      leaf skey {
        type string;
      }
      leaf i {
        type int32;
      }
      container inner {
        leaf  l {
          type string;
        }
      }
    }
  }
}
```

{% endcode %}

Let us now populate the database and look at the Plain CDB Subscriber and how it can use the Java API to react to changes to the data. This component subscribes to changes under the path `/devices/device{ex0}/config` which is configuration changes for the device named `ex0` which is a device connected to NSO via the router NED.

Being an application component in the `cdb` package implies that this component is realized by a Java class that implements the `com.tailf.ncs.ApplicationComponent` Java interface. This interface inherits the Java standard `Runnable` interface which requires the `run()` method to be implemented. In addition to this method, there is a `init()` and a `finish()` method that has to be implemented. When the NSO Java-VM starts this class will be started in a separate thread with an initial call to `init()` before the thread starts. When the package is requested to stop execution a call to `finish()` is performed and this method is expected to end thread execution.

{% code title="Example: Plain CDB Subscriber Java Code" %}

```java
public class PlainCdbSub implements ApplicationComponent  {
    private static final Logger LOGGER
            = LogManager.getLogger(PlainCdbSub.class);

    @Resource(type = ResourceType.CDB, scope = Scope.INSTANCE,
              qualifier = "plain")
    private Cdb cdb;

    private CdbSubscription sub;
    private int subId;
    private boolean requestStop;

    public PlainCdbSub() {
    }

    public void init() {
        try {
            LOGGER.info(" init cdb subscriber ");
            sub = new CdbSubscription(cdb);
            String str = "/devices/device{ex0}/config";
            subId = sub.subscribe(1, new Ncs(), str);
            sub.subscribeDone();
            LOGGER.info("subscribeDone");
            requestStop = false;
        } catch (Exception e) {
            throw new RuntimeException("FAIL in init", e);
        }
    }

    public void run() {
        try {
            while (!requestStop) {
                try {
                    sub.read();
                    sub.diffIterate(subId, new Iter());
                } finally {
                    sub.sync(CdbSubscriptionSyncType.DONE_SOCKET);
                }
            }
        } catch (ConfException e) {
            if (e.getErrorCode() == ErrorCode.ERR_EOF) {
                // Triggered by finish method
                // if we throw further NCS JVM will try to restart
                // the package
                LOGGER.warn(" Socket Closed!");
            } else {
                throw new RuntimeException("FAIL in run", e);
            }
        } catch (Exception e) {
            LOGGER.warn("Exception:" + e.getMessage());
            throw new RuntimeException("FAIL in run", e);
        } finally {
            requestStop = false;
            LOGGER.warn(" run end ");
        }
    }

    public void finish() {
        requestStop = true;
        LOGGER.warn(" PlainSub in finish () =>");
        try {
            // ResourceManager will close the resource (cdb) used by this
            // instance that triggers ConfException with ErrorCode.ERR_EOF
            // in run method
            ResourceManager.unregisterResources(this);
        } catch (Exception e) {
            throw new RuntimeException("FAIL in finish", e);
        }
        LOGGER.warn(" PlainSub in finish () => ok");
    }

    private class Iter implements CdbDiffIterate  {
        public DiffIterateResultFlag iterate(ConfObject[] kp,
                                             DiffIterateOperFlag op,
                                             ConfObject oldValue,
                                             ConfObject newValue,
                                             Object state) {
            try {
                String kpString = Conf.kpToString(kp);
                LOGGER.info("diffIterate: kp= " + kpString + ", OP=" + op
                            + ", old_value=" + oldValue + ", new_value="
                            + newValue);
                return DiffIterateResultFlag.ITER_RECURSE;
            } catch (Exception e) {
                return DiffIterateResultFlag.ITER_CONTINUE;
            }
        }
    }
}
```

{% endcode %}

We will walk through the code and highlight different aspects. We start with how the `Cdb` instance is retrieved in this example. It is always possible to open a socket to NSO and create the `Cdb` instance with this socket. But with this comes the responsibility to manage that socket. In NSO, there is a resource manager that can take over this responsibility. In the code, the field that should contain the `Cdb` instance is simply annotated with a `@Resource` annotation. The resource manager will find this annotation and create the `Cdb` instance as specified. In this example below (Resource Annotation) `Scope.INSTANCE` implies that new instances of this example class should have unique `Cdb` instances (see more in [The Resource Manager](https://nso-docs.cisco.com/guides/development/core-concepts/pages/LKELpQIToC2wE9KokBCU#ncs.ug.javavm.resman)).

{% code title="Example: Resource Annotation" %}

```java
    @Resource(type = ResourceType.CDB, scope = Scope.INSTANCE,
              qualifier = "plain")
    private Cdb cdb;
```

{% endcode %}

The `init()` method (shown in the example below, (Plain Subscriber Init) is called before this application component thread is started. For this subscriber, this is the place to set up the subscription. First, an `CdbSubscription` instance is created and in this instance, the subscription points are registered (one in this case). When all subscription points are registered a call to `CdbSubscriber.subscribeDone()` will indicate that the registration is finished and the subscriber is ready to start.

{% code title="Example: Plain Subscriber Init" %}

```java
    public void init() {
        try {
            LOGGER.info(" init cdb subscriber ");
            sub = new CdbSubscription(cdb);
            String str = "/devices/device{ex0}/config";
            subId = sub.subscribe(1, new Ncs(), str);
            sub.subscribeDone();
            LOGGER.info("subscribeDone");
            requestStop = false;
        } catch (Exception e) {
            throw new RuntimeException("FAIL in init", e);
        }
    }
```

{% endcode %}

The `run()` method comes from the standard Java API Runnable interface and is executed when the application component thread is started. For this subscriber (see example below (Plain CDB Subscriber)) a loop over the `CdbSubscription.read()` method drives the subscription. This call will block until data has changed for some of the subscription points that were registered, and the IDs for these subscription points will then be returned. In our example, since we only have one subscription point, we know that this is the one stored as `subId`. This subscriber chooses to find the changes by calling the `CdbSubscription.diffIterate()` method. Important is to acknowledge the subscription by calling `CdbSubscription.sync()` or else this subscription will block the ongoing transaction.

{% code title="Example: Plain CDB Subscriber" %}

```java
    public void run() {
        try {
            while (!requestStop) {
                try {
                    sub.read();
                    sub.diffIterate(subId, new Iter());
                } finally {
                    sub.sync(CdbSubscriptionSyncType.DONE_SOCKET);
                }
            }
        } catch (ConfException e) {
            if (e.getErrorCode() == ErrorCode.ERR_EOF) {
                // Triggered by finish method
                // if we throw further NCS JVM will try to restart
                // the package
                LOGGER.warn(" Socket Closed!");
            } else {
                throw new RuntimeException("FAIL in run", e);
            }
        } catch (Exception e) {
            LOGGER.warn("Exception:" + e.getMessage());
            throw new RuntimeException("FAIL in run", e);
        } finally {
            requestStop = false;
            LOGGER.warn(" run end ");
        }
    }
```

{% endcode %}

The call to the `CdbSubscription.diffIterate()` requires an object instance implementing an `iterate()` method. To do this, the `CdbDiffIterate` interface is implemented by a suitable class. In our example, this is done by a private inner class called `Iter` (Example below (Plain Subscriber Iterator Implementation)). The `iterate()` method is called for all changes and the path, type of change, and data are provided as arguments. In the end, the `iterate()` should return a flag that controls how further iteration should prolong, or if it should stop. Our example `iterate()` method just logs the changes.

{% code title="Example: Plain Subscriber Iterator Implementation" %}

```java
    private class Iter implements CdbDiffIterate  {
        public DiffIterateResultFlag iterate(ConfObject[] kp,
                                             DiffIterateOperFlag op,
                                             ConfObject oldValue,
                                             ConfObject newValue,
                                             Object state) {
            try {
                String kpString = Conf.kpToString(kp);
                LOGGER.info("diffIterate: kp= " + kpString + ", OP=" + op
                            + ", old_value=" + oldValue + ", new_value="
                            + newValue);
                return DiffIterateResultFlag.ITER_RECURSE;
            } catch (Exception e) {
                return DiffIterateResultFlag.ITER_CONTINUE;
            }
        }
    }
```

{% endcode %}

The `finish()` method (Example below (Plain Subscriber `finish`)) is called when the NSO Java-VM wants the application component thread to stop execution. An orderly stop of the thread is expected. Here the subscription will stop if the subscription socket and underlying `Cdb` instance are closed. This will be done by the `ResourceManager` when we tell it that the resources retrieved for this Java object instance could be unregistered and closed. This is done by a call to the `ResourceManager.unregisterResources()` method.

{% code title="Example: Plain Subscriber finish" %}

```java
    public void finish() {
        requestStop = true;
        LOGGER.warn(" PlainSub in finish () =>");
        try {
            // ResourceManager will close the resource (cdb) used by this
            // instance that triggers ConfException with ErrorCode.ERR_EOF
            // in run method
            ResourceManager.unregisterResources(this);
        } catch (Exception e) {
            throw new RuntimeException("FAIL in finish", e);
        }
        LOGGER.warn(" PlainSub in finish () => ok");
    }
```

{% endcode %}

We will now compile and start the [examples.ncs/sdk-api/cdb-java](https://github.com/NSO-developer/nso-examples/tree/6.7/sdk-api/cdb-java) example, populate some config data, and look at the result. The example below (Plain Subscriber Startup) shows how to do this.

{% code title="Example: Plain Subscriber Startup" %}

```bash
$ make clean all
$ ncs-netsim start
DEVICE ex0 OK STARTED
DEVICE ex1 OK STARTED
DEVICE ex2 OK STARTED

$ ncs
```

{% endcode %}

By far, the easiest way to populate the database with some actual data is to run the CLI (see the example below (Populate Data using CLI)).

{% code title="Example: Populate Data using CLI" %}

```bash
$ ncs_cli -u admin
admin connected from 127.0.0.1 using console on ncs
admin@ncs# config exclusive
Entering configuration mode exclusive
Warning: uncommitted changes will be discarded on exit
admin@ncs(config)# devices sync-from
sync-result {
    device ex0
    result true
}
sync-result {
    device ex1
    result true
}
sync-result {
    device ex2
    result true
}

admin@ncs(config)# devices device ex0 config r:sys syslog server 4.5.6.7 enabled
admin@ncs(config-server-4.5.6.7)# commit
Commit complete.
admin@ncs(config-server-4.5.6.7)# top
admin@ncs(config)# exit
admin@ncs# show devices device ex0 config r:sys syslog
NAME
----------
4.5.6.7
10.3.4.5
```

{% endcode %}

We have now added a server to the Syslog. What remains is to check what our 'Plain CDB Subscriber' `ApplicationComponent` got as a result of this update. In the `logs` directory of the [examples.ncs/sdk-api/cdb-java](https://github.com/NSO-developer/nso-examples/tree/6.7/sdk-api/cdb-java) example there is a file named `PlainCdbSub.out` which contains the log data from this application component. At the beginning of this file, a lot of logging is performed which emanates from the `sync-from` of the device. At the end of this file, we can find the three log rows that come from our update. See the extract in the example below (Plain Subscriber Output) (with each row split over several to fit on the page).

{% code title="Example: Plain Subscriber Output" %}

```
<INFO> 05-Feb-2015::13:24:55,760  PlainCdbSub$Iter
  (cdb-examples:Plain CDB Subscriber) -Run-4: - diffIterate:
  kp= /ncs:devices/device{ex0}/config/r:sys/syslog/server{4.5.6.7},
  OP=MOP_CREATED, old_value=null, new_value=null
<INFO> 05-Feb-2015::13:24:55,761  PlainCdbSub$Iter
  (cdb-examples:Plain CDB Subscriber) -Run-4: - diffIterate:
  kp= /ncs:devices/device{ex0}/config/r:sys/syslog/server{4.5.6.7}/name,
  OP=MOP_VALUE_SET, old_value=null, new_value=4.5.6.7
<INFO> 05-Feb-2015::13:24:55,762  PlainCdbSub$Iter
  (cdb-examples:Plain CDB Subscriber) -Run-4: - diffIterate:
  kp= /ncs:devices/device{ex0}/config/r:sys/syslog/server{4.5.6.7}/enabled,
  OP=MOP_VALUE_SET, old_value=null, new_value=true
```

{% endcode %}

We will turn to look at another subscriber which has a more elaborate diff iteration method. In our example `cdb` package, we have an application component named `CdbCfgSubscriber`. This component consists of a subscriber for the subscription point `/ncs:devices/device/config/r:sys/interfaces/interface`. The iterate() method is here implemented as an inner class called `DiffIterateImpl`.

The code for this subscriber is left out but can be found in the file `ConfigCdbSub.java`.

The example below (Run CdbCfgSubscriber Example) shows how to build and run the example.

{% code title="Example: Run CdbCfgSubscriber Example" %}

```bash
$ make clean all
$ ncs-netsim start
DEVICE ex0 OK STARTED
DEVICE ex1 OK STARTED
DEVICE ex2 OK STARTED

$ ncs

$ ncs_cli -u admin
admin@ncs# devices sync-from suppress-positive-result
admin@ncs# config
admin@ncs(config)# no devices device ex* config r:sys interfaces
admin@ncs(config)# devices device ex0 config r:sys interfaces \
> interface en0 mac 3c:07:54:71:13:09 mtu 1500 duplex half unit 0 family inet \
> address 192.168.1.115 broadcast 192.168.1.255 prefix-length 32
admin@ncs(config-address-192.168.1.115)# commit
Commit complete.
admin@ncs(config-address-192.168.1.115)# top
admin@ncs(config)# exit
```

{% endcode %}

If we look at the file `logs/ConfigCdbSub.out`, we will find log records from the subscriber (see the example below (Subscriber Output)). At the end of this file the last `DUMP DB` will show only one remaining interface.

{% code title="Example: Subscriber Output" %}

```
...
<INFO> 05-Feb-2015::16:10:23,346  ConfigCdbSub
 (cdb-examples:CdbCfgSubscriber)-Run-1: -  Device {ex0}
<INFO> 05-Feb-2015::16:10:23,346  ConfigCdbSub
 (cdb-examples:CdbCfgSubscriber)-Run-1: -     INTERFACE
<INFO> 05-Feb-2015::16:10:23,346  ConfigCdbSub
 (cdb-examples:CdbCfgSubscriber)-Run-1: -       name: {en0}
<INFO> 05-Feb-2015::16:10:23,346  ConfigCdbSub
 (cdb-examples:CdbCfgSubscriber)-Run-1: -       description:null
<INFO> 05-Feb-2015::16:10:23,350  ConfigCdbSub
 (cdb-examples:CdbCfgSubscriber)-Run-1: -       speed:null
<INFO> 05-Feb-2015::16:10:23,354  ConfigCdbSub
 (cdb-examples:CdbCfgSubscriber)-Run-1: -       duplex:half
<INFO> 05-Feb-2015::16:10:23,354  ConfigCdbSub
 (cdb-examples:CdbCfgSubscriber)-Run-1: -       mtu:1500
<INFO> 05-Feb-2015::16:10:23,354  ConfigCdbSub
 (cdb-examples:CdbCfgSubscriber)-Run-1: -       mac:<<60,7,84,113,19,9>>
<INFO> 05-Feb-2015::16:10:23,354  ConfigCdbSub
 (cdb-examples:CdbCfgSubscriber)-Run-1: -       UNIT
<INFO> 05-Feb-2015::16:10:23,354  ConfigCdbSub
 (cdb-examples:CdbCfgSubscriber)-Run-1: -        name: {0}
<INFO> 05-Feb-2015::16:10:23,355  ConfigCdbSub
 (cdb-examples:CdbCfgSubscriber)-Run-1: -        descripton: null
<INFO> 05-Feb-2015::16:10:23,355  ConfigCdbSub
 (cdb-examples:CdbCfgSubscriber)-Run-1: -        vlan-id:null
<INFO> 05-Feb-2015::16:10:23,355  ConfigCdbSub
 (cdb-examples:CdbCfgSubscriber)-Run-1: -         ADDRESS-FAMILY
<INFO> 05-Feb-2015::16:10:23,355  ConfigCdbSub
 (cdb-examples:CdbCfgSubscriber)-Run-1: -           key: {192.168.1.115}
<INFO> 05-Feb-2015::16:10:23,355  ConfigCdbSub
 (cdb-examples:CdbCfgSubscriber)-Run-1: -           prefixLength: 32
<INFO> 05-Feb-2015::16:10:23,355  ConfigCdbSub
 (cdb-examples:CdbCfgSubscriber)-Run-1: -           broadCast:192.168.1.255
<INFO> 05-Feb-2015::16:10:23,356  ConfigCdbSub
 (cdb-examples:CdbCfgSubscriber)-Run-1: -  Device {ex1}
<INFO> 05-Feb-2015::16:10:23,356  ConfigCdbSub
 (cdb-examples:CdbCfgSubscriber)-Run-1: -  Device {ex2}
```

{% endcode %}

### Operational Data

We will look once again at the YANG model for the CDB package in the [examples.ncs/sdk-api/cdb-java](https://github.com/NSO-developer/nso-examples/tree/6.7/sdk-api/cdb-java) example. Inside the `test.yang` YANG model, there is a `test` container. As a child in this container, there is a list `stats-item` (see the example below (CDB Simple Operational Data).

{% code title="Example: CDB Simple Operational Data" %}

```yang
    list stats-item {
      config false;
      tailf:cdb-oper;
      key skey;
      leaf skey {
        type string;
      }
      leaf i {
        type int32;
      }
      container inner {
        leaf  l {
          type string;
        }
      }
    }
```

{% endcode %}

Note the list `stats-item` has the substatement `config false;` and below it, we find a `tailf:cdb-oper;` statement. A standard way to implement operational data is to define a callpoint in the YANG model and write instrumentation callback methods for retrieval of the operational data (see more on data callbacks in [DP API](https://nso-docs.cisco.com/guides/development/core-concepts/pages/Uzy6qvKpLQF47FSwk0S2#ug.java_api_overview.dp)). Here on the other hand we use the `tailf:cdb-oper;` statement which implies that these instrumentation callbacks are automatically provided internally by NSO. The downside is that we must populate this operational data in CDB from the outside.

An example of Java code that creates operational data using the Navu API is shown in the example below (Creating Operational Data using Navu API)).

{% code title="Example: Creating Operational Data using Navu API" %}

```java
    public static void createEntry(String key)
            throws  IOException, ConfException {

        try (Maapi maapi = new Maapi(UnixDomainSocketAddress.of(Conf.NCS_PATH))) {
            maapi.startUserSession("system", InetAddress.getByName(null),
                                   "system", new String[]{},
                                   MaapiUserSessionFlag.PROTO_TCP);
            NavuContext operContext = new NavuContext(maapi);
            int th = operContext.startOperationalTrans(Conf.MODE_READ_WRITE);
            NavuContainer mroot = new NavuContainer(operContext);
            LOGGER.debug("ROOT --> " + mroot);

            ConfNamespace ns = new test();
            NavuContainer testModule = mroot.container(ns.hash());
            NavuList list =  testModule.container("test").list("stats-item");
            LOGGER.debug("LIST: --> " + list);

            List<ConfXMLParam> param = new ArrayList<>();
            param.add(new ConfXMLParamValue(ns, "skey", new ConfBuf(key)));
            param.add(new ConfXMLParamValue(ns, "i",
                    new ConfInt32(key.hashCode())));
            param.add(new ConfXMLParamStart(ns, "inner"));
            param.add(new ConfXMLParamValue(ns, "l", new ConfBuf("test-" + key)));
            param.add(new ConfXMLParamStop(ns, "inner"));
            list.setValues(param.toArray(new ConfXMLParam[0]));
            maapi.applyTrans(th, false);
            maapi.finishTrans(th);
            maapi.endUserSession();
        }
    }
```

{% endcode %}

An example of Java code that deletes operational data using the CDB API is shown in the example below (Deleting Operational Data using CDB API).

{% code title="Example: Deleting Operational Data using CDB API" %}

```java
    public static void deleteEntry(String key)
            throws IOException, ConfException {
        Cdb c = new Cdb("writer", UnixDomainSocketAddress.of(Conf.NCS_PATH));

        CdbSession sess = c.startSession(CdbDBType.CDB_OPERATIONAL,
                                         EnumSet.of(CdbLockType.LOCK_REQUEST,
                                                    CdbLockType.LOCK_WAIT));
        ConfPath path = new ConfPath("/t:test/stats-item{%x}",
                                     new ConfKey(new ConfBuf(key)));
        sess.delete(path);
        sess.endSession();
    }
```

{% endcode %}

In the [examples.ncs/sdk-api/cdb-java](https://github.com/NSO-developer/nso-examples/tree/6.7/sdk-api/cdb-java) example the `cdb` package, there is also an application component with an operational data subscriber that subscribes to data from the path `"/t:test/stats-item"` (see the example below (CDB Operational Subscriber Java code)).

{% code title="Example: CDB Operational Subscriber Java code" %}

```java
public class OperCdbSub implements ApplicationComponent, CdbDiffIterate {
    private static final Logger LOGGER = LogManager.getLogger(OperCdbSub.class);

    // let our ResourceManager inject Cdb sockets to us
    // no explicit creation of creating and opening sockets needed
    @Resource(type = ResourceType.CDB, scope = Scope.INSTANCE,
              qualifier = "sub-sock")
    private Cdb cdbSub;
    @Resource(type = ResourceType.CDB, scope = Scope.INSTANCE,
              qualifier = "data-sock")
    private Cdb cdbData;

    private boolean requestStop;
    private int point;
    private CdbSubscription cdbSubscription;

    public OperCdbSub() {
    }

    public void init() {
        LOGGER.info(" init oper subscriber ");
        try {
            cdbSubscription = cdbSub.newSubscription();
            String path = "/t:test/stats-item";
            point = cdbSubscription.subscribe(
                    CdbSubscriptionType.SUB_OPERATIONAL,
                    1, test.hash, path);
            cdbSubscription.subscribeDone();
            LOGGER.info("subscribeDone");
            requestStop = false;
        } catch (Exception e) {
            LOGGER.error("Fail in init", e);
        }
    }

    public void run() {
        try {
            while (!requestStop) {
                try {
                    int[] points = cdbSubscription.read();
                    CdbSession cdbSession
                            = cdbData.startSession(CdbDBType.CDB_OPERATIONAL);
                    EnumSet<DiffIterateFlags> diffFlags
                            = EnumSet.of(DiffIterateFlags.ITER_WANT_PREV);
                    cdbSubscription.diffIterate(points[0], this, diffFlags,
                                                cdbSession);
                    cdbSession.endSession();
                } finally {
                    cdbSubscription.sync(
                                    CdbSubscriptionSyncType.DONE_OPERATIONAL);
                }
            }
        } catch (Exception e) {
            LOGGER.error("Fail in run shouldrun", e);
        }
        requestStop = false;
    }

    public void finish() {
        requestStop = true;
        try {
            ResourceManager.unregisterResources(this);
        } catch (Exception e) {
            LOGGER.error("Fail in finish", e);
        }
    }

    @Override
    public DiffIterateResultFlag iterate(ConfObject[] kp,
                                         DiffIterateOperFlag op,
                                         ConfObject oldValue,
                                         ConfObject newValue,
                                         Object initstate) {
        LOGGER.info(op + " " + Arrays.toString(kp) + " value: " + newValue);
        switch (op) {
            case MOP_DELETED:
                break;
            case MOP_CREATED:
            case MOP_MODIFIED: {
                break;
            }
            default:
                break;
        }
        return DiffIterateResultFlag.ITER_RECURSE;
    }
}
```

{% endcode %}

Notice that the `CdbOperSubscriber` is very similar to the `CdbConfigSubscriber` described earlier.

In the [examples.ncs/sdk-api/cdb-py](https://github.com/NSO-developer/nso-examples/tree/6.7/sdk-api/cdb-py) and [examples.ncs/sdk-api/cdb-java](https://github.com/NSO-developer/nso-examples/tree/6.7/sdk-api/cdb-java) examples, there are two shell scripts `setoper` and `deloper` that will execute the above `CreateEntry()` and `DeleteEntry()` respectively. We can use these to populate the operational data in CDB for the `test.yang` YANG model (see the example below (Populating Operational Data)).

{% code title="Example: Populating Operational Data" %}

```bash
$ make clean all
$ ncs
$ ./setoper eth0
$ ./setoper ethX
$ ./deloper ethX
$ ncs_cli -u admin

admin@ncs# show test
SKEY  I        L
--------------------------
eth0  3123639  test-eth0
```

{% endcode %}

And if we look at the output from the 'CDB Operational Subscriber' that is found in the `logs/OperCdbSub.out`, we will see output similar to the example below (Operational subscription Output).

{% code title="Example: Operational Subscription Output" %}

```
<INFO> 05-Feb-2015::16:27:46,583  OperCdbSub
 (cdb-examples:OperSubscriber)-Run-0:
 - MOP_CREATED [{eth0}, t:stats-item, t:test] value: null
<INFO> 05-Feb-2015::16:27:46,584  OperCdbSub
 (cdb-examples:OperSubscriber)-Run-0:
 - MOP_VALUE_SET [t:skey, {eth0}, t:stats-item, t:test] value: eth0
<INFO> 05-Feb-2015::16:27:46,584  OperCdbSub
 (cdb-examples:OperSubscriber)-Run-0:
 - MOP_VALUE_SET [t:l, t:inner, {eth0}, t:stats-item, t:test] value: test-eth0
<INFO> 05-Feb-2015::16:27:46,585  OperCdbSub
 (cdb-examples:OperSubscriber)-Run-0:
 - MOP_VALUE_SET [t:i, {eth0}, t:stats-item, t:test] value: 3123639
<INFO> 05-Feb-2015::16:27:52,429  OperCdbSub
 (cdb-examples:OperSubscriber)-Run-0:
 - MOP_CREATED [{ethX}, t:stats-item, t:test] value: null
<INFO> 05-Feb-2015::16:27:52,430  OperCdbSub
 (cdb-examples:OperSubscriber)-Run-0:
 - MOP_VALUE_SET [t:skey, {ethX}, t:stats-item, t:test] value: ethX
<INFO> 05-Feb-2015::16:27:52,430  OperCdbSub
 (cdb-examples:OperSubscriber)-Run-0:
 - MOP_VALUE_SET [t:l, t:inner, {ethX}, t:stats-item, t:test] value: test-ethX
<INFO> 05-Feb-2015::16:27:52,431  OperCdbSub
 (cdb-examples:OperSubscriber)-Run-0:
 - MOP_VALUE_SET [t:i, {ethX}, t:stats-item, t:test] value: 3123679
<INFO> 05-Feb-2015::16:28:00,669  OperCdbSub
 (cdb-examples:OperSubscriber)-Run-0:
 - MOP_DELETED [{ethX}, t:stats-item, t:test] value: null
```

{% endcode %}

## Automatic Schema Upgrades and Downgrades <a href="#ug.cdb.upgrade" id="ug.cdb.upgrade"></a>

Software upgrades and downgrades represent one of the main problems in managing the configuration data of network devices. Each software release for a network device is typically associated with a certain version of configuration data layout, i.e., a schema. In NSO the schema is the data model stored in the `.fxs` files. Once CDB has initialized, it also stores a copy of the schema associated with the data it holds.

Every time NSO starts, CDB will check the current contents of the `.fxs` files with its own copy of the schema files. If CDB detects any changes in the schema, it initiates an upgrade transaction. In the simplest case, CDB automatically resolves the changes and commits the new data before NSO reaches start-phase one.

The CDB upgrade can be followed by checking the `devel.log`. The development log is meant to be used as support while the application is developed. It is enabled in `ncs.conf` as shown in the example below (Enabling Developer Logging).

{% code title="Example: Enabling Developer Logging" %}

```xml
    <developer-log>
      <enabled>true</enabled>
      <file>
        <name>./logs/devel.log</name>
        <enabled>true</enabled>
      </file>
      <syslog>
        <enabled>true</enabled>
      </syslog>
    </developer-log>
    <developer-log-level>trace</developer-log-level>
```

{% endcode %}

CDB can automatically handle the following changes to the schema:

* **Deleted elements**: When an element is deleted from the schema, CDB simply deletes it (and any children) from the database.
* **Added elements**: If a new element is added to the schema it needs to either be optional, dynamic, or have a default value. New elements with a default are added and set to their default value. New dynamic or optional elements are simply noted as a schema change.
* **Re-ordering elements**: An element with the same name, but in a different position on the same level, is considered to be the same element. If its type hasn't changed it will retain its value, but if the type has changed it will be upgraded as described below.
* **Type changes**: If a leaf is still present but its type has changed, automatic coercions are performed, so for example integers may be transformed to their string representation if the type changed from e.g. int32 to string. Automatic type conversion succeeds as long as the string representation of the current value can be parsed into its new type. (Which of course also implies that a change from a smaller integer type, e.g. int8, to a larger type, e.g., int32, succeeds for any value - while the opposite will not hold, but might!).\
  \
  If the coercion fails, any supplied default value will be used. If no default value is present in the new schema, the automatic upgrade will fail and the leaf will be deleted after the CDB upgrade.\
  \
  Note: The conversion between the `empty` and `boolean` types deviate from the aforementioned rule. Let's consider a scenario where a leaf of type `boolean` is being upgraded to a leaf of type `empty`. If the original leaf is set to `true`, it will be upgraded to a `set` empty leaf. Conversely, if the original leaf is set to `false`, it will be deleted after the upgrade. On the other hand, a `set` empty leaf will be upgraded to a leaf of type `boolean` and will be set to `true`.\
  \
  Type changes when user-defined types are used are also handled automatically, provided that some straightforward rules are followed for the type definitions. Read more about user-defined types in the confd\_types(3) manual page, which also describes these rules.
* **Node type changes**: CDB can handle automatic type changes between a container and a list. When converting from a container to a list, the child nodes of the container are mapped to the child nodes of the list, applying type coercion on the nodes when necessary. Conversely, a list can be automatically transformed into a container provided the list contains at most one list entry. Node attributes will remain intact, with the exception of the list key entry. Attributes set on a container will be transferred to the list key entry and vice versa. However, attributes on the container child node corresponding to the list key value will be lost in the upgrade.\
  \
  Additionally, type changes between leaf and leaf-list are allowed, and the data is kept intact if the number of entries in the leaf-list is exactly one. If a leaf-list has more than one entry, all entries will be deleted when upgrading to leaf.\
  \
  Type changes to and from empty leaf are possible to some extent. A type change from any type is allowed to empty leaf, but an empty leaf can only be changed to a presence container. Node attributes will only be preserved for node changes between empty leaf and container.
* **Hash changes**: When a hash value of a particular element has changed (due to an addition of, or a change to, a `tailf:id-value` statement) CDB will update that element.
* **Key changes**: When a key of a list is modified, CDB tries to upgrade the key using the same rules as explained above for adding, deleting, re-ordering, change of type, and change of hash value. If an automatic upgrade of a key fails the entire list entry will be deleted.\
  \
  When individual entries upgrade successfully but result in an invalid list, all list entries will be deleted. This can happen, e.g., when an upgrade removes a leaf from the key, resulting in several entries having the same key.
* **Default values**: If a leaf has a default value, that has not been changed from its default, then the automatic upgrade will use the new default value (if any). If the leaf value has been changed from the old default, then that value will be kept.
* **Adding / Removing namespaces**: If a namespace no longer is present after an upgrade, CDB removes all data in that namespace. When CDB detects a new namespace, it is initialized with default values.
* **Changing to/from operational**: Elements that previously had `config false` set that are changed into database elements will be treated as added elements. In the opposite case, where data elements in the new data model are tagged with `config false`, the elements will be deleted from the database.
* **Callpoint changes**: CDB only considers the part of the data model in YANG modules that do not have external data callpoints. But while upgrading, CDB handles moving subtrees into CDB from a callpoint and vice versa. CDB simply considers these as added and deleted schema elements.\
  \
  Thus an application can be developed using CDB in the first development cycle. When the external database component is ready it can easily replace CDB without changing the schema.

Should the automatic upgrade fail, exit codes and log entries will indicate the reason (see [Disaster Management](https://nso-docs.cisco.com/guides/development/core-concepts/pages/hgWUBFw1TA0R6WyLxOgc#ug.ncs_sys_mgmt.disaster)).

## Using Initialization Files for Upgrade <a href="#d5e3066" id="d5e3066"></a>

As described earlier, when NSO starts with an empty CDB database, CDB will load all instantiated XML documents found in the CDB directory and use these to initialize the database. We can also use this mechanism for CDB upgrade since CDB will again look for files in the CDB directory ending in `.xml` when doing an upgrade.

This allows for handling many of the cases that the automatic upgrade can not do by itself, e.g., the addition of mandatory leaves (without default statements), or multiple instances of new dynamic containers. Most of the time we can probably simply use the XML init file that is appropriate for a fresh install of the new version and also for the upgrade from a previous version.

When using XML files for the initialization of CDB, the complete contents of the files are used. On upgrade, however, doing this could lead to modification of the user's existing configuration - e.g., we could end up resetting data that the user has modified since CDB was first initialized. For this reason, two restrictions are applied when loading the XML files on upgrade:

* Only data for elements that are new as of the upgrade, i.e., elements that did not exist in the previous schema, will be considered.
* The data will only be loaded if all old, i.e., previously existing, optional/dynamic parent elements and instances exist in the current configuration.

To clarify this, let's make up the following example. Some `ServerManager` package was developed and delivered. It was realized that the data model had a serious shortcoming in that there was no way to specify the protocol to use, TCP or UDP. To fix this, in a new version of the package, another leaf was added to the `/servers/server` list, and the new YANG module can be seen in the example below (New YANG module for the ServerManager Package).

{% code title="Example: New YANG Module for the ServerManager Package" %}

```yang
module servers {
  namespace "http://example.com/ns/servers";
  prefix servers;

  import ietf-inet-types {
    prefix inet;
  }

  revision "2007-06-01" {
      description "added protocol.";
  }

  revision "2006-09-01" {
      description "Initial servers data model";
  }

  /*  A set of server structures  */
  container servers {
    list server {
      key name;
      max-elements 64;
      leaf name {
        type string;
      }
      leaf ip {
        type inet:ip-address;
        mandatory true;
      }
      leaf port {
        type inet:port-number;
        mandatory true;
      }
      leaf protocol {
        type enumeration {
            enum tcp;
            enum udp;
        }
        mandatory true;
      }
    }
  }
}
```

{% endcode %}

The differences from the earlier version of the YANG module can be seen in the example below (Difference between YANG Modules).

{% code title="Example: Difference between YANG Modules" %}

```diff
diff ../servers1.4.yang ../servers1.5.yang

9,12d8
>   revision "2007-06-01" {
>       description "added protocol.";
>   }
>
31,37d26
>         mandatory true;
>       }
>       leaf protocol {
>         type enumeration {
>             enum tcp;
>             enum udp;
>         }
```

{% endcode %}

Since it was considered important that the user explicitly specified the protocol, the new leaf was made mandatory. The XML init file must include this leaf, and the result can be seen in the example below (Protocol Upgrade Init File) like this:

{% code title="Example: Protocol Upgrade Init File" %}

```xml
<servers:servers xmlns:servers="http://example.com/ns/servers">
  <servers:server>
    <servers:name>www</servers:name>
    <servers:ip>192.168.3.4</servers:ip>
    <servers:port>88</servers:port>
    <servers:protocol>tcp</servers:protocol>
  </servers:server>
  <servers:server>
    <servers:name>www2</servers:name>
    <servers:ip>192.168.3.5</servers:ip>
    <servers:port>80</servers:port>
    <servers:protocol>tcp</servers:protocol>
  </servers:server>
  <servers:server>
    <servers:name>smtp</servers:name>
    <servers:ip>192.168.3.4</servers:ip>
    <servers:port>25</servers:port>
    <servers:protocol>tcp</servers:protocol>
  </servers:server>
  <servers:server>
    <servers:name>dns</servers:name>
    <servers:ip>192.168.3.5</servers:ip>
    <servers:port>53</servers:port>
    <servers:protocol>udp</servers:protocol>
  </servers:server>
</servers:servers>
```

{% endcode %}

We can then just use this new init file for the upgrade, and the existing server instances in the user's configuration will get the new `/servers/server/protocol` leaf filled in as expected. However some users may have deleted some of the original servers from their configuration, and in those cases, we do not want those servers to get re-created during the upgrade just because they are present in the XML file - the above restrictions make sure that this does not happen. The configuration after the upgrade can be seen in the example below (Configuration After Upgrade).

Here is what the configuration looks like after the upgrade if the `smtp` server has been deleted before the upgrade:

{% code title="Example: Configuration After Upgrade" %}

```xml
    <servers xmlns="http://example.com/ns/servers">
      <server>
        <name>dns</name>
        <ip>192.168.3.5</ip>
        <port>53</port>
        <protocol>udp</protocol>
      </server>
      <server>
        <name>www</name>
        <ip>192.168.3.4</ip>
        <port>88</port>
        <protocol>tcp</protocol>
      </server>
      <server>
        <name>www2</name>
        <ip>192.168.3.5</ip>
        <port>80</port>
        <protocol>tcp</protocol>
      </server>
    </servers>
```

{% endcode %}

This example also implicitly shows a limitation of this method. If the user has created additional servers, the new XML file will not specify what protocol to use for those servers, and the upgrade cannot succeed unless the package upgrade component method is used, see below. However, the example is a bit contrived. In practice, this limitation is rarely a problem. It does not occur for new lists or optional elements, nor for new mandatory elements that are not children of old lists. In fact, correctly adding this `protocol` leaf for user-created servers would require user input; it cannot be done by any fully automated procedure.

{% hint style="info" %}
Since CDB will attempt to load all `*.xml` files in the CDB directory at the time of upgrade, it is important to not leave XML init files from a previous version that are no longer valid there.
{% endhint %}

It is always possible to write a package-specific upgrade component to change the data belonging to a package before the upgrade transaction is committed. This will be explained in the following section.

## New Validation Points <a href="#cdb.upgrade-add-vp" id="cdb.upgrade-add-vp"></a>

One case the system does not handle directly is the addition of new custom validation points using the `tailf:validate` statement during an upgrade. The issue that surfaces is that the schema upgrade is performed before the (new) user code gets deployed and therefore the code required for validation is not yet available. It results in an error similar to `no registration found for callpoint NEW-VALIDATION/validate` or simply `application communication failure`.

One way to solve this problem is to first redeploy the package with the custom validation code and then perform the schema upgrade through the full `packages reload` action. For example, suppose you are upgrading the package `test-svc`. Then you first perform `packages package test-svc redeploy`, followed by `packages reload`. The main downside to this approach is that the new code must work with the old data model, which may require extra effort when there are major data model changes.

An alternative is to temporarily disable the validation by starting the NSO with the `--ignore-initial-validation` option. In this case, you should stop the `ncs` process and start it using `--ignore-initial-validation` and `--with-package-reload` options to perform the schema upgrade without custom validation. However, this may result in data in the CDB that would otherwise not pass custom validation. If you still want to validate the data, you can write an upgrade component to do this one-time validation.

## Writing an Upgrade Package Component <a href="#ncs.cdb.upgrade.comp" id="ncs.cdb.upgrade.comp"></a>

In previous sections, we showed how automatic upgrades and XML initialization files can help in upgrading CDB when YANG models have changed. In some situations, this is not sufficient. For instance, if a YANG model is changed and new mandatory leaves are introduced that need calculations to set the values then a programmatic upgrade is needed. This is when the upgrade component of a package comes into play.

An `upgrade` component is a Java or Python class with a standard `main()` method that becomes a standalone program that is run as part of the package `reload` action.

As with any package component type, the `upgrade` component has to be defined in the `package-meta-data.xml` file for the package (see the example below (Upgrade Package Components)).

{% code title="Example: Upgrade Package Components" %}

```xml
<ncs-package xmlns="http://tail-f.com/ns/ncs-packages">
    ....
  <component>
    <name>do-upgrade</name>
    <upgrade>
      <java-class-name>com.example.DoUpgrade</java-class-name>
    </upgrade>
  </component>
</ncs-package>
```

{% endcode %}

Let's recapitulate how packages are loaded and reloaded. NSO can search the `/ncs-config/load-path` for packages to run and will copy these to a private directory tree under `/ncs-config/state-dir` with root directory `packages-in-use.cur`. However, NSO will only do this search when `packages-in-use.cur` is empty or when a `reload` is requested. This scheme makes package upgrades controlled and predictable, for more on this, see [Loading Packages](https://nso-docs.cisco.com/guides/development/core-concepts/pages/TpxOL8qGnNMkxkp0ZCeB#ug.package_mgmt.loading).

<div data-with-frame="true"><figure><img src="/files/BwvN276QfWEgkGjqRXr1" alt="" width="563"><figcaption><p>NSO Package before Reload</p></figcaption></figure></div>

So in preparation for a package upgrade, the new packages replace the old ones in the load path. In our scenario, the YANG model changes are such that the automatic schema upgrade that CDB performs is not sufficient, therefore the new packages also contain `upgrade` components. At this point, NSO is still running with the old package definitions.

<div data-with-frame="true"><figure><img src="/files/y4ln1Xb3a3ZE2TvwPvgg" alt="" width="563"><figcaption><p>NSO Package at Reload</p></figcaption></figure></div>

When the package reload is requested, the packages in the load path are copied to the state directory. The old state directory is scratched, so that packages that no longer exist in the load path are removed and new packages are added. Unchanged packages will be unchanged. Automatic schema CDB upgrades will be performed, and afterward, for all packages that have an upgrade component and for which at least one YANG model was changed, this upgrade component will be executed. Also for added packages that have an upgrade component, this component will be executed. Hence the upgrade component needs to be programmed in such a way that care is taken for both the `new` and `upgrade` package scenarios.

So how should an upgrade component be implemented? In the previous section, we described how CDB can perform an automatic upgrade. But this means that CDB has deleted all values that are no longer part of the schema. Well, not quite yet. At the initial phase of the NSO startup procedure (called start-phase0), it is possible to use all the CDB Java/Python API calls to access the data using the schema from the database as it looked before the automatic upgrade. That is, the complete database as it stood before the upgrade is still available to the application. It is under this condition that the upgrade components are executed and this is the reason why they are standalone programs and not executed by the NSO Java/Python-VM as all other Java/Python code for components are.

So the CDB Java/Python API can be used to read data defined by the old YANG models. To write new config data Maapi has a specific method `Maapi.attachInit()`. This method attaches a Maapi instance to the upgrade transaction (or init transaction) during `phase0`. This special upgrade transaction is only available during `phase0`. NSO will commit this transaction when the `phase0` is ended, so the user should only write config data (not attempt to commit, etc.).

We take a look at the example [examples.ncs/service-management/upgrade-service](https://github.com/NSO-developer/nso-examples/tree/6.7/service-management/upgrade-service) to see how an upgrade component can be implemented. Here the *vlan* package has an original version which is replaced with a version `vlan_v2`. See the `vlan_v2-py` package for a Python variant. See the `README` and play with examples to get acquainted.

{% hint style="info" %}
The `upgrade-service` is a `service` package upgrade example. But the upgrade components here described work equally well and in the same way for any package type. The only requirement is that the package contain at least one YANG model for the upgrade component to have meaning. If not the upgrade component will never be executed.
{% endhint %}

The complete YANG model for the version 2 of the VLAN service looks as follows:

{% code title="Example: VLAN Service v2 YANG Model" %}

```yang
module vlan-service {
  namespace "http://example.com/vlan-service";
  prefix vl;

  import tailf-common {
    prefix tailf;
  }
  import tailf-ncs {
    prefix ncs;
  }

  description
    "This service creates a vlan iface/unit on all routers in our network. ";

  revision 2013-08-30 {
    description
      "Added mandatory leaf global-id.";
  }
  revision 2013-01-08 {
    description
      "Initial revision.";
  }

  augment /ncs:services {
    list vlan {
      key name;
      leaf name {
        tailf:info "Unique service id";
        tailf:cli-allow-range;
        type string;
      }

      uses ncs:service-data;
      ncs:servicepoint vlanspnt_v2;

      tailf:action self-test {
        tailf:info "Perform self-test of the service";
        tailf:actionpoint vlanselftest;
        output {
          leaf success {
            type boolean;
          }
          leaf message {
            type string;
            description
              "Free format message.";
          }
        }
      }

      leaf global-id {
        type string;
        mandatory true;
      }
      leaf iface {
        type string;
        mandatory true;
      }
      leaf unit {
        type int32;
        mandatory true;
      }
      leaf vid {
        type uint16;
        mandatory true;
      }
      leaf description {
        type string;
        mandatory true;
      }
    }
  }
}
```

{% endcode %}

If we `diff` the changes between the two YANG models for the service, we see that in version 2, a new mandatory leaf has been added (see the example below (YANG Service diff)).

{% code title="Example: YANG Service diff" %}

```bash
$ diff vlan/src/yang/vlan-service.yang \
                     vlan_v2/src/yang/vlan-service.yang
16a18,22
>   revision 2013-08-30 {
>     description
>       "Added mandatory leaf global-id.";
>   }
>
48a55,58
>       leaf global-id {
>         type string;
>         mandatory true;
>       }
68c78
```

{% endcode %}

We need to create a Java class with a `main()` method that connects to CDB and MAAPI. This main will be executed as a separate program and all private and shared jars defined by the package will be in the classpath. To upgrade the VLAN service, the following Java code is needed:

{% code title="Example: VLAN Service Upgrade Component Java Class" %}

```java
public class UpgradeService {

    public UpgradeService() {
    }

    public static void main(String[] args) throws Exception {
        SocketAddress address = UnixDomainSocketAddress.of(Conf.NCS_PATH);
        Cdb cdb = new Cdb("cdb-upgrade-sock", address);
        cdb.setUseForCdbUpgrade();
        CdbUpgradeSession cdbsess =
            cdb.startUpgradeSession(
                    CdbDBType.CDB_RUNNING,
                    EnumSet.of(CdbLockType.LOCK_SESSION,
                               CdbLockType.LOCK_WAIT));


        Maapi maapi = new Maapi(address);
        int th = maapi.attachInit();

        int no = cdbsess.getNumberOfInstances("/services/vlan");
        for(int i = 0; i < no; i++) {
            Integer offset = Integer.valueOf(i);
            ConfBuf name = (ConfBuf)cdbsess.getElem("/services/vlan[%d]/name",
                                                    offset);
            ConfBuf iface = (ConfBuf)cdbsess.getElem("/services/vlan[%d]/iface",
                                                    offset);
            ConfInt32 unit =
                (ConfInt32)cdbsess.getElem("/services/vlan[%d]/unit",
                                           offset);
            ConfUInt16 vid =
                (ConfUInt16)cdbsess.getElem("/services/vlan[%d]/vid",
                                            offset);

            String nameStr = name.toString();
            System.out.println("SERVICENAME = " + nameStr);

            String globId = String.format("%1$s-%2$s-%3$s", iface.toString(),
                                          unit.toString(), vid.toString());
            ConfPath gidpath = new ConfPath("/services/vlan{%s}/global-id",
                                            name.toString());
            maapi.setElem(th, new ConfBuf(globId), gidpath);
        }

    }
}
```

{% endcode %}

Let's go through the code and point out the different aspects of writing an `upgrade` component. First (see the example below (Upgrade Init)) we create a Java API `Cdb` connection to NSO over Local IPC and call `Cdb.setUseForCdbUpgrade()`. This method will prepare `cdb` sessions for reading old data from the CDB database, and it should only be called in this context. At the end of this first code fragment, we start the CDB upgrade session:

{% code title="Example: Upgrade Init" %}

```java
        SocketAddress address = UnixDomainSocketAddress.of(Conf.NCS_PATH);
        Cdb cdb = new Cdb("cdb-upgrade-sock", address);
        cdb.setUseForCdbUpgrade();
        CdbUpgradeSession cdbsess =
            cdb.startUpgradeSession(
                    CdbDBType.CDB_RUNNING,
                    EnumSet.of(CdbLockType.LOCK_SESSION,
                               CdbLockType.LOCK_WAIT));
```

{% endcode %}

We then create a Java API Maapi connection to NSO and call the `Maapi.attachInit()` method to get the init transaction (see the example below (Upgrade Get Transaction)).

{% code title="Example: Upgrade Get Transaction" %}

```java
        Maapi maapi = new Maapi(address);
        int th = maapi.attachInit();
```

{% endcode %}

Using the `CdbSession` instance we read the number of service instance that exists in the CDB database. We will work on all these instances. Also, if the number of instances is zero the loop will not be entered. This is a simple way to prevent the upgrade component from doing any harm in the case of this being a new package that is added to NSO for the first time:

```java
        int no = cdbsess.getNumberOfInstances("/services/vlan");
        for(int i = 0; i < no; i++) {
```

Via the `CdbUpgradeSession`, the old service data is retrieved:

```java
            ConfBuf name = (ConfBuf)cdbsess.getElem("/services/vlan[%d]/name",
                                                    offset);
            ConfBuf iface = (ConfBuf)cdbsess.getElem("/services/vlan[%d]/iface",
                                                    offset);
            ConfInt32 unit =
                (ConfInt32)cdbsess.getElem("/services/vlan[%d]/unit",
                                           offset);
            ConfUInt16 vid =
                (ConfUInt16)cdbsess.getElem("/services/vlan[%d]/vid",
                                            offset);
```

The value for the new leaf introduced in the new version of the YANG model is calculated, and the value is set using Maapi and the init transaction:

```java
            String globId = String.format("%1$s-%2$s-%3$s", iface.toString(),
                                          unit.toString(), vid.toString());
            ConfPath gidpath = new ConfPath("/services/vlan{%s}/global-id",
                                            name.toString());
            maapi.setElem(th, new ConfBuf(globId), gidpath);
```

At the end of the program, the sockets are closed. Important to note is that no commits or other handling of the init transaction is done. This is NSO's responsibility:

```java
        s1.close();
        s2.close();
```

<div data-with-frame="true"><figure><img src="/files/tMj2x1oBWVtbz24yk7JX" alt="" width="563"><figcaption><p>NSO Advanced Service Upgrade</p></figcaption></figure></div>

In the [examples.ncs/service-management/upgrade-service](https://github.com/NSO-developer/nso-examples/tree/6.7/service-management/upgrade-service) example, this more complicated scenario is illustrated with the `tunnel` package. See the `tunnel-py` package for a Python variant. The `tunnel` package YANG model maps the `vlan_v2` package one-to-one but is a complete rename of the model containers and all leafs:

{% code title="Example: Tunnel Service YANG Model" %}

```yang
module tunnel-service {
  namespace "http://example.com/tunnel-service";
  prefix tl;

  import tailf-common {
    prefix tailf;
  }
  import tailf-ncs {
    prefix ncs;
  }

  description
    "This service creates a tunnel assembly on all routers in our network. ";

  revision 2013-01-08 {
    description
      "Initial revision.";
  }

  augment /ncs:services {
    list tunnel {
      key tunnel-name;
      leaf tunnel-name {
        tailf:info "Unique service id";
        tailf:cli-allow-range;
        type string;
      }

      uses ncs:service-data;
      ncs:servicepoint tunnelspnt;

      tailf:action self-test {
        tailf:info "Perform self-test of the service";
        tailf:actionpoint tunnelselftest;
        output {
          leaf success {
            type boolean;
          }
          leaf message {
            type string;
            description
              "Free format message.";
          }
        }
      }

      leaf gid {
        type string;
        mandatory true;
      }
      leaf interface {
        type string;
        mandatory true;
      }
      leaf assembly {
        type int32;
        mandatory true;
      }
      leaf tunnel-id {
        type uint16;
        mandatory true;
      }
      leaf descr {
        type string;
        mandatory true;
      }
    }
  }
}
```

{% endcode %}

To upgrade from the `vlan_v2` to the `tunnel` package, a new upgrade component for the `tunnel` package has to be implemented:

{% code title="Example: Tunnel Service Upgrade Java Class" %}

```java
public class UpgradeService {

    public UpgradeService() {
    }

    public static void main(String[] args) throws Exception {
        ArrayList<ConfNamespace> nsList = new ArrayList<ConfNamespace>();
        nsList.add(new vlanService());
        SocketAddress address = UnixDomainSocketAddress.of(Conf.NCS_PATH);
        Cdb cdb = new Cdb("cdb-upgrade-sock", address);
        cdb.setUseForCdbUpgrade(nsList);
        CdbUpgradeSession cdbsess =
            cdb.startUpgradeSession(
                    CdbDBType.CDB_RUNNING,
                    EnumSet.of(CdbLockType.LOCK_SESSION,
                               CdbLockType.LOCK_WAIT));


        Maapi maapi = new Maapi(address);
        int th = maapi.attachInit();

        int no = cdbsess.getNumberOfInstances("/services/vlan");
        for(int i = 0; i < no; i++) {
            ConfBuf name =(ConfBuf)cdbsess.getElem("/services/vlan[%d]/name",
                                                   Integer.valueOf(i));
            String nameStr = name.toString();
            System.out.println("SERVICENAME = " + nameStr);

            ConfCdbUpgradePath oldPath =
                new ConfCdbUpgradePath("/ncs:services/vl:vlan{%s}",
                                       name.toString());
            ConfPath newPath = new ConfPath("/services/tunnel{%x}", name);
            maapi.create(th, newPath);

            ConfXMLParam[] oldparams = new ConfXMLParam[] {
                new ConfXMLParamLeaf("vl", "global-id"),
                new ConfXMLParamLeaf("vl", "iface"),
                new ConfXMLParamLeaf("vl", "unit"),
                new ConfXMLParamLeaf("vl", "vid"),
                new ConfXMLParamLeaf("vl", "description"),
            };
            ConfXMLParam[] data =
                cdbsess.getValues(oldparams, oldPath);

            ConfXMLParam[] newparams = new ConfXMLParam[] {
                new ConfXMLParamValue("tl", "gid",       data[0].getValue()),
                new ConfXMLParamValue("tl", "interface", data[1].getValue()),
                new ConfXMLParamValue("tl", "assembly",  data[2].getValue()),
                new ConfXMLParamValue("tl", "tunnel-id", data[3].getValue()),
                new ConfXMLParamValue("tl", "descr",     data[4].getValue()),
            };
            maapi.setValues(th, newparams, newPath);

            maapi.ncsMovePrivateData(th, oldPath, newPath);
        }

    }
}
```

{% endcode %}

We will walk through this code also and point out the aspects that differ from the earlier more simple scenario. First, we want to create the `Cdb` instance and get the CdbSession. However, in this scenario, the old namespace is removed and the Java API cannot retrieve it from NSO. To be able to use CDB to read and interpret the old YANG Model, the old generated and removed Java namespace classes have to be temporarily reinstalled. This is solved by adding a jar (Java archive) containing these removed namespaces to the `private-jar` directory of the tunnel package. The removed namespace can then be instantiated and passed to Cdb via an overridden version of the `Cdb.setUseForCdbUpgrade()` method:

```java
        ArrayList<ConfNamespace> nsList = new ArrayList<ConfNamespace>();
        nsList.add(new vlanService());
        SocketAddress address = UnixDomainSocketAddress.of(Conf.NCS_PATH);
        Cdb cdb = new Cdb("cdb-upgrade-sock", address);
        cdb.setUseForCdbUpgrade(nsList);
        CdbUpgradeSession cdbsess =
            cdb.startUpgradeSession(
                    CdbDBType.CDB_RUNNING,
                    EnumSet.of(CdbLockType.LOCK_SESSION,
                               CdbLockType.LOCK_WAIT));
```

As an alternative to including the old namespace file in the package, a `ConfNamespaceStub` can be constructed for each old model that is to be accessed:

```java
nslist.add(new ConfNamespaceStub(500805321,
                                 "http://example.com/vlan-service",
                                 "http://example.com/vlan-service",
                                 "vl"));
```

Since the old YANG model with the service point is removed, the new service container with the new service has to be created before any config data can be written to this position:

```java
            ConfPath newPath = new ConfPath("/services/tunnel{%x}", name);
            maapi.create(th, newPath);
```

The complete config for the old service is read via the `CdbUpgradeSession`. Note in particular that the path `oldPath` is constructed as a `ConfCdbUpgradePath`. These are the paths that allow access to nodes that are not available in the current schema (i.e., nodes in deleted models).

```java
            ConfXMLParam[] oldparams = new ConfXMLParam[] {
                new ConfXMLParamLeaf("vl", "global-id"),
                new ConfXMLParamLeaf("vl", "iface"),
                new ConfXMLParamLeaf("vl", "unit"),
                new ConfXMLParamLeaf("vl", "vid"),
                new ConfXMLParamLeaf("vl", "description"),
            };
            ConfXMLParam[] data =
                cdbsess.getValues(oldparams, oldPath);
```

The new data structure with the service data is created and written to NSO via Maapi and the init transaction:

```java
            ConfXMLParam[] newparams = new ConfXMLParam[] {
                new ConfXMLParamValue("tl", "gid",       data[0].getValue()),
                new ConfXMLParamValue("tl", "interface", data[1].getValue()),
                new ConfXMLParamValue("tl", "assembly",  data[2].getValue()),
                new ConfXMLParamValue("tl", "tunnel-id", data[3].getValue()),
                new ConfXMLParamValue("tl", "descr",     data[4].getValue()),
            };
            maapi.setValues(th, newparams, newPath);
```


# YANG

Learn the working aspects of YANG data modeling language in NSO.

YANG is a data modeling language used to model configuration and state data manipulated by a NETCONF agent. The YANG modeling language is defined in RFC 6020 (version 1) and RFC 7950 (version 1.1). YANG as a language will not be described in its entirety here - rather, we refer to the IETF RFC text at [RFC6020](https://www.ietf.org/rfc/rfc6020.txt) and [RFC7950](https://www.ietf.org/rfc/rfc7950.txt).

## YANG in NSO <a href="#d5e1847" id="d5e1847"></a>

In NSO, YANG is not only used for NETCONF data. On the contrary, YANG is used to describe the data model as a whole and used by all northbound interfaces.

NSO uses YANG for Service Models as well as for specifying device interfaces. Where do these models come from? When it comes to services, the YANG service model is specified as part of the service design activity. NSO ships several examples of service models that can be used as a starting point. For devices, it depends on the underlying device interface how the YANG model is derived. For native NETCONF/YANG devices the YANG model is of course given by the device. For SNMP devices, the NSO tool-chain generates the corresponding YANG modules, (SNMP NED). For CLI devices, the package for the device contains the YANG data model. This is shipped in text and can be modified to cater for upgrades. Customers can also write their own YANG data models to render the CLI integration (CLI NED). The situation for other interfaces is similar to CLI, a YANG model that corresponds to the device interface data model is written and bundled in the NED package.

NSO also relies on the revision statement in YANG modules for revision management of different versions of the same type of managed device, but running different software versions.

A YANG module can be directly transformed into a final schema (.fxs) file that can be loaded into NSO. Currently, all features of the YANG 1.0 language are supported where `anyxml` statement data is treated as a string. Most features of the YANG 1.1 language are supported. For a list of exceptions, please refer to the `YANG 1.1` section of the `ncsc` man page.

The data models including the .fxs file along with any code are bundled into packages that can be loaded to NSO. This is true for service applications as well as for NEDs and other packages. The corresponding YANG can be found in the `src/yang` directory in the package.

## YANG Introduction <a href="#d5e1856" id="d5e1856"></a>

This section is a brief introduction to YANG. The exact details of all language constructs are fully described in RFC 6020 and RFC 7950.

The NSO programmer must know YANG well since all APIs use various paths that are derived from the YANG data model.

### Modules and Submodules <a href="#d5e1860" id="d5e1860"></a>

A module contains three types of statements: module-header statements, revision statements, and definition statements. The module header statements describe the module and give information about the module itself, the revision statements give information about the history of the module, and the definition statements are the body of the module where the data model is defined.

A module may be divided into submodules, based on the needs of the module owner. The external view remains that of a single module, regardless of the presence or size of its submodules.

The `include` statement allows a module or submodule to reference material in submodules, and the `import` statement allows references to material defined in other modules.

### Data Modeling Basics <a href="#d5e1867" id="d5e1867"></a>

YANG defines four types of nodes for data modeling. In each of the following subsections, the example shows the YANG syntax as well as a corresponding NETCONF XML representation.

### Leaf Nodes <a href="#d5e1870" id="d5e1870"></a>

A leaf node contains simple data like an integer or a string. It has exactly one value of a particular type and no child nodes.

```yang
leaf host-name {
    type string;
    description "Hostname for this system";
}
```

With XML value representation for example:

```xml
<host-name>my.example.com</host-name>
```

An interesting variant of leaf nodes is typeless leafs.

```yang
leaf enabled {
    type empty;
    description "Enable the interface";
}
```

With XML value representation for example:

```xml
<enabled/>
```

### Leaf-list Nodes <a href="#d5e1884" id="d5e1884"></a>

A `leaf-list` is a sequence of leaf nodes with exactly one value of a particular type per leaf.

```
leaf-list domain-search {
         type string;
         description "List of domain names to search";
     }
```

With XML value representation for example:

```xml
<domain-search>high.example.com</domain-search>
<domain-search>low.example.com</domain-search>
<domain-search>everywhere.example.com</domain-search>
```

### Container Nodes <a href="#d5e1893" id="d5e1893"></a>

A `container` node is used to group related nodes in a subtree. It has only child nodes and no value and may contain any number of child nodes of any type (including leafs, lists, containers, and leaf-lists).

```yang
container system {
    container login {
        leaf message {
            type string;
            description
                "Message given at start of login session";
        }
    }
}
```

With XML value representation for example:

```xml
<system>
  <login>
    <message>Good morning, Dave</message>
  </login>
</system>
```

### List Nodes <a href="#d5e1902" id="d5e1902"></a>

A `list` defines a sequence of list entries. Each entry is like a structure or a record instance and is uniquely identified by the values of its key leafs. A list can define multiple keys and may contain any number of child nodes of any type (including leafs, lists, containers, etc.).

```yang
list user {
    key "name";
    leaf name {
        type string;
    }
    leaf full-name {
        type string;
    }
    leaf class {
        type string;
    }
}
```

With XML value representation for example:

```xml
<user>
  <name>glocks</name>
  <full-name>Goldie Locks</full-name>
  <class>intruder</class>
</user>
<user>
  <name>snowey</name>
  <full-name>Snow White</full-name>
  <class>free-loader</class>
</user>
<user>
  <name>rzull</name>
  <full-name>Repun Zell</full-name>
  <class>tower</class>
</user>
```

### Example Module <a href="#d5e1911" id="d5e1911"></a>

These statements are combined to define the module:

```
// Contents of "acme-system.yang"
module acme-system {
    namespace "http://acme.example.com/system";
    prefix "acme";

    organization "ACME Inc.";
    contact "joe@acme.example.com";
    description
        "The module for entities implementing the ACME system.";

    revision 2007-06-09 {
        description "Initial revision.";
    }

    container system {
        leaf host-name {
            type string;
            description "Hostname for this system";
        }

        leaf-list domain-search {
            type string;
            description "List of domain names to search";
        }

        container login {
            leaf message {
                type string;
                description
                    "Message given at start of login session";
            }

            list user {
                key "name";
                leaf name {
                    type string;
                }
                leaf full-name {
                    type string;
                }
                leaf class {
                    type string;
                }
            }
        }
    }
}
```

### State Data <a href="#d5e1916" id="d5e1916"></a>

YANG can model state data, as well as configuration data, based on the `config` statement. When a node is tagged with `config false`, its sub-hierarchy is flagged as state data, to be reported using NETCONF's `get` operation, not the `get-config` operation. Parent containers, lists, and key leafs are reported also, giving the context for the state data.

In this example, two leafs are defined for each interface, a configured speed, and an observed speed. The observed speed is not a configuration, so it can be returned with NETCONF `get` operations, but not with `get-config` operations. The observed speed is not configuration data, and cannot be manipulated using `edit-config`.

```yang
list interface {
    key "name";
    config true;

    leaf name {
        type string;
    }
    leaf speed {
        type enumeration {
            enum 10m;
            enum 100m;
            enum auto;
        }
    }
    leaf observed-speed {
        type uint32;
        config false;
    }
}
```

### Built-in Types <a href="#d5e1929" id="d5e1929"></a>

YANG has a set of built-in types, similar to those of many programming languages, but with some differences due to special requirements from the management domain. The following table summarizes the built-in types.

The table below lists YANG built-in types:

| Name                | Type        | Description                                       |
| ------------------- | ----------- | ------------------------------------------------- |
| binary              | Text        | Any binary data                                   |
| bits                | Text/Number | A set of bits or flags                            |
| boolean             | Text        | `true` or `false`                                 |
| decimal64           | Number      | 64-bit fixed point real number                    |
| empty               | Empty       | A leaf that does not have any value               |
| enumeration         | Text/Number | Enumerated strings with associated numeric values |
| identityref         | Text        | A reference to an abstract identity               |
| instance-identifier | Text        | References a data tree node                       |
| int8                | Number      | 8-bit signed integer                              |
| int16               | Number      | 16-bit signed integer                             |
| int32               | Number      | 32-bit signed integer                             |
| int64               | Number      | 64-bit signed integer                             |
| leafref             | Text/Number | A reference to a leaf instance                    |
| string              | Text        | Human readable string                             |
| uint8               | Number      | 8-bit unsigned integer                            |
| uint16              | Number      | 16-bit unsigned integer                           |
| uint32              | Number      | 32-bit unsigned integer                           |
| uint64              | Number      | 64-bit unsigned integer                           |
| union               | Text/Number | Choice of member types                            |

### Derived Types (`typedef`)

YANG can define derived types from base types using the `typedef` statement. A base type can be either a built-in type or a derived type, allowing a hierarchy of derived types. A derived type can be used as the argument for the `type` statement.

```
typedef percent {
    type uint16 {
        range "0 .. 100";
    }
    description "Percentage";
}

leaf completed {
    type percent;
}
```

With XML value representation for example:

```xml
<completed>20</completed>
```

User-defined typedefs are useful when we want to name and reuse a type several times. It is also possible to restrict leafs inline in the data model as in:

```yang
leaf completed {
    type uint16 {
        range "0 .. 100";
    }
    description "Percentage";
}
```

### Reusable Node Groups (`grouping`) <a href="#d5e2029" id="d5e2029"></a>

Groups of nodes can be assembled into the equivalent of complex types using the `grouping` statement. `grouping` defines a set of nodes that are instantiated with the `uses` statement:

```
grouping target {
    leaf address {
        type inet:ip-address;
        description "Target IP address";
    }
    leaf port {
        type inet:port-number;
        description "Target port number";
    }
}

container peer {
    container destination {
        uses target;
    }
}
```

With XML value representation for example:

```xml
<peer>
  <destination>
    <address>192.0.2.1</address>
    <port>830</port>
  </destination>
</peer>
```

The grouping can be refined as it is used, allowing certain statements to be overridden. In this example, the description is refined:

```yang
container connection {
    container source {
        uses target {
            refine "address" {
                description "Source IP address";
            }
            refine "port" {
                description "Source port number";
            }
        }
    }
    container destination {
        uses target {
            refine "address" {
                description "Destination IP address";
            }
            refine "port" {
                description "Destination port number";
            }
        }
    }
}
```

### Choices (`choice`) <a href="#d5e2043" id="d5e2043"></a>

YANG allows the data model to segregate incompatible nodes into distinct choices using the `choice` and `case` statements. The `choice` statement contains a set of `case` statements that define sets of schema nodes that cannot appear together. Each `case` may contain multiple nodes, but each node may appear in only one `case` under a `choice`.

When the nodes from one case are created, all nodes from all other cases are implicitly deleted. The device handles the enforcement of the constraint, preventing incompatibilities from existing in the configuration.

The choice and case nodes appear only in the schema tree, not in the data tree or XML encoding. The additional levels of hierarchy are not needed beyond the conceptual schema.

```yang
container food {
   choice snack {
       mandatory true;
       case sports-arena {
           leaf pretzel {
               type empty;
           }
           leaf beer {
               type empty;
           }
       }
       case late-night {
           leaf chocolate {
               type enumeration {
                   enum dark;
                   enum milk;
                   enum first-available;
               }
           }
       }
   }
}
```

With XML value representation for example:

```xml
<food>
  <chocolate>first-available</chocolate>
</food>
```

### Extending Data Models (`augment`) <a href="#d5e2060" id="d5e2060"></a>

YANG allows a module to insert additional nodes into data models, including both the current module (and its submodules) or an external module. This is useful e.g. for vendors to add vendor-specific parameters to standard data models in an interoperable way.

The `augment` statement defines the location in the data model hierarchy where new nodes are inserted, and the `when` statement defines the conditions when the new nodes are valid.

```yang
augment /system/login/user {
    when "class != 'wheel'";
    leaf uid {
        type uint16 {
            range "1000 .. 30000";
        }
    }
}
```

This example defines a `uid` node that only is valid when the user's `class` is not `wheel`.

If a module augments another model, the XML representation of the data will reflect the prefix of the augmenting model. For example, if the above augmentation were in a module with the prefix `other`, the XML would look like:

```xml
<user>
  <name>alicew</name>
  <full-name>Alice N. Wonderland</full-name>
  <class>drop-out</class>
  <other:uid>1024</other:uid>
</user>
```

### RPC Definitions <a href="#d5e2076" id="d5e2076"></a>

YANG allows the definition of NETCONF RPCs. The method names, input parameters, and output parameters are modeled using YANG data definition statements.

```
rpc activate-software-image {
    input {
        leaf image-name {
            type string;
        }
    }
    output {
        leaf status {
            type string;
        }
    }
}
```

```xml
<rpc message-id="101"
     xmlns="urn:ietf:params:xml:ns:netconf:base:1.0">
  <activate-software-image xmlns="http://acme.example.com/system">
    <name>acmefw-2.3</name>
 </activate-software-image>
</rpc>

<rpc-reply message-id="101"
           xmlns="urn:ietf:params:xml:ns:netconf:base:1.0">
  <status xmlns="http://acme.example.com/system">
    The image acmefw-2.3 is being installed.
  </status>
</rpc-reply>
```

### Notification Definitions <a href="#d5e2083" id="d5e2083"></a>

YANG allows the definition of notifications suitable for NETCONF. YANG data definition statements are used to model the content of the notification.

```
notification link-failure {
    description "A link failure has been detected";
    leaf if-name {
        type leafref {
            path "/interfaces/interface/name";
        }
    }
    leaf if-admin-status {
        type ifAdminStatus;
    }
}
```

```xml
<notification xmlns="urn:ietf:params:netconf:capability:notification:1.0">
  <eventTime>2007-09-01T10:00:00Z</eventTime>
  <link-failure xmlns="http://acme.example.com/system">
    <if-name>so-1/2/3.0</if-name>
    <if-admin-status>up</if-admin-status>
  </link-failure>
</notification>
```

## Working With YANG Modules <a href="#d5e2090" id="d5e2090"></a>

Assume we have a small trivial YANG file `test.yang`:

```yang
module test {
  namespace "http://tail-f.com/test";
  prefix "t";

  container top {
      leaf a {
          type int32;
      }
      leaf b {
          type string;
      }
  }
}
```

{% hint style="success" %}
There is an Emacs mode suitable for YANG file editing in the system distribution. It is called `yang-mode.el`.
{% endhint %}

We can use `ncsc` compiler to compile the YANG module.

```bash
$ ncsc -c test.yang
```

The above command creates an output file `test.fxs` that is a compiled schema that can be loaded into the system. The `ncsc` compiler with all its flags is fully described in [ncsc(1)](/guides/resources/man/ncsc.1) in Manual Pages.

There exist several standards-based auxiliary YANG modules defining various useful data types. These modules, as well as their accompanying `.fxs` files can be found in the `${NCS_DIR}/src/confd/yang` directory in the distribution.

The modules are:

* `ietf-yang-types`: Defining some basic data types such as counters, dates, and times.
* `ietf-inet-types`: Defining several useful types related to IP addresses.

Whenever we wish to use any of those predefined modules we need to not only import the module into our YANG module, but we must also load the corresponding .fxs file for the imported module into the system.

So, if we extend our test module so that it looks like:

```yang
module test {
    namespace "http://tail-f.com/test";
    prefix "t";

    import ietf-inet-types {
        prefix inet;
    }

    container top {
        leaf a {
            type int32;
        }
        leaf b {
            type string;
        }
        leaf ip {
            type inet:ipv4-address;
        }
    }
}
```

Normally when importing other YANG modules we must indicate through the `--yangpath` flag to `ncsc` where to search for the imported module. In the special case of the standard modules, this is not required.

We compile the above as:

```bash
$ ncsc -c test.yang
$ ncsc --get-info test.fxs
fxs file
Ncsc version:           "3.0_2"
uri:                    http://tail-f.com/test
id:                     http://tail-f.com/test
prefix:                 "t"
flags:                  6
type:                   cs
mountpoint:             undefined
exported agents:        all
dependencies:           ['http://www.w3.org/2001/XMLSchema',
                         'urn:ietf:params:xml:ns:yang:inet-types']
source:                 ["test.yang"]
```

We see that the generated `.fxs` file has a dependency on the standard `urn:ietf:params:xml:ns:yang:inet-types` namespace. Thus if we try to start NSO we must also ensure that the fxs file for that namespace is loaded.

Failing to do so gives:

```bash
$ ncs -c ncs.conf --foreground --verbose
The namespace urn:ietf:params:xml:ns:yang:inet-types (referenced by http://tail-f.com/test) could not be found in the loadPath.
Daemon died status=21
```

The remedy is to modify `ncs.conf` so that it contains the proper load path or to provide the directory containing the `fxs` file, alternatively, we can provide the path on the command line. The directory `${NCS_DIR}/etc/ncs` contains pre-compiled versions of the standard YANG modules.

```bash
$ ncs -c ncs.conf --addloadpath ${NCS_DIR}/etc/ncs --foreground --verbose
```

`ncs.conf` is the configuration file for NSO itself. It is described in the [ncs.conf(5)](/guides/resources/man/ncs.conf.5) in Manual Pages.

## Integrity Constraints <a href="#d5e2141" id="d5e2141"></a>

The YANG language has built-in declarative constructs for common integrity constraints. These constructs are conveniently specified as `must` statements.

A `must` statement is an XPath expression that must evaluate to true or a non-empty node-set.

An example is:

```yang
 container interface {
    leaf ifType {
        type enumeration {
            enum ethernet;
            enum atm;
        }
    }
    leaf ifMTU {
        type uint32;
    }
    must "ifType != 'ethernet' or "
      +  "(ifType = 'ethernet' and ifMTU = 1500)" {
        error-message "An ethernet MTU must be 1500";
    }
    must "ifType != 'atm' or "
       + "(ifType = 'atm' and ifMTU <= 17966 and ifMTU >= 64)" {
        error-message "An atm MTU must be  64 .. 17966";
    }
}
```

XPath is a very powerful tool here. It is often possible to express the most realistic validation constraints using XPath expressions. Note that for performance reasons, it is recommended to use the `tailf:dependency` statement in the `must` statement. The compiler gives a warning if a `must` statement lacks a `tailf:dependency` statement, and it cannot derive the dependency from the expression. The options `--fail-on-warnings` or `-E TAILF_MUST_NEED_DEPENDENCY` can be given to force this warning to be treated as an error. See `tailf:dependency` in [tailf\_yang\_extensions(5)](/guides/resources/man/tailf_yang_extensions.5) in Manual Pages for details.

Another useful built-in constraint checker is the `unique` statement.

With the YANG code:

```yang
list server {
      key "name";
      unique "ip port";
      leaf name {
          type string;
      }
      leaf ip {
          type inet:ip-address;
      }
      leaf port {
          type inet:port-number;
      }
  }
```

We specify that the combination of IP and port must be unique. Thus the configuration is not valid:

```xml
<server>
  <name>smtp</name>
  <ip>192.0.2.1</ip>
  <port>25</port>
</server>

<server>
  <name>http</name>
  <ip>192.0.2.1</ip>
  <port>25</port>
</server>
```

The usage of leafrefs (See the YANG specification) ensures that we do not end up with configurations with dangling pointers. Leafrefs are also especially good, since the CLI and Web UI can render a better interface.

If other constraints are necessary, validation callback functions can be programmed in Java, Python, or Erlang. See `tailf:validate` in [tailf\_yang\_extensions(5)](/guides/resources/man/tailf_yang_extensions.5) in Manual Pages for details.

## The `when` statement <a href="#d5e2173" id="d5e2173"></a>

The `when` statement is used to make its parent statement conditional. If the XPath expression specified as the argument to this statement evaluates to false, the parent node cannot be given configured. Furthermore, if the parent node exists, and some other node is changed so that the XPath expression becomes false, the parent node is automatically deleted. For example:

```yang
leaf a {
    type boolean;
}
leaf b {
    type string;
    when "../a = 'true'";
}
```

This data model snippet says that `b` can only exist if `a` is true. If `a` is true, and `b` has a value, and `a` is set to false, `b` will automatically be deleted.

Since the XPath expression in theory can refer to any node in the data tree, it has to be re-evaluated when any node in the tree is modified. But this would have a disastrous performance impact, so to avoid this, NSO keeps track of dependencies for each when expression. In many cases, the **confdc** can figure out these dependencies by itself. In the example above, NSO will detect that `b` is dependent on `a`, and evaluate `b`'s XPath expression only if `a` is modified. If `confdc` cannot detect the dependencies by itself, it requires a `tailf:dependency` statement in the `when` statement. See `tailf:dependency` in [tailf\_yang\_extensions(5)](/guides/resources/man/tailf_yang_extensions.5) in Manual Pages for details.

## Using the Tail-f Extensions with YANG <a href="#d5e2188" id="d5e2188"></a>

Tail-f has an extensive set of extensions to the YANG language that integrates YANG models in NSO. For example, when we have `config false;` data, we may wish to invoke user C code to deliver the statistics data in runtime. To do this we annotate the YANG model with a Tail-f extension called `tailf:callpoint`.

Alternatively, we may wish to invoke user code to validate the configuration, this is also controlled through an extension called `tailf:validate`.

All these extensions are handled as normal YANG extensions. (YANG is designed to be extended) We have defined the Tail-f proprietary extensions in a file `${NCS_DIR}/src/ncs/yang/tailf-common.yang`

Continuing with our previous example, by adding a callpoint and a validation point, we get:

```yang
module test {
   namespace "http://tail-f.com/test";
   prefix "t";

   import ietf-inet-types {
      prefix inet;
   }
   import tailf-common {
      prefix tailf;
   }

   container top {
      leaf a {
          type int32;
          config false;
          tailf:callpoint mycp;
      }
      leaf b {
         tailf:validate myvalcp {
            tailf:dependency "../a";
         }
         type string;
      }
      leaf ip {
         type inet:ipv4-address;
      }
   }
}
```

The above module contains a callpoint and a validation point. The exact syntax for all Tail-f extensions is defined in the `tailf-common.yang` file.

Note the import statement where we import `tailf-common`.

When we are using YANG specifications to generate Java classes for ConfM, these extensions are ignored. They only make sense on the device side. It is worth mentioning them though since EMS developers will certainly get the YANG specifications from the device developers, thus the YANG specifications may contain extensions

The man page [tailf\_yang\_extensions(5)](/guides/resources/man/tailf_yang_extensions.5) in Manual Pages describes all the Tail-f YANG extensions.

### Using a YANG Annotation File <a href="#d5e2207" id="d5e2207"></a>

Sometimes it is convenient to specify all Tail-f extension statements in-line in the original YANG module. But in some cases, e.g. when implementing a standard YANG module, it is better to keep the Tail-f extension statements in a separate annotation file. When the YANG module is compiled to an `fxs` file, the compiler is given the original YANG module and any number of annotation files.

A YANG annotation file is a normal YANG module that imports the module to annotate. Then the `tailf:annotate` statement is used to annotate nodes in the original module. For example, the module test above can be annotated like this:

```yang
module test {
   namespace "http://tail-f.com/test";
   prefix "t";

   import ietf-inet-types {
      prefix inet;
   }

   container top {
      leaf a {
          type int32;
          config false;
      }
      leaf b {
         type string;
      }
      leaf ip {
         type inet:ipv4-address;
      }
   }
}
```

```yang
module test-ann {
   namespace "http://tail-f.com/test-ann";
   prefix "ta";

   import test {
      prefix t;
   }
   import tailf-common {
      prefix tailf;
   }

   tailf:annotate "/t:top/t:a" {
       tailf:callpoint mycp;
   }

   tailf:annotate "/t:top" {
       tailf:annotate "t:b" {  // recursive annotation
           tailf:validate myvalcp {
               tailf:dependency "../t:a";
           }
       }
   }
}
```

To compile the module with annotations, use the `-a` parameter to `confdc`:

```
confdc -c -a test-ann.yang test.yang
```

## Custom Help Texts and Error Messages <a href="#d5e2219" id="d5e2219"></a>

Certain parts of a YANG model are used by northbound agents, e.g. CLI and Web UI, to provide the end-user with custom help texts and error messages.

### Custom Help Texts

A YANG statement can be annotated with a `description` statement which is used to describe the definition for a reader of the module. This text is often too long and too detailed to be useful as help text in a CLI. For this reason, NSO by default does not use the text in the `description` for this purpose. Instead, a tail-f-specific statement, `tailf:info` is used. It is recommended that the standard `description` statement contains a detailed description suitable for a module reader (e.g. NETCONF client or server implementor), and `tailf:info` contains a CLI help text.

As an alternative, NSO can be instructed to use the text in the `description` statement also for CLI help text. See the option `--use-description` in [ncsc(1)](/guides/resources/man/ncsc.1) in Manual Pages.

For example, CLI uses the help text to prompt for a value of this particular type. The CLI shows this information during tab/command completion or if the end-user explicitly asks for help using the `?-`character. The behavior depends on the mode the CLI is running in.

The Web UI uses this information likewise to help the end-user.

The `mtu` definition below has been annotated to enrich the end-user experience:

```yang
leaf mtu {
    type uint16 {
        range "1 .. 1500";
    }
    description
       "MTU is the largest frame size that can be transmitted
        over the network. For example, an Ethernet MTU is 1,500
        bytes. Messages longer than the MTU must be divided
        into smaller frames.";
    tailf:info
       "largest frame size";
}
```

### Custom Help Text in a `typedef` <a href="#d5e2240" id="d5e2240"></a>

Alternatively, we could have provided the help text in a `typedef` statement as in:

```
 typedef mtuType {
    type uint16 {
        range "1 .. 1500";
    }
    description
        "MTU is the largest frame size that can be transmitted over the
         network. For example, an Ethernet MTU is 1,500
         bytes. Messages longer than the MTU must be
         divided into smaller frames.";
    tailf:info
       "largest frame size";
}

leaf mtu {
    type mtuType;
}
```

If there is an explicit help text attached to a leaf, it overrides the help text attached to the type.

### Custom Error Messages <a href="#d5e2247" id="d5e2247"></a>

A statement can have an optional error message statement. The northbound agents, for example, the CLI uses this to inform the end-user about a provided value that is not of the correct type. If no custom error message statement is available NSO generates a built-in error message, e.g. `1505 is too large`.

All northbound agents use the extra information provided by an `error-message` statement.

The `typedef` statement below has been annotated to enrich the end-user experience when it comes to error information:

```
typedef mtuType {
   type uint32 {
       range "1..1500" {
           error-message
              "The MTU must be a positive number not "
            + "larger than 1500";
       }
   }
}
```

## Example: Modeling a List of Interfaces <a href="#d5e2256" id="d5e2256"></a>

Say, for example, that we want to model the interface list on a Linux-based device. Running the `ip link list` command reveals the type of information we have to model

```bash
$ /sbin/ip link list
1: eth0: <BROADCAST,MULTICAST,UP>; mtu 1500 qdisc pfifo_fast qlen 1000
    link/ether 00:12:3f:7d:b0:32 brd ff:ff:ff:ff:ff:ff
2: lo: <LOOPBACK,UP>; mtu 16436 qdisc noqueue
    link/loopback 00:00:00:00:00:00 brd 00:00:00:00:00:00
3: dummy0: <BROADCAST,NOARP> mtu 1500 qdisc noop
    link/ether a6:17:b9:86:2c:04 brd ff:ff:ff:ff:ff:ff
```

And, this is how we want to represent the above in XML:

```xml
<?xml version="1.0"?>
<config xmlns="http://example.com/ns/link">
  <links>
    <link>
      <name>eth0</name>
      <flags>
        <UP/>
        <BROADCAST/>
        <MULTICAST/>
      </flags>
      <addr>00:12:3f:7d:b0:32</addr>
      <brd>ff:ff:ff:ff:ff:ff</brd>
      <mtu>1500</mtu>
    </link>

    <link>
      <name>lo</name>
      <flags>
        <UP/>
        <LOOPBACK/>
      </flags>
      <addr>00:00:00:00:00:00</addr>
      <brd>00:00:00:00:00:00</brd>
      <mtu>16436</mtu>
    </link>
  </links>
</config>
```

An interface or a `link` has data associated with it. It also has a name, an obvious choice to use as the key - the data item that uniquely identifies an individual interface.

The structure of a YANG model is always a header, followed by type definitions, followed by the actual structure of the data. A YANG model for the interface list starts with a header:

```yang
module links {
    namespace "http://example.com/ns/links";
    prefix link;

    revision 2007-06-09 {
      description "Initial revision.";
    }
    ...
```

A number of datatype definitions may follow the YANG module header. Looking at the output from `/sbin/ip` we see that each interface has a number of boolean flags associated with it, e.g. `UP`, and `NOARP`.

One way to model a sequence of boolean flags is as a sequence of statements:

```yang
leaf UP {
    type boolean;
    default false;
}
leaf NOARP {
    type boolean;
    default false;
}
```

A better way is to model this as:

```yang
leaf UP {
    type empty;
}
leaf NOARP {
    type empty;
}
```

We could choose to group these leafs together into a grouping. This makes sense if we wish to use the same set of boolean flags in more than one place. We could thus create a named grouping such as:

```
grouping LinkFlags {
    leaf UP {
        type empty;
    }
    leaf NOARP {
        type empty;
    }
    leaf BROADCAST {
        type empty;
    }
    leaf MULTICAST {
        type empty;
    }
    leaf LOOPBACK {
        type empty;
    }
    leaf NOTRAILERS {
        type empty;
    }
}
```

The output from `/sbin/ip` also contains Ethernet MAC addresses. These are best represented by the `mac-address` type defined in the `ietf-yang-types.yang` file. The `mac-address` type is defined as:

```
typedef mac-address {
    type string {
        pattern '[0-9a-fA-F]{2}(:[0-9a-fA-F]{2}){5}';
    }
    description
       "The mac-address type represents an IEEE 802 MAC address.

       This type is in the value set and its semantics equivalent to
       the MacAddress textual convention of the SMIv2.";
    reference
      "IEEE 802: IEEE Standard for Local and Metropolitan Area
                 Networks: Overview and Architecture
       RFC 2579: Textual Conventions for SMIv2";
}
```

This defines a restriction on the string type, restricting values of the defined type `mac-address` to be strings adhering to the regular expression `[0-9a-fA-F]{2}(:[0-9a-fA-F]{2}){5}` Thus strings such as `a6:17:b9:86:2c:04` will be accepted.

Queue disciplines are associated with each device. They are typically used for bandwidth management. Another string restriction we could do is to define an enumeration of the different queue disciplines that can be attached to an interface.

We could write this as:

```
typedef QueueDisciplineType {
   type enumeration {
      enum pfifo_fast;
      enum noqueue;
      enum noop;
      enum htp;
   }
}
```

There are a large number of queue disciplines and we only list a few here. The example serves to show that by using enumerations we can restrict the values of the data set in a way that ensures that the data entered always is valid from a syntactical point of view.

Now that we have a number of usable datatypes, we continue with the actual data structure describing a list of interface entries:

```yang
container links {
    list link {
        key name;
        unique addr;
        max-elements 1024;
        leaf name {
            type string;
        }
        container flags {
            uses LinkFlags;
        }
        leaf addr {
            type yang:mac-address;
            mandatory true;
        }
        leaf brd {
            type yang:mac-address;
            mandatory true;
        }
        leaf qdisc {
            type QueueDisciplineType;
            mandatory true;
        }
        leaf qlen {
            type uint32;
            mandatory true;
        }
        leaf mtu {
            type uint32;
            mandatory true;
        }
    }
}
```

The `key` attribute on the leaf named "name" is important. It indicates that the leaf is the instance key for the list entry named `link`. All the `link` leafs are guaranteed to have unique values for their `name` leafs due to the key declaration.

If one leaf alone does not uniquely identify an object, we can define multiple keys. At least one leaf must be an instance key - we cannot have lists without a key.

List entries are ordered and indexed according to the value of the key(s).

### Modeling Relationships <a href="#ug.yang.relationships" id="ug.yang.relationships"></a>

A very common situation when modeling a device configuration is that we wish to model a relationship between two objects. This is achieved by means of the `leafref` statements. A `leafref` points to a child of a list entry which either is defined using a `key` or `unique` attribute.

The `leafref` statement can be used to express three flavors of relationships: extensions, specializations, and associations. Below we exemplify this by extending the `link` example from above.

Firstly, assume we want to put/store the queue disciplines from the previous section in a separate container - not embedded inside the `links` container.

We then specify a separate container, containing all the queue disciplines which each refers to a specific `link` entry. This is written as:

```yang
container queueDisciplines {
    list queueDiscipline {
        key linkName;
        max-elements 1024;
        leaf linkName {
            type leafref {
                path "/config/links/link/name";
            }
        }

        leaf type {
            type QueueDisciplineType;
            mandatory true;
        }
        leaf length {
            type uint32;
        }
    }
}
```

The `linkName` statement is both an instance key of the `queueDiscipline` list, and at the same time refers to a specific `link` entry. This way we can extend the amount of configuration data associated with a specific `link` entry.

Secondly, assume we want to express a restriction or specialization on Ethernet `link` entries, e.g. it should be possible to restrict interface characteristics such as 10Mbps and half duplex.

We then specify a separate container, containing all the specializations which each refers to a specific `link`:

```yang
container linkLimitations {
    list LinkLimitation {
        key linkName;
        max-elements 1024;
        leaf linkName {
            type leafref {
                path "/config/links/link/name";
            }
        }
        container limitations {
            leaf only10Mbs { type boolean;}
            leaf onlyHalfDuplex { type boolean;}
        }
    }
}
```

The `linkName` leaf is both an instance key to the `linkLimitation` list, and at the same time refers to a specific `link` leaf. This way we can restrict or specialize a specific `link`.

Thirdly, assume we want to express that one of the `link` entries should be the default link. In that case, we enforce an association between a non-dynamic `defaultLink` and a certain `link` entry:

```yang
leaf defaultLink {
    type leafref {
        path "/config/links/link/name";
    }
}
```

### Ensuring Uniqueness <a href="#d5e2348" id="d5e2348"></a>

Key leafs are always unique. Sometimes we may wish to impose further restrictions on objects. For example, we can ensure that all `link` entries have a unique MAC address. This is achieved through the use of the `unique` statement:

```yang
container servers {
    list server {
        key name;
        unique "ip port";
        unique "index";
        max-elements 64;
        leaf name {
            type string;
        }
        leaf index {
            type uint32;
            mandatory true;
        }
        leaf ip {
            type inet:ip-address;
            mandatory true;
        }
        leaf port {
            type inet:port-number;
            mandatory true;
        }
    }
}
```

In this example, we have two `unique` statements. These two groups ensure that each server has a unique index number as well as a unique IP and port pair.

### Default Values <a href="#d5e2357" id="d5e2357"></a>

A leaf can have a static or dynamic default value. Static default values are defined with the `default` statement in the data model. For example:

```yang
leaf mtu {
    type int32;
    default 1500;
}
```

and:

```yang
leaf UP {
    type boolean;
    default true;
}
```

A dynamic default value means that the default value for the leaf is the value of some other leaf in the data model. This can be used to make the default values configurable by the user. Dynamic default values are defined using the `tailf:default-ref` statement. For example, suppose we want to make the MTU default value configurable:

```yang
container links {
    leaf mtu {
        type uint32;
    }
    list link {
        key name;
        leaf name {
            type string;
        }
        leaf mtu {
            type uint32;
            tailf:default-ref '../../mtu';
        }
    }
}
```

Now suppose we have the following data:

```xml
<links>
  <mtu>1000</mtu>
  <link>
    <name>eth0</name>
    <mtu>1500</mtu>
  </link>
  <link>
    <name>eth1</name>
  </link>
</links>
```

In the example above, link `eth0` has the mtu 1500, and the link `eth1` has the `mtu` 1000. Since `eth1` does not have a `mtu` value set, it defaults to the value of `../../mtu`, which is 1000 in this case.

{% hint style="info" %}
Whenever a leaf has a default value, it implies that the leaf can be left out from the XML document, i.e. mandatory = false.
{% endhint %}

With the default value mechanism an old configuration can be used even after having added new settings.

Another example where default values are used is when a new instance is created. If all leafs within the instance have default values, these need not be specified in, for example, a NETCONF `create` operation.

### The Final Interface YANG Model <a href="#d5e2383" id="d5e2383"></a>

Here is the final interface YANG model with all constructs described above:

```yang
module links {
    namespace "http://example.com/ns/link";
    prefix link;

    import ietf-yang-types {
        prefix yang;
    }


    grouping LinkFlagsType {
        leaf UP {
            type empty;
        }
        leaf NOARP {
            type empty;
        }
        leaf BROADCAST {
            type empty;
        }
        leaf MULTICAST {
            type empty;
        }
        leaf LOOPBACK {
            type empty;
      }
        leaf NOTRAILERS {
            type empty;
        }
    }

    typedef QueueDisciplineType {
        type enumeration {
            enum pfifo_fast;
            enum noqueue;
            enum noop;
            enum htb;
        }
    }
    container config {
        container links {
            list link {
                key name;
                unique addr;
                max-elements 1024;
                leaf name {
                    type string;
                }
                container flags {
                    uses LinkFlagsType;
                }
                leaf addr {
                    type yang:mac-address;
                    mandatory true;
                }
                leaf brd {
                    type yang:mac-address;
                    mandatory true;
                }
                leaf mtu {
                    type uint32;
                    default 1500;
                }
            }
        }
        container queueDisciplines {
            list queueDiscipline {
                key linkName;
                max-elements 1024;
                leaf linkName {
                    type leafref {
                        path "/config/links/link/name";
                    }
                }
                leaf type {
                    type QueueDisciplineType;
                    mandatory true;
                }
                leaf length {
                    type uint32;
                }
            }
        }
        container linkLimitations {
            list linkLimitation {
                key linkName;
                leaf linkName {
                    type leafref {
                        path "/config/links/link/name";
                    }
                }
                container limitations {
                    leaf only10Mbps {
                        type boolean;
                        default false;
                    }
                    leaf onlyHalfDuplex {
                        type boolean;
                        default false;
                    }
                }
            }
        }
        container defaultLink {
            leaf linkName {
                type leafref {
                    path "/config/links/link/name";
                }
            }
        }
    }
}
```

If the above YANG file is saved on disk, as `links.yang`, we can compile and link it using the `confdc` compiler:

```bash
$ confdc -c links.yang
```

We now have a ready-to-use schema file named `links.fxs` on disk. To run this example, we need to copy the compiled `links.fxs` to a directory where NSO can find it.

## More on leafrefs <a href="#ug.yang.leafrefs" id="ug.yang.leafrefs"></a>

A `leafref` is used to model relationships in the data model, as described in [Modeling Relationships](#ug.yang.relationships). In the simplest case, the `leafref` is a single leaf that references a single key in a list:

```yang
list host {
    key "name";
    leaf name {
        type string;
    }
    ...
}

leaf host-ref {
    type leafref {
        path "../host/name";
    }
}
```

But sometimes a list has more than one key, or we need to refer to a list entry within another list. Consider this example:

```yang
list host {
    key "name";
    leaf name {
        type string;
    }

    list server {
        key "ip port";
        leaf ip {
            type inet:ip-address;
        }
        leaf port {
            type inet:port-number;
        }
        ...
    }
}
```

If we want to refer to a specific server on a host, we must provide three values; the host name, the server IP, and the server port. Using leafrefs, we can accomplish this by using three connected leafs:

```yang
leaf server-host {
    type leafref {
        path "/host/name";
    }
}
leaf server-ip {
    type leafref {
        path "/host[name=current()/../server-host]/server/ip";
    }
}
leaf server-port {
    type leafref {
        path "/host[name=current()/../server-host]"
           + "/server[ip=current()/../server-ip]/../port";
    }
}
```

The path specification for `server-ip` means the IP address of the server under the host with the same name as specified in `server-host`.

The path specification for `server-port` means the port number of the server with the same IP as specified in `server-ip`, under the host with the same name as specified in `server-host`.

This syntax quickly gets awkward and error-prone. NSO supports a shorthand syntax, by introducing an XPath function `deref()` (see [XPATH FUNCTIONS](/guides/resources/man/tailf_yang_extensions.5#xpath-functions) in Manual Pages ). Technically, this function follows a `leafref` value and returns all nodes that the `leafref` refers to (typically just one). The example above can be written like this:

```yang
leaf server-host {
    type leafref {
        path "/host/name";
    }
}
leaf server-ip {
    type leafref {
        path "deref(../server-host)/../server/ip";
    }
}
leaf server-port {
    type leafref {
        path "deref(../server-ip)/../port";
    }
}
```

Note that using the `deref` function is syntactic sugar for the basic syntax. The translation between the two formats is trivial. Also note that `deref()` is an extension to YANG, and third-party tools might not understand this syntax. To make sure that only plain YANG constructs are used in a module, the parameter `--strict-yang` can be given to `confdc -c`.

## Using Multiple Namespaces <a href="#d5e2425" id="d5e2425"></a>

There are several reasons for supporting multiple configuration namespaces. Multiple namespaces can be used to group common datatypes and hierarchies to be used by other YANG models. Separate namespaces can be used to describe the configuration of unrelated sub-systems, i.e. to achieve strict configuration data model boundaries between these sub-systems.

As an example, `datatypes.yang` is a YANG module that defines a reusable data type.

```yang
module datatypes {
  namespace "http://example.com/ns/dt";
  prefix dt;

  grouping countersType {
     leaf recvBytes {
        type uint64;
        mandatory true;
     }
     leaf sentBytes {
        type uint64;
        mandatory true;
     }
  }
}
```

We compile and link `datatypes.yang` into a final schema file representing the `http://example.com/ns/dt` namespace:

```bash
$ confdc -c datatypes.yang
```

To reuse our user defined `countersType`, we must import the `datatypes` module.

```yang
module test {
    namespace "http://tail-f.com/test";
    prefix "t";

    import datatypes {
        prefix dt;
    }

    container stats {
        uses dt:countersType;
    }
}
```

When compiling this new module that refers to another module, we must indicate to `confdc` where to search for the imported module:

```bash
$ confdc -c test.yang --yangpath /path/to/dt
```

`confdc` also searches for referred modules in the colon (:) separated path defined by the environment variable `YANG_MODPATH` and . (dot) is implicitly included.

## Module Names, Namespaces, and Revisions <a href="#ug.yang.names_namespaces_and_revisions" id="ug.yang.names_namespaces_and_revisions"></a>

We have three different entities that define our configuration data.

* The module name. A system typically consists of several modules. In the future, we also expect to see standard modules in a manner similar to how we have standard SNMP modules.

  It is highly recommended to have the vendor name embedded in the module name, similar to how vendors have their names in proprietary MIBs today.
* The XML namespace. A module defines a namespace. This is an important part of the module header. For example, we have:

  ```yang
   module acme-system {
       namespace "http://acme.example.com/system";
       .....
  ```

  \
  The namespace string must uniquely define the namespace. It is very important that once we have settled on a namespace we never change it. The namespace string should remain the same between revisions of a product. Do not embed revision information in the namespace string since that breaks manager-side NETCONF scripts.
* The `revision` statement as in:

  ```yang
   module acme-system {
       namespace "http://acme.example.com/system";
       prefix "acme";

       revision 2007-06-09;
       .....
  ```

  \
  The revision is exposed to a NETCONF manager in the capabilities sent from the agent to the NETCONF manager in the initial hello message. The fine details of revision management are being worked on in the IETF NETMOD working group and are not finalized at the time of this writing.

  What is clear though, is that a manager should base its version decisions on the information in the revision string.

  \
  A capabilities reply from a NETCONF agent to the manager may look as:

  ```xml
  <?xml version="1.0" encoding="UTF-8"?>
  <hello xmlns="urn:ietf:params:xml:ns:netconf:base:1.0">
  <capabilities>
    <capability>urn:ietf:params:netconf:base:1.0</capability>
    <capability>urn:ietf:params:netconf:capability:writable-running:1.0</capability>
    <capability>urn:ietf:params:netconf:capability:candidate:1.0</capability>
    <capability>urn:ietf:params:netconf:capability:confirmed-commit:1.0</capability>
    <capability>urn:ietf:params:netconf:capability:xpath:1.0</capability>
    <capability>urn:ietf:params:netconf:capability:validate:1.0</capability>
    <capability>urn:ietf:params:netconf:capability:rollback-on-error:1.0</capability>
    <capability>http://example.com/ns/link?revision=2007-06-09</capability>
    ....
  ```

  where the revision information for the `http://example.com/ns/link` namespace is encoded as `?revision=2007-06-09` using standard URI notation.

  \
  When we change the data model for a namespace, it is recommended to change the revision statement and never make any changes to the data model that are backward incompatible. This means that all leafs that are added must be either optional or have a default value. That way it is ensured that the old NETCONF client code will continue to function on the new data model. Section 10 of RFC 6020 and section 11 of RFC 7950 define exactly what changes can be made to a data model to not break old NETCONF clients.

## Hash Values and the `id-value` Statement <a href="#ug.yang.id_value" id="ug.yang.id_value"></a>

Internally and in the programming APIs, NSO uses integer values to represent YANG node names and the namespace URI. This conserves space and allows for more efficient comparisons (including `switch` statements) in the user application code. By default, `confdc` automatically computes a hash value for the namespace URI and for each string that is used as a node name.

Conflicts can occur in the mapping between strings and integer values - i.e. the initial assignment of integers to strings is unable to provide a unique, bi-directional mapping. Such conflicts are extremely rare (but possible) when the default hashing mechanism is used.

The conflicts are detected either by `confdc` or by the NSO daemon when it loads the `.fxs` files.

If there are any conflicts reported they will pertain to XML tags (or the namespace URI),

There are two different cases:

* Two different strings mapped to the same integer. This is the classical hash conflict - extremely rare due to the high quality of the hash function used. The resolution is to manually assign a unique value to one of the conflicting strings. The value should be greater than 2^31+2 but less than 2^32-1. This way it will be out of the range of the automatic hash values, which are between 0 and 2^31-1. The best way to choose a value is by using a random number generator, as in `2147483649 + rand:uniform(2147483645)`. The `tailf:id-value` should be placed as a substatement to the statement where the conflict occurs, or in the `module` statement in case of namespace URI conflict.
* One string mapped to two different integers. This is even more rare than the previous case - it can only happen if a hash conflict was detected and avoided through the use of `tailf:id-value` on one of the strings, and that string also occurs somewhere else. The resolution is to add the same `tailf:id-value` to the second occurrence of the string.

## NSO Caveats <a href="#ug.yang.caveats" id="ug.yang.caveats"></a>

### The `union` Type and Value Conversion <a href="#d5e2497" id="d5e2497"></a>

When converting a string to an enumeration value, the order of types in the union is important when the types overlap. The first matching type will be used, so we recommend having the narrower (or more specific) types first.

Consider the example below:

```yang
leaf example {
  type union {
    type string; // NOTE: widest type first
    type int32;
    type enumeration {
      enum "unbounded";
    }
  }
}
```

Converting the string `42` to a typed value using the YANG model above, will always result in a string value even though it is the string representation of an `int32`. Trying to convert the string `unbounded` will also result in a string value instead of the enumeration because the enumeration is placed after the string.

Instead, consider the example below where the string (being a wider type) is placed last:

```yang
leaf example {
  type union {
    type enumeration {
      enum "unbounded";
    }
    type int32;
    type string; // NOTE: widest type last
  }
}
```

Converting the string `42` to the corresponding union value will result in a `int32`. Trying to convert the string `unbounded` will also result in the enumeration value as expected. The relative order of the `int32` and enumeration does not matter as they do not overlap.

Using the C and Python APIs to convert a string to a given value is further limited by the lack of restriction matching on the types. Consider the following example:

```yang
leaf example {
  type union {
    type string {
      pattern "[a-z]+[0-9]+";
    }
    type int32;
  }
}
```

Converting the string `42` will result in a string value, even though the pattern requires the string to begin with a character in the "a" to "z" range. This value will be considered invalid by NSO if used in any calls handled by NSO.

To avoid issues when working with unions place wider types at the end. As an example put `string` last, `int8` before `int16` etc.

### User-defined Types <a href="#d5e2524" id="d5e2524"></a>

When using user-defined types together with NSO the compiled schema does not contain the original type as specified in the YANG file. This imposes some limitations on the running system.

High-level APIs are unable to infer the correct type of a value as this information is left out when the schema is compiled. It is possible to work around this issue by specifying the type explicitly whenever setting values of a user-defined type.

### XML Representation: Union of `type` `empty` and `type` `string`

The normal representation of a type `empty` leaf in XML is `<leaf-name/>`. However, there is an exception when a leaf is a union of type `empty` and for example type `string`. Consider the example below:

```yang
leaf example {
  type union {
    type empty;
    type string;
  }
}
```

In this case, both `<example>example</example>` and `</example>` will represent `empty` being set.


# NSO Concurrency Model

Learn how NSO enhances transactional efficiency with parallel transactions.

From version 6.0, NSO uses the so-called 'optimistic concurrency', which greatly improves parallelism. With this approach, NSO avoids the need for serialization and a global lock to run user code which would otherwise limit the number of requests the system can process in a given time unit.

Using this concurrency model, your code, such as a service mapping or custom validation code, can run in parallel, either with another instance of the same service or an entirely different service (or any other provisioning code, for that matter). As a result, the system can take better advantage of available resources, especially the additional CPU cores, making it a lot more performant.

## Optimistic Concurrency <a href="#d5e8425" id="d5e8425"></a>

Transactional systems, such as NSO, must process each request in a way that preserves what are known as the ACID properties, such as atomicity and isolation of requests. A traditional approach to ensure this behavior is by using locking to apply requests or transactions one by one. The main downside is that requests are processed sequentially and may not be able to fully utilize the available resources.

{% hint style="info" %}
Refer to [Transactions](/guides/development/core-concepts/transactions) for more information on what kind of work NSO performs in a transaction.
{% endhint %}

Optimistic concurrency, on the other hand, allows transactions to run in parallel. It works on the premise that data conflicts are rare, so most of the time the transactions can be applied concurrently and will retain the required properties. NSO ensures this by checking that there are no conflicts with other transactions just before each transaction is committed. In particular, NSO will verify that all the data accessed as part of the transaction is still valid when applying changes. Otherwise, the system will reject the transaction.

Such a model makes sense because a lot of the time concurrent transactions deal with separate sets of data. Even if multiple transactions share some data in a read-only fashion, it is fine as they still produce the same result.

<div data-with-frame="true"><figure><img src="/files/y8B2mIR8eB3WY56oIIOe" alt="" width="563"><figcaption><p>Nonconflicting Concurrent Transactions</p></figcaption></figure></div>

In the figure, `svc1` in the `T1` transaction and `svc2` in the `T2` transaction both read (but do not change) the same, shared piece of data and can proceed as usual, unperturbed.

On the other hand, a conflict is when a piece of data, that has been read by one transaction, is changed by another transaction before the first transaction is committed. In this case, at the moment the first transaction completes, it is already working with stale data and must be rejected, as the following figure shows.

<div data-with-frame="true"><figure><img src="/files/XXj9j2rz61OVwpH1EdPa" alt="" width="563"><figcaption><p>Conflicting Concurrent Transactions</p></figcaption></figure></div>

In the figure, the transaction `T1` reads `dns-server` to use in the provisioning of `svc1` but transaction `T2` changes `dns-server` value in the meantime. The two transactions conflict and `T1` is rejected because `T2` completed first.

To be precise, for a transaction to experience a conflict, both of the following has to be true:

1. It reads some data that is changed after being read and before the transaction is completed.
2. It commits a set of changes in NSO.

This means a set of read-only transactions or transactions, where nothing is changed, will never conflict. It is also possible that multiple write-only transactions won't conflict even when they update the same data nodes.

Allowing multiple concurrent transactions to write (and only write, not read) to the same data without conflict may seem odd at first. But from a transaction's standpoint, it does not depend on the current value because it was never read. Suppose the value changed the previous day, the transaction would do the exact same thing and you wouldn't consider it a conflict. So, the last write wins, regardless of the time elapsed between the two transactions.

{% hint style="danger" %}
It is extremely important that you do not mix multiple transactions, because it will prevent NSO from detecting conflicts properly. For example, starting multiple separate transactions and using one to write data, based on what was read from a different one, can result in subtle bugs that are hard to troubleshoot.
{% endhint %}

While the optimistic concurrency model allows transactions to run concurrently most of the time, ultimately some synchronization (a global lock) is still required to perform the conflict checks and serialize data writes to the CDB and devices. The following figure shows everything that happens after a client tries to apply a configuration change, including acquiring and releasing the lock. This process takes place, for example, when you enter the **commit** command on the NSO CLI or when a PUT request of the RESTCONF API is processed.

<div data-with-frame="true"><figure><img src="/files/dsbW59A7h4gVHTBJ6fvB" alt="" width="563"><figcaption><p>Stages of a Transaction Commit</p></figcaption></figure></div>

As the figure shows (and you can also observe it in the progress trace output), service mapping, validation, and transforms all happen in the transaction before taking a (global) transaction lock.

At the same time, NSO tracks all of the data reads and writes from the start of the transaction, right until the lock and conflict check. This includes service mapping callbacks and XML templates, as well as transform and custom validation hooks if you are using any. It even includes reads done as part of the YANG validation and rollback creation that NSO performs automatically.

If reads do not overlap with writes from other transactions, the conflict check passes. The change is written to the CDB and disseminated to the affected network devices, through the *prepare* and *commit* phases. Kickers and subscribers are called and, finally, the global lock can be released.

On the other hand, if there is overlap and the system detects a conflict, the transaction obviously cannot proceed. To recover if this happens, the transaction should be retried. Sometimes the system can do it automatically and sometimes the client itself must be prepared to retry it.

{% hint style="info" %}
An ingenious developer might consider avoiding the need for retries by using explicit locking, in the way the NETCONF `lock` command does. However, be aware that such an approach is likely to significantly degrade the throughput of the whole system and is discouraged. If explicit locking is required, it should be considered with caution and sufficient testing.
{% endhint %}

In general, what affects the chance of conflict is the actual data that is read and written by each transaction. So, if there is more data, the surface for potential conflict is bigger. But you can minimize this chance by accounting for it in the application design.

## Identifying Conflicts <a href="#d5e8475" id="d5e8475"></a>

When a transaction conflict occurs, NSO logs an entry in the developer log, often found at `logs/devel.log` or a similar path. Suppose you have the following code in Python:

```python
with ncs.maapi.single_write_trans('admin', 'system') as t:
    root = ncs.maagic.get_root(t)
    # Read a value that can change during this transaction
    dns_server = root.mysvc_dns
    # Now perform complex work... or time.sleep(10) for testing
    # Finally, write the result
    root.some_data = 'the result'
    t.apply()
```

If the `/mysvc-dns` leaf changes while the code is executing, the `t.apply()` line fails and the developer log contains an entry similar to the following example:

```
<INFO> 23-Aug-2022::03:31:17.029 linux-nso ncs[<0.18350.3>]: ncs writeset collector:
   check conflict tid=3347 min=234 seq=237 wait=0ms against=[3346] elapsed=1ms
   -> conflict on: /mysvc-dns read: <<"10.1.2.2">> (op: get_delem tid: 3347)
   write: <<"10.1.1.138">> (op: write tid: 3346 user: admin) phase(s): work
   write tids: 3346
```

Here, the transaction with id 3347 reads a value of `/mysvc-dns` as “10.1.2.2” but that value was changed by the transaction with id 3346 to “10.1.1.138” by the time the first transaction called `t.apply()`. The entry also contains some additional data, such as the user that initiated the other transaction and the low-level operations that resulted in the conflict.

At the same time, the Python code raises an `ncs.error.Error` exception, with `confd_errno` set to the value of `ncs.ERR_TRANSACTION_CONFLICT` and error text, such as the following:

```
Conflict detected (70): Transaction 3347 conflicts with transaction 3346 started by
   user admin: /mysvc:mysvc-dns read-op get_delem write-op write in work phase(s)
```

In Java code, a matching `com.tailf.conf.ConfException` is thrown, with `errorCode` set to the `com.tailf.conf.ErrorCode.ERR_TRANSACTION_CONFLICT` value.

A thing to keep in mind when examining conflicts is that the transaction that performed the read operations is the one that gets the error and causes the log entry, while the other transaction, performing the write operations to the same path, is already completed successfully.

The error includes a reference to the `work` phase. The phase tells which part of the transaction encountered a conflict. The `work` phase signifies changes in an open transaction before it is applied. In practice, this is a direct read in the code that started the transaction before calling the `apply()` or `applyTrans()` function: the example reads the value of the leaf into `dns_server`.

On the other hand, if two transactions configure two service instances and the conflict arises in the mapping code, then the phase shows `transform` instead. It is also possible for a conflict to occur in more than one place, such as the phase `transform,work` denoting a conflict in both, the service mapping code as well as the initial transaction.

The complete list of conflict sources, that is, the possible values for the phase, is as follows:

* `work`: read in an open transaction before it is applied
* `rollback`: read during rollback file creation
* `pre-transform`: read while validating service input parameters according to the service YANG model
* `transform`: read during service (FASTMAP) or another transform invocation
* `validation`: read while validating the final configuration (YANG validation)

For example, `pre-transform` indicates that the service YANG model validation is the source of the conflict. This can help tremendously when you try to narrow down the conflicting code in complex scenarios. In addition, the phase information is useful when you troubleshoot automatic transaction retries in case of conflict: when the phase includes `work`, automatic retry is not possible.

## Automatic Retries <a href="#d5e8528" id="d5e8528"></a>

In some situations, NSO can retry a transaction that first failed to apply due to a conflict. A prerequisite is that NSO knows which code caused the conflict and that it can run that code again.

Changes done in the work phase are changes made directly by an external agent, such as a Python script connecting to the NSO or a remote NETCONF client. Since NSO is not in control of and is not aware of the logic in the external agent, it can only reject the conflicting transaction.

However, for the phases that follow the work phase, all the logic is implemented in NSO and NSO can run it on demand. For example, NSO is in charge of calling the service mapping code and the code can be run as many times as needed (a requirement for service re-deploy and similar). So, in case of a conflict, NSO can rerun all of the necessary logic to provision or de-provision a service.

NSO keeps checkpoints for each transaction, to restart it from the conflicting phase and save itself from redoing the work from the preceding phases if possible. NSO automatically checks if the transaction checkpoint read- or write-set grows too large. This allows for larger transactions to go through without memory exhaustion. When all checkpoints are skipped, no transaction retries are possible, and the transaction fails. When later-stage checkpoints are skipped, the transaction retry will take more time.

The read-set and write-set size limits that NSO uses for transaction checkpoints are configurable in `ncs.conf` under:

* /ncs-config/checkpoint/max-read-set-size
* /ncs-config/checkpoint/max-write-set-size
* /ncs-config/checkpoint/total-size-limit

See [ncs.conf(5) ](/guides/resources/man#section-5-file-formats-and-syntax)for details.

A transaction checkpoint reaching a size limit will result in a log entry:

```
not creating rollback checkpoint, write-set size limit exceeded
```

If checkpoints are skipped, we might miss retry points/attempts if the transaction fails due to conflicts.

Moreover, in case of conflicts during service mapping, NSO optimizes the process even further. It tracks the conflicting services to not schedule them concurrently in the future. This automatic retry behavior is enabled by default.

For services, retries can be configured further or even disabled under `/services/global-settings`. You can also find the service conflicts NSO knows about by running the `show services scheduling conflict` command. For example:

```
admin@ncs# unhide debug
admin@ncs# show services scheduling conflict | notab
services scheduling conflict mysvc-servicepoint mysvc-servicepoint
 type           dynamic
 first-seen     2022-08-27T17:15:10+00:00
 inactive-after 2022-08-27T17:15:09+00:00
 expires-after  2022-08-27T18:05:09+00:00
 ttl-multiplier 1
admin@ncs#
```

Since a given service may not always conflict and can evolve over time, NSO reverts to default scheduling after expiry time, unless new conflicts occur.

Sometimes, you know in advance that a service will conflict, either with itself or another service. You can encode this information in the service YANG model using the `conflicts-with` parameter under the `servicepoint` definition:

```yang
list mysvc {
  uses ncs:service-data;
  ncs:servicepoint mysvc-servicepoint {
    ncs:conflicts-with "mysvc-servicepoint";
    ncs:conflicts-with "some-other-servicepoint";
  }
  // ...
}
```

The parameter ensures that NSO will never schedule and execute this service concurrently with another service using the specified `servicepoint`. It adds a non-expiring `static` scheduling conflict entry. This way, you can avoid the unnecessary occasional retry when the dynamic scheduling conflict entry expires.

Declaring a conflict with itself is especially useful when you have older, non-thread-safe service code that cannot be easily updated to avoid threading issues.

For the NSO CLI and JSON-RPC (WebUI) interfaces, a commit of a transaction that results in a conflict will trigger an automatic rebase and retry when the resulting configuration is the same despite the conflict. If the rebase does not resolve the conflict, the transaction will fail. The conflict can, in some CLI cases, be resolved manually. A successful automatic rebase and a retry will generate something like the following pseudo-log entries in the developer log (trace log level):

```
<INFO> … check for read-write conflicts: conflict found
<INFO> … rebase transaction
…
<INFO> … rebase transaction: ok
<INFO> … retrying transaction after rebase
```

## Handling Conflicts <a href="#ncs.development.concurrency.handling" id="ncs.development.concurrency.handling"></a>

When a transaction fails to apply due to a read-write conflict in the work phase, NSO rejects the transaction and returns a corresponding error. In such a case, you must start a new transaction and redo all the changes.

Why is this necessary? Suppose you have code, let's say as part of a CDB subscriber or a standalone program, similar to the following Python snippet:

```python
with ncs.maapi.single_write_trans('admin', 'system') as t:
    if t.get_elem('/mysvc-use-dhcp') == True:
        # do something
    else:
        # do something entirely different that breaks
        # your network if mysvc-use-dhcp happens to be true
    t.apply()
```

If `mysvc-use-dhcp` has one value when your code starts provisioning but is changed mid-process, your code needs to restart from the beginning or you can end up with a broken system. To guard against such a scenario, NSO needs to be conservative and return an error.

Since there is a chance of a transaction failing to apply due to a conflict, robust code should implement a retry scheme. You can implement the retry algorithm yourself, or you can use one of the provided helpers.

In Python, `Maapi` class has a `run_with_retry()` method, which creates a new transaction and calls a user-supplied function to perform the work. On conflict, `run_with_retry()` will recreate the transaction and call the user function again. For details, please see the relevant API documentation.

The same functionality is available in Java as well, as the `Maapi.ncsRunWithRetry()` method. Where it differs from the Python implementation is that it expects the function to be implemented inside a `MaapiRetryableOp` object.

As an alternative option, available only in Python, you can use the `retry_on_conflict()` function decorator.

Example code for each of these approaches is shown next. In addition, the [examples.ncs/scaling-performance/conflict-retry](https://github.com/NSO-developer/nso-examples/tree/6.7/scaling-performance/conflict-retry) example showcases this functionality as part of a concrete service.

## Example Retrying Code in Python <a href="#d5e8569" id="d5e8569"></a>

Suppose you have some code in Python, such as the following:

```python
with ncs.maapi.single_write_trans('admin', 'python') as t:
    root = ncs.maagic.get_root(t)
    # First read some data, then write some too.
    # Finally, call apply.
    t.apply()
```

Since the code performs reads and writes of data in NSO through a newly established transaction, there is a chance of encountering a conflict with another, concurrent transaction.

On the other hand, if this was a service mapping code, you wouldn't be creating a new transaction yourself because the system would already provide one for you. You wouldn't have to worry about the retry because, again, the system would handle it for you through the automatic mechanism described earlier.

Yet, you may find such code in CDB subscribers, standalone scripts, or action implementations. As a best practice, the code should handle conflicts.

If you have an existing `ncs.maapi.Maapi` object already available, the simplest option might be to refactor the actual logic into a separate function and call it through `run_with_retry()`. For the current example, this might look like the following:

```python
def do_provisioning(t):
    """Function containing the actual logic"""
    root = ncs.maagic.get_root(t)
    # First read some data, then write some too.
    # ...
    # Finally, return True to signal apply() has to be called.
    return True

# Need to replace single_write_trans() with a Maapi object
with ncs.maapi.Maapi() as m:
    with ncs.maapi.Session(m, 'admin', 'python'):
        m.run_with_retry(do_provisioning)
```

If the new function is not entirely independent and needs additional values passed as parameters, you can wrap it inside an anonymous (lambda) function:

```
m.run_with_retry(lambda t: do_provisioning(t, one_param, another_param))
```

An alternative implementation with a decorator is also possible and might be easier to implement if the code relies on the `single_write_trans()` or similar function. Here, the code does not change unless it has to be refactored into a separate function. The function is then adorned with the `@ncs.maapi.retry_on_conflict()` decorator. For example:

```python
from ncs.maapi import retry_on_conflict

@retry_on_conflict()
def do_provisioning():
    # This is the same code as before but in a function
    with ncs.maapi.single_write_trans('admin', 'python') as t:
        root = ncs.maagic.get_root(t)
        # First read some data, then write some too.
        # ...
        # Finally, call apply().
        t.apply()

do_provisioning()
```

The major benefit of this approach is when the code is already in a function and only a decorator needs to be added. It can also be used with methods of the `Action` class and alike.

```python
class MyAction(ncs.dp.Action):
    @ncs.dp.Action.action
    @retry_on_conflict()
    def cb_action(self, uinfo, name, kp, input, output, trans):
        with ncs.maapi.single_write_trans('admin', 'python') as t:
            ...
```

For actions in particular, please note that the order of decorators is important and the decorator is only useful when you start your own write transaction in the wrapped function. This is what `single_write_trans()` does in the preceding example because the old transaction cannot be used any longer in case of conflict.

## Example Retrying Code in Java <a href="#d5e8591" id="d5e8591"></a>

Suppose you have some code in Java, such as the following:

```java
public class MyProgram {
    public static void main(String[] arg) throws Exception {
        try (Maapi maapi = new Maapi(UnixDomainSocketAddress.of(Conf.NCS_PATH))) {
            maapi.startUserSession("admin", "system");
            NavuContext context = new NavuContext(maapi);
            int tid = context.startRunningTrans(Conf.MODE_READ_WRITE);

            // Your code here that reads and writes data.

            // Finally, call apply.
            context.applyClearTrans();
            maapi.endUserSession();
        }
    }
}
```

To read and write some data in NSO, the code starts a new transaction with the help of `NavuContext.startRunningTrans()` but could have called `Maapi.startTrans()` directly as well. Regardless of the way such a transaction is started, there is a chance of encountering a read-write conflict. To handle those cases, the code can be rewritten to use `Maapi.ncsRunWithRetry()`.

The `ncsRunWithRetry()` call creates and manages a new transaction, then delegates work to an object implementing the `com.tailf.maapi.MaapiRetryableOp` interface. So, you need to move the code that does the work into a new class, let's say `MyProvisioningOp`:

```java
public class MyProvisioningOp implements MaapiRetryableOp {
    public boolean execute(Maapi maapi, int tid)
        throws IOException, ConfException, MaapiException
    {
        // Create context for the provided, managed transaction;
        // note the extra parameter compared to before and no calling
        // context.startRunningTrans() anymore.
        NavuContext context = new NavuContext(maapi, tid);

        // Your code here that reads and writes data.

        // Finally, return true to signal apply() has to be called.
        return true;
    }
}
```

This class does not start its own transaction any more but uses the transaction handle `tid`, provided by the `ncsRunWithRetry()` wrapper.

You can create the `MyProvisioningOp` as an inner or nested class if you wish so but note that, depending on your code, you may need to designate it as a `static class` to use it directly as shown here.

If the code requires some extra parameters when called, you can also define additional properties on the new class and use them for this purpose. With the new class ready, you instantiate and call into it with the `ncsRunWithRetry()` function. For example:

```java
public class MyProgram {
    public static void main(String[] arg) throws Exception {
        try (Maapi maapi = new Maapi(UnixDomainSocketAddress.of(Conf.NCS_PATH))) {
            maapi.startUserSession("admin", "system");
            // Deletegate work to MyProvisioningOp, with retry.
            maapi.ncsRunWithRetry(new MyProvisioningOp());
            // No more calling applyClearTrans() or friends,
            // ncsRunWithRetry() does that for you.
            maapi.endUserSession();
        }
    }
}
```

And what if your use case requires you to customize how the transaction is started or applied? `ncsRunWithRetry()` can take additional parameters that allow you to control those aspects. Please see the relevant API documentation for the full reference.

## Designing for Concurrency <a href="#ncs.development.concurrency.designing" id="ncs.development.concurrency.designing"></a>

In general, transaction conflicts in NSO cannot be avoided altogether, so your code should handle them gracefully with retries. Retries are required to ensure correctness but do take up additional time and resources. Since a high percentage of retries will notably decrease the throughput of the system, you should endeavor to construct your data models and logic in a way that minimizes the chance of conflicts.

A conflict arises when one transaction changes a value that one or more other ongoing transactions rely on. From this, you can make a couple of observations that should help guide your implementation.

First, if the shared data changes infrequently, it will rarely cause a conflict (regardless of the number of reads) because it only affects the transactions happening at the time it is changed. Conversely, a frequent change can clash with other transactions much more often and warrants spending some effort to analyze and possibly make conflict-free.

Next, if a transaction runs a long time, a greater number of other write transactions can potentially run in the meantime, increasing the chances of a conflict. For this reason, you should avoid long-running read-write transactions.

Likewise, the more data nodes and the different parts of the data tree the transaction touches, the more likely it is to run into a conflict. Limiting the scope and the amount of the changes to shared data is an important design aspect.

Also, when considering possible conflicts, you must account for all the changes in the transaction. This includes changes propagated to other parts of the data model through dependencies. For example, consider the following YANG snippet. Changing a single `provision-dns` leaf also changes every `mysvc` list item because of the `when` statement.

```yang
leaf provision-dns {
  type boolean;
}
list mysvc {
  container dns {
    when "../../provision-dns";
    // ...
  }
}
```

Ultimately, what matters is the read-write overlap with other transactions. Thus, you should avoid needless reads in your code: if there are no reads of the changed values, there can't be any conflicts.

### Avoiding Needless Reads

A technique used in some existing projects, in service mapping code and elsewhere, is to first prepare all the provisioning parameters by reading a number of things from the CDB. But some of these parameters, or even most, may not really be needed for that particular invocation.

Consider the following service mapping code:

```python
def cb_create(self, tctx, root, service, proplist):
    device = root.devices.device[service.device]

    # Search device interfaces and CDB for mgmt IP
    device_ip = find_device_ip(device)

    # Find the best server to use for this device
    ntp_servers = root.my_settings.ntp_servers
    use_ntp_server = find_closest_server(device_ip, ntp_servers)

    if service.do_ntp:
        device.ntp.servers.append(use_ntp_server)
```

Here, a service performs NTP configuration when enabled through the `do_ntp` switch. But even if the switch is off, there are still a lot of reads performed. If one of the values changes during provisioning, such as the list of the available NTP servers in `ntp_servers`, it will cause a conflict and a retry.

An improved version of the code only calculates the NTP server value if it is actually needed:

```python
def cb_create(self, tctx, root, service, proplist):
    device = root.devices.device[service.device]

    if service.do_ntp:
        # Search device interfaces and CDB for mgmt IP
        device_ip = find_device_ip(device)

        # Find the best server to use for this device
        ntp_servers = root.my_settings.ntp_servers
        use_ntp_server = find_closest_server(device_ip, ntp_servers)

        device.ntp.servers.append(use_ntp_server)
```

### Handling Dependent Services <a href="#d5e8638" id="d5e8638"></a>

Another thing to consider in addition to the individual service implementation is the placement and interaction of the service within the system. What happens if one service is used to generate input for another service? If the two services run concurrently, writes of the first service will invalidate reads of the other one, pretty much guaranteeing a conflict. Then it is wasteful to run both services concurrently and they should really run serially.

A way to achieve this is through a design pattern called stacked services. You create a third service that instantiates the first service (generating the input data) before the second one (dependent on the generated data).

### Searching and Enumerating Lists <a href="#d5e8642" id="d5e8642"></a>

When there is a need to search or filter a list for specific items, you will often find for-loops or similar constructs in the code. For example, to configure NTP, you might have the following:

```
for ntp_server in root.my_settings.ntp_servers:
    # Only select active servers
    if ntp_server.is_active:
        # Do something
```

This approach is especially prevalent in ordered-by-user lists since the order of the items and their processing is important.

The interesting bit is that such code reads every item in the list. If the list is changed while the transaction is ongoing, you get a conflict with the message identifying the `get_next` operation (which is used for list traversal). This is not very surprising: if another active item is added or removed, it changes the result of your algorithm. So, this behavior is expected and desirable to ensure correctness.

However, you can observe the same conflict behavior in less obvious scenarios. If the list model contains a `unique` YANG statement, NSO performs the same kind of enumeration of list items for you to verify the unique constraint. Likewise, a `must` or `when` statement can also trigger the evaluation of every item during validation, depending on the XPath expression.

NSO knows how to discern between access to specific list items based on the key value, where it tracks reads only to those particular items, and enumerating the list, where no key value is supplied and a list with all elements is treated as a single item. This works for your code as well as for the XPath expressions (in YANG and otherwise). As you can imagine, adding or removing items in the first case doesn't cause conflicts, while in the second one, it does.

In the end, it depends on the situation whether list enumeration can affect throughput or not. In the example, the NTP servers could be configured manually, by the operator, so they would rarely change, making it a non-issue. But your use case might differ.

### Python Assigning to Self

As several service invocations may run in parallel, Python self-assignment in service handling code can cause difficult-to-debug issues. Therefore, NSO checks for such patterns and issues an alarm (default) or a log entry containing a warning and a keypath to the service instance that caused the warning. See [NSO Python VM](/guides/development/core-concepts/nso-virtual-machines/nso-python-vm) for details.

### **Controlling `no-overwrite` Behavior in Concurrent Environments**

The `no-overwrite` commit parameter mechanism prevents NSO from applying configuration changes that conflict with the device's current state. Since NSO 6.4, the `no-overwrite` mechanism has been enhanced to include configurable compare scopes using a new `compare` parameter. This enables fine-grained control over how NSO validates device state consistency before applying changes.

You can choose from the following three `compare` scopes:

<table><thead><tr><th valign="top">Scope</th><th valign="top">Description</th><th valign="top">Use Case and Considerations</th></tr></thead><tbody><tr><td valign="top"><code>write-set-only</code></td><td valign="top">Only modified data is checked (pre-6-4 behavior).</td><td valign="top"><p>Minimizes the amount of data fetched from the device, thus, reducing overhead.</p><p>Suitable for scenarios where performance is critical, and the device is trusted to validate its own configuration constraints.<br><br>Use this scope for devices with simple YANG models or when minimal validation is sufficient.</p></td></tr><tr><td valign="top"><code>write-and-full-read-set</code></td><td valign="top">Both modified and read data are checked (introduced in NSO 6.4).</td><td valign="top"><p>Provides the highest level of consistency by ensuring that all dependent data matches NSO’s CDB. Recommended for critical devices or complex configurations where data integrity is paramount.</p><p>Can be resource-intensive, especially for devices with third-party YANG models containing extensive dependencies (<code>when</code>, <code>must</code>, or <code>leafref</code> expressions).<br><br>Use this scope for devices requiring strict configuration alignment with NSO’s CDB.</p></td></tr><tr><td valign="top"><code>write-and-service-read-set</code></td><td valign="top">Checks only modified data and reads during the transform phase (new default introduced in 6.4).</td><td valign="top"><p>Offers improved performance over <code>write-and-full-read-set</code> by limiting the scope to service-related reads, while still ensuring consistency for service-driven configurations.</p><p>Avoids performance penalties from validation-phase reads in complex device models.<br><br>Balances performance and accuracy for devices with complex YANG models, such as those used in 3PY NEDs, where validation-phase reads can significantly increase the read-set size.<br><br>Use this scope for multi-vendor environments with third-party devices where services drive configuration changes.</p></td></tr></tbody></table>

These compare scopes are critical when designing for concurrency, as they determine the risk of conflicting changes and impact service transaction performance.


# Service Handling of Ambiguous Device Models

Perform handling of ambiguous device models.

When new NED versions with diverging XML namespaces are introduced, adaptations might be needed in the services for these new NEDs. But not necessarily; it depends on where in the specific NED models the ambiguities reside. Existing services might not refer to these parts of the model and in that case, they do not need any adaptations.

Finding out if and where services need adaptations can be non-trivial. An important exception is template services which check and point out ambiguities at load time (NSO startup). In Java or Python code this is harder and essentially falls back to code reviews and testing.

The changes in service code to handle ambiguities are straightforward but different for templates and code.

## Template Services <a href="#d5e8316" id="d5e8316"></a>

In templates, there are new processing instructions `if-ned-id` and `elif-ned-id`. When the template specifies a node in an XML namespace where an ambiguity exists, the `if-ned-id` process instruction is used to resolve that ambiguity.

The processing instruction `else` can be used in conjunction with `if-ned-id` and `elif-ned-id` to capture all other NED IDs.

For the nodes in the XML namespace where no ambiguities occur, this process instruction is not necessary.

```xml
<config-template xmlns="http://tail-f.com/ns/config/1.0">
  <devices xmlns="http://tail-f.com/ns/ncs">
    <device foreach="{apache-device}">
      <name>{current()}</name>
      <config>
        <?if-ned-id apache-nc-1.0:apache-nc-1.0?>
          <vhosts xmlns="urn:apache">
            <vhost>
              <hostname>{/vhost}</hostname>
              <doc-root>/srv/www/{/vhost}</doc-root>
            </vhost>
          </vhosts>
        <?elif-ned-id apache-nc-1.1:apache-nc-1.1?>
          <public xmlns="urn:apache">
            <vhosts>
              <vhost>
                <hostname>{/vhost}</hostname>
                <aliases>{/vhost}.public</aliases>
                <doc-root>/srv/www/{/vhost}</doc-root>
              </vhost>
            </vhosts>
          </public>
        <?end?>
      </config>
    </device>
  </devices>
</config-template>
```

## Java Services <a href="#d5e8330" id="d5e8330"></a>

In Java, the service code must handle the ambiguities by code where the devices' `ned-id` is tested before setting the nodes and values for the diverging paths.

The `ServiceContext` class has a new convenience method, `getNEDIdByDeviceName` which helps retrieve the `ned-id` from the device name string.

```java
    @ServiceCallback(servicePoint="websiteservice",
                     callType=ServiceCBType.CREATE)
    public Properties create(ServiceContext context,
                             NavuNode service,
                             NavuNode root,
                             Properties opaque)
                             throws DpCallbackException {

...

                NavuLeaf elemName = elem.leaf(Ncs._name_);
                NavuContainer md = root.container(Ncs._devices_).
                    list(Ncs._device_).elem(elemName.toKey());

                String ipv4Str = baseIp + ((subnet<<3) + server);
                String ipv6Str = "::ff:ff:" + ipv4Str;
                String ipStr = ipv4Str;
                String nedIdStr =
                    context.getNEDIdByDeviceName(elemName.valueAsString());
                if ("webserver-nc-1.0:webserver-nc-1.0".equals(nedIdStr)) {
                    ipStr = ipv4Str;
                } else if ("webserver2-nc-1.0:webserver2-nc-1.0"
                           .equals(nedIdStr)) {
                    ipStr = ipv6Str;
                }

                md.container(Ncs._config_).
                    container(webserver.prefix, webserver._wsConfig_).
                    list(webserver._listener_).
                    sharedCreate(new String[] {ipStr, ""+8008});

                ms.list(lb._backend_).sharedCreate(
                    new String[]{baseIp + ((subnet<<3) + server++),
                                 ""+8008});
...

            return opaque;
        } catch (Exception e) {
            throw new DpCallbackException("Service create failed", e);
        }

    }
```

## Python Services <a href="#d5e8338" id="d5e8338"></a>

In the Python API, there is also a need to handle ambiguities by checking the `ned-id` before setting the diverging paths. Use `get_ned_id()` from `ncs.application` to resolve NED IDs.

```python
import ncs

def _get_device(service, name):
    dev_path = '/ncs:devices/ncs:device{%s}' % (name, )
    return ncs.maagic.cd(service, dev_path)

class ServiceCallbacks(Service):
    @Service.create
    def cb_create(self, tctx, root, service, proplist):
        self.log.info('Service create(service=', service._path, ')')

        for name in service.apache_device:
            self.create_apache_device(service, name)

        template = ncs.template.Template(service)
        self.log.info(
            'applying web-server-template for device {}'.format(name))
        template.apply('web-server-template')
        self.log.info(
            'applying load-balancer-template for device {}'.format(name))
        template.apply('load-balancer-template')

    def create_apache_device(self, service, name):
        dev = _get_device(service, name)
        if 'apache-nc-1.0:apache-nc-1.0' == ncs.application.get_ned_id(dev):
            self.create_apache1_device(dev)
        elif 'apache-nc-1.1:apache-nc-1.1' == ncs.application.get_ned_id(dev):
            self.create_apache2_device(dev)
        else:
            raise Exception('unknown ned-id {}'.format(get_ned_id(dev)))

    def create_apache1_device(self, dev):
        self.log.info(
            'creating config for apache1 device {}'.format(dev.name))
        dev.config.ap__listen_ports.listen_port.create(("*", 8080))
        dev.config.ap__clash = dev.name

    def create_apache2_device(self, dev):
        self.log.info(
            'creating config for apache2 device {}'.format(dev.name))
        dev.config.ap__system.listen_ports.listen_port.create(("*", 8080))
        dev.config.ap__clash = dev.name
```


# NSO Virtual Machines

Extend product functionality to add custom service code or expose data through data provider mechanism.


# NSO Python VM

Run your Python code using Python Virtual Machine (VM).

NSO is capable of starting one or several Python VMs where Python code in user-provided packages can run.

An NSO package containing a `python` directory will be considered to be a Python Package. By default, a Python VM will be started for each Python package that has a `python-class-name` defined in its `package-meta-data.xml` file. In this Python VM, the `PYTHONPATH` environment variable will be pointing to the `python` directory in the package.

If any required package that is listed in the `package-meta-data.xml` contains a `python` directory, the path to that directory will be added to the `PYTHONPATH` of the started Python VM and thus its accompanying Python code will be accessible.

Several Python packages can be started in the same Python VM if their corresponding `package-meta-data.xml` files contain the same *`python-package/vm-name`*.

A Python package skeleton can be created by making use of the `ncs-make-package` command:

```bash
ncs-make-package --service-skeleton python <package-name>
```

## YANG Model <a href="#d5e1542" id="d5e1542"></a>

The `tailf-ncs-python-vm.yang` defines the `python-vm` container which, along with `ncs.conf`, is the entry point for controlling the NSO Python VM functionality. Study the content of the YANG model in the example below (The Python VM YANG Model). For a full explanation of all the configuration data, look at the YANG file and man `ncs.conf`. Here will follow a description of the most important configuration parameters.

Note that some of the nodes beneath `python-vm` are by default invisible due to a hidden attribute. To make everything under `python-vm` visible in the CLI, two steps are required:

1. First, the following XML snippet must be added to `ncs.conf`:\\

   ```xml
   <hide-group>
      <name>debug</name>
   </hide-group>
   ```
2. Next, the `unhide` command may be used in the CLI session:

   ```cli
   admin@ncs(config)# unhide debug
   admin@ncs(config)#
   ```

The `sanity-checks`/`self-assign-warning` controls the self-assignment warnings for Python services with off, log, and alarm (default) modes. An example of a self-assignment:

```python
class ServiceCallbacks(Service):
    @Service.create
    def cb_create(self, tctx, root, service, proplist):
        self.counter = 42
```

As several service invocations may run in parallel, self-assignment will likely cause difficult-to-debug issues. An alarm or a log entry will contain a warning and a keypath to the service instance that caused the warning. Example log entry:

```xml
<WARNING> ... Assigning to self is not thread safe: /mysrvc:mysrvc{2}
```

With the `logging`/`level`, the amount of logged information can be controlled. This is a global setting applied to all started Python VMs unless explicitly set for a particular VM, see [Debugging of Python packages](#debugging-of-python-packages). The levels correspond to the pre-defined Python levels in the Python `logging` module, ranging from `level-critical` to `level-debug`.

{% hint style="info" %}
Refer to the official Python documentation for the `logging` module for more information about the log levels.
{% endhint %}

The `logging`/`log-file-prefix` define the prefix part of the log file path used for the Python VMs. This prefix will be appended with a Python VM-specific suffix which is based on the Python package name or the *`python-package/vm-name`* from the `package-meta-data.xml` file. The default prefix is `logs/ncs-python-vm` so e.g., if a Python package named `l3vpn` is started, a logfile with the name `logs/ncs-python-vm-l3vpn.log` will be created.

The `status`*/*`start` and `status`*/*`current` contains operational data. The `status`*/*`start` command will show information about what Python classes, as declared in the `package-meta-data.xml` file, were started and whether the outcome was successful or not. The `status`*/*`current` command will show which Python classes that are currently running in a separate thread. The latter assumes that the user-provided code cooperates by informing NSO about any thread(s) started by the user code, see [Structure of the User-provided Code](#structure-of-the-user-provided-code).

The `start` and `stop` actions make it possible to start and stop a particular Python VM.

{% code title="Example: The Python VM YANG Model" %}

```cli
> yanger -f tree tailf-ncs-python-vm.yang
          
submodule: tailf-ncs-python-vm (belongs-to tailf-ncs)
  +--rw python-vm
     +--rw sanity-checks
     |  +--rw self-assign-warning?   enumeration
     +--rw logging
     |  +--rw log-file-prefix?   string
     |  +--rw level?             py-log-level-type
     |  +--rw vm-levels* [node-id]
     |     +--rw node-id    string
     |     +--rw level      py-log-level-type
     +--rw status
     |  +--ro start* [node-id]
     |  |  +--ro node-id     string
     |  |  +--ro packages* [package-name]
     |  |     +--ro package-name    string
     |  |     +--ro components* [component-name]
     |  |        +--ro component-name    string
     |  |        +--ro class-name?       string
     |  |        +--ro status?           enumeration
     |  |        +--ro error-info?       string
     |  +--ro current* [node-id]
     |     +--ro node-id     string
     |     +--ro packages* [package-name]
     |        +--ro package-name    string
     |        +--ro components* [component-name]
     |           +--ro component-name    string
     |           +--ro class-names* [class-name]
     |              +--ro class-name    string
     |              +--ro status?       enumeration
     +---x stop
     |  +---w input
     |  |  +---w name    string
     |  +--ro output
     |     +--ro result?   string
     +---x start
        +---w input
        |  +---w name    string
        +--ro output
           +--ro result?   string
```

{% endcode %}

## Structure of the User-provided Code

The `package-meta-data.xml` file must contain a `component` of type `application` with a `python-class-name` specified as shown in the example below.

{% code title="Example: package-meta-data.xml Excerpt" %}

```xml
<component>
  <name>L3VPN Service</name>
  <application>
    <python-class-name>l3vpn.service.Service</python-class-name>
  </application>
</component>
<component>
  <name>L3VPN Service model upgrade</name>
  <upgrade>
    <python-class-name>l3vpn.upgrade.Upgrade</python-class-name>
  </upgrade>
</component>
```

{% endcode %}

The component name (`L3VPN Service` in the example) is a human-readable name of this application component. It will be shown when doing `show python-vm` in the CLI. The `python-class-name` should specify the Python class that implements the application entry point. Note that it needs to be specified using Python's dot notation and should be fully qualified (given the fact that `PYTHONPATH` is pointing to the package `python` directory).

Study the excerpt of the directory listing from a package named `l3vpn` below.

{% code title="Example: Python Package Directory Structure" %}

```
packages/
+-- l3vpn/
    +-- package-meta-data.xml
    +-- python/
    |   +-- l3vpn/
    |       +-- __init__.py
    |       +-- service.py
    |       +-- upgrade.py
    |       +-- _namespaces/
    |           +-- __init__.py
    |           +-- l3vpn_ns.py
    +-- src
        +-- Makefile
        +-- yang/
            +-- l3vpn.yang
```

{% endcode %}

Look closely at the `python` directory above. Note that directly under this directory is another directory named the package (`l3vpn`) that contains the user code. This is an important structural choice that eliminates the chance of code clashes between dependent packages (only if all dependent packages use this pattern of course).

As you can see, the `service.py` is located according to the description above. There is also a `__init__.py` (which is empty) there to make the `l3vpn` directory considered a module from Python's perspective.

Note the `_namespaces/l3vpn_ns.py` file. It is generated from the `l3vpn.yang` model using the `ncsc --emit-python` command and contains constants representing the namespace and the various components of the YANG model, which the User code can import and make use of.

The `service.py` file should include a class definition named `Service` which acts as the component's entry point. See [The Application Component](#ncs.development.pythonvm.cthread) for details.

Notice that there is also a file named `upgrade.py` present which holds the implementation of the `upgrade` component specified in the `package-meta-data.xml` excerpt above. See [The Upgrade Component](#ncs.development.pythonvm.upgrade) for details regarding `upgrade` components.

### The `application` Component <a href="#ncs.development.pythonvm.cthread" id="ncs.development.pythonvm.cthread"></a>

The Python class specified in the `package-meta-data.xml` file will be started in a Python thread which we call a `component` thread. This Python class should inherit `ncs.application.Application` and should implement the methods `setup()` and `teardown()`.

NSO supports two different modes for executing the implementations of the registered callpoints, `threading` and `multiprocessing`.

The default `threading` mode will use a single thread pool for executing the callbacks for all callpoints.

The `multiprocessing` mode will start a subprocess for each callpoint. Depending on the user code, this can greatly improve the performance on systems with a lot of parallel requests, as a separate worker process will be created for each Service, Nano Service, and Action.

The behavior is controlled by three factors:

* `callpoint-model` setting in the `package-meta-data.xml` file.
* Number of registered callpoints in the `Application`.
* Operating System support for killing child processes when the parent exits.

If the `callpoint-model` is set to `multiprocessing`, more than one callpoint is registered in the `Application` and the Operating System supports killing child processes when the parent exits, NSO will enable multiprocessing mode.

{% code title="Example: Component Class Skeleton" %}

```python
import ncs

class Service(ncs.application.Application):
    def setup(self):
        # The application class sets up logging for us. It is accessible
        # through 'self.log' and is a ncs.log.Log instance.
        self.log.info('Service RUNNING')

        # Service callbacks require a registration for a 'service point',
        # as specified in the corresponding data model.
        #
        self.register_service('l3vpn-servicepoint', ServiceCallbacks)

        # If we registered any callback(s) above, the Application class
        # took care of creating a daemon (related to the service/action point).

        # When this setup method is finished, all registrations are
        # considered done and the application is 'started'.

    def teardown(self):
        # When the application is finished (which would happen if NCS went
        # down, packages were reloaded or some error occurred) this teardown
        # method will be called.

        self.log.info('Service FINISHED')
```

{% endcode %}

The `Service` class will be instantiated by NSO when started or whenever packages are reloaded. Custom initialization, such as registering service and action callbacks should be done in the `setup()` method. If any cleanup is needed when NSO finishes or when packages are reloaded it should be placed in the `teardown()` method.

The existing log functions are named after the standard Python log levels, thus in the example above the `self.log` object contains the functions `debug`*,*`info`*,*`warning`*,*`error`*,*`critical`. Where to log and with what level can be controlled from NSO?

### The `upgrade` Component <a href="#ncs.development.pythonvm.upgrade" id="ncs.development.pythonvm.upgrade"></a>

The Python class specified in the `upgrade` section of `package-meta-data.xml` will be run by NSO in a separately started Python VM. The class must be instantiable using the empty constructor and it must have a method called `upgrade` as in the example below. It should inherit `ncs.upgrade.Upgrade`.

{% code title="Example: Upgrade Class Example" %}

```python
import ncs
import _ncs


class Upgrade(ncs.upgrade.Upgrade):
    """An upgrade 'class' that will be instantiated by NSO.

    This class can be named anything as long as NSO can find it using the
    information specified in <python-class-name> for the <upgrade>
    component in package-meta-data.xml.

    Is should inherit ncs.upgrade.Upgrade.

    NSO will instantiate this class using the empty contructor.
    The class MUST have a method named 'upgrade' (as in the example below)
    which will be called by NSO.
    """

    def upgrade(self, cdbsock, trans):
        """The upgrade 'method' that will be called by NSO.

        Arguments:
        cdbsock -- a connected CDB data socket for reading current (old) data.
        trans -- a ncs.maapi.Transaction instance connected to the init
                 transaction for writing (new) data.

        There is no need to connect a CDB data socket to NSO - that part is
        already taken care of and the socket is passed in the first argument
        'cdbsock'. A session against the DB needs to be started though. The
        session doesn't need to be ended and the socket doesn't need to be
        closed - NSO will do that automatically.

        The second argument 'trans' is already attached to the init transaction
        and ready to be used for writing the changes. It can be used to create a
        maagic object if that is preferred. There's no need to detach or finish
        the transaction, and, remember to NOT apply() the transaction when work
        is finished.

        The method should return True (or None, which means that a return
        statement is not needed) if everything was OK.
        If something went wrong the method should return False or throw an
        error. The northbound client initiating the upgrade will be alerted
        with an error message.

        Anything written to stdout/stderr will end up in the general log file
        for various output from Python VMs. If not configured the file will
        be named ncs-python-vm.log.
        """

        # start a session against running
        _ncs.cdb.start_session2(cdbsock, ncs.cdb.RUNNING,
                                ncs.cdb.LOCK_SESSION | ncs.cdb.LOCK_WAIT)

        # loop over a list and do some work
        num = _ncs.cdb.num_instances(cdbsock, '/path/to/list')
        for i in range(0, num):
            # read the key (which in this example is 'name') as a ncs.Value
            value = _ncs.cdb.get(cdbsock, '/path/to/list[{0}]/name'.format(i))
            # create a mandatory leaf 'level' (enum - low, normal, high)
            key = str(value)
            trans.set_elem('normal', '/path/to/list{{{0}}}/level'.format(key))

        # not really needed
        return True

        # Error return example:
        #
        # This indicates a failure and the string written to stdout below will
        # written to the general log file for various output from Python VMs.
        #
        # print('Error: not implemented yet')
        # return False
```

{% endcode %}

## The NSO client timeouts

The section `/ncs-config/api` in **ncs.conf** contains a number of very important timeouts. See `$NCS_DIR/src/ncs/ncs_config/tailf-ncs-config.yang` and [ncs.conf(5)](/guides/resources/man/ncs.conf.5) for details.

* `new-session-timeout` controls how long NSO will wait for the NSO Python VM to respond to a new session.
* `query-timeout` controls how long NSO will wait for the NSO Python VM to respond to a request to get data.
* `connect-timeout` controls how long NSO will wait for the NSO Python VM to initialize a Dp connection after the initial socket connect.
* `action-timeout` controls how long NSO will wait for the NSO Python VM to respond to an action request callback.

For `new-session-timeout`, `query-timeout` and `connect-timeout`, whenever any of these timeouts trigger, NSO will close the sockets from NSO to the NSO Python VM. The NSO Python VM will detect the closed socket and exit.

For `action-timeout`, whenever this timeout triggers, NSO will only close the sockets from the NSO Python VM to the clients without exiting the Python VM.

## Debugging of Python Packages

Python code packages are not running with an attached console and the standard out from the Python VMs are collected and put into the common log file `ncs-python-vm.log`. Possible Python compilation errors will also end up in this file.

Normally the logging objects provided by the Python APIs are used. They are based on the standard Python `logging` module. This gives the possibility to control the logging if needed, e.g., getting a module local logger to increase logging granularity.

The default logging level is set to `info`. For debugging purposes, it is very useful to increase the logging level:

```bash
    $ ncs_cli -u admin
    admin@ncs> config
    admin@ncs% set python-vm logging level level-debug
    admin@ncs% commit
```

This sets the global logging level and will affect all started Python VMs. It is also possible to set the logging level for a single package (or multiple packages running in the same VM), which will take precedence over the global setting:

```bash
    $ ncs_cli -u admin
    admin@ncs> config
    admin@ncs% set python-vm logging vm-levels pkg_name level level-debug
    admin@ncs% commit
```

The debugging output is printed to separate files for each package and the log file naming is `ncs-python-vm-`*`pkg_name`*`.log`

Log file output example for package `l3vpn`:

```bash
    $ tail -f logs/ncs-python-vm-l3vpn.log
    2016-04-13 11:24:07 - l3vpn - DEBUG - Waiting for Json msgs
    2016-04-13 11:26:09 - l3vpn - INFO - action name: double
    2016-04-13 11:26:09 - l3vpn - INFO - action input.number: 21
```

## Using Non-standard Python <a href="#ncs.development.pythonvm.nonstdpython" id="ncs.development.pythonvm.nonstdpython"></a>

There are occasions where the standard Python installation is incompatible or maybe not preferred to be used together with NSO. In such cases, there are several options to tell NSO to use another Python installation for starting a Python VM.

By default NSO will use the file `$NCS_DIR/bin/ncs-start-python-vm` when starting a new Python VM. The last few lines in that file read:

```
        if [ -x "$(which python3)" ]; then
            echo "Starting python3 -u $main $*"
            exec python3 -u "$main" "$@"
        fi
        echo "Starting python -u $main $*"
        exec python -u "$main" "$@"
```

As seen above NSO first looks for `python3` and if found it will be used to start the VM. If `python3` is not found NSO will try to use the command `python` instead. Here we describe a couple of options for deciding which Python NSO should start.

### Configure NSO to Use a Custom Start Command (recommended) <a href="#d5e1719" id="d5e1719"></a>

NSO can be configured to use a custom start command for starting a Python VM. This can be done by first copying the file `$NCS_DIR/bin/ncs-start-python-vm` to a new file and then changing the last lines of that file to start the desired version of Python. After that, edit `ncs.conf` and configure the new file as the start command for a new Python VM. When the file `ncs.conf` has been changed reload its content by executing the command `ncs --reload`.

Example:

```bash
$ cd $NCS_DIR/bin
$ pwd
/usr/local/nso/bin
$ cp ncs-start-python-vm my-start-python-vm
$ # Use your favourite editor to update the last lines of the new
$ # file to start the desired Python executable.
```

Add the following snippet to `ncs.conf`:

```xml
<python-vm>
    <start-command>/usr/local/nso/bin/my-start-python-vm</start-command>
</python-vm>
```

The new `start-command` will take effect upon the next restart or configuration reload.

### Changing the Path to `python3` or `python` <a href="#d5e1732" id="d5e1732"></a>

Another way of telling NSO to start a specific Python executable is to configure the environment so that executing `python3` or `python` starts the desired Python. This may be done system-wide or can be made specific for the user running NSO.

### Updating the Default Start Command (not recommended) <a href="#d5e1739" id="d5e1739"></a>

Changing the last line of `$NCS_DIR/bin/ncs-start-python-vm` is of course an option but altering any of the installation files of NSO is discouraged.

## Handling Python Dependencies in NSO Packages

### Recommended: Add Dependencies to the `python` Directory

Python package dependencies can be installed in the `packages/<my-package>/python/` directory and loaded when the NSO Python VM is started for the package.

{% code title="Quick Start (Example)" overflow="wrap" %}

```
pip install --target packages/<my-package>/python/ -r packages/<my-package>/python/requirements.txt
```

{% endcode %}

#### Benefits

Installing the NSO package Python dependencies in the package `python` directory provides several advantages:

* Dependency isolation: Prevents Python package version conflicts between different NSO packages.
* Portability: Improves reproducibility across environments.
* System cleanliness: Keeps the host’s Python installation unmodified.
* High Availability: In NSO HA setups, the `packages ha sync action` can copy the self-contained packages, including their Python dependencies, across the cluster.

{% hint style="warning" %}
The Python dependencies must be installed using the same Python version, Python package version, and Linux distribution version as used by the test and production environment where the package runs.
{% endhint %}

#### Best Practices

* Include a `requirements.txt` to document Python dependencies.
* Place the `requirements.txt` file inside the NSO package to make it self-contained.

### Alternative: NSO Python VM in a Virtual Environment

NSO Python VM instances can run in isolated Python virtual environments using Python’s built-in `venv` module. This allows packages to manage their own Python dependencies without conflicts.

#### How It Works

To enable Python virtual environment support for an NSO package:

1. Create a `use_venv` file in the `packages/<my-package>/python/` directory.
2. Add the path to your Python virtual environment in this file.
3. The `$NCS_DIR/bin/ncs-start-python-vm` script will automatically activate the specified virtual environment when starting the Python VM for that package.

{% code title="Example Structure" %}

```none
packages/
└── my-package/
    └── python/
        ├── use_venv          # Contains: path/to/my/venv
        └── my_program.py
```

{% endcode %}

{% code title="Quick Start (Example)" overflow="wrap" %}

```bash
cd $NCS_RUN_DIR  # Or to the project run-time directory
python3 -m venv ./pyvenv
./pyvenv/bin/pip install -r packages/<my-package>/python/requirements.txt
echo "./pyvenv" > packages/<my-package>/python/use_venv
```

{% endcode %}

#### Packages Sharing Python VM Instance

When multiple packages share the same `vm-name` (i.e., Python VM instance) but specify different Python virtual environments, NSO will log an informational message in the developer log. The first Python virtual environment encountered will be used for all packages sharing that `vm-name`. Use unique `vm-name` values for packages requiring different Python virtual environments.

#### Benefits

Using virtual environments with NSO Python packages provides several advantages:

* Dependency isolation: Prevents Python package version conflicts between different NSO packages.
* Portability: Improves reproducibility across environments.
* System cleanliness: Keeps the host’s Python installation unmodified.
* Version flexibility: Enables testing and deployment with different Python versions.
* Reproducible builds: Ensures consistent dependency versions across development and production environments.

#### Best Practices

* Use paths from the NSO run-time directory to where the Python virtual environment is located.
* Include a `requirements.txt` to document Python dependencies.
* Use unique `vm-name` values when packages require different Python virtual environments.
* Ensure the NSO user has read/execute permissions on the venv path.
* Ensure all nodes in a high availability setup has the same copy of the Python virtual environment.
* Check the Python VM log, `ncs-python-vm.log`, for activation messages to verify the Python virtual environment used by the NSO package.

{% hint style="info" %}
The [examples.ncs/misc/py-package-deps](https://github.com/NSO-developer/nso-examples/tree/6.7/misc/py-package-deps) example demonstrates how to either install Python package dependencies in the NSO package `python` directory, or as an alternative, use a Python virtual environment to manage dependencies that automatically activates when the Python VM for a package starts.
{% endhint %}


# NSO Java VM

Run your Java code using Java Virtual Machine (VM).

The NSO Java VM is the execution container for all Java classes supplied by deployed NSO packages.

The classes, and other resources, are structured in `jar` files and the specific use of these classes is described in the `component` tag in the respective `package-meta-data.xml` file. Also as a framework, it starts and controls other utilities for the use of these components. To accomplish this, a main class `com.tailf.ncs.NcsMain`, implementing the `Runnable` interface is started as a thread. This thread can be the main thread (running in a java `main()`) or be embedded into another Java program.

When the `NcsMain` thread starts it establishes a socket connection towards NSO. This is called the NSO Java VM control socket. It is the responsibility of `NcsMain` to respond to command requests from NSO and pass these commands as events to the underlying finite state machine (FSM). The `NcsMain` FSM will execute all actions as requested by NSO. This includes class loading and instantiation as well as registration and start of services, NEDs, etc.

<div data-with-frame="true"><figure><img src="/files/zt9jRbHYNCLh1UxZqhks" alt="" width="563"><figcaption><p>NSO Service Manager</p></figcaption></figure></div>

When NSO detects the control socket connection from the NSO Java VM, it starts an initialization process:

1. First, NSO sends a `INIT_JVM` request to the NSO Java VM. At this point, the NSO Java VM will load schemas i.e. retrieve all known YANG module definitions. The NSO Java VM responds when all modules are loaded.
2. Then, NSO sends a `LOAD_SHARED_JARS` request for each deployed NSO package. This request contains the URLs for the jars situated in the `shared-jar` directory in the respective NSO package. The classes and resources in these jars will be globally accessible for all deployed NSO packages.
3. The next step is to send a `LOAD_PACKAGE` request for each deployed NSO package. This request contains the URLs for the jars situated in the `private-jar` directory in the respective NSO package. These classes and resources will be private to the respective NSO package. In addition, classes that are referenced in a `component` tag in the respective NSO package `package-meta-data.xml` file will be instantiated.
4. NSO will send a `INSTANTIATE_COMPONENT` request for each component in each deployed NSO package. At this point, the NSO Java VM will register a start method for the respective component. NSO will send these requests in a proper start phase order. This implies that the `INSTANTIATE_COMPONENT` requests can be sent in an order that mixes components from different NSO packages.
5. Lastly, NSO sends a `DONE_LOADING` request which indicates that the initialization process is finished. After this, the NSO Java VM is up and running.

See [Debugging Startup](#ug.javavm.debug) for tips on customizing startup behavior and debugging problems when the Java VM fails to start

## YANG Model <a href="#d5e1181" id="d5e1181"></a>

The file `tailf-ncs-java-vm.yang` defines the `java-vm` container which, along with `ncs.conf`, is the entry point for controlling the NSO Java VM functionality. Study the content of the YANG model in the example below (The Java VM YANG model). For a full explanation of all the configuration data, look at the YANG file and man `ncs.conf`.

Many of the nodes beneath `java-vm` are by default invisible due to a hidden attribute. To make everything under `java-vm` visible in the CLI, two steps are required:

1. First, the following XML snippet must be added to `ncs.conf`:\\

   ```xml
   <hide-group>
       <name>debug</name>
   </hide-group>
   ```
2. Next, the `unhide` command may be used in the CLI session:

   ```cli
   admin@ncs(config)# unhide debug
   admin@ncs(config)#
   ```

{% code title="Example: The Java VM YANG Model" %}

```cli
        > yanger -f tree tailf-ncs-java-vm.yang
          submodule: tailf-ncs-java-vm (belongs-to tailf-ncs)
  +--rw java-vm
     +--rw stdout-capture
     |  +--rw enabled?   boolean
     |  +--rw file?      string
     |  +--rw stdout?    empty
     +--rw connect-time?                     uint32
     +--rw initialization-time?              uint32
     +--rw synchronization-timeout-action?   enumeration
     +--rw exception-error-message
     |  +--rw verbosity?   error-verbosity-type
     +--rw java-logging
     |  +--rw logger* [logger-name]
     |     +--rw logger-name    string
     |     +--rw level          log-level-type
     +--ro start-status?                     enumeration
     +--ro status?                           enumeration
     +---x stop
     |  +--ro output
     |     +--ro result?   string
     +---x start
     |  +--ro output
     |     +--ro result?   string
     +---x restart
        +--ro output
           +--ro result?   string
```

{% endcode %}

## Java Packages and the Class Loader

Each NSO package will have a specific java classloader instance that loads its private jar classes. These package classloaders will refer to a single shared classloader instance as its parent. The shared classloader will load all shared jar classes for all deployed NSO packages.

{% hint style="info" %}
The `jar`'s in the `shared-jar` and `private-jar` directories should NOT be part of the Java classpath.
{% endhint %}

The purpose of this is first to keep integrity between packages which should not have access to each other's classes, other than the ones that are contained in the shared jars. Secondly, this way it is possible to hot redeploy the private jars and classes of a specific package while keeping other packages in a run state.

Should this class loading scheme not be desired, it is possible to suppress it by starting the NSO Java VM with the system property `TAILF_CLASSLOADER` set to false.

```
java -DTAILF_CLASSLOADER=false ...
```

This will force NSO Java VM to use the standard Java system classloader. For this to work, all `jar`'s from all deployed NSO packages need to be part of the classpath. The drawback of this is that all classes will be globally accessible and hot redeploy will have no effect.

There are four types of components that the NSO Java VM can handle:

* The `ned` type. The NSO Java VM will handle NEDs of sub-type `cli` and `generic` which are the ones that have a Java implementation.
* The `callback` type. These are any forms of callbacks that are defined by the DP API.
* The `application` type. These are user-defined daemons that implement a specific `ApplicationComponent` Java interface.
* The `upgrade` type. This component type is activated when deploying a new version of a NSO package and the NSO automatic CDB data upgrade is not sufficient. See [Writing an Upgrade Package Component](https://nso-docs.cisco.com/guides/development/core-concepts/nso-virtual-machines/pages/FxpCNgv5QKnfWrJw4nXf#ncs.cdb.upgrade.comp) for more information.

In some situations, several NSO packages are expected to use the same code base, e.g. when third-party libraries are used or the code is structured with some common parts. Instead of duplicate jars in several NSO packages, it is possible to create a new NSO package, add these jars to the `shared-jar` directory, and let the `package-meta-data.xml` file contains no component definitions at all. The NSO Java VM will load these shared jars and these will be accessible from all other NSO packages.

Inside the NSO Java VM, each component type has a specific Component Manager. The responsibility of these Managers is to manage a set of component classes for each NSO package. The Component Manager acts as an FSM that controls when a component should be registered, started, stopped, etc.

<div data-with-frame="true"><figure><img src="/files/0KEuZzmQ4TTs5F8cDQIt" alt="" width="563"><figcaption><p>Component Managers</p></figcaption></figure></div>

For instance, the `DpMuxManager` controls all callback implementations (services, actions, data providers, etc). It can load, register, start, and stop such callback implementations.

## The NED Component Type <a href="#d5e1240" id="d5e1240"></a>

NEDs can be of type `netconf`, `snmp`, `cli`*,* or `generic`. Only the `cli` and `generic` types are relevant for the NSO Java VM because these are the ones that have a Java implementation. Normally these NED components come in self-contained and prefabricated NSO packages for some equipment or class of equipment. It is however possible to tailor make NEDs for any protocol. For more information on this see [Network Element Drivers (NEDs)](/guides/development/advanced-development/developing-neds) and [Writing a data model for a CLI NED](/guides/development/advanced-development/developing-neds#writing-a-data-model-for-a-cli-ned) in NED Development

### The Callback Component Type <a href="#d5e1251" id="d5e1251"></a>

Callbacks are the collective name for a number of different functions that can be implemented in Java. One of the most important is the service callbacks, but also actions, transaction control, and data provision callbacks are in common use in an NSO implementation. For more on how to program callback using the DP API, see [DP API](https://nso-docs.cisco.com/guides/development/core-concepts/nso-virtual-machines/pages/Uzy6qvKpLQF47FSwk0S2#ug.java_api_overview.dp).

### The Application Component Type <a href="#d5e1255" id="d5e1255"></a>

For programs that are none of the above types but still need to access NSO as a daemon process, it is possible to use the `ApplicationComponent` Java interface. The `ApplicationComponent` interface expects the implementing classes to implement a `init()`, `finish()` and a `run()` method.

The NSO Java VM will start each class in a separate thread. The `init()` is called before the thread is started. The `run()` runs in a thread similar to the `run()` method in the standard Java `Runnable` interface. The `finish()` method is called when the NSO Java VM wants the application thread to stop. It is the responsibility of the programmer to stop the application thread i.e., stop the execution in the `run()` method when `finish()` is called. Note, that making the thread stop when `finish()` is called is important so that the NSO Java VM will not be hanging at a `STOP_VM` request.

{% code title="Example: ApplicationComponent Interface" %}

```java
package com.tailf.ncs;

/**
 * User defined Applications should implement this interface that
 * extends Runnable, hence also the run() method has to be implemented.
 * These applications are registered as components of type
 * "application" in a Ncs packages.
 *
 * Ncs Java VM will start this application in a separate thread.
 * The init() method is called before the thread is started.
 * The finish() method is expected to stop the thread. Hence stopping
 * the thread is user responsibility
 *
 */
public interface ApplicationComponent extends Runnable {

    /**
     * This method is called by the Ncs Java vm before the
     * thread is started.
     */
    public void init();

    /**
     * This method is called by the Ncs Java vm when the thread
     * should be stopped. Stopping the thread is the responsibility of
     * this method.
     */
    public void finish();

}
```

{% endcode %}

An example of an application component implementation is found in [SNMP Notification Receiver](/guides/development/connected-topics/snmp-notification-receiver).

## The Resource Manager <a href="#ncs.ug.javavm.resman" id="ncs.ug.javavm.resman"></a>

User Implementations typically need resources like Maapi, Maapi Transaction, Cdb, Cdb Session, etc. to fulfill their tasks. These resources can be instantiated and used directly in the user code. This implies that the user code needs to handle connection and close of additional sockets used by these resources. There is however another recommended alternative, and that is to use the Resource manager. The Resource manager is capable of injecting these resources into the user code. The principle is that the programmer will annotate the field that should refer to the resource rather than instantiate it.

{% code title="Example: Resource Injection" %}

```java
@Resource(type=ResourceType.MAAPI, scope=Scope.INSTANCE)
public Maapi m;
```

{% endcode %}

This way the NSO Java VM and the Resource manager can keep control over used resources and also can intervene e.g. close sockets at forced shutdowns.

The Resource manager can handle two types of resources: `MAAPI` and `CDB`.

{% code title="Example: Resource Types" %}

```java
package com.tailf.ncs.annotations;

/**
 * ResourceType set by the Ncs ResourceManager
 */
public enum ResourceType {

    MAAPI(1),
    CDB(2);
}
```

{% endcode %}

For both the Maapi and Cdb resource types a socket connection is opened towards NSO by the Resource manager. At a stop, the Resource manager will disconnect these sockets before ending the program. User programs can also tell the resource manager when its resources are no longer needed with a call to `ResourceManager.unregisterResources()`.

The resource annotation has three attributes:

* `type` defines the resource type.
* `scope` defines if this resource should be unique for each instance of the Java class (`Scope.INSTANCE`) or shared between different instances and classes (`Scope.CONTEXT`). For CONTEXT scope the sharing is confined to the defining NSO package, i.e., a resource cannot be shared between NSO packages.
* `qualifier` is an optional string to identify the resource as a unique resource. All instances that share the same context-scoped resource need to have the same qualifier. If the qualifier is not given it defaults to the value `DEFAULT` i.e., shared between all instances that have the `DEFAULT` qualifier.

{% code title="Example: Resource Annotation" %}

```java
package com.tailf.ncs.annotations;

/**
 * Annotation class for Action Callbacks Attributes are callPoint and callType
 */
@Retention(RetentionPolicy.RUNTIME)
@Target(ElementType.FIELD)
public @interface Resource {

    public ResourceType type();

    public Scope scope();

    public String qualifier() default "DEFAULT";

}
```

{% endcode %}

{% code title="Example: Scopes" %}

```java
package com.tailf.ncs.annotations;

/**
 * Scope for resources managed by the Resource Manager
 */
public enum Scope {

    /**
     * Context scope implies that the resource is
     * shared for all fields having the same qualifier in any class.
     * The resource is shared also between components in the package.
     * However sharing scope is confined to the package i.e sharing cannot
     * be extended between packages.
     * If the qualifier is not given it becomes "DEFAULT"
     */
    CONTEXT(1),
    /**
     * Instance scope implies that all instances will
     * get new resource instances. If the instance needs
     * several resources of the same type they need to have
     * separate qualifiers.
     */
    INSTANCE(2);
}
```

{% endcode %}

When the NSO Java VM starts it will receive component classes to load from NSO. Note, that the component classes are the classes that are referred to in the `package-meta-data.xml` file. For each component class, the Resource Manager will scan for annotations and inject resources as specified.

However, the package jars can contain lots of classes in addition to the component classes. These will be loaded at runtime and will be unknown by the NSO Java VM and therefore not handled automatically by the Resource Manager. These classes can also use resource injection but need a specific call to the Resource Manager for the mechanism to take effect. Before the resources are used for the first time the resource should be used, a call of `ResourceManager.registerResources(...)` will force the injection of the resources. If the same class is registered several times the Resource manager will detect this and avoid multiple resource injections.

{% code title="Example: Force Resource Injection" %}

```
MyClass myclass = new MyClass();
try {
    ResourceManager.registerResources(myclass);
} catch (Exception e) {
    LOGGER.error("Error injecting Resources", e);
}
```

{% endcode %}

## The Alarm Centrals

The `AlarmSourceCentral` and `AlarmSinkCentral`, which is part of the NSO Alarm API, can be used to simplify reading and writing alarms. The NSO Java VM will start these centrals at initialization. User implementations can therefore expect this to be set up without having to handle the start and stop of either the `AlarmSinkCentral` or the `AlarmSourceCentral`. For more information on the alarm API, see [Alarm Manager](/guides/operation-and-usage/operations/alarm-manager).

## Embedding the NSO Java VM <a href="#d5e1325" id="d5e1325"></a>

As stated above the NSO Java VM is executed in a thread implemented by the `NcsMain`. This implies that somewhere a java `main()` must be implemented that launches this thread. For NSO this is provided by the `NcsJVMLauncher` class. In addition to this, there is a script named `ncs-start-java-vm` that starts Java with the `NcsJVMLauncher.main()`. This is the recommended way of launching the NSO Java VM and how it is set up in a default installation. If there is a need to run the NSO Java VM as an embedded thread inside another program. This can be done simply by instantiating the class `NcsMain` and starting this instance in a new thread.

{% code title="Example: Starting NcsMain" %}

```
NcsMain ncsMain   = NcsMain.getInstance(host);
Thread  ncsThread = new Thread(ncsMain);

ncsThread.start();
```

{% endcode %}

However, with the embedding of the NSO Java VM comes the responsibility to manage the life cycle of the NSO Java VM thread. This thread cannot be started before NSO has started and is running or else the NSO Java VM control socket connection will fail. Also, running NSO without the NSO Java VM being launched will render runtime errors as soon as NSO needs NSO Java VM functionality.

## Logging

NSO has extensive logging functionality. Log settings are typically very different for a production system compared to a development system. Furthermore, the logging of the NSO daemon and the NSO Java VM is controlled by different mechanisms. During development, we typically want to turn on the `developer-log`. The sample `ncs.conf` that comes with the NSO release has log settings suitable for development, while the `ncs.conf` created by a System Install are suitable for production deployment.

The NSO Java VM uses Log4j for logging and will read its default log settings from a provided `log4j2.xml` file in the `ncs.jar`. Following that, NSO itself has `java-vm` log settings that are directly controllable from the NSO CLI. We can do:

```cli
admin@ncs(config)# java-vm java-logging logger com.tailf.maapi level level-trace
admin@ncs(config-logger-com.tailf.maapi)# commit
Commit complete.
```

This will dynamically reconfigure the log level for package `com.tailf.maapi` to be at the level `trace`. Where the Java logs end up is controlled by the `log4j2.xml` file. By default, the NSO Java VM writes to stdout. If the NSO Java VM is started by NSO, as controlled by the `ncs.conf` parameter `/java-vm/auto-start`, NSO will pick up the stdout of the service manager and write it to:

```cli
admin@ncs(config)# show full-configuration java-vm stdout-capture
java-vm stdout-capture file /var/log/ncs/ncs-java-vm.log
```

(The `details` pipe command also displays default values)

## The NSO Java VM Timeouts <a href="#d5e1406" id="d5e1406"></a>

The section `/ncs-config/api` in `ncs.conf` contains a number of very important timeouts. See `$NCS_DIR/src/ncs/ncs_config/tailf-ncs-config.yang` and [ncs.conf(5)](/guides/resources/man/ncs.conf.5) in Manual Pages for details.

* `new-session-timeout` controls how long NSO will wait for the NSO Java VM to respond to a new session.
* `query-timeout` controls how long NSO will wait for the NSO Java VM to respond to a request to get data.
* `connect-timeout` controls how long NSO will wait for the NSO Java VM to initialize a DP connection after the initial socket connect.
* `action-timeout` controls how long NSO will wait for the NSO Java VM to respond to an action request callback.

For `new-session-timeout`, `query-timeout`, and `connect-timeout`, whenever any of these timeouts trigger, NSO will close the sockets from NSO to the NSO Java VM. The NSO Java VM will detect the socket close and exit.

For `action-timeout`, whenever this timeout triggers, NSO will close only the sockets from NSO Java VM to the clients without exiting the Java VM.

If NSO is configured to start (and restart) the NSO Java VM, the NSO Java VM will be automatically restarted. If the NSO Java VM is started by some external entity, if it runs within an application server, it is up to that entity to restart the NSO Java VM.

## Debugging Startup <a href="#ug.javavm.debug" id="ug.javavm.debug"></a>

When using the `auto-start` feature (the default), NSO will start the NSO Java VM (as outlined in the start of this section), there are a number of different settings in the `java-vm` YANG model (see `$NCS_DIR/src/ncs/yang/tailf-ncs-java-vm.yang`) that controls what happens when something goes wrong during the startup.

The two timeout configurations `connect-time` and `initialization-time` are most relevant during startup. If the Java VM fails during the initial stages (during `INIT_JVM`, `LOAD_SHARED_JARS`, or `LOAD_PACKAGE`) either because of a timeout or because of a crash, NSO will log `The NCS Java VM synchronization failed` in `ncs.log`.

{% hint style="info" %}
The synchronization error message in the log will also have a hint as to what happened:

* `closed` usually means that the Java VM crashed (and closed the socket connected to NSO)
* `timeout` means that it failed to start (or respond) within the time limit. For example, if the Java VM runs out of memory and crashes, this will be logged as `closed`.
  {% endhint %}

After logging, NSO will take action based on the `synchronization-timeout-action` setting:

* `log`: NSO will log the failure, and if `auto-restart` is set to true NSO will try to restart the Java VM
* `log-stop` (default): NSO will log the failure, and if the Java VM has not stopped already NSO will also try to stop it. No restart action is taken.
* `exit`: NSO will log the failure, and then stop NSO itself.

If you have problems with the Java VM crashing during startup, a common pitfall is running out of memory (either total memory on the machine, or heap in the JVM). If you have a lot of Java code (or a loaded system) perhaps the Java VM did not start in time. Try to determine the root cause, check ncs.log and `ncs-java-vm.log`, and if needed increase the timeout.

For complex problems, for example with the class loader, try logging the internals of the startup:

```cli
admin@ncs(config)# java-vm java-logging logger com.tailf.ncs level level-all
admin@ncs(config-logger-com.tailf.maapi)# commit
Commit complete.
```

Setting this will result in a lot more detailed information in `ncs-java-vm.log` during startup.

When the `auto-restart` setting is `true` (the default), it means that NSO will try to restart the Java VM when it fails (at any point in time, not just during startup). NSO will at most try three restarts within 30 seconds, i.e., if the Java VM crashes more than three times within 30 seconds NSO gives up. You can check the status of the Java VM using the `java-vm` YANG model. For example in the CLI:

```cli
admin@ncs# show java-vm
java-vm start-status started
java-vm status running
```

The `start-status` can have the following values:

* `auto-start-not-enabled`: Autostart is not enabled.
* `stopped`: The Java VM has been stopped or is not yet started.
* `started`: The Java VM has been started. See the leaf 'status' to check the status of the Java application code.
* `failed`: The Java VM has been terminated. If `auto-restart` is enabled, the Java VM restart has been disabled due to too frequent restarts.

The `status` can have the following values:

* `not-connected`: The Java application code is not connected to NSO.
* `initializing`: The Java application code is connected to NSO, but not yet initialized.
* `running`: The Java application code is connected and initialized.
* `timeout`: The Java application connected to NSO, but failed to initialize within the stipulated timeout 'initialization-time'.


# Embedded Erlang Applications

Start user-provided Erlang applications.

NSO is capable of starting user-provided Erlang applications embedded in the same Erlang VM as NSO.

The Erlang code is packaged into applications which are automatically started and stopped by NSO if they are located at the proper place. NSO will search all packages for top-level directories called `erlang-lib`. The structure of such a directory is the same as a standard `lib` directory in Erlang. The directory may contain multiple Erlang applications. Each one must have a valid `.app` file. See the Erlang documentation of `application` and `app` for more info.

An Erlang package skeleton can be created by making use of the `ncs-make-package` command:

```bash
ncs-make-package --erlang-skeleton --erlang-application-name <appname> <package-name>
```

Multiple applications can be generated by using the option `--erlang-application-name NAME` multiple times with different names.

All application code should use the prefix `ec_` for module names, application names, registered processes (if any), and named `ets` tables (if any), to avoid conflict with existing or future names used by NSO itself.

## Erlang API <a href="#d5e1761" id="d5e1761"></a>

The Erlang API to NSO is implemented as an Erlang/OTP application called `econfd`. This application comes in two flavors. One is built into NSO to support applications running in the same Erlang VM as NSO. The other is a separate library which is included in source form in the NSO release, in the `$NCS_DIR/erlang` directory. Building `econfd` as described in the `$NCS_DIR/erlang/econfd/README` file will compile the Erlang code and generate the documentation.

This API can be used by applications written in Erlang in much the same way as the C and Java APIs are used, i.e. code running in an Erlang VM can use the `econfd` API functions to make socket connections to NSO for the data provider, MAAPI, CDB, etc. access. However, the API is also available internally in NSO, which makes it possible to run Erlang application code inside the NSO daemon, without the overhead imposed by the socket communication.

When the application is started, one of its processes should make initial connections to the NSO subsystems, register callbacks, etc. This is typically done in the `init/1` function of a `gen_server` or similar. While the internal connections are made using the exact same API functions (e.g. `econfd_maapi:connect/2`) as for an application running in an external Erlang VM, any `Address` and `Port` arguments are ignored, and instead, standard Erlang inter-process communication is used.

There is little or no support for testing and debugging Erlang code executing internally in NSO since NSO provides a very limited runtime environment for Erlang to minimize disk and memory footprints. Thus the recommended method is to develop Erlang code targeted for this by using `econfd` in a separate Erlang VM, where an interactive Erlang shell and all the other development support included in the standard Erlang/OTP releases are available. When development and testing are completed, the code can be deployed to run internally in NSO without changes.

For information about the Erlang programming language and development tools, refer to [www.erlang.org](https://www.erlang.org/) and the available books about Erlang (some are referenced on the website).

The `--printlog` option to `ncs`, which prints the contents of the NSO error log, is normally only useful for Cisco support and developers, but it may also be relevant for debugging problems with application code running inside NSO. The error log collects the events sent to the OTP error\_logger, e.g. crash reports as well as info generated by calls to functions in the error\_logger(3) module. Another possibility for primitive debugging is to run `ncs` with the `--foreground` option, where calls to `io:format/2` etc will print to standard output. Printouts may also be directed to the developer log by using `econfd:log/3`.

While Erlang application code running in an external Erlang VM can use basically any version of Erlang/OTP, this is not the case for code running inside NSO, since the Erlang VM is evolving and provides limited backward/forward compatibility. To avoid incompatibility issues when loading the `beam` files, the Erlang compiler `erlc` should be of the same version as was used to build the NSO distribution.

NSO provides the VM, `erlc` and the `kernel`, `stdlib`, and `crypto` OTP applications.

{% hint style="info" %}
Application code running internally in the NSO daemon can have an impact on the execution of the standard NSO code. Thus, it is critically important that the application code is thoroughly tested and verified before being deployed for production in a system using NSO.
{% endhint %}

## Application Configuration <a href="#d5e1797" id="d5e1797"></a>

Applications may have dependencies to other applications. These dependencies affect the start order. If the dependent application resides in another package, this should be expressed by using the required package in the `package-meta-data.xml` file. Application dependencies within the same package should be expressed in the `.app`. See below.

The following config settings in the `.app` file are explicitly treated by NSO:

<table data-header-hidden><thead><tr><th width="275.65130615234375"></th><th></th></tr></thead><tbody><tr><td><code>applications</code></td><td>A list of applications that need to be started before this application can be started. This info is used to compute a valid start order.</td></tr><tr><td><code>included_applications</code></td><td>A list of applications that are started on behalf of this application. This info is used to compute a valid start order.</td></tr><tr><td><code>env</code></td><td>A property list, containing <code>[{Key,Val}]</code> tuples. Besides other keys, used by the application itself, a few predefined keys are used by NSO. The key <code>ncs_start_phase</code> is used by NSO to determine which start phase the application is to be started in. Valid values are <code>early_phase0</code>, <code>phase0</code>, <code>phase1</code>, <code>phase1_delayed</code> and <code>phase2</code>. Default is <code>phase1</code>. If the application is not required in the early phases of startup, set <code>ncs_start_phase</code> to <code>phase2</code> to avoid issues with NSO services being unavailable to the application. The key <code>ncs_restart_type</code> is used by NSO to determine what impact a restart of the application will have. This is the same as the <code>restart_type()</code> type in <code>application</code>. Valid values are <code>permanent</code>, <code>transient</code> and <code>temporary</code>. Default is <code>temporary</code>.</td></tr></tbody></table>

## Example <a href="#d5e1835" id="d5e1835"></a>

The [examples.ncs/service-management/rfs-service-erlang](https://github.com/NSO-developer/nso-examples/tree/6.7/service-management/rfs-service-erlang) example in the bundled collection shows how to create a service written in Erlang and execute it internally in NSO. This Erlang example is a subset of the Java example [examples.ncs/service-management/rfs-service](https://github.com/NSO-developer/nso-examples/tree/6.7/service-management/rfs-service).


# API Overview

Overview of NSO APIs.

NSO uses socket communication to coordinate work with applications, such as a Python or Java service. In addition to the control socket, NSO uses a number of worker sockets to process individual requests: performing service mapping or executing an action, for example. We collectively call these programs data provider applications, since the data provider protocol underpins all of them.

## Application Timeouts

The communication with data provider applications is subject to timeouts in order to manage the execution time of requests. These are defined in section `/ncs-config/api` in `ncs.conf`:

* `ncs-config/api/action-timeout`
* `ncs-config/api/query-timeout`
* `ncs-config/api/new-session-timeout`
* `ncs-config/api/connect-timeout`

For executing actions invoked by the clients, NSO uses `action-timeout` to ensures the response from data provider is received within the given time. If the data provider fails to do so within the stipulated timeout, NSO will kill the worker sockets executing the actions and trigger the abort action defined in `cb_abort()` without restarting the NSO VMs. The following code shows a trivial implementation of an abort action callback:

```python
class MyTestAction(Action):
    def cb_abort(self, uinfo):
        self.log.info('Action aborted: ')
```

There are some important points worth noting for action timeout:

* An action callback that times out in one user instance will not affect the result of an action callback in another user instance. This is because NSO executes actions using multiple worker sockets, and an action timeout will only terminate the worker socket executing that specific action.
* Implementing your own abort action callback in `cb_abort` allows you to handle actions that are timing out. If `cb_abort` is not defined, NSO cannot trigger the abort action during a timeout, preventing it from unlocking the action for a user session. Consequently, you must wait for the action callback to finish before attempting it again.

{% hint style="info" %}
See [examples.ncs/sdk-api/action-abort-py](https://github.com/NSO-developer/nso-examples/tree/6.7/sdk-api/action-abort-py) for an example of how to implement an abortable Python action that spawns a separate worker process using the multiprocessing library and returns the worker's outcome via a result queue or terminates the worker if the action is aborted.
{% endhint %}

For NSO operational data queries, NSO uses `query-timeout` to ensure the data provider return operational data within the given time. If the data provider fails to do so within the stipulated timeout, NSO will close its end of the control socket to the data provider. The NSO VMs will detect the socket close and exit.

For connection initiation requests between NSO and data providers, NSO uses `connect-timeout` to ensure the data provider send the initial message after connecting the socket to NSO within the given time. If the data provider fails to do so within the stipulated timeout, NSO will close its end of the control socket to the data provider. The NSO VMs will detect the socket close and exit.

For requests invoked by NSO, NSO uses `new-session-timeout` to ensure the data provider respond to the control socket request within the given time. If the data provider fails to do so within the stipulated timeout, NSO will close its end of the control socket to the data provider. The NSO VMs will detect the socket close and exit.


# Python API Overview

Learn about the NSO Python API and its usage.

The NSO Python library contains a variety of APIs for different purposes. In this section, we introduce these and explain their usage. The NSO Python module deliverables are found in two variants, the low-level APIs and the high-level APIs.

The low-level APIs are a direct mapping of the NSO C APIs, CDB, and MAAPI. These will follow the evolution of the C APIs. See `man confd_lib_lib` for further information.

The high-level APIs are an abstraction layer on top of the low-level APIs to make them easier to use and to improve code readability and development rate for common use cases. E.g. services and action callbacks and common scripting towards NSO.

## Python API Overview <a href="#d5e4354" id="d5e4354"></a>

<table data-header-hidden data-full-width="false"><thead><tr><th width="350"></th><th></th></tr></thead><tbody><tr><td><strong>MAAPI (Management Agent API)</strong><br>Northbound interface that is transactional and user session-based. Using this interface, both configuration and operational data can be read. Configuration and operational data can be written and committed as one transaction. The API is complete in the way that it is possible to write a new northbound agent using only this interface. It is also possible to attach to ongoing transactions to read uncommitted changes and/or modify data in these transactions.</td><td><img src="/files/sCt1fr2cZH59OFoC8wRO" alt="" data-size="original"></td></tr><tr><td><strong>Python low-level CDB API</strong><br>The Southbound interface provides access to the CDB configuration database. Using this interface, configuration data can be read. In addition, operational data that is stored in CDB can be read and written. This interface has a subscription mechanism to subscribe to changes. A subscription is specified on a path that points to an element in a YANG model or an instance in the instance tree. Any change under this point will trigger the subscription. CDB also has functions to iterate through the configuration changes when a subscription has been triggered.</td><td><img src="/files/qkTURcbkVJ8SOVhPb0iE" alt="" data-size="original"></td></tr><tr><td><strong>Python low-level DP API</strong><br>Southbound interface that enables callbacks, hooks, and transforms. This API makes it possible to provide the service callbacks that handle service-to-device mapping logic. Other usual cases are external data providers for operational data or action callback implementations. There are also transaction and validation callbacks, etc. Hooks are callbacks that are fired when certain data is written and the hook is expected to do additional modifications of data. Transforms are callbacks that are used when complete mediation between two different models is necessary.</td><td><img src="/files/4WbyLbznPMPqze3NUzGZ" alt="" data-size="original"></td></tr><tr><td><strong>Python high-level API</strong>: API that resides on top of the MAAPI, CDB, and DP APIs. It provides schema model navigation and instance data handling (read/write). Uses a MAAPI context as data access and incorporates its functionality. It is used in service implementations, action handlers, and Python scripting.</td><td><img src="/files/g4kDYYlhwhLZWfQATDj9" alt="" data-size="original"></td></tr></tbody></table>

## Python scripting <a href="#d5e4389" id="d5e4389"></a>

Scripting in Python is a very easy and powerful way of accessing NSO. This document has several examples of scripts showing various ways of accessing data and requesting actions in NSO.

The examples are directly executable with the Python interpreter after sourcing the `ncsrc` file in the NSO installation directory. This sets up the `PYTHONPATH` environment variable, which enables access to the NSO Python modules.

Edit a file and execute it directly on the command line like this:

```bash
$ python3 script.py
```

## High-level MAAPI API <a href="#d5e4398" id="d5e4398"></a>

The Python high-level MAAPI API provides an easy-to-use interface for accessing NSO. Its main targets are to encapsulate the sockets, transaction handles, data type conversions, and the possibility of using the Python `with` statement for proper resource cleanup.

The simplest way to access NSO is to use the `single_transaction` helper. It creates a MAAPI context and a transaction in one step.

This example shows its usage, connecting as user `admin` and `python` in the AAA context:

{% code title="Example: Single Transaction Helper" %}

```python
import ncs

with ncs.maapi.single_write_trans('admin', 'python') as t:
    t.set_elem2('Kilroy was here', '/ncs:devices/device{ce0}/description')
    t.apply()

with ncs.maapi.single_read_trans('admin', 'python') as t:
    desc = t.get_elem('/ncs:devices/device{ce0}/description')
    print("Description for device ce0 = %s" % desc)
```

{% endcode %}

{% hint style="danger" %}
The example code here shows how to start a transaction but does not properly handle the case of concurrency conflicts when writing data. See [Handling Conflicts](https://nso-docs.cisco.com/guides/development/core-concepts/api-overview/pages/8PJW3J1uuhcu5ttTWYkl#ncs.development.concurrency.handling) for details.
{% endhint %}

{% hint style="warning" %}
When only reading data, always start a `read` transaction to read directly from the CDB datastore and data providers. `write` transactions cache repeated reads done by the same transaction.
{% endhint %}

A common use case is to create a MAAPI context and reuse it for several transactions. This reduces the latency and increases the transaction throughput, especially for backend applications. For scripting the lifetime is shorter and there is no need to keep the MAAPI contexts alive.

This example shows how to keep a MAAPI connection alive between transactions:

{% code title="Example: Reading of Configuration Data using High-level MAAPI" %}

```python
import ncs

with ncs.maapi.Maapi() as m:
    with ncs.maapi.Session(m, 'admin', 'python'):

        # The first transaction
        with m.start_read_trans() as t:
            address = t.get_elem('/ncs:devices/device{ce0}/address')
            print("First read: Address = %s" % address)

        # The second transaction
        with m.start_read_trans() as t:
            address = t.get_elem('/ncs:devices/device{ce1}/address')
            print("Second read: Address = %s" % address)
```

{% endcode %}

## Maagic API

Maagic is a module provided as part of the NSO Python APIs. It reduces the complexity of programming towards NSO, is used on top of the MAAPI high-level API, and addresses areas that require more programming. First, it helps in navigating the model, using standard Python object dot notation, giving very clear and easily read code. The context handlers remove the need to close sockets, user sessions, and transactions and the problems when they are forgotten and kept open. Finally, it removes the need to know the data types of the leafs, helping you to focus on the data to be set.

When using Maagic, you still do the same procedure of starting a transaction.

```python
with ncs.maapi.Maapi() as m:
  with ncs.maapi.Session(m, 'admin', 'python'):
    with m.start_write_trans() as t:
      # Read/write/request ...
```

To use the Maagic functionality, you get access to a Maagic object either pointing to the root of the CDB:

```
root = ncs.maagic.get_root(t)
```

In this case, it is a `ncs.maagic.Node` object with a `ncs.maapi.Transaction` backend.

From here, you can navigate in the model. In the table, you can see examples of how to navigate.

The table below lists Maagic object navigation.

| Action                                       | Returns             |
| -------------------------------------------- | ------------------- |
| `root.devices`                               | `Container`         |
| `root.devices.device`                        | `List`              |
| `root.devices.device['ce0']`                 | `ListElement`       |
| `root.devices.device['ce0'].device_type.cli` | `PresenceContainer` |
| `root.devices.device['ce0'].address`         | `str`               |
| `root.devices.device['ce0'].port`            | `int`               |

You can also get a Maagic object from a keypath:

```
node = ncs.maagic.get_node(t, '/ncs:devices/device{ce0}')
```

### Namespaces <a href="#d5e4469" id="d5e4469"></a>

Maagic handles namespaces by a prefix to the names of the elements. This is optional but recommended to avoid future side effects.

The syntax is to prefix the names with the namespace name followed by two underscores, e.g., `ns_name__ name`.

Examples of how to use namespaces:

```bash
# The examples are equal unless there is a namespace collision.
# For the ncs namespace it would look like this:

root.ncs__devices.ncs__device['ce0'].ncs__address
# equals
root.devices.device['ce0'].address
```

In cases where there is a name collision, the namespace prefix is required to access an entity from a module, except for the module that was first loaded. A namespace is always required for root entities when there is a collision. The module load order is found in the NCS log file: `logs/ncs.log`.

```bash
# This example have three namespaces referring to a leaf, value, with the same
# name and this load order: /ex/a:value=11, /ex/b:value=22 and /ex/c:value=33

root.ex.value # returns 11
root.ex.a__value # returns 11
root.ex.b__value # returns 22
root.ex.c__value # returns 33
```

### Reading Data <a href="#d5e4480" id="d5e4480"></a>

Reading data using Maagic is straightforward. You will just specify the leaf you are interested in and the data is retrieved. The data is returned in the nearest available Python data type.

For non-existing leafs, `None` is returned.

```
dev_name = root.devices.device['ce0'].name # 'ce0'
dev_address = root.devices.device['ce0'].address # '127.0.0.1'
dev_port = root.devices.device['ce0'].port # 10022
```

### Writing Data <a href="#d5e4486" id="d5e4486"></a>

Writing data using Maagic is straightforward. You will just specify the leaf you are interested in and assign a value. Any data type can sent as input, as the `str` function is called, converting it to a string. The format depends on the data type. If the type validation fails, an `Error` exception is thrown.

```
root.devices.device['ce0'].name  = 'ce0'
root.devices.device['ce0'].address  = '127.0.0.1'
root.devices.device['ce0'].port = 10022
root.devices.device['ce0'].port = '10022' # Also valid

# This will raise an Error exception
root.devices.device['ce0'].port = 'netconf'
```

### Deleting Data <a href="#d5e4492" id="d5e4492"></a>

Data is deleted the Python way of using the `del` function:

```
del root.devices.device['ce0'] # List element
del root.devices.device['ce0'].name # Leaf
del root.devices.device['ce0'].device_type.cli # Presence container
```

Some entities have a delete method, this is explained under the corresponding type.

### Object Deletion

The delete mechanism in Maagic is implemented using the `__delattr__` method on the `Node` class. This means that executing the del function on a local or global variable will only delete the object from the Python local or global namespaces. E.g., `del obj`.

### Containers <a href="#d5e4504" id="d5e4504"></a>

Containers are addressed using standard Python dot notation: `root.container1.container2`.

### Presence Containers <a href="#d5e4508" id="d5e4508"></a>

A presence container is created using the `create` method:

```
pc = root.container.presence_container.create()
```

Existence is checked with the `exists` or `bool` functions:

```
root.container.presence_container.exists() # Returns True or False
bool(root.container.presence_container) # Returns True or False
```

A presence container is deleted with the `del` or `delete` functions:

```
del root.container.presence_container
root.container.presence_container.delete()
```

### Choices <a href="#d5e4516" id="d5e4516"></a>

The case of a choice is checked by addressing the name of the choice in the model:

```
ne_type = root.devices.device['ce0'].device_type.ne_type
if ne_type == 'cli':
  # Handle CLI
elif ne_type == 'netconf':
  # Handle NETCONF
elif ne_type == 'generic':
  # Handle generic
else:
  # Don't handle
```

Changing a choice is done by setting a value in any of the other cases:

```
root.devices.device['ce0'].device_type.netconf.create()
str(root.devices.device['ce0'].device_type.ne_type) # Returns 'netconf'
```

### Lists and List Elements <a href="#d5e4522" id="d5e4522"></a>

List elements are created using the create method on the `List` class:

```bash
# Single value key
ce5 = root.devices.device.create('ce5')

# Multiple values key
o = root.container.list.create('foo', 'bar')
```

The objects `ce5` and *`o`* above are of type `ListElement` which is actually an ordinary `container` object with a different name.

Existence is checked with the `exists` or `bool` functions `List` class:

```
'ce0' in root.devices.device # Returns True or False
```

A list element is deleted with the Python `del` function:

```bash
# Single value key
del root.devices.device['ce5']

# Multiple values key
del root.container.list['foo', 'bar']
```

To delete the whole list, use the Python `del` function or `delete()` on the list.

```bash
# use Python's del function
del root.devices.device

# use List's delete() method
root.container.list.delete()
```

### Unions <a href="#d5e4539" id="d5e4539"></a>

Unions are not handled in any specific way - you just read or write to the leaf and the data is validated according to the model.

### Enumeration <a href="#d5e4542" id="d5e4542"></a>

Enumerations are returned as an `Enum` object, giving access to both the integer and string values.

```
str(root.devices.device['ce0'].state.admin_state) # May return 'unlocked'
root.devices.device['ce0'].state.admin_state.string # May return 'unlocked'
root.devices.device['ce0'].state.admin_state.value # May return 1
```

Writing values to enumerations accepts both the string and integer values.

```
root.devices.device['ce0'].state.admin_state = 'locked'
root.devices.device['ce0'].state.admin_state = 0

# This will raise an Error exception
root.devices.device['ce0'].state.admin_state = 3 # Not a valid enum
```

### Leafref <a href="#d5e4549" id="d5e4549"></a>

Leafrefs are read as regular leafs and the returned data type corresponds to the referred leaf.

```bash
# /model/device is a leafref to /devices/device/name

dev = root.model.device # May return 'ce0'
```

Leafrefs are set as the leaf they refer to. The data type is validated as it is set. The reference is validated when the transaction is committed.

```bash
# /model/device is a leafref to /devices/device/name

root.model.device = 'ce0'
```

### Identityref <a href="#d5e4555" id="d5e4555"></a>

Identityrefs are read and written as string values. Writing an identityref without a prefix is possible, but doing so is error-prone and may stop working if another model is added which also has an identity with the same name. The recommendation is to always use a prefix when writing identityrefs. Reading an identityref will always return a prefixed string value.

```bash
# Read
root.devices.device['ce0'].device_type.cli.ned_id # May return 'ios-id:cisco-ios'

# Write when identity cisco-ios is unique throughout the system (not recommended)
root.devices.device['ce0'].device_type.cli.ned_id = 'cisco-ios'

# Write with unique identity
root.devices.device['ce0'].device_type.cli.ned_id = 'ios-id:cisco-ios'
```

### Instance Identifier <a href="#d5e4559" id="d5e4559"></a>

Instance identifiers are read as xpath formatted string values.

```bash
# /model/iref is an instance-identifier

root.model.iref # May return "/ncs:devices/ncs:device[ncs:name='ce0']"
```

Instance identifiers are set as xpath formatted strings. The string is validated as it is set. The reference is validated when the transaction is committed.

```bash
# /model/iref is an instance-identifier

root.devices.device['ce0'].device_type.cli.ned_id = "/ncs:devices/ncs:device[ncs:name='ce0']"
```

### Leaf-list <a href="#d5e4565" id="d5e4565"></a>

A leaf-list is represented by a `LeafList` object. This object behaves very much like a Python list. You may iterate it, check for the existence of a specific element using `in`, or remove specific items using the `del` operator. See examples below.

{% hint style="info" %}
From NSO version 4.5 and onwards, a Yang leaf-list is represented differently than before. Reading a leaf-list using Maagic used to result in an ordinary Python list (or None if the leaf-list was non-existent). Now, reading a leaf-list will give back a `LeafList` object whether it exists or not. The `LeafList` object may be iterated like a Python list and you may check for existence using the `exists()` method or the `bool()` operator. A Maagic leaf-list node may be assigned using a Python list, just like before, and you may convert it to a Python list using the `as_list()` method or by doing `list(my_leaf_list_node)`.
{% endhint %}

```bash
# /model/ll is a leaf-list with the type string

# read a LeafList object
ll = root.model.ll

# iteration
for item in root.model.ll:
    do_stuff(item)

# check if the leaf-list exists (i.e. is non-empty)
if root.model.ll:
    do_stuff()
if root.model.ll.exists():
    do_stuff()

# check the leaf-list contains a specific item
if 'foo' in root.model.ll:
    do_stuff()

# length
len(root.model.ll)

# create a new item in the leaf-list
root.model.ll.create('bar')

# set the whole leaf-list in one operation
root.model.ll = ['foo', 'bar', 'baz']

# remove a specific item from the list
del root.model.ll['bar']
root.model.ll.remove('baz')

# delete the whole leaf-list
del root.model.ll
root.model.ll.delete()

# get the leaf-list as a Python list
root.model.ll.as_list()
```

### Binary <a href="#d5e4581" id="d5e4581"></a>

Binary values are read and written as byte strings.

```bash
# Read
root.model.bin # May return '\x00foo\x01bar'

# Write
root.model.bin = b'\x00foo\x01bar'
```

### Bits <a href="#d5e4585" id="d5e4585"></a>

Reading a `bits` leaf will give a Bits object back (or None if the `bits` leaf is non-existent). To get some useful information out of the Bits object, you can either use the `bytearray()` method to get a Python byte array object in return or the Python `str()` operator to get a space-separated string containing the bit names.

```bash
# read a bits leaf - a Bits object may be returned (None if non-existent)
root.model.bits

# get a bytearray
root.model.bits.bytearray()

# get a space separated string with bit names
str(root.model.bits)
```

There are four ways of setting a `bits` leaf: One is to set it using a string with space-separated bit names, the other one is to set it using a byte array, the third by using a Python binary string, and as a last option is it may be set using a Bits object. Note that updating a Bits object does not change anything in the database - for that to happen, you need to assign it to the Maagic node.

```bash
# set a bits leaf using a string of space separated bit names
root.model.bits = 'turboMode enableEncryption'

# set a bits leaf using a Python bytearray
root.model.bits = bytearray(b'\x11')

# set a bits leaf using a Python binary string
root.model.bits = b'\x11'

# read a bits leaf, update the Bits object and set it
b = x.model.bits
b.clr_bit(0)
x.model.bits = b
```

### Empty Leaf <a href="#d5e4591" id="d5e4591"></a>

An empty leaf is created using the `create` method. If the type empty leaf is part of a union, the leaf must be set to the `C_EMPTY` value instead.

```
pc = root.container.empty_leaf.create()
```

If the type empty leaf is part of a union, then you read the leaf to see if `empty` is the current value. Otherwise, existence is checked with the `exists` or `bool` functions:

```
root.container.empty_leaf.exists() # Returns True or False
bool(root.container.empty_leaf) # Returns True or False
```

An empty leaf is deleted with the `del` or `delete` functions:

```
del root.container.empty_leaf
root.container.empty_leaf.delete()
```

## Maagic Examples <a href="#d5e4599" id="d5e4599"></a>

### Action Requests <a href="#d5e4601" id="d5e4601"></a>

Requesting an action may not require an ongoing transaction and this example shows how to use Maapi as a transactionless back-end for Maagic.

{% code title="Example: Action Request without Transaction" %}

```python
import ncs

with ncs.maapi.Maapi() as m:
    with ncs.maapi.Session(m, 'admin', 'python'):
        root = ncs.maagic.get_root(m)

        output = root.devices.check_sync()

        for result in output.sync_result:
            print('sync-result {')
            print('    device %s' % result.device)
            print('    result %s' % result.result)
            print('}')
```

{% endcode %}

This example shows how to request an action that requires an ongoing transaction. It is also valid to request an action that does not require an ongoing transaction.

{% code title="Example: Action Request with Transaction" %}

```python
import ncs

with ncs.maapi.Maapi() as m:
    with ncs.maapi.Session(m, 'admin', 'python'):
        with m.start_read_trans() as t:
            root = ncs.maagic.get_root(t)

            output = root.devices.check_sync()

            for result in output.sync_result:
                print('sync-result {')
                print('    device %s' % result.device)
                print('    result %s' % result.result)
                print('}')
```

{% endcode %}

Providing parameters to an action with Maagic is very easy: You request an input object, with `get_input` from the Maagic action object, and set the desired (or required) parameters as defined in the model specification.

{% code title="Example: Action Request with Input Parameters" %}

```python
import ncs

with ncs.maapi.Maapi() as m:
    with ncs.maapi.Session(m, 'admin', 'python'):
        root = ncs.maagic.get_root(m)

        input = root.action.double.get_input()
        input.number = 21
        output = root.action.double(input)

        print(output.result)
```

{% endcode %}

If you have a leaf-list, you need to prepare the input parameters

{% code title="Example: Action Request with leaf-list Input Parameters" %}

```python
import ncs

with ncs.maapi.Maapi() as m:
    with ncs.maapi.Session(m, 'admin', 'python'):
        root = ncs.maagic.get_root(m)

        input = root.leaf_list_action.llist.get_input()
        input.args = ['testing action']
        output = root.leaf_list_action.llist(input)

        print(output.result)
```

{% endcode %}

A common use case is to script the creation of devices. With the Python APIs, this is easily done without the need to generate set commands and execute them in the CLI.

{% code title="Example: Create Device, Fetch Host Keys, and Synchronize Configuration" %}

```python
import argparse
import ncs


def parseArgs():
    parser = argparse.ArgumentParser()
    parser.add_argument('--name', help="device name", required=True)
    parser.add_argument('--address', help="device address", required=True)
    parser.add_argument('--port', help="device address", type=int, default=22)
    parser.add_argument('--desc', help="device description",
                        default="Device created by maagic_create_device.py")
    parser.add_argument('--auth', help="device authgroup", default="default")
    return parser.parse_args()


def main(args):
    with ncs.maapi.Maapi() as m:
        with ncs.maapi.Session(m, 'admin', 'python'):
            with m.start_write_trans() as t:
                root = ncs.maagic.get_root(t)

                print("Setting device '%s' configuration..." % args.name)

                # Get a reference to the device list
                device_list = root.devices.device

                device = device_list.create(args.name)
                device.address = args.address
                device.port = args.port
                device.description = args.desc
                device.authgroup = args.auth
                dev_type = device.device_type.cli
                dev_type.ned_id = 'cisco-ios-cli-3.0'
                device.state.admin_state = 'unlocked'

                print('Committing the device configuration...')
                t.apply()
                print("Committed")

                # This transaction is no longer valid

            #
            # fetch-host-keys and sync-from does not require a transaction
            # continue using the Maapi object
            #
            root = ncs.maagic.get_root(m)
            device = root.devices.device[args.name]

            print("Fetching SSH keys...")
            output = device.ssh.fetch_host_keys()
            print("Result: %s" % output.result)

            print("Syncing configuration...")
            output = device.sync_from()
            print("Result: %s" % output.result)
            if not output.result:
                print("Error: %s" % output.info)


if __name__ == '__main__':
    main(parseArgs())
```

{% endcode %}

## PlanComponent

This class is a helper to support service progress reporting using `plan-data` as part of a Reactive FASTMAP nano service. More info about `plan-data` is found in [Nano Services for Staged Provisioning](/guides/development/core-concepts/nano-services).

The interface of the `PlanComponent` is identical to the corresponding Java class and supports the setup of plans and setting the transition states.

```python
class PlanComponent(object):
    """Service plan component.

    The usage of this class is in conjunction with a nano service that
    uses a reactive FASTMAP pattern.
    With a plan the service states can be tracked and controlled.

    A service plan can consist of many PlanComponent's.
    This is operational data that is stored together with the service
    configuration.
    """

    def __init__(self, service, name, component_type):
        """Initialize a PlanComponent."""

    def append_state(self, state_name):
        """Append a new state to this plan component.

        The state status will be initialized to 'ncs:not-reached'.
        """

    def set_reached(self, state_name):
        """Set state status to 'ncs:reached'."""

    def set_failed(self, state_name):
        """Set state status to 'ncs:failed'."""

    def set_status(self, state_name, status):
        """Set state status."""
```

See `pydoc3 ncs.application.PlanComponent` for further information about the Python class.

The pattern is to add an overall plan (self) for the service and separate plans for each component that builds the service.

```
self_plan = PlanComponent(service, 'self', 'ncs:self')
self_plan.append_state('ncs:init')
self_plan.append_state('ncs:ready')
self_plan.set_reached('ncs:init')

route_plan = PlanComponent(service, 'router', 'myserv:router')
route_plan.append_state('ncs:init')
route_plan.append_state('myserv:syslog-initialized')
route_plan.append_state('myserv:ntp-initialized')
route_plan.append_state('myserv:dns-initialized')
route_plan.append_state('ncs:ready')
route_plan.set_reached('ncs:init')
```

When appending a new state to a plan the initial state is set to `ncs:not-reached`. At the completion of a plan the state is set to `ncs:ready`. In this case when the service is completely setup:

```
self_plan.set_reached('ncs:ready')
```

## Python Packages <a href="#d5e4639" id="d5e4639"></a>

### Action Handler <a href="#d5e4641" id="d5e4641"></a>

The Python high-level API provides an easy way to implement an action handler for your modeled actions. The easiest way to create a handler is to use the `ncs-make-package` command. It creates some ready-to-use skeleton code.

```bash
$ cd packages
$ ncs-make-package --service-skeleton python pyaction --component-class
 action.Action \
 --action-example
```

The generated package skeleton:

```bash
$ tree pyaction
pyaction/
+-- README
+-- doc/
+-- load-dir/
+-- package-meta-data.xml
+-- python/
|   +-- pyaction/
|       +-- __init__.py
|       +-- action.py
+-- src/
|   +-- Makefile
|   +-- yang/
|       +-- action.yang
+-- templates/
```

This example action handler takes a number as input, doubles it, and returns the result.

When debugging Python packages refer to [Debugging of Python Packages](/guides/development/core-concepts/nso-virtual-machines/nso-python-vm#debugging-of-python-packages).

{% code title="Example: Action Server Implementation" %}

```bash
# -*- mode: python; python-indent: 4 -*-

from ncs.application import Application
from ncs.dp import Action

# ---------------
# ACTIONS EXAMPLE
# ---------------
class DoubleAction(Action):
    @Action.action
    def cb_action(self, uinfo, name, kp, input, output):
        self.log.info('action name: ', name)
        self.log.info('action input.number: ', input.number)

        output.result = input.number * 2

class LeafListAction(Action):
    @Action.action
    def cb_action(self, uinfo, name, kp, input, output):
        self.log.info('action name: ', name)
        self.log.info('action input.args: ', input.args)
        output.result = [ w.upper() for w in input.args]

# ---------------------------------------------
# COMPONENT THREAD THAT WILL BE STARTED BY NCS.
# ---------------------------------------------
class Action(Application):
    def setup(self):
        self.log.info('Worker RUNNING')
        self.register_action('action-action', DoubleAction)
        self.register_action('llist-action', LeafListAction)

    def teardown(self):
        self.log.info('Worker FINISHED')
```

{% endcode %}

Test the action by doing a request from the NSO CLI:

```
admin@ncs> request action double number 21
result 42
[ok][2016-04-22 10:30:39]
```

The input and output parameters are the most commonly used parameters of the action callback method. They provide the access objects to the data provided to the action request and the returning result.

They are `maagic.Node` objects, which provide easy access to the modeled parameters.

The table below lists the action handler callback parameters:

<table><thead><tr><th width="141">Parameter</th><th width="201">Type</th><th>Description</th></tr></thead><tbody><tr><td><code>self</code></td><td><code>ncs.dp.Action</code></td><td>The action object.</td></tr><tr><td><code>uinfo</code></td><td><code>ncs.UserInfo</code></td><td>User information of the requester.</td></tr><tr><td><code>name</code></td><td><code>string</code></td><td>The tailf:action name.</td></tr><tr><td><code>kp</code></td><td><code>ncs.HKeypathRef</code></td><td>The keypath of the action.</td></tr><tr><td><code>input</code></td><td><code>ncs.maagic.Node</code></td><td>An object containing the parameters of the input section of the action yang model.</td></tr><tr><td><code>output</code></td><td><code>ncs.maagic.Node</code></td><td>The object where to put the output parameters as defined in the output section of the action yang model.</td></tr></tbody></table>

### Service Handler

The Python high-level API provides an easy way to implement a service handler for your modeled services. The easiest way to create a handler is to use the `ncs-make-package` command. It creates some skeleton code.

```bash
$ cd packages
$ ncs-make-package --service-skeleton python pyservice \
 --component-class service.Service
```

The generated package skeleton:

```bash
$ tree pyservice
pyservice/
+-- README
+-- doc/
+-- load-dir/
+-- package-meta-data.xml
+-- python/
|   +-- pyservice/
|       +-- __init__.py
|       +-- service.py
+-- src/
|   +-- Makefile
|   +-- yang/
|       +-- service.yang
+-- templates/
```

This example has some code added for the service logic, including a service template.

When debugging Python packages, refer to [Debugging of Python Packages](/guides/development/core-concepts/nso-virtual-machines/nso-python-vm#debugging-of-python-packages).

Add some service logic to the `cb_create`:

{% code title="Example: High-level Python Service Implementation" %}

```bash
# -*- mode: python; python-indent: 4 -*-

from ncs.application import Application
from ncs.application import Service
import ncs.template

# ------------------------
# SERVICE CALLBACK EXAMPLE
# ------------------------
class ServiceCallbacks(Service):
    @Service.create
    def cb_create(self, tctx, root, service, proplist):
        self.log.info('Service create(service=', service._path, ')')

        # Add this service logic >>>>>>>
        vars = ncs.template.Variables()
        vars.add('MAGIC', '42')
        vars.add('CE', service.device)
        vars.add('INTERFACE', service.unit)
        template = ncs.template.Template(service)
        template.apply('pyservice-template', vars)

        self.log.info('Template is applied')

        dev = root.devices.device[service.device]
        dev.description = "This device was modified by %s" % service._path
        # <<<<<<<<< service logic

    @Service.pre_modification
    def cb_pre_modification(self, tctx, op, kp, root, proplist):
        self.log.info('Service premod(service=', kp, ')')

    @Service.post_modification
    def cb_post_modification(self, tctx, op, kp, root, proplist):
        self.log.info('Service premod(service=', kp, ')')


# ---------------------------------------------
# COMPONENT THREAD THAT WILL BE STARTED BY NCS.
# ---------------------------------------------
class Service(Application):
    def setup(self):
        self.log.info('Worker RUNNING')
        self.register_service('service-servicepoint', ServiceCallbacks)

    def teardown(self):
        self.log.info('Worker FINISHED')
```

{% endcode %}

Add a template to `packages/pyservice/templates/service.template.xml`:

```xml
<config-template xmlns="http://tail-f.com/ns/config/1.0">
  <devices xmlns="http://tail-f.com/ns/ncs">
    <device tags="nocreate">
      <name>{$CE}</name>
      <config tags="merge">
      <interface xmlns="urn:ios">
        <FastEthernet>
          <name>0/{$INTERFACE}</name>
          <description>The maagic: {$MAGIC}</description>
        </FastEthernet>
      </interface>
      </config>
    </device>
  </devices>
</config-template>
```

The table below lists the service handler callback parameters:

<table><thead><tr><th width="160">Parameter</th><th width="272">Type</th><th>Description</th></tr></thead><tbody><tr><td><code>self</code></td><td><code>ncs.application.Service</code></td><td>The service object.</td></tr><tr><td><code>tctx</code></td><td><code>ncs.TransCtxRef</code></td><td>Transaction context.</td></tr><tr><td><code>root</code></td><td><code>ncs.maagic.Node</code></td><td>An object pointing to the root with the current transaction context, using shared operations (<code>create</code>, <code>set_elem</code>, ...) for configuration modifications.</td></tr><tr><td><code>service</code></td><td><code>ncs.maagic.Node</code></td><td>An object pointing to the service with the current transaction context, using shared operations (<code>create</code>, <code>set_elem</code>, ...) for configuration modifications.</td></tr><tr><td><code>proplist</code></td><td>list(tuple(str, str))</td><td>The opaque object for the service configuration used to store hidden state information between invocations. It is updated by returning a modified list.</td></tr></tbody></table>

### Validation Point Handler

The Python high-level API provides an easy way to implement a validation point handler. The easiest way to create a handler is to use the `ncs-make-package` command. It creates ready-to-use skeleton code.

```bash
$ cd packages
$ ncs-make-package --service-skeleton python pyvalidation --component-class
 validation.ValidationApplication \
 --disable-service-example --validation-example
```

The generated package skeleton:

```bash
$ tree pyaction
pyaction/
+-- README
+-- doc/
+-- load-dir/
+-- package-meta-data.xml
+-- python/
|   +-- pyaction/
|       +-- __init__.py
|       +-- validation.py
+-- src/
|   +-- Makefile
|   +-- yang/
|       +-- validation.yang
+-- templates/
```

This example validation point handler accepts all values except `invalid`.

When debugging Python packages refer to [Debugging of Python Packages](/guides/development/core-concepts/nso-virtual-machines/nso-python-vm#debugging-of-python-packages).

{% code title="Example: Validation Implementation" %}

```bash
# -*- mode: python; python-indent: 4 -*-
import ncs
from ncs.dp import ValidationError, ValidationPoint


# ---------------
# VALIDATION EXAMPLE
# ---------------
class Validation(ValidationPoint):
    @ValidationPoint.validate
    def cb_validate(self, tctx, keypath, value, validationpoint):
        self.log.info('validate: ', str(keypath), '=', str(value))
        if value == 'invalid':
            raise ValidationError('invalid value')
        return ncs.CONFD_OK


# ---------------------------------------------
# COMPONENT THREAD THAT WILL BE STARTED BY NCS.
# ---------------------------------------------
class ValidationApplication(ncs.application.Application):
    def setup(self):
        # The application class sets up logging for us. It is accessible
        # through 'self.log' and is a ncs.log.Log instance.
        self.log.info('ValidationApplication RUNNING')

        # When using actions, this is how we register them:
        #
        self.register_validation('pyvalidation-valpoint', Validation)

        # If we registered any callback(s) above, the Application class
        # took care of creating a daemon (related to the service/action point).

        # When this setup method is finished, all registrations are
        # considered done and the application is 'started'.

    def teardown(self):
        # When the application is finished (which would happen if NCS went
        # down, packages were reloaded or some error occurred) this teardown
        # method will be called.

        self.log.info('ValidationApplication FINISHED')
```

{% endcode %}

Test the validation by setting the value to invalid and validating the transaction from the NSO CLI:

```cli
admin@ncs% set validation validate-value invalid
admin@ncs% validate
Failed: 'validation validate-value': invalid value
[ok][2016-04-22 10:30:39]
```

The table below lists the validation point handler callback parameters:

<table><thead><tr><th width="203">Parameter</th><th width="256">Type</th><th>Description</th></tr></thead><tbody><tr><td><code>self</code></td><td><code>ncs.dp.ValidationPoint</code></td><td>The validation point object.</td></tr><tr><td><code>tctx</code></td><td><code>ncs.TransCtxRef</code></td><td>Transaction context.</td></tr><tr><td><code>kp</code></td><td><code>ncs.HKeypathRef</code></td><td>The keypath of the node being validated.</td></tr><tr><td><code>value</code></td><td><code>ncs.Value</code></td><td>Current value of the node being validated.</td></tr><tr><td><code>validationpoint</code></td><td><code>string</code></td><td>The validation point that triggered the validation.</td></tr></tbody></table>

## Low-level APIs

The Python low-level APIs are a direct mapping of the C-APIs. A C call has a corresponding Python function entry. From a programmer's point of view, it wraps the C data structures into Python objects and handles the related memory management when requested by the Python garbage collector. Any errors are reported as `error.Error`.

The low-level APIs will not be described in detail in this document, but you will find a few examples showing their usage in the coming sections.

See `pydoc3 _ncs` and `man confd_lib_lib` for further information.

### Low-level MAAPI API <a href="#d5e4852" id="d5e4852"></a>

This API is a direct mapping of the NSO MAAPI C API. See `pydoc3 _ncs.maapi` and `man confd_lib_maapi` for further information.

Note that additional care must be taken when using this API in service code, as it also exposes functions that do not perform reference counting (see [Reference Counting Overlapping Configuration](https://nso-docs.cisco.com/guides/development/core-concepts/api-overview/pages/Bw8TviXCSEsM9XjXBC1d#ch_svcref.refcount)).

In the service code, you should use the `shared_*` set of functions, such as:

```
shared_apply_template
shared_copy_tree
shared_create
shared_insert
shared_set_elem
shared_set_elem2
shared_set_values
```

And, avoid the non-shared variants:

```
load_config()
load_config_cmds()
load_config_stream()
apply_template()
copy_tree()
create()
insert()
set_elem()
set_elem2()
set_object
set_values()
```

The following example is a script to read and de-crypt a password using the Python low-level MAAPI API.

<pre data-title="Example: Setting of Configuration Data using MAAPI"><code>import socket
import _ncs
from _ncs import maapi

sock_maapi = socket.socket(family=socket.AF_UNIX)

maapi.connect(sock_maapi,
              path=_ncs.PATH)

maapi.load_schemas(sock_maapi)

maapi.start_user_session(
                  sock_maapi,
                  'admin',
                  'python',
                  [],
                  '127.0.0.1',
                  _ncs.PROTO_TCP)

maapi.install_crypto_keys(sock_maapi)


th = maapi.start_trans(sock_maapi, _ncs.RUNNING, _ncs.READ)

<strong>path = "/devices/authgroups/group{default}/umap{admin}/remote-password"
</strong>encrypted_password = maapi.get_elem(sock_maapi, th, path)

decrypted_password = _ncs.decrypt(str(encrypted_password))

maapi.finish_trans(sock_maapi, th)
maapi.end_user_session(sock_maapi)
sock_maapi.close()

print("Default authgroup admin password = %s" % decrypted_password)
</code></pre>

This example is a script to do a `check-sync` action request using the low-level MAAPI API.

{% code title="Example: Action Request" %}

```python
import socket
import _ncs
from _ncs import maapi

sock_maapi = socket.socket(family=socket.AF_UNIX)

maapi.connect(sock_maapi,
              path=_ncs.PATH)

maapi.load_schemas(sock_maapi)

_ncs.maapi.start_user_session(
                  sock_maapi,
                  'admin',
                  'python',
                  [],
                  '127.0.0.1',
                  _ncs.PROTO_TCP)

ns_hash = _ncs.str2hash("http://tail-f.com/ns/ncs")

results = maapi.request_action(sock_maapi, [], ns_hash, "/devices/check-sync")
for result in results:
    v = result.v
    t = v.confd_type()
    if t == _ncs.C_XMLBEGIN:
        print("sync-result {")
    elif t == _ncs.C_XMLEND:
        print("}")
    elif t == _ncs.C_BUF:
        tag = result.tag
        print("    %s %s" % (_ncs.hash2str(tag), str(v)))
    elif t == _ncs.C_ENUM_HASH:
        tag = result.tag
        text = v.val2str((ns_hash, '/devices/check-sync/sync-result/result'))
        print("    %s %s" % (_ncs.hash2str(tag), text))

maapi.end_user_session(sock_maapi)
sock_maapi.close()
```

{% endcode %}

### Low-level CDB API

This API is a direct mapping of the NSO CDB C API. See `pydoc3 _ncs.cdb` and `man confd_lib_cdb` for further information.

Setting of operational data has historically been done using one of the CDB APIs (Python, Java, C). This example shows how to set a value and trigger subscribers for operational data using the Python low-level API. API.

{% code title="Example: Setting of Operational Data using CDB API" %}

```python
import socket
import _ncs
from _ncs import cdb

sock_cdb = socket.socket(family=socket.AF_UNIX)

cdb.connect(
    sock_cdb,
    type=cdb.DATA_SOCKET,
    path=_ncs.PATH)

cdb.start_session2(sock_cdb, cdb.OPERATIONAL, cdb.LOCK_WAIT | cdb.LOCK_REQUEST)

path = "/operdata/value"
cdb.set_elem(sock_cdb, _ncs.Value(42, _ncs.C_UINT32), path)

new_value = cdb.get(sock_cdb, path)

cdb.end_session(sock_cdb)
sock_cdb.close()

print("/operdata/value is now %s" % new_value)
```

{% endcode %}

### Low-level Event Notification API

The Python `_ncs.events` low-level module provides an API for subscribing to and processing NSO event notifications. Typically, the event notification API is used by applications that manage NSO using the SDK API using, for example, MAAPI or for debug purposes. In addition to subscribing to the various events, streams available over other northbound interfaces, such as NETCONF, RESTCONF, etc., can be subscribed to as well.

See [`examples.ncs/sdk-api/event-notifications`](https://github.com/NSO-developer/nso-examples/tree/6.7/sdk-api/event-notifications) for an example. The [`examples.ncs/common/event_notifications.py`](https://github.com/NSO-developer/nso-examples/tree/6.7/common/event_notifications.py) Python script used by the example can also be used as a standalone application to, for example, debug any NSO instance.

## Error Codes

The low-level APIs raise `_ncs.error.Error`. The high-level APIs raise `ncs.error.Error`, which uses the same underlying error information. The exception object contains:

* `confd_errno`: the numeric error code.
* `confd_strerror`: the library error message for the code.
* `confd_lasterr`: optional additional error text from NSO or application code.

In an exception string such as `external error (19): application communication failure`, `external error` is the `confd_strerror` text, `19` is `confd_errno`, and the text after the colon is `confd_lasterr`.

Use the Python constants from `ncs` or `_ncs` when checking `confd_errno`. Most C `CONFD_ERR_*` constants are exposed as Python `ERR_*` constants. The NSO-specific `NCS_ERR_*` constants keep their names. For codes without a dedicated `confd_strerror()` entry, the base error message is `Unknown error`; the detailed cause is typically carried in `confd_lasterr`.

| Code | Python constant                     | C constant                            | Error message                                               | Description                                                                                                                    |
| ---- | ----------------------------------- | ------------------------------------- | ----------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------ |
| 1    | `ncs.ERR_NOEXISTS`                  | `CONFD_ERR_NOEXISTS`                  | `item does not exist`                                       | A requested value or object does not exist, typically when reading through CDB or MAAPI.                                       |
| 2    | `ncs.ERR_ALREADY_EXISTS`            | `CONFD_ERR_ALREADY_EXISTS`            | `item already exists`                                       | An attempt was made to create an object that already exists.                                                                   |
| 3    | `ncs.ERR_ACCESS_DENIED`             | `CONFD_ERR_ACCESS_DENIED`             | `access denied`                                             | AAA authorization rules denied access to the object.                                                                           |
| 4    | `ncs.ERR_NOT_WRITABLE`              | `CONFD_ERR_NOT_WRITABLE`              | `item is not writable`                                      | An attempt was made to write an object that is not writable.                                                                   |
| 5    | `ncs.ERR_BADTYPE`                   | `CONFD_ERR_BADTYPE`                   | `item has a bad/wrong type`                                 | An object was created or written with a value whose type does not match the YANG type.                                         |
| 6    | `ncs.ERR_NOTCREATABLE`              | `CONFD_ERR_NOTCREATABLE`              | `item is not creatable`                                     | An attempt was made to create an object that cannot be created.                                                                |
| 7    | `ncs.ERR_NOTDELETABLE`              | `CONFD_ERR_NOTDELETABLE`              | `item is not deletable`                                     | An attempt was made to delete an object that cannot be deleted.                                                                |
| 8    | `ncs.ERR_BADPATH`                   | `CONFD_ERR_BADPATH`                   | `badly formatted or nonexistent path`                       | A path argument was malformed, referred to a nonexistent node, or was invalid in a printf-style path function.                 |
| 9    | `ncs.ERR_NOSTACK`                   | `CONFD_ERR_NOSTACK`                   | `cannot pop an empty stack`                                 | A path stack pop was attempted without a preceding push.                                                                       |
| 10   | `ncs.ERR_LOCKED`                    | `CONFD_ERR_LOCKED`                    | `locked`                                                    | An attempt was made to lock an object that is already locked.                                                                  |
| 11   | `ncs.ERR_INUSE`                     | `CONFD_ERR_INUSE`                     | `in use`                                                    | A commit or operation could not proceed because another user or process holds a lock.                                          |
| 12   | `ncs.ERR_NOTSET`                    | `CONFD_ERR_NOTSET`                    | `notset`                                                    | A mandatory leaf has no value, either because it was deleted or never set after create.                                        |
| 13   | `ncs.ERR_NON_UNIQUE`                | `CONFD_ERR_NON_UNIQUE`                | `non unique`                                                | A set of leaves covered by a YANG `unique` statement is not unique.                                                            |
| 14   | `ncs.ERR_BAD_KEYREF`                | `CONFD_ERR_BAD_KEYREF`                | `bad keyref`                                                | A key reference points to a target that does not exist.                                                                        |
| 15   | `ncs.ERR_TOO_FEW_ELEMS`             | `CONFD_ERR_TOO_FEW_ELEMS`             | `too few elems`                                             | A `min-elements` constraint was violated; the node has too few elements or entries.                                            |
| 16   | `ncs.ERR_TOO_MANY_ELEMS`            | `CONFD_ERR_TOO_MANY_ELEMS`            | `too many elems`                                            | A `max-elements` constraint was violated; the node has too many elements or entries.                                           |
| 17   | `ncs.ERR_BADSTATE`                  | `CONFD_ERR_BADSTATE`                  | `operation in wrong state`                                  | A function was called out of order, for example in a MAAPI commit sequence.                                                    |
| 18   | `ncs.ERR_INTERNAL`                  | `CONFD_ERR_INTERNAL`                  | `internal error`                                            | An internal NSO or libconfd error occurred. This normally indicates a product bug or missing specific error code.              |
| 19   | `ncs.ERR_EXTERNAL`                  | `CONFD_ERR_EXTERNAL`                  | `external error`                                            | An error originated in user code, such as a service, action, validation, or data provider callback.                            |
| 20   | `ncs.ERR_MALLOC`                    | `CONFD_ERR_MALLOC`                    | `Failed to allocate`                                        | Memory allocation failed.                                                                                                      |
| 21   | `ncs.ERR_PROTOUSAGE`                | `CONFD_ERR_PROTOUSAGE`                | `Bad protocol usage or unexpected retval`                   | An API function or callback was used incorrectly, or a callback returned an unexpected value.                                  |
| 22   | `ncs.ERR_NOSESSION`                 | `CONFD_ERR_NOSESSION`                 | `A session must be established prior to this command`       | The function requires an established user session before it is called.                                                         |
| 23   | `ncs.ERR_TOOMANYTRANS`              | `CONFD_ERR_TOOMANYTRANS`              | `Too many transactions`                                     | A new MAAPI transaction was rejected because the transaction limit was reached.                                                |
| 24   | `ncs.ERR_OS`                        | `CONFD_ERR_OS`                        | `system call failed`                                        | An operating system call failed. Inspect the OS error string or errno for the underlying cause.                                |
| 25   | `ncs.ERR_HA_CONNECT`                | `CONFD_ERR_HA_CONNECT`                | `Failed to connect to remote HA node`                       | Connecting to a remote HA node failed.                                                                                         |
| 26   | `ncs.ERR_HA_CLOSED`                 | `CONFD_ERR_HA_CLOSED`                 | `Remote HA node closed socket to us`                        | A remote HA node closed the connection, or a sync response from the primary timed out.                                         |
| 27   | `ncs.ERR_HA_BADFXS`                 | `CONFD_ERR_HA_BADFXS`                 | `Remote HA node has incompatible config to us`              | The remote HA node has a different set or version of FXS files.                                                                |
| 28   | `ncs.ERR_HA_BADTOKEN`               | `CONFD_ERR_HA_BADTOKEN`               | `Remote HA node has bad credentials`                        | The remote HA node presented a different HA token.                                                                             |
| 29   | `ncs.ERR_HA_BADNAME`                | `CONFD_ERR_HA_BADNAME`                | `remote HA node has other/wrong name`                       | The remote HA node name differs from the expected name.                                                                        |
| 30   | `ncs.ERR_HA_BIND`                   | `CONFD_ERR_HA_BIND`                   | `failed to bind HA socket`                                  | Binding the HA socket for incoming HA connections failed.                                                                      |
| 31   | `ncs.ERR_HA_NOTICK`                 | `CONFD_ERR_HA_NOTICK`                 | `Remote HA node doesn't send us any ticks`                  | A remote HA node did not produce the expected live ticks.                                                                      |
| 32   | `ncs.ERR_VALIDATION_WARNING`        | `CONFD_ERR_VALIDATION_WARNING`        | `Validation warnings`                                       | `maapi_validate()` returned validation warnings.                                                                               |
| 33   | `ncs.ERR_SUBAGENT_DOWN`             | `CONFD_ERR_SUBAGENT_DOWN`             | `Subagent down`                                             | An operation towards a mounted NETCONF subagent failed because the subagent is not up.                                         |
| 34   | `ncs.ERR_LIB_NOT_INITIALIZED`       | `CONFD_ERR_LIB_NOT_INITIALIZED`       | `Library not initialized`                                   | The libconfd library was not initialized properly before use.                                                                  |
| 35   | `ncs.ERR_TOO_MANY_SESSIONS`         | `CONFD_ERR_TOO_MANY_SESSIONS`         | `Maximum number of sessions reached`                        | The maximum number of sessions has been reached.                                                                               |
| 36   | `ncs.ERR_BAD_CONFIG`                | `CONFD_ERR_BAD_CONFIG`                | `Error in a configuration`                                  | A configuration contains an error.                                                                                             |
| 37   | `ncs.ERR_RESOURCE_DENIED`           | `CONFD_ERR_RESOURCE_DENIED`           | `Data provider returned CONFD_ERRCODE_RESOURCE_DENIED`      | A data provider callback returned `CONFD_ERRCODE_RESOURCE_DENIED`.                                                             |
| 38   | `ncs.ERR_INCONSISTENT_VALUE`        | `CONFD_ERR_INCONSISTENT_VALUE`        | `Data provider returned CONFD_ERRCODE_INCONSISTENT_VALUE`   | A data provider callback returned `CONFD_ERRCODE_INCONSISTENT_VALUE`.                                                          |
| 39   | `ncs.ERR_APPLICATION_INTERNAL`      | `CONFD_ERR_APPLICATION_INTERNAL`      | `Data provider returned CONFD_ERRCODE_APPLICATION_INTERNAL` | A data provider callback returned `CONFD_ERRCODE_APPLICATION_INTERNAL`.                                                        |
| 40   | `ncs.ERR_UNSET_CHOICE`              | `CONFD_ERR_UNSET_CHOICE`              | `No case selected for mandatory choice`                     | No `case` has been selected for a mandatory YANG `choice`.                                                                     |
| 41   | `ncs.ERR_MUST_FAILED`               | `CONFD_ERR_MUST_FAILED`               | `Unsatisfied must constraint`                               | A YANG `must` constraint is not satisfied.                                                                                     |
| 42   | `ncs.ERR_MISSING_INSTANCE`          | `CONFD_ERR_MISSING_INSTANCE`          | `Required instance does not exist`                          | An `instance-identifier` leaf with `require-instance true` points to a missing instance.                                       |
| 43   | `ncs.ERR_INVALID_INSTANCE`          | `CONFD_ERR_INVALID_INSTANCE`          | `Instance does not conform to path filters`                 | An `instance-identifier` leaf does not conform to its specified path filters.                                                  |
| 44   | `ncs.ERR_UNAVAILABLE`               | `CONFD_ERR_UNAVAILABLE`               | `Unavailable functionality`                                 | The requested functionality is unavailable, for example getting or setting attributes on an operational data element.          |
| 45   | `ncs.ERR_EOF`                       | `CONFD_ERR_EOF`                       | `Lost connection to ConfD`                                  | Used when an API function returns `CONFD_EOF`; the connection to NSO was closed or no normal success value was returned.       |
| 46   | `ncs.ERR_NOTMOVABLE`                | `CONFD_ERR_NOTMOVABLE`                | `item is not movable`                                       | An attempt was made to move an object that cannot be moved.                                                                    |
| 47   | `ncs.ERR_HA_WITH_UPGRADE`           | `CONFD_ERR_HA_WITH_UPGRADE`           | `HA secondary not allowed during upgrade`                   | An in-service data model upgrade was attempted on an HA node, or HA primary/secondary mode was changed during such an upgrade. |
| 48   | `ncs.ERR_TIMEOUT`                   | `CONFD_ERR_TIMEOUT`                   | `Operation timed out`                                       | The operation did not complete within the specified timeout.                                                                   |
| 49   | `ncs.ERR_ABORTED`                   | `CONFD_ERR_ABORTED`                   | `Operation was aborted`                                     | The operation was aborted.                                                                                                     |
| 50   | `ncs.ERR_XPATH`                     | `CONFD_ERR_XPATH`                     | `XPath compilation/evaluation failed`                       | Compilation or evaluation of an XPath expression failed.                                                                       |
| 51   | `ncs.ERR_NOT_IMPLEMENTED`           | `CONFD_ERR_NOT_IMPLEMENTED`           | `Operation not implemented`                                 | The requested operation is not implemented, often because the library is newer than the NSO daemon.                            |
| 52   | `ncs.ERR_HA_BADVSN`                 | `CONFD_ERR_HA_BADVSN`                 | `Remote HA node has incompatible version`                   | The remote HA node uses an incompatible protocol version.                                                                      |
| 53   | `ncs.ERR_POLICY_FAILED`             | `CONFD_ERR_POLICY_FAILED`             | `Policy expression evaluated to false`                      | A user-defined policy expression evaluated to false.                                                                           |
| 54   | `ncs.ERR_POLICY_COMPILATION_FAILED` | `CONFD_ERR_POLICY_COMPILATION_FAILED` | `Policy XPath expression could not be compiled`             | A user-defined policy XPath expression could not be compiled.                                                                  |
| 55   | `ncs.ERR_POLICY_EVALUATION_FAILED`  | `CONFD_ERR_POLICY_EVALUATION_FAILED`  | `Policy expression failed XPath evaluation`                 | A user-defined policy expression failed XPath evaluation.                                                                      |
| 56   | `ncs.NCS_ERR_CONNECTION_REFUSED`    | `NCS_ERR_CONNECTION_REFUSED`          | `Unknown error`                                             | NSO failed to connect to a device.                                                                                             |
| 57   | `ncs.ERR_START_FAILED`              | `CONFD_ERR_START_FAILED`              | `Failed to proceed to next start phase`                     | The NSO daemon failed to proceed to the next start phase.                                                                      |
| 58   | `ncs.ERR_DATA_MISSING`              | `CONFD_ERR_DATA_MISSING`              | `Data provider returned CONFD_ERRCODE_DATA_MISSING`         | A data provider callback returned `CONFD_ERRCODE_DATA_MISSING`.                                                                |
| 59   | `ncs.ERR_CLI_CMD`                   | `CONFD_ERR_CLI_CMD`                   | `CLI command error`                                         | Execution of a CLI command failed.                                                                                             |
| 60   | `ncs.ERR_UPGRADE_IN_PROGRESS`       | `CONFD_ERR_UPGRADE_IN_PROGRESS`       | `Not allowed during upgrade`                                | The requested operation is not allowed while an in-service data model upgrade is in progress.                                  |
| 61   | `ncs.ERR_NOTRANS`                   | `CONFD_ERR_NOTRANS`                   | `Transaction not found`                                     | An invalid transaction handle was passed to a MAAPI function.                                                                  |
| 62   | `ncs.NCS_ERR_SERVICE_CONFLICT`      | `NCS_ERR_SERVICE_CONFLICT`            | `Unknown error`                                             | A service invocation outside the transaction lock modified data that was also modified by another service invocation.          |
| 63   | `ncs.NCS_ERR_CONNECTION_TIMEOUT`    | `NCS_ERR_CONNECTION_TIMEOUT`          | `Unknown error`                                             | NSO reported a connection timeout for a device operation.                                                                      |
| 64   | `ncs.NCS_ERR_CONNECTION_CLOSED`     | `NCS_ERR_CONNECTION_CLOSED`           | `Unknown error`                                             | NSO reported that the connection for a device operation was closed.                                                            |
| 65   | `ncs.NCS_ERR_DEVICE`                | `NCS_ERR_DEVICE`                      | `Unknown error`                                             | NSO reported a device-specific error during a device operation.                                                                |
| 66   | `ncs.NCS_ERR_TEMPLATE`              | `NCS_ERR_TEMPLATE`                    | `Unknown error`                                             | NSO reported a template processing error.                                                                                      |
| 67   | `ncs.ERR_NO_MOUNT_ID`               | `CONFD_ERR_NO_MOUNT_ID`               | `Path is ambiguous due to traversing a mount point`         | A path is ambiguous because it traverses a schema mount point without a mount id.                                              |
| 68   | `ncs.ERR_STALE_INSTANCE`            | `CONFD_ERR_STALE_INSTANCE`            | `Required instance does not exist`                          | An `instance-identifier` leaf with `require-instance true` contains stale data after upgrade.                                  |
| 69   | `ncs.ERR_HA_BADCONFIG`              | `CONFD_ERR_HA_BADCONFIG`              | `Remote HA node has bad configuration`                      | A remote HA node has bad HA application configuration that prevents it from functioning properly.                              |
| 70   | `ncs.ERR_TRANSACTION_CONFLICT`      | `CONFD_ERR_TRANSACTION_CONFLICT`      | `Conflict detected`                                         | A transaction conflict was detected, such as data read by one transaction being modified by another before apply.              |
| 71   | `ncs.ERR_HA_ABORT`                  | `CONFD_ERR_HA_ABORT`                  | `Transaction aborted`                                       | A transaction was aborted during an HA transition.                                                                             |
| 72   | `ncs.ERR_BAD_PAYLOAD`               | `CONFD_ERR_BAD_PAYLOAD`               | `Error in payload data`                                     | Payload data was invalid or could not be processed.                                                                            |
| 73   | `ncs.ERR_CHECKSUM_MISMATCH`         | `CONFD_ERR_CHECKSUM_MISMATCH`         | `Checksum mismatch`                                         | A dry-run changeset did not match the commit changeset.                                                                        |

## Advanced Topics

### Schema Loading - Internals <a href="#ncs.development.python_api_overview.advanced.schema_loading" id="ncs.development.python_api_overview.advanced.schema_loading"></a>

When schemas are loaded, either upon direct request or automatically by methods and classes in the `maapi` module, they are statically cached inside the Python VM. This fact presents a problem if one wants to connect to several different NSO nodes with diverging schemas from the same Python VM.

Take for example the following program that connects to two different NSO nodes (with diverging schemas) and shows their ned-id's.

{% code title="Example: Reading NED-IDs (read\_nedids.py)" %}

```python
import ncs


def print_ned_ids(path):
    with ncs.maapi.single_read_trans('admin', 'system', db=ncs.OPERATIONAL, path=path) as t:
        dev_ned_id = ncs.maagic.get_node(t, '/devices/ned-ids/ned-id')
        for id in dev_ned_id.keys():
            print(id)


if __name__ == '__main__':
    print('=== lsa-1 ===')
    print_ned_ids('/tmp/nso-lsa-1/nso-ipc')
    print('=== lsa-2 ===')
    print_ned_ids('/tmp/nso-lsa-2/nso-ipc')
```

{% endcode %}

Running this program may produce output like this:

```bash
            $ python3 read_nedids.py
            === lsa-1 ===
            {ned:lsa-netconf}
            {ned:netconf}
            {ned:snmp}
            {cisco-nso-nc-5.5:cisco-nso-nc-5.5}
            === lsa-2 ===
            {ned:lsa-netconf}
            {ned:netconf}
            {ned:snmp}
            {"[<_ncs.Value type=C_IDENTITYREF(44) value='idref<211668964'...>]"}
            {"[<_ncs.Value type=C_IDENTITYREF(44) value='idref<151824215'>]"}
            {"[<_ncs.Value type=C_IDENTITYREF(44) value='idref<208856485'...>]"}
```

The output shows identities in string format for the active NEDs on the different nodes. Note that for `lsa-2`, the last three lines do not show the name of the identity but instead the representation of a `_ncs.Value`. The reason for this is that `lsa-2` has different schemas which do not include these identities. Schemas for this Python VM were loaded and cached during the first call to `ncs.maapi.single_read_trans()` so no schema loading occurred during the second call.

The way to make the program above work as expected is to force the reloading of schemas by passing an optional argument to `single_read_trans()` like so:

```python
with ncs.maapi.single_read_trans('admin', 'system', db=ncs.OPERATIONAL, port=port,
    load_schemas=ncs.maapi.LOAD_SCHEMAS_RELOAD) as t:
```

Running the program with this change may produce something like this:

```
          === lsa-1 ===
          {ned:lsa-netconf}
          {ned:netconf}
          {ned:snmp}
          {cisco-nso-nc-5.5:cisco-nso-nc-5.5}
          === lsa-2 ===
          {ned:lsa-netconf}
          {ned:netconf}
          {ned:snmp}
          {cisco-asa-cli-6.13:cisco-asa-cli-6.13}
          {cisco-ios-cli-6.72:cisco-ios-cli-6.72}
          {router-nc-1.0:router-nc-1.0}
```

Now, this was just an example of what may happen when wrong schemas are loaded. Implications may be more severe though, especially if maagic nodes are kept between reloads. In such cases, accessing an "invalid" maagic object may in the best case result in undefined behavior making the program not work, but might even crash the program. So care needs to be taken to not reload schemas in a Python VM if there are dependencies to other parts in the same VM that need previous schemas.

Functions and methods that accept the `load_schemas` argument:

* `ncs.maapi.Maapi() constructor`
* `ncs.maapi.single_read_trans()`
* `ncs.maapi.single_write_trans()`

### The way of using `multiprocessing.Process`

When using multiprocessing in NSO, the default start method is now `spawn` instead of `fork`. With the `spawn` method, a new Python interpreter process is started, and all arguments passed to `multiprocessing.Process` must be picklable.

If you pass Python objects that reference low-level C structures (for example `_ncs.dp.DaemonCtxRef` or `_ncs.UserInfo`), Python will raise an error like:

```python
TypeError: cannot pickle '<object>' object
```

{% code title="Example: using multiprocessing.Process" %}

```python
import ncs
import _ncs
from ncs.dp import Action
from multiprocessing import Process
import multiprocessing

def child(uinfo, self):
    print(f"uinfo: {uinfo}, self: {self}")

class DoAction(Action):
    @Action.action
    def cb_action(self, uinfo, name, kp, input, output, trans):
          t1 = multiprocessing.Process(target=child, args=(uinfo, self))
          t1.start()

class Main(ncs.application.Application):
    def setup(self):
        self.log.info('Main RUNNING')
        self.register_action('sleep', DoAction)

    def teardown(self):
        self.log.info('Main FINISHED')
```

{% endcode %}

This happens because `self` and `uinfo` contain low-level C references that cannot be serialized (pickled) and sent to the child process.

To fix this, avoid passing entire objects such as `self` or `uinfo` to the process. Instead, pass only simple or primitive data types (like strings, integers, or dictionaries) that can be pickled.

{% code title="Example: using multiprocessing.Process with primitive data" %}

```python
import ncs
import _ncs
from ncs.dp import Action
from multiprocessing import Process
import multiprocessing

def child(usid, th, action_point):
    print(f"uinfo: {usid}, th: {th}, action_point: {action_point}")

class DoAction(Action):
    @Action.action
    def cb_action(self, uinfo, name, kp, input, output, trans):
          usid = uinfo.usid
          th = uinfo.actx_thandle
          action_point = self.actionpoint
          t1 = multiprocessing.Process(target=child, args=(usid,th,action_point,))
          t1.start()

class Main(ncs.application.Application):
    def setup(self):
        self.log.info('Main RUNNING')
        self.register_action('sleep', DoAction)

    def teardown(self):
        self.log.info('Main FINISHED')
```

{% endcode %}

### Fetch bulk live-status via MAAPI

In MAAPI, when reading live-status data with `get_elem`, `get_case`, `get_values`, and similar calls, each read call may trigger its own request to the device. To improve fetch performance, you can use `set_read_intent` to fetch live-status data in bulk.

{% code title="Example: Fetch bulk live-status via MAAPI" %}

```python
import ncs

with ncs.maapi.single_read_trans('admin', 'python') as t:
    t.set_read_intent(["/ncs:devices/device/live-status/foo:foo",
                       "/ncs:devices/device/live-status/bar:bar"])

    t.get_elem("/ncs:devices/device{dev0}/live-status/foo:foo/a")
    t.get_elem("/ncs:devices/device{dev0}/live-status/foo:foo/b")

    t.get_elem("/ncs:devices/device{dev0}/live-status/bar:bar/baz")
```

{% endcode %}

With `set_read_intent`, it only requires one read request to get all the live-status data under `foo:foo` and `bar:bar`, and the data will be cached for later usage. To check the current read-intent, use `get_read_intent`, and `clear_read_intent` to clear it.


# Java API Overview

Learn about the NSO Java API and its usage.

The NSO Java library contains a variety of APIs for different purposes. In this section, we introduce these and explain their usage. The Java library deliverables are found as two jar files (`ncs.jar` and `conf-api.jar`). The jar files and their dependencies can be found under `$NCS_DIR/java/jar/`.

For convenience, the Java build tool Apache ant (<https://ant.apache.org/>) is used to run all of the examples. However, this tool is not a requirement for NSO.

For same-host applications, Java APIs should use Local IPC by default. TCP is still available for applications that need to connect to NSO remotely.

The following APIs are included in the library:

<table data-header-hidden data-full-width="false"><thead><tr><th width="345"></th><th></th></tr></thead><tbody><tr><td><strong>MAAPI (Management Agent API)</strong><br>Northbound interface that is transactional and user session-based. Using this interface both configuration and operational data can be read. Configuration data can be written and committed as one transaction. The API is complete in the way that it is possible to write a new northbound agent using only this interface. It is also possible to attach to ongoing transactions in order to read uncommitted changes and/or modify data in these transactions.</td><td><img src="/files/sCt1fr2cZH59OFoC8wRO" alt="" data-size="original"></td></tr><tr><td><strong>CDB API</strong><br>The southbound interface provides access to the CDB configuration database. Using this interface configuration data can be read. In addition, operational data that is stored in CDB can be read and written. This interface has a subscription mechanism to subscribe to changes. A subscription is specified on a path that points to an element in a YANG model or an instance in the instance tree. Any change under this point will trigger the subscription. CDB has also functions to iterate through the configuration changes when a subscription has been triggered.</td><td><img src="/files/qkTURcbkVJ8SOVhPb0iE" alt="" data-size="original"></td></tr><tr><td><strong>DP API</strong><br>Southbound interface that enables callbacks, hooks, and transforms. This API makes it possible to provide the service callbacks that handle service-to-device mapping logic. Other usual cases are external data providers for operational data or action callback implementations. There are also transaction and validation callbacks, etc. Hooks are callbacks that are fired when certain data is written and the hook is expected to do additional modifications of data. Transforms are callbacks that are used when complete mediation between two different models is necessary.</td><td><img src="/files/4WbyLbznPMPqze3NUzGZ" alt="" data-size="original"></td></tr><tr><td><strong>NED API (Network Element Driver)</strong><br>Southbound interface that mediates communication for devices that do not speak either NETCONF or SNMP. All prepackaged NEDs for different devices are written using this interface. It is possible to use the same interface to write your own NED. There are two types of NEDs, CLI NEDs and Generic NEDs. CLI NEDs can be used for devices that can be controlled by a Cisco-style CLI syntax, in this case the NED is developed primarily by building a YANG model and a relatively small part in Java. In other cases the Generic NED can be used for any type of communication protocol.</td><td><img src="/files/p0cOEPRE84owAeoZH4kB" alt="" data-size="original"></td></tr><tr><td><strong>NAVU API (Navigation Utilities)</strong><br>API that resides on top of the MAAPI and CDB APIs. It provides schema model navigation and instance data handling (read/write). Uses either a MAAPI or CDB context as data access and incorporates a subset of functionality from these (navigational and data read/write calls). Its major use is in service implementations which normally is about navigating device models and setting device data.</td><td><img src="/files/YDpFTSkuWEg9PcyFzkd0" alt="" data-size="original"></td></tr><tr><td><strong>ALARM API</strong><br>Eastbound API that is used both to consume and produce alarms in alignment with the NSO Alarm model. To consume alarms the AlarmSource interface is used. To produce a new alarm the AlarmSink interface is used. There is also a possibility to buffer produced alarms and make asynchronous writes to CDB to improve alarm performance.</td><td><img src="/files/xorWBz0ELnHzPYpJ9wL4" alt="" data-size="original"></td></tr><tr><td><strong>NOTIF API (Notification API)</strong><br>Northbound API that is used to subscribe to system events from NSO. These events are generated for audit log events, for different transaction states, for HA state changes, upgrade events, user sessions, etc.</td><td><img src="/files/OqVnfn0OnBO3vrpQ3WVv" alt="" data-size="original"></td></tr><tr><td><strong>HA API (High Availability)</strong><br>Northbound api used to manage a High Availability cluster of NSO instances. An NSO instance can be in one of three states NONE, PRIMARY or SECONDARY. With the HA API the state can be queried and changed for NSO instances in the cluster.</td><td><img src="/files/5oWPoETR1gSIVIAHcmBJ" alt="" data-size="original"></td></tr></tbody></table>

In addition, the Conf API framework contains utility classes for data types, keypaths, etc.

## MAAPI <a href="#ug.ncs.api.maapi" id="ug.ncs.api.maapi"></a>

The Management Agent API (MAAPI) provides an interface to the Transaction engine in NSO. As such it is very versatile. Here are some examples of how the MAAPI interface can be used.

* Read and write configuration data stored by NSO or in an external database.
* Write our own northbound interface.
* We could access data inside a not yet committed transaction, e.g. as validation logic where our Java code can attach itself to a running transaction and read through the not yet committed transaction, and validate the proposed configuration change.
* During database upgrade we can access and write data to a special upgrade transaction.

The first step of a typical sequence of MAAPI API calls when writing a management application would be to create a user session. Creating a user session is the equivalent of establishing an SSH connection from a NETCONF manager. It is up to the MAAPI application to authenticate users. When TCP is used, the connection between MAAPI and NSO is neither encrypted, nor authenticated. The Maapi Java package does however include an `authenticate()` method that can be used by the application to hook into the AAA framework of NSO and let NSO authenticate the user.

{% code title="Example: Establish a MAAPI Connection" %}

```
    Maapi maapi = new Maapi(UnixDomainSocketAddress.of(Conf.NCS_PATH));
```

{% endcode %}

When a Maapi socket has been created the next step is to create a user session and supply the relevant information about the user for authentication.

{% code title="Example: Starting a User Session" %}

```
    maapi.startUserSession("admin", "maapi", new String[] {"admin"});
```

{% endcode %}

When the user has been authenticated and a user session has been created the Maapi reference is now ready to establish a new transaction toward a data store. The following code snippet starts a read/write transaction towards the running data store.

{% code title="Example: Start a Read/Write transaction Towards Running" %}

```
    int th = maapi.startTrans(Conf.DB_RUNNING,
                              Conf.MODE_READ_WRITE);
```

{% endcode %}

\\

The `startTrans(int db,int mode)` method of the Maapi class returns an integer that represents a transaction handler. This transaction handler is used when invoking the various Maapi methods.

An example of a typical transactional method is the `getElem()` method:

{% code title="Example: Maapi.getElem()" %}

```java
    public ConfValue getElem(int tid,
                             String fmt,
                             Object... arguments)
```

{% endcode %}

The `getElem(int th, String fmt, Object ... arguments)` first parameter is the transaction handle which is the integer that was returned by the `startTrans()` method. The *`fmt`* is a path that leads to a leaf in the data model. The path is expressed as a format string that contain fixed text with zero to many embedded format specifiers. For each specifier, one argument in the variable argument list is expected.

The currently supported format specifiers in the Java API are:

* `%d` - requiring an integer parameter (type int) to be substituted.
* `%s` - requiring a `java.lang.String` parameter to be substituted.
* `%x` - requiring subclasses of type `com.tailf.conf.ConfValue` to be substituted.

```
    ConfValue val = maapi.getElem(th,
                                  "/hosts/host{%x}/interfaces{%x}/ip",
                                  new ConfBuf("host1"),
                                  new ConfBuf("eth0"));
```

The return value *`val`* contains a reference to a `ConfValue` which is a superclass of all the `ConfValues` that maps to the specific yang data type. If the Yang data type `ip` in the Yang model is `ietf-inet-types:ipv4-address`, we can narrow it to the subclass which is the corresponding `com.tailf.conf.ConfIPv4`.

```
    ConfIPv4 ipv4addr = (ConfIPv4)val;
```

The opposite operation of the `getElem()` is the `setElem()` method which set a leaf with a specific value.

```
    maapi.setElem(th ,
                  new ConfUInt16(1500),
                  "/hosts/host{%x}/interfaces{%x}/ip/mtu",
                  new ConfBuf("host1"),
                  new ConfBuf("eth0"));
```

We have not yet committed the transaction so no modification is permanent. The data is only visible inside the current transaction. To commit the transaction we call:

```
    maapi.applyTrans(th)
```

The method `applyTrans()` commits the current transaction to the running datastore.

{% code title="Example: Commit a Transaction" %}

```
    int th = maapi.startTrans(Conf.DB_RUNNING, Conf.MODE_READ_WRITE);
    try {
        maapi.lock(Conf.DB_RUNNING);
        /// make modifications to th
        maapi.setElem(th, .....);
        maapi.applyTrans(th);
        maapi.finishTrans(th);
    } catch(Exception e) {
        maapi.finishTrans(th);
    }  finally {
        maapi.unLock(Conf.DB_RUNNING);
    }
```

{% endcode %}

It is also possible to run the code above without `lock(Conf.DB_RUNNING)`.

Calling the `applyTrans()` method also performs additional validation of the new data as required by the data model and may fail if the validation fails. You can perform the validation beforehand, using the `validateTrans()` method.

Additionally, applying transaction can fail in case of a conflict with another, concurrent transaction. The best course of action in this case is to retry the transaction. Please see [Handling Conflicts](https://nso-docs.cisco.com/guides/development/core-concepts/api-overview/pages/8PJW3J1uuhcu5ttTWYkl#ncs.development.concurrency.handling) for details.

The MAAPI is also intended to attach to already existing NSO transaction to inspect not yet committed data for example if we want to implement validation logic in Java. See the example below (Attach Maapi to the Current Transaction).

## CDB API <a href="#d5e3408" id="d5e3408"></a>

This API provides an interface to the CDB Configuration database which stores all configuration data. With this API the user can:

* Start a CDB Session to read configuration data.
* Subscribe to changes in CDB - The subscription functionality makes it possible to receive events/notifications when changes occur in CDB.

CDB can also be used to store operational data, i.e., data which is designated with a "config false" statement in the YANG data model. Operational data is read/write trough the CDB API. NETCONF and the other northbound agents can only read operational data.

Java CDB API is intended to be fast and lightweight and the CDB read Sessions are expected to be short lived and fast. The NSO transaction manager is surpassed by CDB and therefore write operations on configurational data is prohibited. If operational data is stored in CDB both read and write operations on this data is allowed.

CDB is always locked for the duration of the session. It is therefore the responsibility of the programmer to make CDB interactions short in time and assure that all CDB sessions are closed when interaction has finished.

To initialize the CDB API a CDB socket has to be created and passed into the API base class `com.tailf.cdb.Cdb`:

{% code title="Example: Establish a Connection to CDB" %}

```
    Cdb cdb = new Cdb("MyCdbSock", UnixDomainSocketAddress.of(Conf.NCS_PATH));
```

{% endcode %}

After the `cdb` socket has been established, a user could either start a CDB Session or start a subscription of changes in CDB:

{% code title="Example: Establish a CDB Session" %}

```
    CdbSession session = cdb.startSession(CdbDBType.RUNNING);

    /*
     * Retrieve the number of children in the list and
     * loop over these children
     */
    for(int i = 0; i < session.getNumberOfInstances("/servers/server"); i++) {
        ConfBuf name =
           (ConfBuf) session.getElem("/servers/server[%d]/hostname", i);
        ConfIPv4 ip =
           (ConfIPv4) session.getElem("/servers/server[%d]/ip", i);
    }
```

{% endcode %}

We can refer to an element in a model with an expression like `/servers/server`. This type of string reference to an element is called keypath or just path. To refer to element underneath a list, we need to identify which instance of the list elements that is of interest.

This can be performed either by pinpointing the sequence number in the ordered list, starting from 0. For instance the path: `/servers/server[2]/port` refers to the `port` leaf of the third server in the configuration. This numbering is only valid during the current CDB session. Note, the database is locked during this session.

We can also refer to list instances using the key values for the list. Remember that we specify in the data model which leaf or leafs in list that constitute the key. In our case, a server has the `name` leaf as key. The syntax for keys is a space-separated list of key values enclosed within curly brackets: `{ Key1 Key2 ...}`. So, `/servers/server{www}/ip` refers to the `ip` leaf of the server whose name is `www`.

A YANG list may have more than one key for example the keypath: `/dhcp/subNets/subNet{192.168.128.0 255.255.255.0}/routers` refers to the routers list of the subnet which has key `192.168.128.0`, `255.255.255.0`.

The keypath syntax allows for formatting characters and accompanying substitution arguments. For example, `getElem("server[%d]/ifc{%s}/mtu",2,"eth0")` is using a keypath with a mix of sequence number and keyvalues with formatting characters and argument. Expressed in text the path will reference the MTU of the third server instance's interface named `eth0`.

The `CdbSession` Java class have a number of methods to control current position in the model.

* `CdbSession.cwd()` to get current position.
* `CdbSession.cd()` to change current position.
* `CdbSession.pushd()` to change and push a new position to a stack.
* `CdbSession.popd()` to change back to an stacked position.

Using relative paths and e.g. `CdbSession.pushd()`, it is possible to write code that can be re-used for common sub-trees.

The current position also includes the namespace. If an element of another namespace should be read, then the prefix of that namespace should be set in the first tag of the keypath, like: `/smp:servers/server` where `smp` is the prefix of the namespace. It is also possible to set the default namespace for the CDB session with the method `CdbSession.setNamespace(ConfNamespace)`.

{% code title="Example: Establish a CDB Subscription" %}

```
    CdbSubscription sub = cdb.newSubscription();
    int subid = sub.subscribe(1, new servers(), "/servers/server/");

    // tell CDB we are ready for notifications
    sub.subscribeDone();

    // now do the blocking read
    while (true) {
        int[] points = sub.read();
        // now do something here like diffIterate
        .....
    }
```

{% endcode %}

The CDB subscription mechanism allows an external Java program to be notified when different parts of the configuration changes. For such a notification, it is also possible to iterate through the change set in CDB for that notification.

Subscriptions are primarily to the running data store. Subscriptions towards the operational data store in CDB is possible, but the mechanism is slightly different see below.

The first thing to do is to register in CDB which paths should be subscribed to. This is accomplished with the `CdbSubscription.subscribe(...)` method. Each registered path returns a subscription point identifier. Each subscriber can have multiple subscription points, and there can be many different subscribers.

Every point is defined through a path - similar to the paths we use for read operations, with the difference that instead of fully instantiated paths to list instances we can choose to use tag paths i.e. leave out key value parts to be able to subscribe on all instances. We can subscribe either to specific leaves, or entire sub trees. Assume a YANG data model on the form of:

```yang
   container servers {
     list server {
       key name;
       leaf name { type string;}
       leaf ip { type inet:ip-address; }
       leaf port type inet:port-number; }
       .....
```

Explaining this by example we get:

```
/servers/server/port
```

A subscription on a leaf. Only changes to this leaf will generate a notification.

```
    /servers
```

Means that we subscribe to any changes in the subtree rooted at `/servers`. This includes additions or removals of server instances, as well as changes to already existing server instances.

```
    /servers/server{www}/ip
```

Means that we only want to be notified when the server "www" changes its ip address.

```
    /servers/server/ip
```

Means we want to be notified when the leaf ip is changed in any server instance.

When adding a subscription point the client must also provide a priority, which is an integer. As CDB is changed, the change is part of a transaction. For example, the transaction is initiated by a commit operation from the CLI or an edit-config operation in NETCONF resulting in the running database being modified. As the last part of the transaction, CDB will generate notifications in lock-step priority order. First, all subscribers at the lowest numbered priority are handled; once they all have replied and synchronized by calling `sync(CdbSubscriptionSyncType synctype)`, the next set - at the next priority level - is handled by CDB. Not until all subscription points have been acknowledged, is the transaction complete.

This implies that if the initiator of the transaction was, for example, a commit command in the CLI, the command will hang until notifications have been acknowledged.

Note that even though the notifications are delivered within the transaction, a subscriber can't reject the changes (since this would break the two-phase commit protocol used by the NSO backplane towards all data providers).

When a client is done subscribing, it needs to inform NSO it is ready to receive notifications. This is done by first calling `subscribeDone()`, after which the subscription socket is ready to be polled.

As a subscriber has read its subscription notifications using `read()`, it can iterate through the changes that caused the particular subscription notification using the `diffIterate()` method.

It is also possible to start a new read-session to the `CDB_PRE_COMMIT_RUNNING` database to read the running database as it was before the pending transaction.

Subscriptions towards the operational data in CDB are similar to the above, but because the operational data store is designed for light-weight access (and thus, does not have transactions and normally avoids the use of any locks), there are several differences, in particular:

* Subscription notifications are only generated if the writer obtains the subscription lock, by using the `startSession()` with the `CdbLockType.LOCKREQUEST`. In addition, when starting a session towards the operation data, we need to pass the `CdbDBType.CDB_OPERATIONAL` when starting a CDB session:\\

  ```
      CdbSession sess =
           cdb.startSession(CdbDBType.CDB_OPERATIONAL,
                            EnumSet.of(CdbLockType.LOCK_REQUEST));
  ```
* No priorities are used.
* Neither the writer that generated the subscription notifications nor other writers to the same data are blocked while notifications are being delivered. However, the subscription lock remains in effect until notification delivery is complete.
* The previous value for modified leaf is not available when using the `diffIterate()` method.

Essentially a write operation towards the operational data store, combined with the subscription lock, takes on the role of a transaction for configuration data as far as subscription notifications are concerned. This means that if operational data updates are done with many single-element write operations, this can potentially result in a lot of subscription notifications. Thus, it is a good idea to use the multi-element `setObject()` taking an array of ConfValues which sets a complete container or `setValues()` taking an array of `ConfXMLParam` and potent of setting an arbitrary part of the model. This to keep down notifications to subscribers when updating operational data.

Write operations that do not attempt to obtain the subscription lock, are allowed to proceed even during notification delivery. Therefore, it is the responsibility of the programmer to obtain the lock as needed when writing to the operational data store. E.g. if subscribers should be able to reliably read the exact data that resulted from the write that triggered their subscription, the subscription lock must always be obtained when writing that particular set of data elements. One possibility is of course to obtain the lock for all writes to operational data, but this may have an unacceptable performance impact.

To view registered subscribers, use the `ncs --status` command. For details on how to use the different subscription functions, see the Javadoc for NSO Java API.

The code in the [examples.ncs/sdk-api/cdb-java](https://github.com/NSO-developer/nso-examples/tree/6.7/sdk-api/cdb-java) example illustrates three different types of CDB subscribers:

* A simple CDB config subscriber that utilizes the low-level CDB API directly to subscribe to changes in the subtree of the configuration.
* Two Navu CDB subscribers, one subscribing to configuration changes, and one subscribing to changes in operational data.

## DP API <a href="#ug.java_api_overview.dp" id="ug.java_api_overview.dp"></a>

The DP API makes it possible to create callbacks which are called when certain events occur in NSO. As the name of the API indicates, it is possible to write data provider callbacks that provide data to NSO that is stored externally. However, this is only one of several callback types provided by this API. There exist callback interfaces for the following types:

* Service Callbacks - invoked for service callpoints in the YANG model. Implements service to device information mappings. See, for example, [examples.ncs/service-management/rfs-service](https://github.com/NSO-developer/nso-examples/tree/6.7/service-management/rfs-service).
* Action Callbacks - invoked for a certain action in the YANG model which is defined with a callpoint directive.
* Authentication Callbacks - invoked for external authentication functions.
* Authorization Callbacks - invoked for external authorization of operations and data. Note, avoid this callback if possible since performance will otherwise be affected.
* Data Callbacks - invoked for data provision and manipulation for certain data elements in the YANG model which is defined with a callpoint directive.
* DB Callbacks - invoked for external database stores.
* Range Action Callbacks - A variant of action callback where ranges are defined for the key values.
* Range Data Callbacks - A variant of data callback where ranges are defined for the data values.
* Snmp Inform Response Callbacks - invoked for response on Snmp inform requests on a certain element in the Yang model which is defined by a callpoint directive.
* Transaction Callbacks - invoked for external participants in the two-phase commit protocol.
* Transaction Validation Callbacks - invoked for external transaction validation in the validation phase of a two-phase commit.
* Validation Callbacks - invoked for validation of certain elements in the YANG Model which is designed with a callpoint directive.

The callbacks are methods in ordinary java POJOs. These methods are adorned with a specific Java Annotations syntax for that callback type. The annotation makes it possible to add metadata information to NSO about the supplied method. The annotation includes information about which `callType` and, when necessary, which `callpoint` the method should be invoked for.

{% hint style="info" %}
Only one Java object can be registered on one and the same `callpoint`. Therefore, when a new Java object registers on a `callpoint` that already has been registered, the earlier registration (and Java object) will be silently removed.
{% endhint %}

### Transaction and Data Callbacks <a href="#d5e3544" id="d5e3544"></a>

By default, NSO stores all configuration data in its CDB data store. We may wish to store and configure other data in NSO than what is defined by the NSO built-in YANG models, alternatively, we may wish to store parts of the NSO tree outside NSO (CDB) i.e. in an external database. Say, for example, that we have our customer database stored in a relational database disjunct from NSO. To implement this, we must do a number of things: We must define a callpoint somewhere in the configuration tree, and we must implement what is referred to as a data provider. Also, NSO executes all configuration changes inside transactions and if we want NSO (CDB) and our external database to participate in the same two-phase commit transactions, we must also implement a transaction callback. Altogether, it will appear as if the external data is part of the overall NSO configuration, thus the service model data can refer directly to this external data - typically to validate service instances.

The basic idea for a data provider is that it participates entirely in each NSO transaction, and it is also responsible for reading and writing all data in the configuration tree below the callpoint. Before explaining how to write a data provider and what the responsibilities of a data provider are, we must explain how the NSO transaction manager drives all participants in a lock-step manner through the phases of a transaction.

A transaction has a number of phases, the external data provider gets called in all the different phases. This is done by implementing a transaction callback class and then registering that class. We have the following distinct phases of an NSO transaction:

* `init()`: In this phase, the transaction callback class `init()` methods get invoked. We use annotation on the method to indicate that it's the `init()` method as in:\\

  ```java
      public  class MyTransCb {

          @TransCallback(callType=TransCBType.INIT)
          public void init(DpTrans trans) throws DpCallbackException {
              return;
          }
  ```

  \
  Each different callback method we wish to register must be annotated with an annotation from `TransCBType`.

  \
  The callback is invoked when a transaction starts, but NSO delays the actual invocation as an optimization. For a data provider providing configuration data, `init()` is invoked just before the first data-reading callback, or just before the `transLock()` callback (see below), whichever comes first. When a transaction has started, it is in a state we refer to as `READ`. NSO will, while the transaction is in the `READ` state, execute a series of read operations towards (possibly) different callpoints in the data provider.

  \
  Any write operations performed by the management station are accumulated by NSO and the data provider doesn't see them while in the `READ` state.
* `transLock()`: This callback gets invoked by NSO at the end of the transaction. NSO has accumulated a number of write operations and will now initiate the final write phases. Once the `transLock()` callback has returned, the transaction is in the `VALIDATE`state. In the `VALIDATE` state, NSO will (possibly) execute a number of read operations to validate the new configuration. Following the read operations for validations comes the invocation of one of the `writeStart()` or `transUnlock()` callbacks.
* `transUnlock()`: This callback gets invoked by NSO if the validation fails or if the validation was done separately from the commit (e.g. by giving a `validate` command in the CLI). Depending on where the transaction originated, the behavior after a call to `transUnlock()` differs. If the transaction originated from the CLI, the CLI reports to the user that the configuration is invalid and the transaction remains in the `READ` state whereas if the transaction originated from a NETCONF client, the NETCONF operation fails and a NETCONF `rpc` error is reported to the NETCONF client/manager.
* `writeStart()`: If the validation succeeded, the `writeStart()` callback will be called and the transaction will enter the `WRITE` state. While in `WRITE` state, a number of calls to the write data callbacks `setElem()`, `create()` and `remove()` will be performed.

  \
  If the underlying database supports real atomic transactions, this is a good place to start such a transaction.

  \
  The application should not modify the real running data here. If, later, the `abort()` callback is called, all write operations performed in this state must be undone.
* `prepare()`: Once all write operations are executed, the `prepare()` callback is executed. This callback ensures that all participants have succeeded in writing all elements. The purpose of the callback is merely to indicate to NSO that the data provider is ok, and has not yet encountered any errors.
* `abort()`: If any of the participants die or fail to reply in the `prepare()` callback, the remaining participants all get invoked in the `abort()` callback. All data written so far in this transaction should be disposed of.
* `commit()`: If all participants successfully replied in their respective `prepare()` callbacks, all participants get invoked in their respective `commit()` callbacks. This is the place to make all data written by the write callbacks in `WRITE` state permanent.
* `finish()`: And finally, the `finish()` callback gets invoked at the end. This is a good place to deallocate any local resources for the transaction. The `finish()` callback can be called from several different states.

The following picture illustrates the conceptual state machine an NSO transaction goes through.

<div data-with-frame="true"><figure><img src="/files/LXfow7HEwpRzqjUoebIg" alt="" width="375"><figcaption><p>NSO Transaction State Machine</p></figcaption></figure></div>

All callback methods are optional. If a callback method is not implemented, it is the same as having an empty callback which simply returns.

Similar to how we have to register transaction callbacks, we must also register data callbacks. The transaction callbacks cover the life span of the transaction, and the data callbacks are used to read and write data inside a transaction. The data callbacks have access to what is referred to as the transaction context in the form of a `DpTrans` object.

We have the following data callbacks:

* `getElem()`: This callback is invoked by NSO when NSO needs to read the actual value of a leaf element. We must also implement the `getElem()` callback for the keys. NSO invokes `getElem()` on a key as an existence test.\\

  We define the `getElem` callback inside a class as:\\

  ```java
  public static class DataCb {

      @DataCallback(callPoint="foo", callType=DataCBType.GET_ELEM)
          public ConfValue getElem(DpTrans trans, ConfObject[] kp)
          throws DpCallbackException {
             .....
  ```
* `existsOptional()`: This callback is called for all type less and optional elements, i.e. `presence` containers and leafs of type `empty` (unless in a union). If we have presence containers or leafs of type `empty` (unless in a union), we cannot use the `getElem()` callback to read the value of such a node, since it does not have a type. Type `empty` leafs in a union are instead read using `getElem()` callback.
* An example of a data model could be:\\

  ```yang
    container bs {
      presence "";
      tailf:callpoint bcp;
      list b {
        key name;
        max-elements 64;
        leaf name {
          type string;
        }
        container opt {
          presence "";
          leaf ii {
            type int32;
          }
        }
        leaf foo {
          type empty;
        }
      }
    }
  ```

  The above YANG fragment has three nodes that may or may not exist and that do not have a type. If we do not have any such elements, nor any operational data lists without keys (see below), we do not need to implement the `existsOptional()` callback.

  \
  If we have the above data model, we must implement the `existsOptional()`, and our implementation must be prepared to reply to calls of the function for the paths `/bs`, `/bs/b/opt`, and `/bs/b/foo`. The leaf `/bs/b/opt/ii` is not mandatory, but it does have a type namely `int32`, and thus the existence of that leaf will be determined through a call to the `getElem()` callback.

  \
  The `existsOptional()` callback may also be invoked by NSO as an "existence test" for an entry in an operational data list without keys. Normally this existence test is done with a `getElem()` request for the first key, but since there are no keys, this callback is used instead. Thus, if we have such lists, we must also implement this callback, and handle a request where the keypath identifies a list entry.
* `iterator()` and `getKey()`: This pair of callbacks is used when NSO wants to traverse a YANG list. The job of the `iterator()` callback is to return an `Iterator` object that is invoked by the library. For each `Object` returned by the `iterator`, the NSO library will invoke the `getKey()` callback on the returned object. The `getkey` callback shall return a `ConfKey` value.

  \
  An alternative to the `getKey()` callback is to register the optional `getObject()` callback whose job it is to return not just the key, but the entire YANG list entry. It is possible to register both `getKey()` and `getObject()` or either. If the `getObject()` is registered, NSO will attempt to use it only when bulk retrieval is executed.

We also have two additional optional callbacks that may be implemented for efficiency reasons.

* `getObject()`: If this optional callback is implemented, the work of the callback is to return an entire `object`, i.e., a list instance. This is not the same `getObject()` as the one that is used in combination with the `iterator()`
* `numInstances()`: When NSO needs to figure out how many instances we have of a certain element, by default NSO will repeatedly invoke the `iterator()` callback. If this callback is installed, it will be called instead.

The following example illustrates an external data provider. The example is possible to run from the examples collection. It resides under [examples.ncs/sdk-api/external-db](https://github.com/NSO-developer/nso-examples/tree/6.7/sdk-api/external-db).

The example comes with a tailor-made database - `MyDb`. That source code is provided with the example but not shown here. However, the functionality will be obvious from the method names like `newItem()`, `lock()`, `save()`, etc.

Two classes are implemented, one for the transaction callbacks and another for the data callbacks.

The data model we wish to incorporate into NSO is a trivial list of work items. It looks like:

{% code title="Example: work.yang" %}

```yang
        module work {
  namespace "http://example.com/work";
  prefix w;
  import ietf-yang-types {
    prefix yang;
  }
  import tailf-common {
    prefix tailf;
  }
  description "This model is used as a simple example model
               illustrating how to have NCS configuration data
               that is stored outside of NCS - i.e not in CDB";

  revision 2010-04-26 {
    description "Initial revision.";
  }

  container work {
    tailf:callpoint workPoint;
    list item {
      key key;
      leaf key {
        type int32;
      }
      leaf title {
        type string;
      }
      leaf responsible {
        type string;
      }
      leaf comment {
        type string;
      }
    }
  }
}
```

{% endcode %}

Note the callpoint directive in the model, it indicates that an external Java callback must register itself using that name. That callback will be responsible for all data below the callpoint.

To compile the `work.yang` data model and then also to generate Java code for the data model, we invoke `make all` in the example package src directory. The Makefile will compile the yang files in the package, generate Java code for those data models, and then also invoke ant in the Java src directory.

The Data callback class looks as follows:

{% code title="Example: DataCb Class" %}

```java
    @DataCallback(callPoint=work.callpoint_workPoint,
                  callType=DataCBType.ITERATOR)
    public Iterator<Object> iterator(DpTrans trans,
                                     ConfObject[] keyPath)
        throws DpCallbackException {
        return MyDb.iterator();
    }

    @DataCallback(callPoint=work.callpoint_workPoint,
                  callType=DataCBType.GET_NEXT)
    public ConfKey getKey(DpTrans trans, ConfObject[] keyPath,
                          Object obj)
        throws DpCallbackException {
        Item i = (Item) obj;
        return new ConfKey( new ConfObject[] { new ConfInt32(i.key) });
    }


    @DataCallback(callPoint=work.callpoint_workPoint,
                  callType=DataCBType.GET_ELEM)
    public ConfValue getElem(DpTrans trans, ConfObject[] keyPath)
        throws DpCallbackException {

        ConfInt32 kv = (ConfInt32) ((ConfKey) keyPath[1]).elementAt(0);
        Item i = MyDb.findItem( kv.intValue() );
        if (i == null) return null; // not found

        // switch on xml elem tag
        ConfTag leaf = (ConfTag) keyPath[0];
        switch (leaf.getTagHash()) {
        case work._key:
            return new ConfInt32(i.key);
        case work._title:
            return new ConfBuf(i.title);
        case work._responsible:
            return new ConfBuf(i.responsible);
        case work._comment:
            return new ConfBuf(i.comment);
        default:
            throw new DpCallbackException("xml tag not handled");
        }
    }

    @DataCallback(callPoint=work.callpoint_workPoint,
                  callType=DataCBType.SET_ELEM)
    public int setElem(DpTrans trans, ConfObject[] keyPath,
                       ConfValue newval)
        throws DpCallbackException {
        return Conf.REPLY_ACCUMULATE;
    }

    @DataCallback(callPoint=work.callpoint_workPoint,
                  callType=DataCBType.CREATE)
    public int create(DpTrans trans, ConfObject[] keyPath)
        throws DpCallbackException {
        return Conf.REPLY_ACCUMULATE;
    }

    @DataCallback(callPoint=work.callpoint_workPoint,
                  callType=DataCBType.REMOVE)
    public int remove(DpTrans trans, ConfObject[] keyPath)
        throws DpCallbackException {
        return Conf.REPLY_ACCUMULATE;
    }

    @DataCallback(callPoint=work.callpoint_workPoint,
                  callType=DataCBType.NUM_INSTANCES)
    public int numInstances(DpTrans trans, ConfObject[] keyPath)
        throws DpCallbackException {
        return MyDb.numItems();
    }


    @DataCallback(callPoint=work.callpoint_workPoint,
                  callType=DataCBType.GET_OBJECT)
    public ConfValue[] getObject(DpTrans trans, ConfObject[] keyPath)
        throws DpCallbackException {
        ConfInt32 kv = (ConfInt32) ((ConfKey) keyPath[0]).elementAt(0);
        Item i = MyDb.findItem( kv.intValue() );
        if (i == null) return null; // not found
        return getObject(trans, keyPath, i);
    }

    @DataCallback(callPoint=work.callpoint_workPoint,
                  callType=DataCBType.GET_NEXT_OBJECT)
    public ConfValue[] getObject(DpTrans trans, ConfObject[] keyPath,
                                 Object obj)
        throws DpCallbackException {
        Item i = (Item) obj;
        return new ConfValue[] {
            new ConfInt32(i.key),
            new ConfBuf(i.title),
            new ConfBuf(i.responsible),
            new ConfBuf(i.comment)
        };
    }
```

{% endcode %}

First, we see how the Java annotations are used to declare the type of callback for each method. Secondly, we see how the `getElem()` callback inspects the `keyPath` parameter passed to it to figure out exactly which element NSO wants to read. The `keyPath` is an array of `ConfObject` values. Keypaths are central to the understanding of the NSO Java library since they are used to denote objects in the configuration. A keypath uniquely identifies an element in the instantiated configuration tree.

Furthermore, the `getElem()` switches on the tag `keyPath[0]` which is a `ConfTag` using symbolic constants from the class "work". The "work" class was generated through the call to `ncsc --emit-java ...`.

The three write callbacks, `setElem()`, `create()` and `remove()` all return the value `Conf.REPLY_ACCUMULATE`. If our backend database has real support to abort transactions, it is a good idea to initiate a new backend database transaction in the Transaction callback `init()` (more on that later), whereas if our backend database doesn't support proper transactions, we can fake real transactions by returning `Conf.REPLY_ACCUMULATE` instead of actually writing the data. Since the final verdict of the NSO transaction as a whole may very well be to abort the transaction, we must be prepared to undo all write operations. The `Conf.REPLY_ACCUMULATE` return value means that we ask the library to cache the write for us.

The transaction callback class looks like this:

{% code title="Example: TransCb Class" %}

```java
    @TransCallback(callType=TransCBType.INIT)
    public void init(DpTrans trans) throws DpCallbackException {
        return;
    }

    @TransCallback(callType=TransCBType.TRANS_LOCK)
    public void transLock(DpTrans trans) throws DpCallbackException {
        MyDb.lock();
    }

    @TransCallback(callType=TransCBType.TRANS_UNLOCK)
    public void transUnlock(DpTrans trans) throws DpCallbackException {
        MyDb.unlock();
    }

    @TransCallback(callType=TransCBType.PREPARE)
    public void prepare(DpTrans trans) throws DpCallbackException {
        Item i;
        ConfInt32 kv;
        for (Iterator<DpAccumulate> it = trans.accumulated();
             it.hasNext(); ) {
            DpAccumulate ack= it.next();
            // check op
            switch (ack.getOperation()) {
            case DpAccumulate.SET_ELEM:
                kv = (ConfInt32)  ((ConfKey) ack.getKP()[1]).elementAt(0);
                if ((i = MyDb.findItem( kv.intValue())) == null)
                    break;
                // check leaf tag
                ConfTag leaf = (ConfTag) ack.getKP()[0];
                switch (leaf.getTagHash()) {
                case work._title:
                    i.title = ack.getValue().toString();
                    break;
                case work._responsible:
                    i.responsible = ack.getValue().toString();
                    break;
                case work._comment:
                    i.comment = ack.getValue().toString();
                    break;
                }
                break;
            case DpAccumulate.CREATE:
                kv = (ConfInt32)  ((ConfKey) ack.getKP()[0]).elementAt(0);
                MyDb.newItem(new Item(kv.intValue()));
                break;
            case DpAccumulate.REMOVE:
                kv = (ConfInt32)  ((ConfKey) ack.getKP()[0]).elementAt(0);
                MyDb.removeItem(kv.intValue());
                break;
            }
        }
        try {
            MyDb.save("running.prep");
        } catch (Exception e) {
            throw
              new DpCallbackException("failed to save file: running.prep",
                                      e);
        }
    }

    @TransCallback(callType=TransCBType.ABORT)
    public void abort(DpTrans trans) throws DpCallbackException {
        MyDb.restore("running.DB");
        MyDb.unlink("running.prep");
    }

    @TransCallback(callType=TransCBType.COMMIT)
    public void commit(DpTrans trans) throws DpCallbackException {
        try {
            MyDb.rename("running.prep","running.DB");
        } catch (DpCallbackException e) {
            throw new DpCallbackException("commit failed");
        }
    }

    @TransCallback(callType=TransCBType.FINISH)
    public void finish(DpTrans trans) throws DpCallbackException {
        ;
    }
}
```

{% endcode %}

We can see how the `prepare()` callback goes through all write operations and actually executes them towards our database `MyDb`.

### Service and Action Callbacks <a href="#d5e3716" id="d5e3716"></a>

Both service and action callbacks are fundamental in NSO.

Implementing a service callback is one way of creating a service type. This and other ways of creating service types are in-depth described in the [Package Development](/guides/development/advanced-development/developing-packages) section.

Action callbacks are used to implement arbitrary operations in Java. These operations can be basically anything, e.g. downloading a file, performing some test, resetting alarms, etc, but they should not modify the modeled configuration.

The actions are defined in the YANG model by means of `rpc` or `tailf:action` statements. Input and output parameters can optionally be defined via `input` and `output` statements in the YANG model. To specify that the `rpc` or `action` is implemented by a callback, the model uses a `tailf:actionpoint` statement.

The action callbacks are:

* `init()` Similar to the transaction `init()` callback. However note that, unlike the case with transaction and data callbacks, both `init()` and `action()` are registered for each `actionpoint` (i.e. different action points can have different `init()` callbacks), and there is no `finish()` callback - the action is completed when the `action()` callback returns.
* `action()` This callback is invoked to actually execute the `rpc` or `action`. It receives the input parameters (if any) and returns the output parameters (if any).

In the [examples.ncs/service-management/mpls-vpn-java](https://github.com/NSO-developer/nso-examples/tree/6.7/service-management/mpls-vpn-java) example, we can define a `self-test` action. In the `packages/l3vpn/src/yang/l3vpn.yang`, we locate the service callback definition:

```
uses ncs:service-data;
ncs:servicepoint vlanspnt;
```

Beneath the service callback definition, we add an action callback definition so the resulting YANG looks like the following:

```
uses ncs:service-data;
ncs:servicepoint vlanspnt;

tailf:action self-test {
  tailf:info "Perform self-test of the service";
  tailf:actionpoint vlanselftest;
  output {
    leaf success {
      type boolean;
    }
    leaf message {
      type string;
      description
        "Free format message.";
    }
  }
}
```

The `packages/l3vpn/src/java/src/com/example/l3vpnRFS.java` already contains an action implementation but it has been suppressed since no `actionpoint` with the corresponding name has been defined in the YANG model, before now.

```java
/**
 * Init method for selftest action
 */
@ActionCallback(callPoint="l3vpn-self-test",
callType=ActionCBType.INIT)
public void init(DpActionTrans trans) throws DpCallbackException {
}

/**
 * Selftest action implementation for service
 */
@ActionCallback(callPoint="l3vpn-self-test", callType=ActionCBType.ACTION)
public ConfXMLParam[] selftest(DpActionTrans trans, ConfTag name,
                               ConfObject[] kp, ConfXMLParam[] params)
throws DpCallbackException {
    try {
        // Refer to the service yang model prefix
        String nsPrefix = "l3vpn";
        // Get the service instance key
        String str = ((ConfKey)kp[0]).toString();

        return new ConfXMLParam[] {
              new ConfXMLParamValue(nsPrefix, "success", new ConfBool(true)),
              new ConfXMLParamValue(nsPrefix, "message", new ConfBuf(str))};
        } catch (Exception e) {
            throw new DpCallbackException("self-test failed", e);
        }
    }
}
```

### Validation Callbacks <a href="#d5e3761" id="d5e3761"></a>

In the `VALIDATE` state of a transaction, NSO will validate the new configuration. This consists of verification that specific YANG constraints such as `min-elements`, `unique`, etc, as well as arbitrary constraints specified by `must` expressions, are satisfied. The use of `must` expressions is the recommended way to specify constraints on relations between different parts of the configuration, both due to its declarative and concise form and due to performance considerations, since the expressions are evaluated internally by the NSO transaction engine.

In some cases, it may still be motivated to implement validation logic via callbacks in code. The YANG model will then specify a validation point by means of a `tailf:validate` statement. By default, the callback registered for a validation point will be invoked whenever a configuration is validated, since the callback logic will typically be dependent on data in other parts of the configuration, and these dependencies are not known by NSO. Thus it is important from a performance point of view to specify the actual dependencies by means of `tailf:dependency` substatements to the `validate` statement.

Validation callbacks use the MAAPI API to attach to the current transaction. This makes it possible to read the configuration data that is to be validated, even though the transaction is not committed yet. The view of the data is effectively the pre-existing configuration "shadowed" by the changes in the transaction, and thus exactly what the new configuration will look like if it is committed.

Similar to the case of transaction and data callbacks, there are transaction validation callbacks that are invoked when the validation phase starts and stops, and validation callbacks that are invoked for the specific validation points in the YANG model.

The transaction validation callbacks are:

* `init()`: This callback is invoked when the validation phase starts. It will typically attach to the current transaction:

{% code title="Example: Attach MAAPI to the Current Transaction" overflow="wrap" %}

```java
public class SimpleValidator implements DpTransValidateCallback { 
    ... 
    @TransValidateCallback(callType=TransValidateCBType.INIT) 
    public void init(DpTrans trans) throws DpCallbackException{ 
        try { 
            th = trans.thandle; 
            maapi.attach(th, new MyNamesapce().hash(), trans.uinfo.usid); 
            .. 
            } catch(Exception e) { 
            throw new DpCallbackException("failed to attach via maapi: "+ e.getMessage()); 
            } 
        } 
    }
```

{% endcode %}

* `stop()`: This callback is invoked when the validation phase ends. If `init()` attached to the transaction, `stop()` should detach from it.

The actual validation logic is implemented in a validation callback:

* `validate()`: This callback is invoked for a specific validation point.

#### Transforms <a href="#d5e3794" id="d5e3794"></a>

Transforms implement a mapping between one part of the data model - the front-end of the transform - and another part - the back-end of the transform. Typically the front-end is visible to northbound interfaces, while the back-end is not, but for operational data (`config false` in the data model), a transform may implement a different view (e.g. aggregation) of data that is also visible without going through the transform.

The implementation of a transform uses techniques already described in this section: Transaction and data callbacks are registered and invoked when the front-end data is accessed, and the transform uses the MAAPI API to attach to the current transaction and accesses the back-end data within the transaction.

To specify that the front-end data is provided by a transform, the data model uses the `tailf:callpoint` statement with a `tailf:transform true` substatement. Since transforms do not participate in the two-phase commit protocol, they only need to register the `init()` and `finish()` transaction callbacks. The `init()` callback attaches to the transaction and `finish()` detaches from it. Also, a transform for operational data only needs to register the data callbacks that read data, i.e. `getElem()`, `existsOptional()`, etc.

#### Hooks <a href="#d5e3808" id="d5e3808"></a>

Hooks make it possible to have changes to the configuration trigger additional changes. In general, this should only be done when the data that is written by the hook is not visible to northbound interfaces since otherwise, the additional changes will make it difficult e.g. EMS or NMS systems to manage the configuration - the complete configuration resulting from a given change cannot be predicted. However, one use case in NSO for hooks that trigger visible changes is precisely to model-managed devices that have this behavior: hooks in the device model can emulate what the device does on certain configuration changes, and thus the device configuration in NSO remains in sync with the actual device configuration.

The implementation technique for a hook is very similar to that for a transform. Transaction and data callbacks are registered, and the MAAPI API is used to attach to the current transaction and write the additional changes into the transaction. As for transforms, only the `init()` and `finish()` transaction callbacks need to be registered, to do the MAAPI attach and detach. However only data callbacks that write data, i.e. `setElem()`, `create()`, etc need to be registered, and depending on which changes should trigger the hook invocation, it is possible to register only a subset of those. For example, if the hook is registered for a leaf in the data model, and only changes to the value of that leaf should trigger invocation of the hook, it is sufficient to register `setElem()`.

To specify that changes to some part of the configuration should trigger a hook invocation, the data model uses the `tailf:callpoint` statement with a `tailf:set-hook` or `tailf:transaction-hook` sub-statement. A set-hook is invoked immediately when a northbound agent requests a write operation on the data, while a transaction-hook is invoked when the transaction is committed. For the NSO-specific use case mentioned above, a `set-hook` should be used. The `tailf:set-hook` and `tailf:transaction-hook` statements take an argument specifying the extent of the data model the hook applies to.

### NED API <a href="#d5e3823" id="d5e3823"></a>

NSO can speak southbound to an arbitrary management interface. This is of course not entirely automatic like with NETCONF or SNMP, and depending on the type of interface the device has for configuration, this may involve some programming. Devices with a Cisco-style CLI can however be managed by writing YANG models describing the data in the CLI, and a relatively thin layer of Java code to handle the communication to the devices. Refer to Network Element Drivers (NEDs) for more information.

### NAVU API <a href="#ug.java_api_overview.navu" id="ug.java_api_overview.navu"></a>

The NAVU API provides a DOM-driven approach to navigate the NSO service and device models. The main features of the NAVU API are dynamic schema loading at start-up and lazy loading of instance data. The navigation model is based on the YANG language structure. In addition to navigation and reading of values, NAVU also provides methods to modify the data model. Furthermore, it supports the execution of actions modeled in the service model.

By using NAVU, it is easy to drill down through tree structures with minimal effort using the node-by-node navigation primitives. Alternatively, we can use the NAVU search feature. This feature is especially useful when we need to find information deep down in the model structures.

NAVU requires all models i.e. the complete NSO service model with all its augmented sub-models. This is loaded at runtime from NSO. NSO has in turn acquired these from loaded `.fxs` files. The `.fxs` files are a product from the `ncsc` tool with compiles these from the `.yang` files.

The `ncsc` tool can also generate Java classes from the .yang files. These files, extending the `ConfNamespace` base class, are the Java representation of the models and contain all defined nametags and their corresponding hash values. These Java classes can, optionally, be used as help classes in the service applications to make NAVU navigation type-safe, e.g. eliminating errors from misspelled model container names.

<div data-with-frame="true"><figure><img src="/files/wGpw2zveaCB0mAlGUx7w" alt="" width="563"><figcaption><p>NAVU Design Support</p></figcaption></figure></div>

The service models are loaded at start-up and are always the latest version. The models are always traversed in a lazy fashion i.e. data is only loaded when it is needed. This is to minimize the amount of data transferred between NSO and the service applications.

The most important classes of NAVU are the classes implementing the YANG node types. These are used to navigate the DOM. These classes are as follows.

* `NavuContainer`: the NavuContainer is a container representing either the root of the model, a YANG module root, or a YANG container.
* `NavuList`: the NavuList represents a YANG list node.
* `NavuListEntry`: list node entry.
* `NavuLeaf`: the NavuLeaf represents a YANG leaf node.

<div data-with-frame="true"><figure><img src="/files/3ngNuOktB9fF6iDwJ46C" alt="" width="563"><figcaption><p>NAVU YANG Structure</p></figcaption></figure></div>

The remaining part of this section will guide us through the most useful features of the NAVU. Should further information be required, please refer to the corresponding Javadoc pages.

NAVU relies on MAAPI as the underlying interface to access NSO. The starting point in NAVU configuration is to create a `NavuContext` instance using the `NavuContext(Maapi maapi)` constructor. To read and/or write data a transaction has to be started in Maapi. There are methods in the `NavuContext` class to start and handle this transaction.

If data has to be written, the Navu transaction has to be started differently depending on the data being the configuration or operational data. Such a transaction is started by the methods `NavuContext.startRunningTrans()` or `NavuContext.startOperationalTrans()` respectively. The Javadoc describes this in more detail.

When navigating using NAVU we always start by creating a `NavuContainer` and passing in the `NavuContext` instance, this is a base container from which navigation can be started. Furthermore, we need to create a root `NavuContainer` which is the top of the YANG module in which to navigate down. This is done by using the `NavuContainer.container(int hash)` method. Here the argument is the hash value for the module namespace.

{% code title="Example: NSO Module" %}

```yang
module tailf-ncs {
  namespace "http://tail-f.com/ns/ncs";
  ...
}
```

{% endcode %}

{% code title="Example: NSO NavuContainer Instance" %}

```java
    .....
      NavuContext context = new NavuContext(maapi);
      context.startRunningTrans(Conf.MODE_READ);
      // This will be the base container "/"
      NavuContainer base = new NavuContainer(context);

      // This will be the ncs root container "/ncs"
      NavuContainer root = base.container(new Ncs().hash());
      .....
      // This method finishes the started read transaction and
      // clears the context from this transaction.
      context.finishClearTrans();
```

{% endcode %}

NAVU maps the YANG node types; `container`, `list`, `leaf`, and `leaf-list` into its own structure. As mentioned previously `NavuContainer` is used to represent both the `module` and the `container` node type. The `NavuListEntry` is also used to represent a `list` node instance (actually `NavuListEntry` extends `NavuContainer`). i.e. an element of a list node.

Consider the YANG excerpt below.

{% code title="Example: NSO List Element" %}

```yang
submodule tailf-ncs-devices {
  ...
  container devices {
    .....

      list device {

        key name;

        leaf name {
          type string;
        }
        ....
      }
    }

    .......
  }
}
```

{% endcode %}

If the purpose is to directly access a list node, we would typically do a direct navigation to the list element using the NAVU primitives.

{% code title="Example: NAVU List Direct Element Access" %}

```java
    .....
    NavuContext context = new NavuContext(maapi);
    context.startRunningTrans(Conf.MODE_READ);

    NavuContainer base = new NavuContainer(context);
    NavuContainer ncs = base.container(new Ncs().hash());
    NavuContainer dev = ncs.container("devices").
                             list("device").
                             elem( key);

    NavuListEntry devEntry = (NavuListEntry)dev;
    .....
    context.finishClearTrans();
```

{% endcode %}

Or if we want to iterate over all elements of a list we could do as follows.

{% code title="Example: NAVU List Element Iterating" %}

```java
    .....
    NavuContext context = new NavuContext(maapi);
    context.startRunningTrans(Conf.MODE_READ);

    NavuContainer base = new NavuContainer(context);
    NavuContainer ncs = base.container(new Ncs().hash());
    NavuList listOfDevs = ncs.container("devices").
                             list("device");

    for (NavuContainer dev: listOfDevs.elements()) {
        .....
    }
    .....
    context.finishClearTrans();
```

{% endcode %}

The above example uses the `select()` which uses a recursive regexp match against its children.

Alternatively, if the purpose is to drill down deep into a structure we should use `select()`. The `select()` offers a wild card-based search. The search is relative and can be performed from any node in the structure.

{% code title="Example: NAVU Leaf Access" %}

```java
    .....
    NavuContext context = new NavuContext(maapi);
    context.startRunningTrans(Conf.MODE_READ);

    NavuContainer base = new NavuContainer(context);
    NavuContainer ncs = base.container(new Ncs().hash());

    for (NavuNode node: ncs.container("devices").select("dev.*/.*")) {
        NavuContainer dev = (NavuContainer)node;
        .....
    }
    .....
    context.finishClearTrans();
```

{% endcode %}

All of the above are valid ways of traversing the lists depending on the purpose. If we know what we want, we use direct access. If we want to apply something to a large amount of nodes, we use `select()`.

An alternative method is to use the `xPathSelect()` where an XPath query could be issued instead.

{% code title="Example: NAVU Leaf Access" %}

```java
    .....
    NavuContext context = new NavuContext(maapi);
    context.startRunningTrans(Conf.MODE_READ);

    NavuContainer base = new NavuContainer(context);
    NavuContainer ncs = base.container(new Ncs().hash());

    for (NavuNode node: ncs.container("devices").xPathSelect("device/*")) {
        NavuContainer devs = (NavuContainer)node;
        .....
    }
    .....
    context.finishClearTrans();
```

{% endcode %}

`NavuContainer` and `NavuList` are structural nodes with NAVU. i.e. they have no values. Values are always kept by `NavuLeaf`. A `NavuLeaf` represents the YANG node types `leaf`. A `NavuLeaf` can be both read and set. `NavuLeafList` represents the YANG node type `leaf-list` and has some features in common with both `NavuLeaf` (which it inherits from) and `NavuList`.

{% code title="Example: NSO Leaf" %}

```yang
module tailf-ncs {
  namespace "http://tail-f.com/ns/ncs";
  ...
  container ncs {
    .....

      list service {

        key object-id;

        leaf object-id {
          type string;
        }
        ....

        leaf reference {
          type string;
        }
        ....

      }
    }

    .......
  }
}
```

{% endcode %}

To read and update a leaf, we simply navigate to the leaf and request the value. And in the same manner, we can update the value.

{% code title="Example: NAVU List Element Iterating" %}

```java
    .....
    NavuContext context = new NavuContext(maapi);
    context.startRunningTrans(Conf.MODE_READ);

    NavuContainer base = new NavuContainer(context);
    NavuContainer ncs = base.container(new Ncs().hash());

    for (NavuNode node: ncs.select("sm/ser.*/.*")) {
        NavuContainer rfs = (NavuContainer)node;
        if (rfs.leaf(Ncs._description_).value()==null) {
            /*
             * Setting dummy value.
             */
            rfs.leaf(Ncs._description_).set(new ConfBuf("Dummy value"));
        }
    }
    .....
    context.finishClearTrans();
```

{% endcode %}

In addition to the YANG standard node types, NAVU also supports the Tailf proprietary node type `action`. An action is considered being a `NavuAction`. It differs from an ordinary container in that it can be executed using the `call()` primitive. Input and output parameters are represented as ordinary nodes. The action extension of YANG allows an arbitrary structure to be defined both for input and output parameters.

Consider the excerpt below. It represents a module on a managed device. When connected and synchronized to the NSO, the module will appear in the `/devices/device/config` container.

{% code title="Example: YANG Action" %}

```yang
module interfaces {
  namespace "http://router.com/interfaces";
  prefix i;
  .....

  list interface {
    key name;
    max-elements 64;

    tailf:action ping-test {
      description "ping a machine ";
      tailf:exec "/tmp/mpls-ping-test.sh" {
        tailf:args "-c $(context) -p $(path)";
      }

      input {
        leaf ttl {
            type int8;
        }
      }

      output {
        container rcon {
          leaf result {
            type string;
          }
          leaf ip {
            type inet:ipv4-address;
          }
          leaf ival {
            type int8;
          }
        }
      }
    }

   .....

  }

  .....
}
```

{% endcode %}

To execute the action below we need to access a device with this module loaded. This is done in a similar way to non-action nodes.

{% code title="Example: NAVU Action Execution (1)" %}

```java
    .....
    NavuContext context = new NavuContext(maapi);
    context.startRunningTrans(Conf.MODE_READ);

    NavuContainer base = new NavuContainer(context);
    NavuContainer ncs = base.container(new Ncs().hash());

    /*
     * Execute ping on all devices with the interface module.
     */
    for (NavuNode node: ncs.container(Ncs._devices_).
                   select("device/.*/config/interface/.*")) {
        NavuContainer if = (NavuContainer)node;

        NavuAction ping = if.action(interfaces.i_ping_test_);


        /*
         * Execute action.
         */
        ConfXMLParamResult[] result = ping.call(new ConfXMLParam[] {
                new ConfXMLParamValue(new interfaces().hash(),
                                      interfaces._ttl,
                                      new ConfInt64(64))};

        //or we could execute it with XML-String

        result = ping.call("<if:ttl>64</if:ttl>");
        /*
         * Output the result of the action.
         */
         System.out.println("result_ip: "+
         ((ConfXMLParamValue)result[1]).getValue().toString());

         System.out.println("result_ival:" +
         ((ConfXMLParamValue)result[2]).getValue().toString());
    }
    .....
    context.finishClearTrans();
```

{% endcode %}

Or, we could do it with `xPathSelect()`.

{% code title="Example: NAVU Action Execution (2)" %}

```java
    .....
    NavuContext context = new NavuContext(maapi);
    context.startRunningTrans(Conf.MODE_READ);

    NavuContainer base = new NavuContainer(context);
    NavuContainer ncs = base.container(new Ncs().hash());

    /*
     * Execute ping on all devices with the interface module.
     */
    for (NavuNode node: ncs.container(Ncs._devices_).
                   xPathSelect("device/config/interface")) {
        NavuContainer if = (NavuContainer)node;

        NavuAction ping = if.action(interfaces.i_ping_test_);


        /*
         * Execute action.
         */
        ConfXMLParamResult[] result = ping.call(new ConfXMLParam[] {
                new ConfXMLParamValue(new interfaces().hash(),
                                      interfaces._ttl,
                                      new ConfInt64(64))};

        //or we could execute it with XML-String

        result = ping.call("<if:ttl>64</if:ttl>");
        /*
         * Output the result of the action.
         */
         System.out.println("result_ip: "+
         ((ConfXMLParamValue)result[1]).getValue().toString());

         System.out.println("result_ival:" +
         ((ConfXMLParamValue)result[2]).getValue().toString());
    }
    .....
    context.finishClearTrans();
```

{% endcode %}

The examples above have described how to attach to the NSO module and navigate through the data model using the NAVU primitives. When using NAVU in the scope of the NSO Service manager, we normally don't have to worry about attaching the `NavuContainer` to the NSO data model. NSO does this for us providing `NavuContainer` nodes pointing at the nodes of interest.

## ALARM API <a href="#d5e3951" id="d5e3951"></a>

Since this API is potent for both producing and consuming alarms, this becomes an API that can be used both north and eastbound. It adheres to the NSO Alarm model.

For more information see [Alarm Manager](/guides/operation-and-usage/operations/alarm-manager)*.*

The `com.tailf.ncs.alarmman.consumer.AlarmSource` class is used to subscribe to alarms. This class establishes a listener towards an alarm subscription server called `com.tailf.ncs.alarmman.consumer.AlarmSourceCentral`. The `AlarmSourceCentral` needs to be instantiated and started prior to the instantiation of the `AlarmSource` listener. The NSO Java VM takes care of starting the `AlarmSourceCentral` so any use of the ALARM API inside the NSO Java VM can expect this server to be running.

For situations where alarm subscription outside of the NSO Java VM is desired, starting the `AlarmSourceCentral` is performed by creating a `Cdb` connection, passing this `Cdb` to the `AlarmSourceCentral` class, and then calling the `start()` method.

```
    Cdb cdb = new Cdb("my-alarm-source-socket",
                      UnixDomainSocketAddress.of(Conf.NCS_PATH));

    // Get and start alarm source - this must only be done once per JVM
    AlarmSourceCentral source = new AlarmSourceCentral(10000, cdb);
    source.start();
```

To retrieve alarms from the `AlarmSource` listener, either a blocking `takeAlarm()` or a timeout based `pollAlarm()` can be used. The first method will wait indefinitely for new alarms to arrive while the second will timeout if an alarm has not arrived in the stipulated time. When a listener no longer is needed then a `stopListening()` call should be issued to deactivate it, or the `AlarmSource` can be used in a try-with-resources statement.

{% code title="Consuming alarms inside NSO Java VM" %}

```
        try (AlarmSource mySource = new AlarmSource()) {
            mySource.startListening();
            // Get an alarms.
            Alarm alarm = mySource.takeAlarm();

            while (alarm != null){
                System.out.println(alarm);

                for (Attribute attr: alarm.getCustomAttributes()){
                    System.out.println(attr);
                }

                alarm = mySource.takeAlarm();
            }

        } catch (Exception e) {
            e.printStackTrace();
        }
```

{% endcode %}

{% code title="Consuming alarms outside NSO Java VM" %}

```
        try (AlarmSource mySource = new AlarmSource(source)) {
            mySource.startListening();
            // Get an alarms.
            Alarm alarm = mySource.takeAlarm();

            while (alarm != null){
                System.out.println(alarm);

                for (Attribute attr: alarm.getCustomAttributes()){
                    System.out.println(attr);
                }

                alarm = mySource.takeAlarm();
            }

        } catch (Exception e) {
            e.printStackTrace();
        }
```

{% endcode %}

Both the `takeAlarm()` and the `pollAlarm()` method returns a `Alarm` object from which all alarm information can be retrieved.

The `com.tailf.ncs.alarmman.producer.AlarmSink` is used to persistently store alarms in NSO. This can be performed either directly or by the use of an alarm storage server called `com.tailf.ncs.alarmman.producer.AlarmSinkCentral`.

To directly store alarms an AlarmSink instance is created using the `AlarmSink(Maapi maapi)` constructor.

```
        Maapi maapi = new Maapi(UnixDomainSocketAddress.of(Conf.NCS_PATH));
        maapi.startUserSession("system", "system");

        AlarmSink sink = new AlarmSink(maapi);
```

On the other hand, if the alarms are to be stored using the `AlarmSinkCentral` then the `AlarmSink()` constructor without arguments is used.

```
        AlarmSink sink = new AlarmSink();
```

However, this case requires that the `AlarmSinkCentral` is started prior to the instantiation of the `AlarmSink`. The NSO Java VM will take care of starting this server so any use of the ALARM API inside the Java VM can expect this server to be running. If it is desired to store alarms in an application outside of the NSO java VM, the `AlarmSinkCentral` needs to be started like the following example:

```
       Maapi maapi = new Maapi(UnixDomainSocketAddress.of(Conf.NCS_PATH));
       maapi.startUserSession("system", "system");

       AlarmSinkCentral sinkCentral = new AlarmSinkCentral(1000, maapi);
       sinkCentral.start();
```

The alarm sink can then be started with the `AlarmSink(AlarmSinkCentral central)` constructor, i.e.:

```
       AlarmSink sink = new AlarmSink(sinkCentral);
```

To store an alarm using the `AlarmSink`, an `Alarm` instance must be created. This alarm alarm instance is then stored by a call to the `submitAlarm()` method.

```
    ArrayList<AlarmId> idList = new ArrayList<AlarmId>();

    ConfIdentityRef alarmType =
        new ConfIdentityRef(NcsAlarms.hash,
                               NcsAlarms._ncs_dev_manager_alarm);

    ManagedObject managedObject1 =
        new ManagedObject("/ncs:devices/device{device0}/config/root1");
    ManagedObject managedObject2 =
        new ManagedObject("/ncs:devices/device{device0}/config/root2");

    idList.add(new AlarmId(new ManagedDevice("device0"),
                           alarmType,
                           managedObject1));
    idList.add(new AlarmId(new ManagedDevice("device0"),
                           alarmType,
                           managedObject2));

    ManagedObject managedObject3 =
        new ManagedObject("/ncs:devices/device{device0}/config/root3");

    Alarm myAlarm =
        new Alarm(new ManagedDevice("device0"),
                  managedObject3,
                  alarmType,
                  PerceivedSeverity.WARNING,
                  false,
                  "This is a warning",
                  null,
                  idList,
                  null,
                  ConfDatetime.getConfDatetime(),
                  new AlarmAttribute(myAlarm.hash,
                                     myAlarm._custom_alarm_attribute_,
                                     new ConfBuf("An alarm attribute")),
                  new AlarmAttribute(myAlarm.hash,
                                     myAlarm._custom_status_change_,
                                     new ConfBuf("A status change")));

     sink.submitAlarm(myAlarm);
```

## NOTIF API (Notification API) <a href="#ug.java_api_overview.notif" id="ug.java_api_overview.notif"></a>

Applications can subscribe to certain events generated by NSO. The event types are defined by the `com.tailf.notif.NotificationType` enumeration. The following notification can be subscribed to:

* `NotificationType.NOTIF_AUDIT`: all audit log events are sent from NSO on the event notification socket.
* `NotificationType.NOTIF_COMMIT_SIMPLE`: an event indicating that a user has somehow modified the configuration.
* `NotificationType.NOTIF_COMMIT_DIFF`: an event indicating that a user has somehow modified the configuration. The main difference between this event and the above-mentioned `NOTIF_COMMIT_SIMPLE` is that this event is synchronous, i.e. the entire transaction hangs until we have explicitly called `Notif.diffNotificationDone()`. The purpose of this event is to give the applications a chance to read the configuration diffs from the transaction before it commits. A user subscribing to this event can use the MAAPI API to attach `Maapi.attach()` to the running transaction and use `Maapi.diffIterate()` to iterate through the diff.
* `NotificationType.NOTIF_COMMIT_FAILED`: This event is generated when a data provider fails in its commit callback. NSO executes a two-phase commit procedure towards all data providers when committing transactions. When a provider fails to commit, the system is an unknown state. If the provider is "external", the name of the failing daemon is provided. If the provider is another NETCONF agent, the IP address and port of that agent is provided.
* `NotificationType.NOTIF_COMMIT_PROGRESS`: This event provides progress information about the commit of a transaction.
* `NotificationType.NOTIF_PROGRESS`: This event provides progress information about the commit of a transaction or an action being applied. Subscribing to this notification type means that all notifications of the type `NotificationType.NOTIF_COMMIT_PROGRESS` are subscribed to as well.
* `NotificationType.NOTIF_CONFIRMED_COMMIT`: This event is generated when a user has started a confirmed commit, when a confirming commit is issued, or when a confirmed commit is aborted; represented by `ConfirmNotification.confirm_type`. For a confirmed commit, the timeout value is also present in the notification.
* `NotificationType.NOTIF_FORWARD_INFO`: This event is generated whenever the server forwards (proxies) a northbound agent.
* `NotificationType.NOTIF_HA_INFO`: an event related to NSO's perception of the current cluster configuration.
* `NotificationType.NOTIF_HEARTBEAT`: This event can be used by applications that wish to monitor the health and liveness of the server itself. It needs to be requested through a Notif instance which has been constructed with a heartbeat\_interval. The server will continuously generate heartbeat events on the notification socket. If the server fails to do so, the server is hung. The timeout interval is measured in milliseconds. The recommended value is 10000 milliseconds to cater for truly high load situations. Values less than 1000 are changed to 1000.
* `NotificationType.NOTIF_SNMPA`: This event is generated whenever an SNMP PDU is processed by the server. The application receives an `SnmpaNotification` with a list of all varbinds in the PDU. Each varbind contains subclasses that are internal to the SnmpaNotification.
* `NotificationType.NOTIF_SUBAGENT_INFO`: Only sent if NSO runs as a primary agent with subagents enabled. This event is sent when the subagent connection is lost or reestablished. There are two event types, defined in `SubagentNotification.subagent_info_type}`: "subagent up" and "subagent down".
* `NotificationType.NOTIF_DAEMON`: all log events that also go to the `/NCSConf/logs/NSCLog` log are sent from NSO on the event notification socket.
* `NotificationType.NOTIF_NETCONF`: All log events that also go to the `/NCSConf/logs/netconfLog` log are sent from NSO on the event notification socket.
* `NotificationType.NOTIF_DEVEL`: All log events that also go to the `/NCSConf/logs/develLog` log are sent from NSO on the event notification socket.
* `NotificationType.NOTIF_TAKEOVER_SYSLOG`: If this flag is present, NSO will stop Syslogging. The idea behind the flag is that we want to configure Syslogging for NSO to let NSO log its startup sequence. Once NSO is started we wish to subsume the syslogging done by NSO. Typical applications that use this flag want to pick up all log messages, reformat them, and use some local logging method. Once all subscriber sockets with this flag set are closed, NSO will resume to syslog.
* `NotificationType.NOTIF_UPGRADE_EVENT`: This event is generated for the different phases of an in-service upgrade, i.e. when the data model is upgraded while the server is running. The application receives an `UpgradeNotification` where the `UpgradeNotification.event_type` gives the specific upgrade event. The events correspond to the invocation of the Maapi functions that drive the upgrade.
* `NotificationType.NOTIF_COMPACTION`: This event is generated after each CDB compaction performed by NSO. The application receives a `CompactionNotification` where `CompactionNotification.dbfile` indicates which datastore was compacted, and `CompactionNotification.compaction_type` indicates whether the compaction was triggered manually or automatically by the system.
* `NotificationType.NOTIF_USER_SESSION`: An event related to user sessions. There are 6 different user session-related event types, defined in `UserSessNotification.user_sess_type`: session starts/stops, session locks/unlocks database, and session starts/stop database transaction.

To receive events from the NSO the application opens a socket and passes it to the notification base class `com.tailf.notif.Notif` together with an EnumSet of NotificationType for all types of notifications that should be received. Looping over the `Notif.read()` method will read and deliver notifications which are all subclasses of the `com.tailf.notif.Notification` base class.

```
    SocketAddress address = UnixDomainSocketAddress.of(Conf.NCS_PATH);
    EnumSet notifSet = EnumSet.of(NotificationType.NOTIF_COMMIT_SIMPLE,
                                  NotificationType.NOTIF_AUDIT);
    try (Notif notif = new Notif(address, notifSet)) {
        while (true) {
            Notification n = notif.read();

            if (n instanceof CommitNotification) {
                // handle NOTIF_COMMIT_SIMPLE case
                .....
            } else if (n instanceof AuditNotification) {
                 // handle NOTIF_AUDIT case
                .....
            }
        }
    }
```

## HA API <a href="#d5e4061" id="d5e4061"></a>

The HA API is used to set up and control High-Availability cluster nodes. This package is used to connect to the High Availability (HA) subsystem. Configuration data can then be replicated on several nodes in a cluster. (see [High Availability](/guides/administration/management/high-availability))

The following example configures three nodes in a HA cluster. One is set as primary and the other two as secondaries.

{% code title="Example: HA Cluster Setup" %}

```
  ....

  Socket s0 = new Socket("host1", Conf.NCS_PORT);
  Socket s1 = new Socket("host2", Conf.NCS_PORT);
  Socket s2 = new Socket("host3", Conf.NCS_PORT);

  // For local IPC (only works on the same host) use:
  // SocketAddress node1 = UnixDomainSocketAddress.of("/tmp/nso/nso-ipc1");
  // Socket s0 = SocketFactory.getSocket(node1);

  Ha ha0 = new Ha(s0, "clus0");
  Ha ha1 = new Ha(s1, "clus0");
  Ha ha2 = new Ha(s2, "clus0");

  ConfHaNode primary =
      new ConfHaNode(new ConfBuf("node0"),
                     new ConfIPv4(InetAddress.getByName("localhost")));


  ha0.bePrimary(primary.nodeid);

  ha1.beSecondary(new ConfBuf("node1"), primary, true);

  ha2.beSecondary(new ConfBuf("node2"), primary, true);

  HaStatus status0 = ha0.status();
  HaStatus status1 = ha1.status();
  HaStatus status2 = ha2.status();

  ....
```

{% endcode %}

## Java API Conf Package

This section describes the types and how these types map to various YANG types and Java classes.

All types inherit the base class `com.tailf.conf.ConfObject`.

Following the type hierarchy of `ConfObject` subclasses are distinguished by:

* `Value`: A concrete value classes which inherits `ConfValue` that in turn is a subclass of `ConfObject`.
* `TypeDescriptor`: a class representing the type of a ConfValue. A type-descriptor is represented as an instance of `ConfTypeDescriptor`. Usage is primarily to be able to map a ConfValue to its internal integer value representation or vice versa.
* `Tag`: A tag is a representation of an element in the YANG model. A Tag is represented as an instance of `com.tailf.conf.Tag`. The primary usage of tags are in the representation of keypaths.
* `Key`: a key is a representation of the instance key for an element instance. A key is represented as an instance of `com.tailf.conf.ConfKey`. A ConfKey is constructed from an array of values (ConfValue\[]). The primary usage of keys is in the representation of keypaths.
* `XMLParam`: subclasses of ConfXMLParam which are used to represent a, possibly instantiated, subtree of a YANG model. Useful in several APIs where multiple values can be set or retrieved in one function call.

The class `ConfObject` defines public int constants for the different value types. Each value type is mapped to a specific YANG type and is also represented by a specific subtype of `ConfValue`. Having a ConfValue instance it is possible to retrieve its integer representation by the use of the static method `getConfTypeDescriptor()` in class `ConfTypeDescriptor`. This function returns a `ConfTypeDescriptor` instance representing the value from which the integer representation can be retrieved. The values represented as integers are:

The table lists `ConfValue` types.

| Constant                | YANG type                        | ConfValue                 | Description             |
| ----------------------- | -------------------------------- | ------------------------- | ----------------------- |
| `J_STR`                 | string                           | `ConfBuf`                 | Human readable string   |
| `J_BUF`                 | string                           | `ConfBuf`                 | Human readable string   |
| `J_INT8`                | int8                             | `ConfInt8`                | 8-bit signed integer    |
| `J_INT16`               | int16                            | `ConfInt16`               | 16-bit signed integer   |
| `J_INT32`               | int32                            | `ConfInt32`               | 32-bit signed integer   |
| `J_INT64`               | int64                            | `ConfInt64`               | 64-bit signed integer   |
| `J_UINT8`               | uint8                            | `ConfUInt8`               | 8-bit unsigned integer  |
| `J_UINT16`              | uint16                           | `ConfUInt16`              | 16-bit unsigned integer |
| `J_UINT32`              | uint32                           | `ConfUInt32`              | 32-bit unsigned integer |
| `J_UINT64`              | uint64                           | `ConfUInt64`              | 64-bit unsigned integer |
| `J_IPV4`                | inet:ipv4-address                | `ConfIPv4`                | 64-bit unsigned         |
| `J_IPV6`                | inet:ipv6-address                | `ConfIPv6`                | IP v6 Address           |
| `J_BOOL`                | boolean                          | `ConfBoolean`             | Boolean value           |
| `J_QNAME`               | xs:QName                         | `ConfQName`               | A namespace/tag pair    |
| `J_DATETIME`            | yang:date-and-time               | `ConfDateTime`            | Date and Time Value     |
| `J_DATE`                | xs:date                          | `ConfDate`                | XML schema Date         |
| `J_ENUMERATION`         | enum                             | `ConfEnumeration`         | An enumeration value    |
| `J_BIT32`               | bits                             | `ConfBit32`               | 32 bit value            |
| `J_BIT64`               | bits                             | `ConfBit64`               | 64 bit value            |
| `J_LIST`                | leaf-list                        | `-`                       | -                       |
| `J_INSTANCE_IDENTIFIER` | instance-identifier              | `ConfObjectRef`           | yang builtin            |
| `J_OID`                 | tailf:snmp-oid                   | `ConfOID`                 | -                       |
| `J_BINARY`              | tailf:hex-list, tailf:octet-list | `ConfBinary, ConfHexList` | -                       |
| `J_IPV4PREFIX`          | inet:ipv4-prefix                 | `ConfIPv4Prefix`          | -                       |
| `J_IPV6PREFIX`          | -                                | `ConfIPv6Prefix`          | -                       |
| `J_IPV6PREFIX`          | inet:ipv6-prefix                 | `ConfIPv6Prefix`          | -                       |
| `J_DEFAULT`             | -                                | `ConfDefault`             | default value indicator |
| `J_NOEXISTS`            | -                                | `ConfNoExists`            | no value indicator      |
| `J_DECIMAL64`           | decimal64                        | `ConfDecimal64`           | yang builtin            |
| `J_IDENTITYREF`         | identityref                      | `ConfIdentityRef`         | yang builtin            |

An important class in the `com.tailf.conf` package, not inheriting `ConfObject`, is `ConfPath`. ConfPath is used to represent a keypath that can point to any element in an instantiated model. As such it is constructed from an array of `ConfObject[]` instances where each element is expected to be either a `ConfTag` or a `ConfKey`.

As an example take the keypath `/ncs:devices/device{d1}/iosxr:interface/Loopback{lo0}`. The following code snippets show the instantiating of a `ConfPath` object representing this keypath:

```
    ConfPath keyPath = new ConfPath(new ConfObject[] {
                                    new ConfTag("ncs","devices"),
                                    new ConfTag("ncs","device"),
                                    new ConfKey(new ConfObject[] {
                                                new ConfBuf("d1")}),
                                    new ConfTag("iosxr","interface"),
                                    new ConfTag("iosxr","Loopback"),
                                    new ConfKey(new ConfObject[] {
                                                new ConfBuf("lo0")})
                                    });
```

Another more commonly used option is to use the format string + arguments constructor from `ConfPath`. Where `ConfPath` parsers and creates the `ConfTag`/`ConfKey` representation from the string representation instead.

```
    // either this way
    ConfPath key1 = new ConfPath("/ncs:devices/device{d1}"+
                                 "/iosxr:interface/Loopback{lo0}"
    // or this way
    ConfPath key2 = new ConfPath("/ncs:devices/device{%s}"+
                                 "/iosxr:interface/Loopback{%s}",
                                 new ConfBuf("d1"),
                                 new ConfBuf("lo0"));
```

The usage of `ConfXMLParam` is in tagged value arrays `ConfXMLParam[]` of subtypes of `ConfXMLParam`. These can in collaboration represent an arbitrary YANG model subtree. It does not view a node as a path but instead, it behaves as an XML instance document representation. We have 4 subtypes of `ConfXMLParam`:

* `ConfXMLParamStart`: Represents an opening tag. Opening node of a container or list entry.
* `ConfXMLParamStop`: Represents a closing tag. The closing tag of a container or a list entry.
* `ConfXMLParamValue`: Represent a value and a tag. Leaf tag with the corresponding value.
* `ConfXMLParamLeaf`: Represents a leaf tag without the leafs value.

Each element in the array is associated with the node in the data model.

The array corresponding to the `/servers/server{www}` which is a representation of the instance XML document:

```xml
    <servers>
      <server>
        <name>www</name>
      </server>
    </servers>
```

The list entry above could be populated as:

```
    ConfXMLParam[] tree = new ConfXMLParam[] {
        new ConfXMLParamStart(ns.hash(),ns._servers),
        new ConfXMLParamStart(ns.hash(),ns._server),
        new ConfXMLParamValue(ns.hash(),ns._name),
        new ConfXMLParamStop(ns.hash(),ns._server),
        new ConfXMLParamStop(ns.hash,ns._servers)};
```

## Namespace Classes and the Loaded Schema <a href="#d5e4328" id="d5e4328"></a>

A namespace class represents the namespace for a YANG module. As such it maps the symbol name of each element in the YANG module to its corresponding hash value.

A namespace class is a subclass of `ConfNamespace` and comes in one of two shapes. Either created at compile time using the `ncsc` compiler or created at runtime with the use of `Maapi.loadSchemas`. These two types also indicate two main usages of namespace classes. The first is in programming where the symbol names are used e.g. in Navu navigation. This is where the compiled namespaces are used. The other is for internal mapping between symbol names and hash values. This is where the runtime type normally is used, however, compiled namespace classes can be used for these mappings too.

The compiled namespace classes are generated from compiled .fxs files through `ncsc`,(`ncsc --emit-java`).

```bash
ncsc --java-disable-prefix --java-package \
       com.example.app.namespaces \
       --emit-java \
       java/src/com/example/app/namespaces/foo.java \
       foo.fxs
```

Runtime namespace classes are created by calling `Maapi.loadschema()`. That's it, the rest is dynamic. All namespaces known by NSO are downloaded and runtime namespace classes are created. these can be retrieved by calling `Maapi.getAutoNsList()`.

```
    Maapi maapi = new Maapi(UnixDomainSocketAddress.of(Conf.NCS_PATH));
    maapi.loadSchemas();

    ArrayList<ConfNamespace> nsList = maapi.getAutoNsList();
```

The schema information is loaded automatically at the first connect of the NSO server, so no manual method call to `Maapi.loadSchemas()` is needed.

With all schemas loaded, the Java engine can make mappings between hash codes and symbol names on the fly. Also, the `ConfPath` class can find and add namespace information when parsing keypaths provided that the namespace prefixes are added in the start element for each namespace.

```
    ConfPath key1 = new ConfPath("/ncs:devices/device{d1}/iosxr:interface");
```

As an option, several APIs e.g. MAAPI can set the default namespace which will be the expected namespace for paths without prefixes. For example, if the namespace class `smp` is generated with the legal path `/smp:servers/server` an option in MAAPI could be the following:

```
    Maapi maapi = new Maapi(UnixDomainSocketAddress.of(Conf.NCS_PATH));
    int th =  maapi.startTrans(Conf.DB_CANDIDATE,
                               Conf.MODE_READ_WRITE);

    // Because we will use keypaths without prefixes
    maapi.setNamespace(th, new smp().uri());


    ConfValue val = maapi.getElem(th, "/devices/device{d1}/address");
```

## Advanced Topics

### Fetch bulk live-status via MAAPI

In MAAPI, when reading live-status data with `getElem`, `getCase`, `getValues`, and similar calls, each read call may trigger its own request to the device. To improve fetch performance, you can use `setReadIntent` to fetch live-status data in bulk.

{% code title="Example: Fetch bulk live-status via MAAPI" %}

```java
    Maapi maapi = new Maapi(UnixDomainSocketAddress.of(Conf.NCS_PATH));
    int th =  maapi.startTrans(Conf.DB_RUNNING, Conf.MODE_READ);

    maapi.setReadIntent(th, Arrays.asList(
                                "/ncs:devices/device/live-status/foo:foo",
                                "/ncs:devices/device/live-status/bar:bar"
                            ));
    maapi.getElem(th, "/ncs:devices/device{dev0}/live-status/foo:foo/a");
    maapi.getElem(th, "/ncs:devices/device{dev0}/live-status/foo:foo/b");

    maapi.getElem(th, "/ncs:devices/device{dev0}/live-status/bar:bar/baz");
```

{% endcode %}

With `setReadIntent`, it only requires one read request to get all the live-status data under `foo:foo` and `bar:bar`, and the data will be cached for later usage. To check the current read-intent, use `getReadIntent`, and `clearReadIntent` to clear it.


# Northbound APIs

Understand different types of northbound APIs and their working mechanism.

This section describes the various northbound programmatic APIs in NSO NETCONF, REST, and SNMP. These APIs are used by external systems that need to communicate with NSO, such as portals, OSS, or BSS systems.

NSO has two northbound interfaces intended for human usage, the CLI and the WebUI. These interfaces are described in [NSO CLI](/guides/operation-and-usage/operations) and [Web User Interface](/guides/operation-and-usage/webui) respectively.

There are also programmatic Java, Python, and Erlang APIs intended to be used by applications integrated with NSO itself. See [Running Application Code](https://nso-docs.cisco.com/guides/development/core-concepts/pages/sXgWFsCfe9tG1xRpMdkO#ncs.development.applications.running) for more information about these APIs.

## Integrating an External System with NSO <a href="#d5e48" id="d5e48"></a>

There are two APIs to choose from when an external system should communicate with NSO:

* NETCONF
* REST

Which one to choose is mostly a subjective matter. REST may, at first sight, appear to be simpler to use, but is not as feature-rich as NETCONF. By using a NETCONF client library such as the open source Java library [JNC](https://github.com/tail-f-systems/JNC) or Python library [ncclient](https://github.com/ncclient/ncclient), the integration task is significantly reduced.

Both NETCONF and REST provide functions for manipulating the configuration (including creating services) and reading the operational state from NSO. NETCONF provides more powerful filtering functions than REST.

NETCONF and SNMP can be used to receive alarms as notifications from NSO. NETCONF provides a reliable mechanism to receive notifications over SSH, whereas SNMP notifications are sent over UDP.

Regardless of the protocol you choose for integration, keep in mind all of them communicate with the NSO server over network sockets, which may be unreliable. Additionally, write transactions in NSO can fail if they conflict with another, concurrent transaction. As a best practice, the client implementation should be able to gracefully handle such errors and be prepared to retry requests. For details on the NSO concurrency, refer to the [NSO Concurrency Model.](/guides/development/core-concepts/nso-concurrency-model)


# NSO NETCONF Server

Description of northbound NETCONF implementation in NSO.

This section describes the northbound NETCONF implementation in NSO. As of this writing, the server supports the following specifications:

* [RFC 4741](https://www.ietf.org/rfc/rfc4741.txt): NETCONF Configuration Protocol
* [RFC 4742](https://www.ietf.org/rfc/rfc4742.txt): Using the NETCONF Configuration Protocol over Secure Shell (SSH)
* [RFC 5277](https://www.ietf.org/rfc/rfc5277.txt): NETCONF Event Notifications
* [RFC 5717](https://www.ietf.org/rfc/rfc5717.txt): Partial Lock Remote Procedure Call (RPC) for NETCONF
* [RFC 6020](https://www.ietf.org/rfc/rfc6020.txt): YANG - A Data Modeling Language for the Network Configuration Protocol (NETCONF)
* [RFC 6021](https://www.ietf.org/rfc/rfc6021.txt): Common YANG Data Types
* [RFC 6022](https://www.ietf.org/rfc/rfc6022.txt): YANG Module for NETCONF Monitoring
* [RFC 6241](https://www.ietf.org/rfc/rfc6241.txt): Network Configuration Protocol (NETCONF)
* [RFC 6242](https://www.ietf.org/rfc/rfc4742.txt): Using the NETCONF Configuration Protocol over Secure Shell (SSH)
* [RFC 6243](https://www.ietf.org/rfc/rfc6243.txt): With-defaults capability for NETCONF
* [RFC 6470](https://www.ietf.org/rfc/rfc6470.txt): NETCONF Base Notifications
* [RFC 6536](https://www.ietf.org/rfc/rfc6536.txt): NETCONF Access Control Model
* [RFC 6991](https://www.ietf.org/rfc/rfc6991.txt): Common YANG Data Types
* [RFC 7895](https://www.ietf.org/rfc/rfc7895.txt): YANG Module Library
* [RFC 7950](https://www.ietf.org/rfc/rfc7950.txt): The YANG 1.1 Data Modeling Language
* [RFC 8071](https://www.ietf.org/rfc/rfc8071.txt): NETCONF Call Home and RESTCONF Call Home
* [RFC 8342](https://www.ietf.org/rfc/rfc8342.txt): Network Management Datastore Architecture (NMDA)
* [RFC 8525](https://www.ietf.org/rfc/rfc8525.txt): YANG Library
* [RFC 8528](https://www.ietf.org/rfc/rfc8528.txt): YANG Schema Mount
* [RFC 8526](https://www.ietf.org/rfc/rfc8526.txt): NETCONF Extensions to Support the Network Management Datastore Architecture
* [RFC 8639](https://www.ietf.org/rfc/rfc8639.txt): Subscription to YANG Notifications
* [RFC 8640](https://www.ietf.org/rfc/rfc8640.txt): Dynamic Subscription to YANG Events and Datastores over NETCONF
* [RFC 8641](https://www.ietf.org/rfc/rfc8641.txt): Subscription to YANG Notifications for Datastore Updates

{% hint style="info" %}
For the `<delete-config>` operation specified in RFC 4741 / RFC 6241, only `<url>` with scheme `file` is supported for the `<target>` parameter - i.e. no data stores can be deleted. The concept of deleting a data store is not well defined and is at odds with the transaction-based configuration management of NSO. To delete the entire contents of a data store, with full transactional support, a `<copy-config>` with an empty `<config/>` element for the `<source>` parameter can be used.
{% endhint %}

{% hint style="info" %}
For the `<partial-lock>` operation, RFC 5717, section 2.4.1 says that if a node in the scope of the lock is deleted by the session owning the lock, it is removed from the scope of the lock. In NSO this is not true; the deleted node is kept in the scope of the lock.
{% endhint %}

NSO NETCONF northbound API can be used by arbitrary NETCONF clients. A simple Python-based NETCONF client called `netconf-console` is shipped as source code in the distribution. See [Using netconf-console](#ug.netconf_agent.netconf_console) for details. Other NETCONF clients will work too, as long as they adhere to the NETCONF protocol. If you need a Java client, the open-source client [JNC](https://github.com/tail-f-systems/JNC) can be used.

When integrating NSO into larger OSS/NMS environments, the NETCONF API is a good choice of integration point.

## Protocol Capabilities <a href="#d5e142" id="d5e142"></a>

The NETCONF server in NSO supports the following capabilities in both NETCONF 1.0 ([RFC 4741](https://www.ietf.org/rfc/rfc4741.txt)) and NETCONF 1.1 ([RFC 6241](https://www.ietf.org/rfc/rfc6241.txt)).

<table data-full-width="false"><thead><tr><th width="234" valign="top">Capability</th><th valign="top">Description</th></tr></thead><tbody><tr><td valign="top"><code>:writable-running</code></td><td valign="top">This capability is always advertised.</td></tr><tr><td valign="top"><code>:candidate</code></td><td valign="top">Not supported by NSO.</td></tr><tr><td valign="top"><code>:confirmed-commit</code></td><td valign="top">Not supported by NSO.</td></tr><tr><td valign="top"><code>:rollback-on-error</code></td><td valign="top">This capability allows the client to set the <code>&#x3C;error-option></code> parameter to <code>rollback-on-error</code>. The other permitted values are <code>stop-on-error</code> (default) and <code>continue-on-error</code>. Note that the meaning of the word "error" in this context is not defined in the specification. Instead, the meaning of this word must be defined by the data model. Also, note that if <code>stop-on-error</code> or <code>continue-on-error</code> is triggered by the server, it means that some parts of the edit operation succeeded, and some parts didn't. The error <code>partial-operation</code> must be returned in this case. <code>partial-operation</code> is obsolete and should not be returned by a server. If some other error occurs (i.e. an error not covered by the meaning of "error" above), the server generates an appropriate error message, and the data store is unaffected by the operation.<br><br>The NSO server never allows partial configuration changes, since it might result in inconsistent configurations, and recovery from such a state can be very difficult for a client. This means that regardless of the value of the <code>&#x3C;error-option></code> parameter, NSO will always behave as if it had the value <code>rollback-on-error</code>. So in NSO, the meaning of the word "error" in <code>stop-on-error</code> and <code>continue-on-error</code>, is something that never can happen.<br><br>It is possible to configure the NETCONF server to generate an <code>operation-not-supported</code> error if the client asks for the <code>error-option</code> <code>continue-on-error</code>. See <a href="/pages/Tis8ciGgCuwsxHHFpNud">ncs.conf(5)</a> in Manual Pages.</td></tr><tr><td valign="top"><code>:validate</code></td><td valign="top">NSO supports both version 1.0 and 1.1 of this capability.</td></tr><tr><td valign="top"><code>:startup</code></td><td valign="top">Not supported by NSO.</td></tr><tr><td valign="top"><code>:url</code></td><td valign="top"><p>The URL schemes supported are <code>file</code>, <code>ftp</code>, and <code>sftp</code> (SSH File Transfer Protocol). There is no standard URL syntax for the <em><code>sftp</code></em> scheme, but NSO supports the syntax used by <code>curl</code>:</p><pre><code>sftp://&#x3C;user>:&#x3C;password>@&#x3C;host>/&#x3C;path>
</code></pre><p>Note that user name and password must be given for <code>sftp</code> URLs. NSO does not support <code>validate</code> from a URL.</p></td></tr><tr><td valign="top"><code>:xpath</code></td><td valign="top">The NETCONF server supports XPath according to the W3C XPath 1.0 specification (<a href="https://www.w3.org/TR/xpath">https://www.w3.org/TR/xpath</a>).</td></tr></tbody></table>

The following list of optional standard capabilities is also supported:

<table><thead><tr><th width="237" valign="top">Capability</th><th valign="top">Description</th></tr></thead><tbody><tr><td valign="top"><code>:notification</code></td><td valign="top">NSO implements the <code>urn:ietf:params:netconf:capability:notification:1.0</code> capability, including support for the optional replay feature. See <a href="#ug.netconf_agent.notif">Notification Capability</a> for details.</td></tr><tr><td valign="top"><code>:with-defaults</code></td><td valign="top"><p>NSO implements the <code>urn:ietf:params:netconf:capability:with-defaults:1.0</code> capability, which is used by the server to inform the client how default values are handled by the server, and by the client to control whether default values should be generated to replies or not.</p><p>If the capability is enabled, NSO also implements the <code>urn:ietf:params:netconf:capability:with-operational-defaults:1.0</code> capability, which targets the operational state datastore while the <code>:with-defaults</code> capability targets configuration data stores.</p></td></tr><tr><td valign="top"><code>:yang-library:1.0</code></td><td valign="top">NSO implements the <code>urn:ietf:params:netconf:capability:yang-library:1.0</code> capability, which informs the client that the server implements the YANG module library <a href="https://www.ietf.org/rfc/rfc7895.txt">RFC 7895</a>, and informs the client about the current <code>module-set-id</code>.</td></tr><tr><td valign="top"><code>:yang-library:1.1</code></td><td valign="top">NSO implements the <code>urn:ietf:params:netconf:capability:yang-library:1.1</code> capability, which informs the client that the server implements the YANG library <a href="https://www.ietf.org/rfc/rfc8525.txt">RFC 8525</a>, and informs the client about the current <code>content-id</code>.</td></tr></tbody></table>

## Protocol YANG Modules <a href="#d5e255" id="d5e255"></a>

In addition to the protocol capabilities listed above, NSO also implements a set of YANG modules that are closely related to the protocol.

* `ietf-netconf-nmda`: This module from [RFC 8526](https://www.ietf.org/rfc/rfc8526.txt) defines the NMDA extension to NETCONF. It defines the following features:
* `origin`: Indicates that the server supports the origin annotation. It is not advertised by default. The support for `origin` can be enabled in `ncs.conf` (see [ncs.conf(5)](/guides/resources/man/ncs.conf.5) in Manual Pages ). If it is enabled, the `origin` feature is advertised.
* `with-defaults`: Advertised if the server supports the `:with-defaults` capability, which NSO does.
* `ietf-subscribed-notifications`: This module from [RFC 8639](https://www.ietf.org/rfc/rfc8639.txt) defines operations, configuration data nodes, and operational state data nodes related to notification subscriptions. It defines the following features:
* `configured`: Indicates that the server supports configured subscriptions. This feature is not advertised.
* `dscp`: Indicates that the server supports the ability to set the Differentiated Services Code Point (DSCP) value in outgoing packets. This feature is not advertised.
* `encode-json`: Indicates that the server supports JSON encoding of notifications. This is not applicable to NETCONF, and this feature is not advertised.
* `encode-xml`: Indicates that the server supports XML encoding of notifications. This feature is advertised by NSO.
* `interface-designation`: Indicates that a configured subscription can be configured to send notifications over a specific interface. This feature is not advertised.
* `qos`: Indicates that a publisher supports absolute dependencies of one subscription's traffic over another as well as weighted bandwidth sharing between subscriptions. This feature is not advertised.
* `replay`: Indicates that historical event record replay is supported. This feature is advertised by NSO.
* `subtree`: Indicates that the server supports subtree filtering of notifications. This feature is advertised by NSO.
* `supports-vrf`: Indicates that a configured subscription can be configured to send notifications from a specific VRF. This feature is not advertised.
* `xpath`: Indicates that the server supports XPath filtering of notifications. This feature is advertised by NSO.

In addition to this, NSO does not support pre-configuration or monitoring of subtree filters, and thus advertises a deviation module that deviates `/filters/stream-filter/filter-spec/stream-subtree-filter` and `/subscriptions/subscription/target/stream/stream-filter/within-subscription/filter-spec/stream-subtree-filter` as "not-supported".

NSO does not generate `subscription-modified` notifications when the parameters of a subscription change, and there is currently no mechanism to suspend notifications, so `subscription-suspended` and `subscription-resumed` notifications are never generated.

There is basic support for monitoring subscriptions via the `/subscriptions` container. Currently, it is possible to view dynamic subscriptions' attributes: `subscription-id`, `stream`, `encoding`, `receiver`, `stop-time`, and `stream-xpath-filter`. Unsupported attributes are: `stream-subtree-filter`, `receiver/sent-event-records`, `receiver/excluded-event-records`, and `receiver/state`.

* `ietf-yang-push`: This module from [RFC 8641](https://www.ietf.org/rfc/rfc8641.txt) extends operations, data nodes, and operational state defined in `ietf-subscribed-notifications;` and also introduces continuous and customizable notification subscriptions for updates from running and operational datastores. It defines the same features as `ietf-subscribed-notifications` and also the following feature:
  * `on-change`: Indicates that on-change triggered notifications are supported. This feature is advertised by NSO.
    * `dampening-period`: Indicates that dampening-period for on-change subscriptions is supported. This feature is advertised by NSO.
    * `sync-on-start`: Indicates that sync-on-start for on-change subscriptions is supported. This feature is advertised by NSO.
    * `excluded-change`: Indicates that excluded-change for on-change subscription is supported. This feature is advertised by NSO.
  * `periodic`: Indicates that periodic notifications are supported. This feature is advertised by NSO.
    * `period`: Indicates that period for periodic notifications are supported. This feature is advertised by NSO.
    * `anchor-time`: Indicates that anchor-time for periodic subscriptions is supported. This feature is advertised by NSO.

In addition to this, NSO does not support pre-configuration or monitoring of subtree filters and thus advertises a deviation module that deviates `/filters/selection-filter/filter-spec/datastore-subtree-filter` and `/subscriptions/subscription/target/datastore/selection-filter/within-subscription/filter-spec/datastore-subtree-filter` as "not-supported".

The monitoring of subscriptions via the `subscriptions` container currently does not support the attribute `/subscriptions/receivers/receiver/state` .

## Advertising Capabilities and YANG Modules <a href="#d5e376" id="d5e376"></a>

All enabled NETCONF capabilities are advertised in the hello message that the server sends to the client.

A YANG module is supported by the NETCONF server if its fxs file is found in NSO's loadPath, and if the fxs file is exported to NETCONF.

The following YANG modules are built-in, which means that their `fxs` files need not be present in the loadPath. If they are found in the loadPath they are skipped.

* `ietf-netconf`
* `ietf-netconf-with-defaults`
* `ietf-yang-library`
* `ietf-yang-types`
* `ietf-inet-types`
* `ietf-restconf`
* `ietf-datastores`
* `ietf-yang-patch`

All built-in modules are always supported by the server.

All YANG version 1 modules supported by the server are advertised in the hello message, according to the rules defined in [RFC 6020](https://www.ietf.org/rfc/rfc6020.txt).

All YANG version 1 and version 1.1 modules supported by the server are advertised in the YANG library.

If a YANG module (any version) is supported by the server, and its .yang or .yin file is found in the `fxs` file or in the loadPath, then the module is also advertised in the `schema` list defined in `ietf-netconf-monitoring`, made available for download with the RPC operation `get-schema`, and if RESTCONF is enabled, also advertised in the `schema` leaf in `ietf-yang-library`. See [Monitoring of the NETCONF Server](#ug.netconf_agent.monitoring).

### Advertising Device YANG Modules <a href="#d5e417" id="d5e417"></a>

NSO uses [YANG Schema Mount](https://www.ietf.org/rfc/rfc8528.txt) to mount the data models for the devices. There are two mount points, one for the configuration (in `/devices/device/config`), and one for operational state data (in `/devices/device/live-status`). As defined in [YANG Schema Mount](https://www.ietf.org/rfc/rfc8528.txt), a client can read the `module` list from the YANG library in each of these mount points to learn which YANG models each device supports via NSO.

For example, to get the YANG library data for the device `x0`, we can do:

```
$ netconf-console --get -x '/devices/device[name="x0"]/config/yang-library'
<?xml version="1.0" encoding="UTF-8"?>
<rpc-reply xmlns="urn:ietf:params:xml:ns:netconf:base:1.0" message-id="1">
  <data>
    <devices xmlns="http://tail-f.com/ns/ncs">
      <device>
        <name>x0</name>
        <config>
          <yang-library xmlns="urn:ietf:params:xml:ns:yang:ietf-yang-library">
            <module-set>
              <name>common</name>
              <module>
                <name>a</name>
                <namespace>urn:a</namespace>
              </module>
              <module>
                <name>b</name>
                <namespace>urn:b</namespace>
              </module>
            </module-set>
            <schema>
              <name>common</name>
              <module-set>common</module-set>
            </schema>
            <datastore>
              <name xmlns:ds="urn:ietf:params:xml:ns:yang:ietf-datastores">\
                 ds:running\
              </name>
              <schema>common</schema>
            </datastore>
            <datastore>
              <name xmlns:ds="urn:ietf:params:xml:ns:yang:ietf-datastores">\
                ds:intended\
              </name>
              <schema>common</schema>
            </datastore>
            <datastore>
              <name xmlns:ds="urn:ietf:params:xml:ns:yang:ietf-datastores">\
                ds:operational\
              </name>
              <schema>common</schema>
            </datastore>
            <content-id>f0071b28c1e586f2e8609da036379a58</content-id>
          </yang-library>
        </config>
      </device>
    </devices>
  </data>
</rpc-reply>
```

The set of modules reported for a device is the set of modules that NSO knows, i.e., the set of modules compiled for the specific device type. This means that all devices of the same device type will report the same set of modules. Also, note that the device may support other modules that are not known to NSO. Such modules are not reported here.

## NETCONF Transport Protocols <a href="#ug.netconf_agent.transport" id="ug.netconf_agent.transport"></a>

The NETCONF server natively supports the mandatory SSH transport, i.e., SSH is supported without the need for an external SSH daemon (such as `sshd`). It also supports integration with OpenSSH.

### Using OpenSSH <a href="#d5e432" id="d5e432"></a>

NSO is delivered with a program **netconf-subsys** which is an OpenSSH subsystem program. It is invoked by the OpenSSH daemon after successful authentication. It functions as a relay between the ssh daemon and NSO; it reads data from the ssh daemon from standard input and writes the data to NSO over a socket connection, and vice versa. This program is delivered as source code in `$NCS_DIR/src/ncs/netconf/netconf-subsys.c`. It can be modified to fit the needs of the application. For example, it could be modified to read the group names for a user from an external LDAP server.

When using OpenSSH, the users are authenticated by OpenSSH, i.e., the user names are not stored in NSO. To use OpenSSH, compile the `netconf-subsys` program, and put the executable in e.g. `/usr/local/bin`. Then add the following line to the ssh daemon's config file, `sshd_config`:

```
Subsystem     netconf   /usr/local/bin/netconf-subsys
```

The connection from `netconf-subsys` to NSO can be arranged in one of two different ways:

1. Make sure NSO is configured to listen to TCP traffic on localhost, port 2023, and disable SSH in `ncs.conf` (see [ncs.conf(5)](/guides/resources/man/ncs.conf.5) in Manual Pages ). (Re)start `sshd` and NSO. Or:
2. Compile `netconf-subsys` to use a connection to the IPC socket instead of the NETCONF TCP transport (see the `netconf-subsys.c` source for details), and disable both TCP and SSH in `ncs.conf`. (Re)start `sshd` and NSO. This method may be preferable since it makes it possible to use the IPC Access Check (see [Restricting Access to the IPC Socket](/guides/administration/advanced-topics/ipc-connection#restricting-access-to-the-ipc-socket)) to restrict the unauthenticated access to NSO that is needed by `netconf-subsys`.

By default, the `netconf-subsys` program sends the names of the UNIX groups the authenticated user belongs to. To test this, make sure that NSO is configured to give access to the group(s) the user belongs to. The easiest for test is to give access to all groups.

## Configuring the NETCONF Server <a href="#d5e461" id="d5e461"></a>

NSO itself is configured through a configuration file called `ncs.conf`. For a description of the parameters in this file, please see the [ncs.conf(5)](/guides/resources/man/ncs.conf.5) in Manual Pages man page.

### Error Handling <a href="#d5e466" id="d5e466"></a>

When NSO processes `<get>`, `<get-config>`, and `<copy-config>` requests, the resulting data set can be very large. To avoid buffering huge amounts of data, NSO streams the reply to the client as it traverses the data tree and calls data provider functions to retrieve the data.

If a data provider fails to return the data it is supposed to return, NSO can take one of two actions. Either it simply closes the NETCONF transport (default), or it can reply with an inline RPC error and continue to process the next data element. This behavior can be controlled with the `/ncs-config/netconf/rpc-errors` configuration parameter (see [ncs.conf(5)](/guides/resources/man/ncs.conf.5) in Manual Pages).

An inline error is always generated as a child element to the parent of the faulty element. For example, if an error occurs when retrieving the leaf element `mac-address` of an `interface` the error might be:

```xml
<interface>
  <name>atm1</name>
  <rpc-error xmlns="urn:ietf:params:xml:ns:netconf:base:1.0">
    <error-type>application</error-type>
    <error-tag>operation-failed</error-tag>
    <error-severity>error</error-severity>
    <error-message xml:lang="en">Failed to talk to hardware</error-message>
    <error-info>
      <bad-element>mac-address</bad-element>
    </error-info>
  </rpc-error>
  ...
</interface>
```

If a `get_next` call fails in the processing of a list, a reply might look like this:

```xml
<interface>
  <!-- successfully retrieved list entry -->
  <name>eth0</name>
  <mtu>1500</mtu>
  <!-- more leafs here -->
</interface>
<rpc-error xmlns="urn:ietf:params:xml:ns:netconf:base:1.0">
  <error-type>application</error-type>
  <error-tag>operation-failed</error-tag>
  <error-severity>error</error-severity>
  <error-message xml:lang="en">Failed to talk to hardware</error-message>
  <error-info>
    <bad-element>interface</bad-element>
  </error-info>
</rpc-error>
```

## Using `netconf-console` <a href="#ug.netconf_agent.netconf_console" id="ug.netconf_agent.netconf_console"></a>

The `netconf-console` program is a simple NETCONF client. It is delivered as Python source code and can be used as-is or modified.

When NSO has been started, we can use `netconf-console` to query the configuration of the NETCONF Access Control groups:

```
$ netconf-console --get-config -x /nacm/groups
<?xml version="1.0" encoding="UTF-8"?>
<rpc-reply xmlns="urn:ietf:params:xml:ns:netconf:base:1.0" message-id="1">
  <data>
    <nacm xmlns="urn:ietf:params:xml:ns:yang:ietf-netconf-acm">
      <groups>
        <group>
          <name>admin</name>
          <user-name>admin</user-name>
          <user-name>private</user-name>
        </group>
        <group>
          <name>oper</name>
          <user-name>oper</user-name>
          <user-name>public</user-name>
        </group>
      </groups>
    </nacm>
  </data>
</rpc-reply>
```

With the `-x` flag an XPath expression can be specified, to retrieve only data matching that expression. This is a very convenient way to extract portions of the configuration from the shell or from shell scripts.

## Monitoring the NETCONF Server <a href="#ug.netconf_agent.monitoring" id="ug.netconf_agent.monitoring"></a>

[RFC 6022 - YANG Module for NETCONF Monitoring](https://www.ietf.org/rfc/rfc6022.txt) defines a YANG module, `ietf-netconf-monitoring`for monitoring of the NETCONF server. It contains statistics objects such as the number of RPCs received, status objects such as user sessions, and an operation to retrieve data models from the NETCONF server.

This data model defines an RPC operation, `get-schema`, which is used to retrieve YANG modules from the NETCONF server. NSO will report the YANG modules for all fxs files that are reported as capabilities, and for which the corresponding YANG or YIN file is stored in the fxs file or found in the loadPath. If a file is found in the loadPath, it has priority over a file stored in the `fxs` file. Note that by default, the module and its submodules are stored in the `fxs` file by the compiler.

If the YANG (or YIN files) are copied into the loadPath, they can be stored as is or compressed with gzip. The filename extension MUST be `.yang`, `.yin`, `.yang.gz`, or `.yin.gz`.

Also available is a Tail-f-specific data model, `tailf-netconf-monitoring`, which augments `ietf-netconf-monitoring` with additional data about files available for usage with the `<copy-config>` command with a `file` `<url>` source or target. `/ncs-config/netconf-north-bound/capabilities/url/enabled` and `/ncs-config/netconf-north-bound/capabilities/url/file/enabled` must both be set to true. If rollbacks are enabled, those files are listed as well, and they can be loaded using `<copy-config>`.

This data model also adds data about which notification streams are present in the system and data about sessions that subscribe to the streams.

## Notification Capability <a href="#ug.netconf_agent.notif" id="ug.netconf_agent.notif"></a>

This section describes how NETCONF notifications are implemented within NSO, and how the applications generate these events.

Central to NETCONF notifications is the concept of a stream. The stream serves two purposes. It works like a high-level filtering mechanism for the client. For example, if the client subscribes to notifications on the `security` stream, it can expect to get security-related notifications only. Second, each stream may have its own log mechanism. For example, by keeping all debug notifications in a `debug` stream, they can be logged separately from the `security` stream.

### Built-in Notification Streams <a href="#d5e521" id="d5e521"></a>

NSO has built-in support for the well-known stream `NETCONF`, defined in [RFC 5277](https://www.ietf.org/rfc/rfc5277.txt) and [RFC 8639](https://www.ietf.org/rfc/rfc8639.txt). NSO supports the notifications defined in [RFC 6470 - NETCONF Base Notifications](https://www.ietf.org/rfc/rfc6470.txt) on this stream. If the application needs to send any additional notifications on this stream, it can do so.

NSO can be configured to listen to notifications from devices and send those notifications to northbound NETCONF clients. The stream `device-notifications` is used for this purpose. To enable this, the stream `device-notifications` must be configured in `ncs.conf`, and additionally, subscriptions must be created in `/ncs:devices/device/notifications`.

### Defining Notification Streams <a href="#d5e533" id="d5e533"></a>

It is up to the application to define which streams it supports. In NSO, this is done in `ncs.conf` (see [ncs.conf(5)](/guides/resources/man/ncs.conf.5) in Manual Pages). Each stream must be listed, and whether it supports replay or not. The following example enables the built-in stream `device-notifications` with replay support, and an additional, application-specific stream `debug` without replay support:

```xml
<notifications>
  <event-streams>
    <stream>
      <name>device-notifications</name>
      <description>Notifications received from devices</description>
      <replay-support>true</replay-support>
      <builtin-replay-store>
        <enabled>true</enabled>
        <dir>/var/log</dir>
        <max-size>S10M</max-size>
        <max-files>50</max-files>
      </builtin-replay-store>
    </stream>
    <stream>
      <name>debug</name>
      <description>Debug notifications</description>
      <replay-support>false</replay-support>
    </stream>
  </event-streams>
</notifications>
```

The well-known stream `NETCONF` does not have to be listed, but if it isn't listed, it will not support replay.

### Automatic Replay <a href="#d5e544" id="d5e544"></a>

NSO has built-in support for logging of notifications, i.e., if replay support has been enabled for a stream, NSO automatically stores all notifications on disk ready to be replayed should a NETCONF client ask for logged notifications. In the `ncs.conf` fragment above the security stream has been set up to use the built-in notification log/replay store. The replay store uses a set of wrapping log files on a disk (of a certain number and size) to store the security stream notifications.

The reason for using a wrap log is to improve replay performance whenever a NETCONF client asks for notifications in a certain time range. Any problems with log files not being properly closed due to hard power failures etc. are also kept to a minimum, i.e., automatically taken care of by NSO.

## Subscribed Notifications <a href="#ug.netconf_agent.subscribed_notif" id="ug.netconf_agent.subscribed_notif"></a>

This section describes how Subscribed Notifications are implemented for NETCONF within NSO.

Subscribed Notifications is defined in [RFC 8639](https://www.ietf.org/rfc/rfc8639.txt) and the NETCONF transport binding is defined in [RFC 8640](https://www.ietf.org/rfc/rfc8640.txt). Subscribed Notifications build upon NETCONF notifications defined in [RFC 5277](https://www.ietf.org/rfc/rfc5277.txt) and have a number of key improvements:

* Multiple subscriptions on a single transport session
* Support for dynamic and configured subscriptions
* Modification of an existing subscription in progress
* Per-subscription operational counters
* Negotiation of subscription parameters (through the use of hints returned as part of declined subscription requests)
* Subscription state change notifications (e.g., publisher-driven suspension, parameter modification)
* Independence from transport

### Compatibility with NETCONF Notifications <a href="#d5e571" id="d5e571"></a>

Both NETCONF notifications and Subscribed Notifications can be used at the same time and are configured the same way in `ncs.conf`. However, there are some differences and limitations.

For Subscribed Notifications, a new subscription is requested by invoking the RPC `establish-subscription`. For NETCONF notifications, the corresponding RPC is `create-subscription`.

A NETCONF session can only have either the subscribers started with `create-subscription` or `establish-subscription` simultaneously.

* If a session has subscribers established with `establish-subscription` and receives a request to create subscriptions with `create-subscription`, an `<rpc-error>` is sent containing `<error-tag>` `operation-not-supported`.

  If a session has subscribers created with `create-subscription` and receives a request to establish subscriptions with `establish-subscription`, an `<rpc-error>` is sent containing `<error-tag>` `operation-not-supported`.

Dynamic subscriptions send all notifications on the transport session where they were established.

### Monitoring Subscriptions <a href="#ug.netconf_agent.subscribed_notif.monitoring" id="ug.netconf_agent.subscribed_notif.monitoring"></a>

Existing subscriptions and their configuration can be found in the `/subscriptions` container.

For example, for viewing all established subscriptions, we can do:

```
$ netconf-console --get -x /subscriptions
<?xml version="1.0" encoding="UTF-8"?>
<rpc-reply xmlns="urn:ietf:params:xml:ns:netconf:base:1.0" message-id="1">
  <data>
    <subscriptions xmlns="urn:ietf:params:xml:ns:yang:ietf-subscribed-notifications">
      subscription>
       <id>3</id>
       <stream-xpath-filter>/if:interfaces/interface[name='eth0']/enabled</stream-xpath-filter>
       <stream>interface</stream>
       <stop-time>2030-10-04T14:00:00+02:00</stop-time>
       <encoding>encode-xml</encoding>
       <receivers>
         <receiver>
           <name>127.0.0.1:57432</name>
           <state>active</state>
         </receiver>
       </receivers>
      /subscription>
    </subsrcriptions>
  </data>
</rpc-reply>
```

### **Limitations**

It is not possible to establish a subscription with a stored filter from `/filters`.

The support for monitoring subscriptions has basic functionality. It is possible to read `subscription-id`, `stream`, `stream-xpath-filter`, `replay-start-time`, `stop-time`, `encoding`, `receivers/receiver/name`, and `receivers/receiver/state`.

The leaf `stream-subtree-filter` is deviated as "not-supported", hence can not be read.

The unsupported leafs in the subscriptions container are the following: `stream-subtree-filter`, `receiver/sent-event-records`, and `receiver/excluded-event-records`.

## YANG-Push <a href="#ug.netconf_agent.yang_push" id="ug.netconf_agent.yang_push"></a>

This section describes how YANG-Push is implemented for NETCONF within NSO.

YANG-Push is defined in [RFC 8641](https://www.ietf.org/rfc/rfc8641.txt) and the NETCONF transport binding is defined in [RFC 8640](https://www.ietf.org/rfc/rfc8640.txt). YANG-Push implementation in NSO introduces a subscription service that provides updates from a datastore. This implementation supports dynamic subscriptions on updates of datastore nodes. A subscribed receiver is provided with update notifications according to the terms of the subscription. There are two types of notification messages defined to provide updates and these are used according to subscription terms.

* `push-update` notification is a complete, filtered update that reflects the data of the subscribed datastore. It is the type of notification that is used for `periodic` subscriptions. A `push-update` notification can also be used for the `on-change` subscriptions in case of a receiver asks for synchronization, either at the start of a new subscription or by sending a resync request for an established subscription.

  An example `push-update` notification:

  ```xml
  <notification xmlns="urn:ietf:params:xml:ns:netconf:notification:1.0">
    <eventTime>2020-06-10T10:00:00.00Z</eventTime>
    <push-update xmlns="urn:ietf:params:xml:ns:yang:ietf-yang-push">
      <id>1</id>
      <datastore-contents>
        <interfaces xmlns="urn:ietf:params:xml:ns:yang:ietf-interfaces">
          <interface>
            <name>eth0</name>
            <oper-status>up</oper-status>
          </interface>
        </interfaces>
      </datastore-contents>
    </push-update>
  </notification>
  ```
* `push-change-update` notification is the most common type of notification that is used for `on-change` subscriptions. It provides a set of filtered changes that happened on the subscribed datastore since the last update notification. The update records are constructed in the form of `YANG-Patch Media Type` that is defined in [RFC 8072](https://www.ietf.org/rfc/rfc8072.txt).

  \
  An example `push-change-update` notification:

  ```xml
  <notification xmlns="urn:ietf:params:xml:ns:netconf:notification:1.0">
    <eventTime>2020-06-10T10:05:00.00Z</eventTime>
    <push-change-update
      xmlns="urn:ietf:params:xml:ns:yang:ietf-yang-push">
      <id>2</id>
      <datastore-changes>
        <yang-patch>
          <patch-id>s2-p4</patch-id>
          <edit>
            <edit-id>edit1</edit-id>
            <operation>merge</operation>
            <target>/ietf-interfaces:interfaces</target>
            <value>
              <interfaces xmlns="urn:ietf:params:xml:ns:yang:ietf-interfaces">
                <interface>
                  <name>eth0</name>
                  <oper-status>down</oper-status>
                </interface>
              </interfaces>
            </value>
          </edit>
        </yang-patch>
      </datastore-changes>
    </push-change-update>
  </notification>
  ```

### Periodic Subscriptions <a href="#d5e649" id="d5e649"></a>

For periodic subscriptions, updates are triggered periodically according to specified time interval. Optionally a reference `anchor-time` can be provided for a specified `period`.

### On-Change Subscriptions <a href="#d5e654" id="d5e654"></a>

For on-change subscriptions, updates are triggered whenever a change is detected on the subscribed information. In the case of rapidly changing data, instead of receiving frequent notifications for every change, a receiver may specify a `dampening-period` to receive update notifications in a lower frequency. A receiver may request for synchronization at the start of a subscription by using `sync-on-start` option. A receiver may filter out specific types of changes by providing a list of `excluded-change` parameters.

To provide updates for `on-change` subscriptions on `operational` datastore, data provider applications are required to implement push-on-change callbacks. For more details, see the [PUSH ON-CHANGE CALLBACKS](/guides/resources/man/confd_lib_dp.3#push-on-change-callbacks) in the Manual Pages section of [confd\_lib\_dp(3)](/guides/resources/man/confd_lib_dp.3) in Manual Pages.

### YANG-Push Operations <a href="#d5e665" id="d5e665"></a>

In addition to RPCs defined in subscribed notifications, YANG-Push defines `resync-subscription` RPC. Upon receipt of `resync-subscription`, if the subscription is an on-change triggered type, a `push-update` notification is sent to the receiver according to the terms of the subscription. Otherwise, an appropriate error response is sent.

* `resync-subscription`

### Monitoring the YANG-Push Subscriptions <a href="#d5e675" id="d5e675"></a>

YANG-Push subscriptions can be monitored in a similar way to Subscribed Notifications through /subscriptions container. For more information, see [Monitoring Subscriptions](#ug.netconf_agent.subscribed_notif.monitoring).

YANG-Push filters differ from the filters of Subscribed Notifications and they are specified as `datastore-xpath-filter` and `datastore-subtree-filter`. The leaf `datastore-subtree-filter` is deviated as "not-supported", and hence can not be monitored. Also, YANG-Push specific update trigger parameters `periodic/period`, `periodic/anchor-time`, `on-change/dampening-period`, `on-change/sync-on-start` and `on-change/excluded-change` are not supported for monitoring.

### Limitations <a href="#d5e688" id="d5e688"></a>

* `modify-subscriptions` operation does not support changing a subscriptions update trigger type from `periodic` to `on-change` or vice versa.
* `on-change` subscriptions do not work for changes that are made through the CDB-API.
* `on-change` subscriptions do not work on internal callpoints such as `ncs-state`, `ncs-high-availability`, and `live-status`.

## Actions Capability <a href="#ug.netconf_agent.actions_ncs" id="ug.netconf_agent.actions_ncs"></a>

{% hint style="info" %}
This capability is deprecated since actions are now supported in standard YANG 1.1. It is recommended to use standard YANG 1.1 for actions.
{% endhint %}

This capability introduces a new RPC operation that is used to invoke actions defined in the data model. When an action is invoked, the instance on which the action is invoked is explicitly identified by a hierarchy of configuration or state data.

Here is a simple example that invokes the action `sync-from` on the device `ce1`. It uses the `netconf-console` command:

```
$ cat ./sync-from-ce1.xml
<action xmlns="http://tail-f.com/ns/netconf/actions/1.0">
  <data>
    <devices xmlns="http://tail-f.com/ns/ncs">
      <device>
        <name>ce1</name>
        <sync-from/>
      </device>
    </devices>
  </data>
</action>
$ netconf-console --rpc sync-from-ce1.xml
<?xml version="1.0" encoding="UTF-8"?>
<rpc-reply xmlns="urn:ietf:params:xml:ns:netconf:base:1.0" message-id="1">
  <data>
    <devices xmlns="http://tail-f.com/ns/ncs">
      <device>
        <name>ce1</name>
        <sync-from>
          <result>true</result>
        </sync-from>
      </device>
    </devices>
  </data>
</rpc-reply>
```

### Capability Identifier <a href="#d5e717" id="d5e717"></a>

The actions capability is identified by the following capability string:

```
  http://tail-f.com/ns/netconf/actions/1.0
```

## Transactions Capability <a href="#ug.netconf_agent.transactions" id="ug.netconf_agent.transactions"></a>

This capability introduces four new RPC operations that are used to control a two-phase commit transaction on the NETCONF server. The normal `<edit-config>` operation is used to write data in the transaction, but the modifications are not applied until an explicit `<commit-transaction>` is sent.

This capability is formally defined in the YANG module `tailf-netconf-transactions`. It is recommended that this module be enabled.

A typical sequence of operations looks like this:

```
               C                           S
               |                           |
               |  capability exchange      |
               |-------------------------->|
               |<------------------------->|
               |                           |
               |   <start-transaction>     |
               |-------------------------->|
               |<--------------------------|
               |         <ok/>             |
               |                           |
               |     <edit-config>         |
               |-------------------------->|
               |<--------------------------|
               |         <ok/>             |
               |                           |
               |  <prepare-transaction>    |
               |-------------------------->|
               |<--------------------------|
               |         <ok/>             |
               |                           |
               |   <commit-transaction>    |
               |-------------------------->|
               |<--------------------------|
               |         <ok/>             |
               |                           |
```

### Dependencies <a href="#d5e731" id="d5e731"></a>

None.

### Capability Identifier <a href="#d5e734" id="d5e734"></a>

The transactions capability is identified by the following capability string:

```
  http://tail-f.com/ns/netconf/transactions/1.0
```

### New Operation: `<start-transaction>` <a href="#d5e739" id="d5e739"></a>

#### **Description**

Starts a transaction towards a configuration datastore. There can be a single ongoing transaction per session at any time.

When a transaction has been started, the client can send any NETCONF operation, but any `<edit-config>` or `<copy-config>` operation sent from the client must specify the same `<target>` as the `<start-transaction>`, and any `<get-config>` must specify the same \<source> as `<start-transaction>`.

If the server receives an `<edit-config>` or `<copy-config>` with another `<target>`, or a `<get-config>` with another `<source>`, an error must be returned with an `<error-tag>` set to `invalid-value`.

The modifications sent in the `<edit-config>` operations are not immediately applied to the configuration datastore. Instead, they are kept in the transaction state of the server. The transaction state is only applied when a `<commit-transaction>` is received.

The client sends a `<prepare-transaction>` when all modifications have been sent.

#### **Parameters**

* `target:`\
  Name of the configuration datastore towards which the transaction is started.
* `with-inactive:`\
  If this parameter is given, the transaction will handle the `inactive` and `active` attributes. If given, it must also be given in the `<edit-config>` and `<get-config>` invocations in the transaction.

#### **Positive Response**

If the device can satisfy the request, an `<rpc-reply>` is sent that contains an `<ok>` element.

#### **Negative Response**

An `<rpc-error>` element is included in the `<rpc-reply>` if the request cannot be completed for any reason.

If there is an ongoing transaction for this session already, an error must be returned with `<error-app-tag>` set to `bad-state`.

#### **Example**

```xml
  <rpc message-id="101"
       xmlns="urn:ietf:params:xml:ns:netconf:base:1.0">
    <start-transaction xmlns="http://tail-f.com/ns/netconf/transactions/1.0">
      <target>
       <running/>
      </target>
    </start-transaction>
  </rpc>

  <rpc-reply message-id="101"
       xmlns="urn:ietf:params:xml:ns:netconf:base:1.0">
    <ok/>
  </rpc-reply>
```

### New Operation: `<prepare-transaction>` <a href="#d5e772" id="d5e772"></a>

#### **Description**

Prepares the transaction state for commit. The server may reject the prepare request for any reason, for example, due to lack of resources or if the combined changes would result in an invalid configuration datastore.

After a successful `<prepare-transaction>`, the next transaction-related RPC operation must be `<commit-transaction>` or `<abort-transaction>`. Note that an `<edit-config>` cannot be sent before the transaction is either committed or aborted.

Care must be taken by the server to make sure that if `<prepare-transaction>` succeeds then the `<commit-transaction>` should not fail, since this might result in an inconsistent distributed state. Thus, `<prepare-transaction>` should allocate any resources needed to make sure the `<commit-transaction>` will succeed.

#### **Parameters**

None.

#### **Positive Response**

If the device was able to satisfy the request, an `<rpc-reply>` is sent that contains an `<ok>` element.

#### **Negative Response**

An `<rpc-error>` element is included in the `<rpc-reply>` if the request cannot be completed for any reason.

If there is no ongoing transaction in this session, or if the ongoing transaction already has been prepared, an error must be returned with `<error-app-tag>` set to `bad-state`.

#### **Example**

```xml
  <rpc message-id="103"
       xmlns="urn:ietf:params:xml:ns:netconf:base:1.0">
    <prepare-transaction
       xmlns="http://tail-f.com/ns/netconf/transactions/1.0"/>
  </rpc>

  <rpc-reply message-id="103"
       xmlns="urn:ietf:params:xml:ns:netconf:base:1.0">
    <ok/>
  </rpc-reply>
```

### New Operation: `<commit-transaction>` <a href="#d5e793" id="d5e793"></a>

#### **Description**

Applies the changes made in the transaction to the configuration datastore. The transaction is closed after a `<commit-transaction>`.

#### **Parameters**

None.

#### **Positive Response**

If the device was able to satisfy the request, an `<rpc-reply>` is sent that contains an `<ok>` element.

#### **Negative Response**

An `<rpc-error>` element is included in the `<rpc-reply>` if the request cannot be completed for any reason.

If there is no ongoing transaction in this session, or if the ongoing transaction already has not been prepared, an error must be returned with `<error-app-tag>` set to `bad-state`.

#### **Example**

```xml
  <rpc message-id="104"
       xmlns="urn:ietf:params:xml:ns:netconf:base:1.0">
    <commit-transaction
       xmlns="http://tail-f.com/ns/netconf/transactions/1.0"/>
  </rpc>

  <rpc-reply message-id="104"
       xmlns="urn:ietf:params:xml:ns:netconf:base:1.0">
    <ok/>
  </rpc-reply>
```

### New Operation: `<abort-transaction>` <a href="#d5e812" id="d5e812"></a>

#### **Description**

Aborts the ongoing transaction, and all pending changes are discarded. `<abort-transaction>` can be given at any time during an ongoing transaction.

#### **Parameters**

None.

#### **Positive Response**

If the device was able to satisfy the request, an `<rpc-reply>` is sent that contains an `<ok>` element.

#### **Negative Response**

An `<rpc-error>` element is included in the `<rpc-reply>` if the request cannot be completed for any reason.

If there is no ongoing transaction in this session, an error must be returned with `<error-app-tag>` set to `bad-state`.

#### **Example**

```xml
  <rpc message-id="104"
       xmlns="urn:ietf:params:xml:ns:netconf:base:1.0">
    <abort-transaction
       xmlns="http://tail-f.com/ns/netconf/transactions/1.0"/>
  </rpc>

  <rpc-reply message-id="104"
       xmlns="urn:ietf:params:xml:ns:netconf:base:1.0">
    <ok/>
  </rpc-reply>
```

### Modifications to Existing Operations <a href="#d5e831" id="d5e831"></a>

The `<edit-config>` operation is modified so that if it is received during an ongoing transaction, the modifications are not immediately applied to the configuration target. Instead, they are kept in the transaction state of the server. The transaction state is only applied when a `<commit-transaction>` is received.

Note that it doesn't matter if the `<test-option>` is 'set' or 'test-then-set' in the `<edit-config>`, since nothing is actually set when the `<edit-config>` is received.

## Inactive Capability <a href="#ug.netconf_agent.inactive" id="ug.netconf_agent.inactive"></a>

This capability is used by the NETCONF server to indicate that it supports marking nodes as being inactive. A node that is marked as inactive exists in the data store but is not used by the server. Any node can be marked as inactive.

To not confuse clients who do not understand this attribute, the client has to instruct the server to display and handle the inactive nodes. An inactive node is marked with an `inactive` XML attribute, and to make it active, the `active` XML attribute is used.

This capability is formally defined in the YANG module `tailf-netconf-inactive`.

### Dependencies <a href="#d5e842" id="d5e842"></a>

None.

### Capability Identifier

The inactive capability is identified by the following capability string:

```
  http://tail-f.com/ns/netconf/inactive/1.0
```

### New Operations <a href="#d5e850" id="d5e850"></a>

None.

### Modifications to Existing Operations <a href="#d5e853" id="d5e853"></a>

A new parameter, `<with-inactive>`, is added to the `<get>`, `<get-config>`, `<edit-config>`, `<copy-config>`, and `<start-transaction>` operations.

The `<with-inactive>` element is defined in the <http://tail-f.com/ns/netconf/inactive/1.0> namespace, and takes no value.

If this parameter is present in `<get>`, `<get-config>`, or `<copy-config>`, the NETCONF server will mark inactive nodes with the `inactive` attribute.

If this parameter is present in `<edit-config>` or `<copy-config>`, the NETCONF server will treat inactive nodes as existing so that an attempt to create a node that is inactive will fail, and an attempt to delete a node that is inactive will succeed. Further, the NETCONF server accepts the `inactive` and `active` attributes in the data hierarchy, to make nodes inactive or active, respectively.

If the parameter is present in `<start-transaction>`, it must also be present in any `<edit-config>`, `<copy-config>`, `<get>`, or `<get-config>` operations within the transaction. If it is not present in `<start-transaction>`, it must not be present in any `<edit-config>` operation within the transaction.

The `inactive` and `active` attributes are defined in the <http://tail-f.com/ns/netconf/inactive/1.0> namespace. The `inactive` attribute's value is the string `inactive`, and the `active` attribute's value is the string `active`.

#### **Example**

This request creates an `inactive` interface:

```xml
  <rpc message-id="101"
       xmlns="urn:ietf:params:xml:ns:netconf:base:1.0">
    <edit-config>
      <target>
        <running/>
      </target>
      <with-inactive
         xmlns="http://tail-f.com/ns/netconf/inactive/1.0"/>
      <config>
        <top xmlns="http://example.com/schema/1.2/config">
          <interface inactive="inactive">
            <name>Ethernet0/0</name>
            <mtu>1500</mtu>
          </interface>
        </top>
      </config>
    </edit-config>
  </rpc>

  <rpc-reply message-id="101"
       xmlns="urn:ietf:params:xml:ns:netconf:base:1.0">
    <ok/>
  </rpc-reply>
```

This request shows the `inactive` interface:

```xml
  <rpc message-id="102"
       xmlns="urn:ietf:params:xml:ns:netconf:base:1.0">
    <get-config>
      <source>
        <running/>
      </source>
      <with-inactive
         xmlns="http://tail-f.com/ns/netconf/inactive/1.0"/>
    </get-config>
  </rpc>

  <rpc-reply message-id="102"
       xmlns="urn:ietf:params:xml:ns:netconf:base:1.0">
    <data>
      <top xmlns="http://example.com/schema/1.2/config">
        <interface inactive="inactive">
          <name>Ethernet0/0</name>
          <mtu>1500</mtu>
        </interface>
      </top>
    </data>
  </rpc-reply>
```

This request shows that inactive data is not returned unless the client asks for it:

```xml
  <rpc message-id="103"
       xmlns="urn:ietf:params:xml:ns:netconf:base:1.0">
    <get-config>
      <source>
        <running/>
      </source>
    </get-config>
  </rpc>

  <rpc-reply message-id="103"
       xmlns="urn:ietf:params:xml:ns:netconf:base:1.0">
    <data>
    </data>
  </rpc-reply>
```

This request activates the interface:

This request creates an `inactive` interface:

```xml
  <rpc message-id="104"
       xmlns="urn:ietf:params:xml:ns:netconf:base:1.0">
    <edit-config>
      <target>
        <running/>
      </target>
      <with-inactive
         xmlns="http://tail-f.com/ns/netconf/inactive/1.0"/>
      <config>
        <top xmlns="http://example.com/schema/1.2/config">
          <interface active="active">
            <name>Ethernet0/0</name>
          </interface>
        </top>
      </config>
    </edit-config>
  </rpc>

  <rpc-reply message-id="104"
       xmlns="urn:ietf:params:xml:ns:netconf:base:1.0">
    <ok/>
  </rpc-reply>
```

## Rollback ID Capability <a href="#ug.netconf_agent.with-rollback-id" id="ug.netconf_agent.with-rollback-id"></a>

This module extends existing operations with a with-rollback-id parameter which will, when set, extend the result with information about the rollback that was generated for the operation if any.

The rollback ID returned is the ID from within the rollback file which is stable with regards to new rollbacks being created.

### Dependencies <a href="#d5e882" id="d5e882"></a>

None.

### Capability Identifier <a href="#d5e885" id="d5e885"></a>

The transactions capability is identified by the following capability string:

```
  http://tail-f.com/ns/netconf/with-rollback-id
```

### Modifications to Existing Operations <a href="#d5e890" id="d5e890"></a>

This module adds a parameter `with-rollback-id` to the following RPCs:

```
  o  edit-config
  o  copy-config
  o  commit
  o  commit-transaction
```

If `with-rollback-id` is given, rollbacks are enabled, and the operation results in a rollback file being created the response will contain a rollback reference.

## Trace Context

NETCONF supports the IETF standard draft [I-D.draft-ietf-netconf-trace-ctx-extension-00](https://www.ietf.org/archive/id/draft-ietf-netconf-trace-ctx-extension-00.html), that is an adaption of the [W3C Trace Context](https://www.w3.org/TR/2021/REC-trace-context-1-20211123/) standard. Trace Context standardizes the format of `trace-id`, `parent-id`, and key-value pairs sent between distributed entities. The `parent-id` will become the `parent-span-id` for the next generated `span-id` in NSO.

Trace Context consists of two XML attributes `traceparent` and `tracestate` corresponding to the capabilities `urn:ietf:params:xml:ns:yang:traceparent:1.0` and `urn:ietf:params:xml:ns:yang:tracestate:1.0` respectively. The attributes belong to the start XML element `rpc` in a NETCONF request.

Attribute `traceparent` must be of the format:

```
traceparent = <version>-<trace-id>-<parent-id>-<flags>
```

where `version` = "00" and `flags` = "01". The support for the values of `version` and `flags` may change in the future depending on the extension of the standard or functionality.

Attribute `tracestate` is a vendor-specific list of key-value pairs and must be of the format:

```
tracestate = key1=value1,key2=value2
```

Where a value may contain space characters but not end with a space.

Here is an example of the usage of the attributes `traceparent` and `tracestate`:

{% code title="Example: Attributes traceparent and tracestate in NETCONF Request" %}

```xml
<rpc message-id="101"
     xmlns="urn:ietf:params:xml:ns:netconf:base:1.0"
     xmlns:w3ctc="urn:ietf:params:xml:ns:netconf:w3ctc:1.0"
     w3ctc:traceparent="00-100456789abcde10123456789abcde10-001006789abcdef0-01"
     w3ctc:tracestate="key1=value1,key2=value2">
  <edit-config>
    <target>
      <running/>
    </target>
    <config>
      <interfaces xmlns="http://example.com/ns/if">
        <interface>
          <name>eth0</name>
          ...
        </interface>
      </interfaces>
    </config>
  </edit-config>
</rpc>
```

{% endcode %}

If Trace Context is absent in a request, a Trace Context will be generated internally in NSO.

NETCONF also lets LSA clusters to be part of Trace Context handling. A top LSA node will pass down the Trace Context to all LSA nodes beneath. As Trace Context is handled by the progress trace functionality, see also [Progress Trace](/guides/development/advanced-development/progress-trace).

## NETCONF Extensions in NSO <a href="#d5e896" id="d5e896"></a>

The YANG module `tailf-netconf-ncs.yang` augments NETCONF operations with commit parameters and results that are defined in the shared `tailf-ncs-commit-params.yang` module. The parameter semantics are documented centrally in [Commit Parameters](/guides/operation-and-usage/operations/lifecycle-operations#d5e5048).

The shared commit-parameter model is augmented into the following NETCONF operations:

These optional input parameters are augmented into the following NETCONF operations:

* `commit`
* `edit-config`
* `copy-config`
* `prepare-transaction`

In particular, the `prepare-transaction` operation is augmented with the shared `dry-run` parameters and result model from `tailf-ncs-commit-params.yang`, so the same `outformat`, `reverse`, and `with-service-meta-data` leaves are available there as well.

The [examples.ncs/northbound-interfaces/commit-parameters](https://github.com/NSO-developer/nso-examples/tree/6.7/northbound-interfaces/commit-parameters) example includes a NETCONF walkthrough in `demo_nc.py` that uses the standard RFC 6241 `edit-config` operation with the shared commit-parameter model.

FASTMAP attributes such as back pointers and reference counters are typically internal to NSO and are not shown by default. The optional parameter `with-service-meta-data` can be used to include these in the NETCONF reply. The parameter is augmented into the following NETCONF operations:

* `get`
* `get-config`
* `get-data`

## The Query API <a href="#d5e1063" id="d5e1063"></a>

The Query API consists of several RPC operations to start queries, fetch chunks of the result from a query, restart a query, and stop a query.

In the installed release there are two YANG files named `tailf-netconf-query.yang` and `tailf-common-query.yang` that defines these operations. An easy way to find the files is to run the following command from the top directory of the release installation:

```bash
$ find . -name tailf-netconf-query.yang
```

The API consists of the following operations:

* `start-query`: Start a query and return a query handle.
* `fetch-query-result`: Use a query handle to repeatedly fetch chunks of the result.
* `immediate-query`: Start a query and return the entire result immediately.
* `reset-query`: (Re)set where the next fetched result will begin from.
* `stop-query`: Stop (and close) the query.

In the following examples, the following data model is used:

```yang
container x {
  list host {
    key number;
    leaf number {
      type int32;
    }
    leaf enabled {
      type boolean;
    }
    leaf name {
      type string;
    }
    leaf address {
      type inet:ip-address;
    }
  }
}
```

Here is an example of a `start-query` operation:

```xml
<start-query xmlns="http://tail-f.com/ns/netconf/query">
  <foreach>
    /x/host[enabled = 'true']
  </foreach>
  <select>
    <label>Host name</label>
    <expression>name</expression>
    <result-type>string</result-type>
  </select>
  <select>
    <expression>address</expression>
    <result-type>string</result-type>
  </select>
  <sort-by>name</sort-by>
  <limit>100</limit>
  <offset>1</offset>
</start-query>
```

An informal interpretation of this query is:

For each `/x/host` where `enabled` is true, select its `name`, and `address`, and return the result sorted by `name`, in chunks of 100 results at the time.

Let us discuss the various pieces of this request.

The actual XPath query to run is specified by the `foreach` element. The example below will search for all `/x/host` nodes that have the `enabled` node set to `true`:

```xml
<foreach>
  /x/host[enabled = 'true']
</foreach>
```

Now we need to define what we want to have returned from the node set by using one or more `select` sections. What to actually return is defined by the XPath `expression`.

We must also choose how the result should be represented. Basically, it can be the actual value or the path leading to the value. This is specified per select chunk The possible result types are: `string` , `path` , `leaf-value` and `inline`.

The difference between `string` and `leaf-value` is somewhat subtle. In this case of `string` the result will be processed by the XPath function `string()` (which if the result is a node-set will concatenate all the values). The `leaf-value` will return the value of the first node in the result. As long as the result is a leaf node, `string` and `leaf-value` will return the same result. In the example above, we are using `string` as shown below. At least one `result-type` must be specified.

The result-type `inline` makes it possible to return the full sub-tree of data in XML format. The data will be enclosed with a tag: `data`.

Finally, we can specify an optional `label` for a convenient way of labeling the returned data. In the example we have the following:

```xml
<select>
  <label>Host name</label>
  <expression>name</expression>
  <result-type>string</result-type>
</select>
<select>
  <expression>address</expression>
  <result-type>string</result-type>
</select>
```

The returned result can be sorted. This is expressed as XPath expressions, which in most cases are very simple and refer to the found node-set. In this example, we sort the result by the content of the `name` node:

```xml
<sort-by>name</sort-by>
```

To limit the maximum amount of results in each chunk that `fetch-query-result` will return we can set the `limit` element. The default is to get all results in one chunk.

```xml
<limit>100</limit>
```

With the `offset` element we can specify at which node we should start to receive the result. The default is 1, i.e., the first node in the resulting node set.

```xml
<offset>1</offset>
```

Now, if we continue by putting the operation above in a file `query.xml` we can send a request, using the command `netconf-console`, like this:

```bash
$ netconf-console --rpc query.xml
```

The result would look something like this:

```xml
<start-query-result>
  <query-handle>12345</query-handle>
</start-query-result>
```

The query handle (in this example `12345`) must be used in all subsequent calls. To retrieve the result, we can now send:

```xml
<fetch-query-result xmlns="http://tail-f.com/ns/netconf/query">
  <query-handle>12345</query-handle>
</fetch-query-result>
```

Which will result in something like the following:

```xml
<query-result xmlns="http://tail-f.com/ns/netconf/query">
  <result>
    <select>
      <label>Host name</label>
      <value>One</value>
    </select>
    <select>
      <value>10.0.0.1</value>
    </select>
  </result>
  <result>
    <select>
      <label>Host name</label>
      <value>Three</value>
    </select>
    <select>
      <value>10.0.0.1</value>
    </select>
  </result>
</query-result>
```

If we try to get more data with the `fetch-query-result` we might get more `result` entries in return until no more data exists and we get an empty query result back:

```xml
<query-result xmlns="http://tail-f.com/ns/netconf/query">
</query-result>
```

If we want to send the query and get the entire result with only one request, we can do this by using `immediate-query`. This function takes similar arguments as `start-query` and returns the entire result analogous `fetch-query-result`. Note that it is not possible to paginate or set an offset start node for the result list; i.e. the options `limit` and `offset` are ignored.

An example request and response:

```xml
<immediate-query xmlns="http://tail-f.com/ns/netconf/query">
  <foreach>
    /x/host[enabled = 'true']
  </foreach>
  <select>
    <label>Host name</label>
    <expression>name</expression>
    <result-type>string</result-type>
  </select>
  <select>
    <expression>address</expression>
    <result-type>string</result-type>
  </select>
  <sort-by>name</sort-by>
  <timeout>600</timeout>
</immediate-query>
```

```xml
<query-result xmlns="http://tail-f.com/ns/netconf/query">
  <result>
    <select>
      <label>Host name</label>
      <value>One</value>
    </select>
    <select>
      <value>10.0.0.1</value>
    </select>
  </result>
  <result>
    <select>
      <label>Host name</label>
      <value>Three</value>
    </select>
    <select>
      <value>10.0.0.3</value>
    </select>
  </result>
</query-result>
```

If we want to go back in the "stream" of received data chunks and have them repeated, we can do that with the `reset-query` operation. In the example below, we ask to get results from the 42nd result entry:

```xml
<reset-query xmlns="http://tail-f.com/ns/netconf/query">
  <query-handle>12345</query-handle>
  <offset>42</offset>
</reset-query>
```

Finally, when we are done we stop the query:

```xml
<stop-query xmlns="http://tail-f.com/ns/netconf/query">
  <query-handle>12345</query-handle>
</stop-query>
```

## Meta-data in Attributes <a href="#ug.netconf_agent.attributes" id="ug.netconf_agent.attributes"></a>

NSO supports three pieces of meta-data data nodes: tags, annotations, and inactive.

An annotation is a string that acts as a comment. Any data node present in the configuration can get an annotation. An annotation does not affect the underlying configuration but can be set by a user to comment what the configuration does.

An annotation is encoded as an XML attribute `annotation` on any data node. To remove an annotation, set the `annotation` attribute to an empty string.

Any configuration data node can have a set of tags. Tags are set by the user for data organization and filtering purposes. A tag does not affect the underlying configuration.

All tags on a data node are encoded as a space-separated string in an XML attribute `tags`. To remove all tags, set the `tags` attribute to an empty string.

Annotation, tags, and inactive attributes can be present in `<edit-config>`, `<copy-config>`, `<get-config>`, and `<get>`. For example:

```xml
<rpc message-id="101"
     xmlns="urn:ietf:params:xml:ns:netconf:base:1.0">
  <edit-config>
    <target>
      <running/>
    </target>
    <config>
      <interfaces xmlns="http://example.com/ns/if">
        <interface annotation="this is the management interface"
                   tags=" important ethernet ">
          <name>eth0</name>
          ...
        </interface>
      </interfaces>
    </config>
  </edit-config>
</rpc>
```

## Namespace for Additional Error Information <a href="#d5e1189" id="d5e1189"></a>

NSO adds an additional namespace which is used to define elements that are included in the `<error-info>` element. This namespace also describes which `<error-app-tag/>` elements the server might generate, as part of an `<rpc-error/>`.

```xml
<?xml version="1.0" encoding="UTF-8"?>
<xs:schema targetNamespace="http://tail-f.com/ns/netconf/params/1.1"
           xmlns:xs="http://www.w3.org/2001/XMLSchema"
           xml:lang="en">

  <xs:annotation>
    <xs:documentation>
      Tail-f's namespace for additional error information.
      This namespace is used to define elements which are included
      in the 'error-info' element.

      The following are the app-tags used by the NETCONF agent:

        o  not-writable

          Means that an edit-config or copy-config operation was
          attempted on an element which is read-only
          (i.e. non-configuration data).

        o  missing-element-in-choice

          Like the standard error missing-element, but generated when
          one of a set of elements in a choice is missing.

        o  pending-changes

          Means that a lock operation was attempted on the candidate
          database, and the candidate database has uncommitted
          changes. This is not allowed according to the protocol
          specification.

        o  url-open-failed

          Means that the URL given was correct, but that it could not
          be opened. This can e.g. be due to a missing local file, or
          bad ftp credentials. An error message string is provided in
          the &lt;error-message&gt; element.

        o  url-write-failed

          Means that the URL given was opened, but write failed. This
          could e.g. be due to lack of disk space. An error message
          string is provided in the &lt;error-message&gt; element.

        o  bad-state

          Means that an rpc is received when the session is in a state
          which don't accept this rpc.  An example is
          &lt;prepare-transaction&gt; before &lt;start-transaction&gt;

    </xs:documentation>
  </xs:annotation>

  <xs:element name="bad-keyref">
    <xs:annotation>
      <xs:documentation>
        This element will be present in the 'error-info' container when
        'error-app-tag' is "instance-required".
      </xs:documentation>
    </xs:annotation>
    <xs:complexType>
      <xs:sequence>
        <xs:element name="bad-element" type="xs:string">
          <xs:annotation>
            <xs:documentation>
              Contains an absolute XPath expression pointing to the element
              which value refers to a non-existing instance.
            </xs:documentation>
          </xs:annotation>
        </xs:element>
        <xs:element name="missing-element" type="xs:string">
          <xs:annotation>
            <xs:documentation>
              Contains an absolute XPath expression pointing to the missing
              element referred to by 'bad-element'.
            </xs:documentation>
          </xs:annotation>
        </xs:element>
      </xs:sequence>
    </xs:complexType>
  </xs:element>

  <xs:element name="bad-instance-count">
    <xs:annotation>
      <xs:documentation>
        This element will be present in the 'error-info' container when
        'error-app-tag' is "too-few-elements" or "too-many-elements".
      </xs:documentation>
    </xs:annotation>
    <xs:complexType>
      <xs:sequence>
        <xs:element name="bad-element" type="xs:string">
          <xs:annotation>
            <xs:documentation>
              Contains an absolute XPath expression pointing to an
              element which exists in too few or too many instances.
            </xs:documentation>
          </xs:annotation>
        </xs:element>
        <xs:element name="instances" type="xs:unsignedInt">
          <xs:annotation>
            <xs:documentation>
              Contains the number of existing instances of the element
              referd to by 'bad-element'.
            </xs:documentation>
          </xs:annotation>
        </xs:element>
        <xs:choice>
          <xs:element name="min-instances" type="xs:unsignedInt">
            <xs:annotation>
              <xs:documentation>
                Contains the minimum number of instances that must
                exist in order for the configuration to be consistent.
                This element is present only if 'app-tag' is
                'too-few-elems'.
              </xs:documentation>
            </xs:annotation>
          </xs:element>
          <xs:element name="max-instances" type="xs:unsignedInt">
            <xs:annotation>
              <xs:documentation>
                Contains the maximum number of instances that can
                exist in order for the configuration to be consistent.
                This element is present only if 'app-tag' is
                'too-many-elems'.
              </xs:documentation>
            </xs:annotation>
          </xs:element>
        </xs:choice>
      </xs:sequence>
    </xs:complexType>
  </xs:element>

  <xs:attribute name="annotation" type="xs:string">
    <xs:annotation>
      <xs:documentation>
        This attribute can be present on any configuration data node.  It
        acts as a comment for the node.  The annotation does not affect the
        underlying configuration data.
      </xs:documentation>
    </xs:annotation>
  </xs:attribute>

  <xs:attribute name="tags" type="xs:string">
    <xs:annotation>
      <xs:documentation>
        This attribute can be present on any configuration data node.  It
        is a space separated string of tags for the node.  The tags of a
        node does not affect the underlying configuration data, but can
        be used by a user for data organization, and data filtering.
      </xs:documentation>
    </xs:annotation>
  </xs:attribute>

</xs:schema>
```


# RESTCONF API

Description of the RESTCONF API.

RESTCONF is an HTTP-based protocol as defined in [RFC 8040](https://www.ietf.org/rfc/rfc8040.txt). RESTCONF standardizes a mechanism to allow Web applications to access the configuration data, state data, data-model-specific Remote Procedure Call (RPC) operations, and event notifications within a networking device.

RESTCONF uses HTTP methods to provide Create, Read, Update, Delete (CRUD) operations on a conceptual datastore containing YANG-defined data, which is compatible with a server that implements NETCONF datastores as defined in [RFC 6241](https://www.ietf.org/rfc/rfc6241.txt).

Configuration data and state data are exposed as resources that can be retrieved with the GET method. Resources representing configuration data can be modified with the DELETE, PATCH, POST, and PUT methods. Data is encoded with either XML ([W3C.REC-xml-20081126](https://www.w3.org/TR/2008/REC-xml-20081126)) or JSON ([RFC 7951](https://www.ietf.org/rfc/rfc7951.txt)).

This section describes the NSO implementation and extension to or deviation from [RFC 8040](https://www.ietf.org/rfc/rfc8040.txt) respectively.

As of this writing, the server supports the following specifications:

* [RFC 6020](https://www.ietf.org/rfc/rfc6020.txt) - YANG - A Data Modeling Language for the Network Configuration Protocol (NETCONF)
* [RFC 6021](https://www.ietf.org/rfc/rfc6021.txt) - Common YANG Data Types
* [RFC 6470](https://www.ietf.org/rfc/rfc6470.txt) - NETCONF Base Notifications
* [RFC 6536](https://www.ietf.org/rfc/rfc6536.txt) - NETCONF Access Control Model
* [RFC 6991](https://www.ietf.org/rfc/rfc6991.txt) - Common YANG Data Types
* [RFC 7950](https://www.ietf.org/rfc/rfc7950.txt) - The YANG 1.1 Data Modeling Language
* [RFC 7951](https://www.ietf.org/rfc/rfc7951.txt) - JSON Encoding of Data Modeled with YANG
* [RFC 7952](https://www.ietf.org/rfc/rfc7952.txt) - Defining and Using Metadata with YANG
* [RFC 8040](https://www.ietf.org/rfc/rfc8040.txt) - RESTCONF Protocol
* [RFC 8072](https://www.ietf.org/rfc/rfc8072.txt) - YANG Patch Media Type
* [RFC 8341](https://www.ietf.org/rfc/rfc8341.txt) - Network Configuration Access Control Model
* [RFC 8525](https://www.ietf.org/rfc/rfc8525.txt) - YANG Library
* [RFC 8528](https://www.ietf.org/rfc/rfc8528.txt) - YANG Schema Mount
* [RFC 8639](https://www.ietf.org/rfc/rfc8639.txt) - Subscription to YANG Notifications
* [RFC 8641](https://www.ietf.org/rfc/rfc8641.txt) - Subscription to YANG Notifications for Datastore Updates
* [RFC 8650](https://www.ietf.org/rfc/rfc8650.txt) - Dynamic Subscription to YANG Events and Datastores over RESTCONF
* [I-D.draft-ietf-netconf-restconf-trace-ctx-headers-00](https://www.ietf.org/archive/id/draft-ietf-netconf-restconf-trace-ctx-headers-00.html) - RESTCONF Extension to support Trace Context Headers

## Getting Started <a href="#ncs.northbound.restconf.getting_started" id="ncs.northbound.restconf.getting_started"></a>

To enable RESTCONF in NSO, RESTCONF must be enabled in the `ncs.conf` configuration file. The web server configuration for RESTCONF is shared with the WebUI's config, but you may define a separate RESTCONF transport section. The WebUI does not have to be enabled for RESTCONF to work.

Here is a minimal example of what is needed in the `ncs.conf`.

{% code title="Example: NSO Configuration for RESTCONF" %}

```xml
<restconf>
  <enabled>true</enabled>
</restconf>

<webui>
  <transport>
    <tcp>
      <enabled>true</enabled>
      <ip>0.0.0.0</ip>
      <port>8080</port>
    </tcp>
  </transport>
</webui>
```

{% endcode %}

RESTCONF and the WebUI can use separate transport configurations while both remain enabled. The following example configures RESTCONF on TCP port 8090 and the WebUI on TCP port 8080.

{% code title="Example: NSO Separate Transport Configurations for RESTCONF and WebUI" %}

```xml
<restconf>
  <enabled>true</enabled>
  <transport>
    <tcp>
      <enabled>true</enabled>
      <ip>0.0.0.0</ip>
      <port>8090</port>
    </tcp>
  </transport>
</restconf>

<webui>
  <enabled>true</enabled>
  <transport>
    <tcp>
      <enabled>true</enabled>
      <ip>0.0.0.0</ip>
      <port>8080</port>
    </tcp>
  </transport>
</webui>
```

{% endcode %}

The remaining examples use the shared transport configuration on TCP port 8080. You can send RESTCONF requests to NSO using any HTTP client; the following examples use curl. The example below shows what a typical RESTCONF request looks like.

{% code title="Example: A RESTCONF Request using curl " %}

```bash
# Note that the command is wrapped in several lines in order to fit.
#
# The switch '-i' will include any HTTP reply headers in the output
# and the '-s' will suppress some superflous output.
#
# The '-u' switch specify the User:Password for login authentication.
#
# The '-H' switch will add a HTTP header to the request; in this case
# an 'Accept' header is added, requesting the preferred reply format.
#
# Finally, the complete URL to the wanted resource is specified,
# in this case the top of the configuration tree.
#
curl -is -u admin:admin \
-H "Accept: application/yang-data+xml" \
http://localhost:8080/restconf/data
```

{% endcode %}

In the rest of the document, in order to simplify the presentation, the example above will be expressed as:

{% code title="Example: A RESTCONF Request, Simplified" %}

```http
GET /restconf/data
Accept: application/yang-data+xml

# Any reply with relevant headers will be displayed here!
HTTP/1.1 200 OK
```

{% endcode %}

Note the HTTP return code (200 OK) in the example, which will be displayed together with any relevant HTTP headers returned and a possible body of content.

### Top-level GET request <a href="#d5e1282" id="d5e1282"></a>

Send a RESTCONF query to get a representation of the top-level resource, which is accessible through the path: `/restconf`.

{% code title="Example: A Top-level RESTCONF Request" %}

```http
GET /restconf
Accept: application/yang-data+xml

HTTP/1.1 200 OK
<restconf xmlns="urn:ietf:params:xml:ns:yang:ietf-restconf">
  <data/>
  <operations/>
  <yang-library-version>2019-01-04</yang-library-version>
</restconf>
```

{% endcode %}

As can be seen from the result, the server exposes three additional resources:

* `data`: This mandatory resource represents the combined configuration and state data resources that can be accessed by a client.
* `operations`: This optional resource is a container that provides access to the data-model-specific RPC operations supported by the server.
* `yang-library-version`: This mandatory leaf identifies the revision date of the `ietf-yang-library` YANG module that is implemented by this server. This resource exposes which YANG modules are in use by the NSO system.

### Get Resources Under the `data` Resource <a href="#d5e1302" id="d5e1302"></a>

To fetch configuration, operational data, or both, from the server, a request to the `data` resource is made. To restrict the amount of returned data, the following example will prune the amount of output to only consist of the topmost nodes. This is achieved by using the `depth` query argument as shown in the example below:

{% code title="Example: Get the Top-most Resources Under data  " %}

```http
GET /restconf/data?depth=1
Accept: application/yang-data+xml

<data xmlns="urn:ietf:params:xml:ns:yang:ietf-restconf">
  <yang-library xmlns="urn:ietf:params:xml:ns:yang:ietf-yang-library"/>
  <modules-state xmlns="urn:ietf:params:xml:ns:yang:ietf-yang-library"/>
  <dhcp xmlns="http://yang-central.org/ns/example/dhcp"/>
  <nacm xmlns="urn:ietf:params:xml:ns:yang:ietf-netconf-acm"/>
  <netconf-state xmlns="urn:ietf:params:xml:ns:yang:ietf-netconf-monitoring"/>
  <restconf-state xmlns="urn:ietf:params:xml:ns:yang:ietf-restconf-monitoring"/>
  <aaa xmlns="http://tail-f.com/ns/aaa/1.1"/>
  <confd-state xmls="http://tail-f.com/yang/confd-monitoring"/>
  <last-logins xmlns="http://tail-f.com/yang/last-login"/>
</data>
```

{% endcode %}

### Manipulating config data with RESTCONF

Let's assume we are interested in the `dhcp/subnet` resource in our configuration. In the following examples, assume that it is defined by a corresponding Yang module that we have named `dhcp.yang`, looking like this:

{% code title="Example: The dhcp.yang Resource" %}

```cli
> yanger -f tree examples.confd/restconf/basic/dhcp.yang
module: dhcp
  +--rw dhcp
  +--rw max-lease-time?       uint32
  +--rw default-lease-time?   uint32
  +--rw subnet* [net]
  |  +--rw net               inet:ip-prefix
  |  +--rw range!
  |  |  +--rw dynamic-bootp?   empty
  |  |  +--rw low              inet:ip-address
  |  |  +--rw high             inet:ip-address
  |  +--rw dhcp-options
  |  |  +--rw router*        inet:host
  |  |  +--rw domain-name?   inet:domain-name
  |  +--rw max-lease-time?   uint32
```

{% endcode %}

We can issue an HTTP GET request to retrieve the value content of the resource. In this case, we find that there is no such data, which is indicated by the HTTP return code `204 No Content`.

Note also how we have prefixed the `dhcp:dhcp` resource. This is how RESTCONF handles namespaces, where the prefix is the YANG module name and the namespace is as defined by the namespace statement in the YANG module.

{% code title="Example: Get the dhcp/subnet Resource" %}

```http
GET /restconf/data/dhcp:dhcp/subnet

HTTP/1.1 204 No Content
```

{% endcode %}

We can now create the `dhcp/subnet` resource by sending an HTTP POST request + the data that we want to store. Note the `Content-Type` HTTP header, which indicates the format of the provided body. Two formats are supported: XML or JSON. In this example, we are using XML, which is indicated by the `Content-Type` value: `application/yang-data+xml`.

{% code title="Example: Create a New dhcp/subnet Resource" %}

```http
POST /restconf/data/dhcp:dhcp
Content-Type: application/yang-data+xml

<subnet xmlns="http://yang-central.org/ns/example/dhcp"
          xmlns:dhcp="http://yang-central.org/ns/example/dhcp">
  <net>10.254.239.0/27</net>
  <range>
    <dynamic-bootp/>
    <low>10.254.239.10</low>
    <high>10.254.239.20</high>
  </range>
  <dhcp-options>
    <router>rtr-239-0-1.example.org</router>
    <router>rtr-239-0-2.example.org</router>
  </dhcp-options>
  <max-lease-time>1200</max-lease-time>
</subnet>

# If the resource is created, the server might respond as follows:

HTTP/1.1 201 Created
Location: http://localhost:8080/restconf/data/dhcp:dhcp/subnet=10.254.239.0%2F27
```

{% endcode %}

Note the HTTP return code (`201 Created`) indicating that the resource was successfully created. We also got a Location header, which always is returned in a reply to a successful creation of a resource, stating the resulting URI leading to the created resource.

If we now want to modify a part of our `dhcp/subnet` config, we can use the HTTP `PATCH` method, as shown below. Note that the URI used in the request needs to be URL-encoded, such that the key value: `10.254.239.0/27` is URL-encoded as: `10.254.239.0%2F27`.

Also, note the difference of the `PATCH` URI compared to the earlier `POST` request. With the latter, since the resource does not yet exist, we `POST` to the parent resource (`dhcp:dhcp`), while with the `PATCH` request we address the (existing) resource (`10.254.239.0%2F27`).

{% code title="Example: Modify a Part of the dhcp/subnet Resource" %}

```http
PATCH /restconf/data/dhcp:dhcp/subnet=10.254.239.0%2F27

<subnet>
  <max-lease-time>3333</max-lease-time>
</subnet>

# If our modification is successful, the server might respond as follows:

HTTP/1.1 204 No Content
```

{% endcode %}

We can also replace the subnet with some new configuration. To do this, we make use of the `PUT` HTTP method as shown below. Since the operation was successful and no body was returned, we will get a `204 No Content` return code.

{% code title="Example: Replace a dhcp/subnet Resource" %}

```http
PUT /restconf/data/dhcp:dhcp/subnet=10.254.239.0%2F27
Content-Type: application/yang-data+xml

<subnet xmlns="http://yang-central.org/ns/example/dhcp"
          xmlns:dhcp="http://yang-central.org/ns/example/dhcp">
  <net>10.254.239.0/27</net>

  <!-- ...config left out here... -->

</subnet>

# At success, the server will respond as follows:

HTTP/1.1 204 No Content
```

{% endcode %}

To delete the subnet, we make use of the `DELETE` HTTP method as shown below. Since the operation was successful and no body was returned, we will get a `204 No Content` return code.

{% code title="Example: Delete a dhcp/subnet Resource" %}

```http
DELETE /restconf/data/dhcp:dhcp/subnet=10.254.239.0%2F27

HTTP/1.1 204 No Content
```

{% endcode %}

## Protocol YANG Modules <a href="#d5e255" id="d5e255"></a>

In addition to the protocol capabilities listed above, NSO also implements a set of YANG modules that are closely related to the protocol.

* `ietf-subscribed-notifications`: This module from [RFC 8639](https://www.ietf.org/rfc/rfc8639.txt) defines operations, configuration data nodes, and operational state data nodes related to notification subscriptions. It defines the following features:
* `configured`: Indicates that the server supports configured subscriptions. This feature is not advertised.
* `dscp`: Indicates that the server supports the ability to set the Differentiated Services Code Point (DSCP) value in outgoing packets. This feature is not advertised.
* `encode-json`: Indicates that the server supports JSON encoding of notifications. This is not yet implemented for RESTCONF, and this feature is not advertised.
* `encode-xml`: Indicates that the server supports XML encoding of notifications. This feature is advertised by NSO.
* `interface-designation`: Indicates that a configured subscription can be configured to send notifications over a specific interface. This feature is not advertised.
* `qos`: Indicates that a publisher supports absolute dependencies of one subscription's traffic over another as well as weighted bandwidth sharing between subscriptions. This feature is not advertised.
* `replay`: Indicates that historical event record replay is supported. This feature is advertised by NSO.
* `subtree`: Indicates that the server supports subtree filtering of notifications. This is not yet supported for RESTCONF, and this feature is not advertised.
* `supports-vrf`: Indicates that a configured subscription can be configured to send notifications from a specific VRF. This feature is not advertised.
* `xpath`: Indicates that the server supports XPath filtering of notifications. This feature is advertised by NSO.

In addition to this, NSO does not support pre-configuration or monitoring of subtree filters, and thus advertises a deviation module that deviates `/filters/stream-filter/filter-spec/stream-subtree-filter` and `/subscriptions/subscription/target/stream/stream-filter/within-subscription/filter-spec/stream-subtree-filter` as "not-supported".

NSO does not generate `subscription-modified` notifications when the parameters of a subscription change, and there is currently no mechanism to suspend notifications, so `subscription-suspended` and `subscription-resumed` notifications are never generated.

There is basic support for monitoring subscriptions via the `/subscriptions` container. Currently, it is possible to view dynamic subscriptions' attributes: `subscription-id`, `stream`, `encoding`, `receiver`, `stop-time`, and `stream-xpath-filter`. Unsupported attributes are: `stream-subtree-filter`, `receiver/sent-event-records`, `receiver/excluded-event-records`, and `receiver/state`.

* `ietf-yang-push`: This module from [RFC 8641](https://www.ietf.org/rfc/rfc8641.txt) extends operations, data nodes, and operational state defined in `ietf-subscribed-notifications;` and also introduces continuous and customizable notification subscriptions for updates from running and operational datastores. It defines the same features as `ietf-subscribed-notifications` and also the following feature:
  * `on-change`: Indicates that on-change triggered notifications are supported. This feature is advertised by NSO.
    * `dampening-period`: Indicates that dampening-period for on-change subscriptions is supported. This feature is advertised by NSO.
    * `sync-on-start`: Indicates that sync-on-start for on-change subscriptions is supported. This feature is advertised by NSO.
    * `excluded-change`: Indicates that excluded-change for on-change subscription is supported. This feature is advertised by NSO.
  * `periodic`: Indicates that periodic notifications are supported. This feature is advertised by NSO.
    * `period`: Indicates that period for periodic notifications are supported. This feature is advertised by NSO.
    * `anchor-time`: Indicates that anchor-time for periodic subscriptions is supported. This feature is advertised by NSO.

In addition to this, NSO does not support pre-configuration or monitoring of subtree filters and thus advertises a deviation module that deviates `/filters/selection-filter/filter-spec/datastore-subtree-filter` and `/subscriptions/subscription/target/datastore/selection-filter/within-subscription/filter-spec/datastore-subtree-filter` as "not-supported".

The monitoring of subscriptions via the `subscriptions` container currently does not support the attribute `/subscriptions/receivers/receiver/state` .

## Root Resource Discovery

RESTCONF makes it possible to specify where the RESTCONF API is located, as described in the RESTCONF [RFC 8040](https://www.ietf.org/rfc/rfc8040.txt#section-3.1).

As per default, the RESTCONF API root is `/restconf`. Typically there is no need to change the default value although it is possible to change this by configuring the RESTCONF API root in the `ncs.conf` file as:

{% code title="Example: NSO Configuration for RESTCONF" %}

```xml
<restconf>
  <enabled>true</enabled>
  <root-resource>my_own_restconf_root</root-resource>
</restconf>
```

{% endcode %}

The RESTCONF API root will now be `/my_own_restconf_root`.

A client may discover the root resource by getting the `/.well-known/host-meta` resource as shown in the example below:

{% code title="Example: Example Returning /restconf" %}

```
   The client might send the following:

      GET /.well-known/host-meta
      Accept: application/xrd+xml

   The server might respond as follows:

      HTTP/1.1 200 OK

      <XRD xmlns='http://docs.oasis-open.org/ns/xri/xrd-1.0'>
          <Link rel='restconf' href='/restconf'/>
      </XRD>
```

{% endcode %}

{% hint style="info" %}
In this guide, all examples will assume the RESTCONF API root to be `/restconf`.
{% endhint %}

## Capabilities <a href="#d5e1399" id="d5e1399"></a>

A RESTCONF capability is a set of functionality that supplements the base RESTCONF specification. The capability is identified by a uniform resource identifier [(URI)](https://www.ietf.org/rfc/rfc3986.txt). The RESTCONF server includes a `capability` URI leaf-list entry identifying each supported protocol feature. This includes the `basic-mode` default-handling mode, optional query parameters, and may also include other, NSO-specific, capability URIs.

### How to View the Capabilities of the RESTCONF Server <a href="#ncs.northbound.restconf.capabilities" id="ncs.northbound.restconf.capabilities"></a>

To view currently enabled capabilities, use the `ietf-restconf-monitoring` YANG model, which is available as: `/restconf/data/ietf-restconf-monitoring:restconf-state`.

{% code title="Example: NSO RESTCONF Capabilities" %}

```http
GET /restconf/data/ietf-restconf-monitoring:restconf-state
Host: example.com
Accept: application/yang-data+xml

<restconf-state xmlns="urn:ietf:params:xml:ns:yang:ietf-restconf-monitoring"
  xmlns:rcmon="urn:ietf:params:xml:ns:yang:ietf-restconf-monitoring">
<capabilities>
  <capability>
    urn:ietf:params:restconf:capability:defaults:1.0?basic-mode=explicit
  </capability>
  <capability>urn:ietf:params:restconf:capability:depth:1.0</capability>
  <capability>urn:ietf:params:restconf:capability:fields:1.0</capability>
  <capability>urn:ietf:params:restconf:capability:with-defaults:1.0</capability>
  <capability>urn:ietf:params:restconf:capability:filter:1.0</capability>
  <capability>urn:ietf:params:restconf:capability:replay:1.0</capability>
  <capability>http://tail-f.com/ns/restconf/collection/1.0</capability>
  <capability>http://tail-f.com/ns/restconf/query-api/1.0</capability>
  <capability>http://tail-f.com/ns/restconf/partial-response/1.0</capability>
  <capability>http://tail-f.com/ns/restconf/unhide/1.0</capability>
  <capability>urn:ietf:params:xml:ns:yang:traceparent:1.0</capability>
  <capability>urn:ietf:params:xml:ns:yang:tracestate:1.0</capability>
</capabilities>
</restconf-state>
```

{% endcode %}

### The `defaults` Capability

This Capability identifies the `basic-mode` default-handling mode that is used by the server for processing default leafs in requests for data resources.

{% code title="Example: The Default Capability URI" %}

```
          urn:ietf:params:restconf:capability:defaults:1.0
```

{% endcode %}

The `capability` URL will contain a query parameter named `basic-mode` which value tells us what the default behavior of the RESTCONF server is when it returns a leaf. The possible values are shown in the table below (`basic-mode` values):

<table><thead><tr><th width="163" valign="top">Value</th><th valign="top">Description</th></tr></thead><tbody><tr><td valign="top"><code>report-all</code></td><td valign="top">Values set to the YANG default value are reported.</td></tr><tr><td valign="top"><code>trim</code></td><td valign="top">Values set to the YANG default value are not reported.</td></tr><tr><td valign="top"><code>explicit</code></td><td valign="top">Values that has been set by a client to the YANG default value will be reported.</td></tr></tbody></table>

The values presented in the table above can also be used by the Client together with the `with-defaults` query parameter to override the default RESTCONF server behavior. Added to these values, the Client can also use the `report-all-tagged` value.

The table below lists additional `with-defaults` value.

<table><thead><tr><th width="231" valign="top">Value</th><th valign="top">Description</th></tr></thead><tbody><tr><td valign="top"><code>report-all-tagged</code></td><td valign="top">Works as the <code>report-all</code> but a default value will include an XML/JSON attribute to indicate that the value is in fact a default value.</td></tr></tbody></table>

Referring back to the example: Example: NSO RESTCONF Capabilities, where the RESTCONF server returned the default capability:

```
urn:ietf:params:restconf:capability:defaults:1.0?basic-mode=explicit
```

It tells us that values that have been set by a client to the YANG default value will be reported but default values that have not been set by the Client will not be returned. Again, note that this is the default RESTCONF server behavior which can be overridden by the Client by using the `with-defaults` query argument.

### Query Parameter Capabilities <a href="#d5e1469" id="d5e1469"></a>

A set of optional RESTCONF Capability URIs are defined to identify the specific query parameters that are supported by the server. They are defined as:

The table shows query parameter capabilities.

<table><thead><tr><th width="184" valign="top">Name</th><th valign="top">URI</th></tr></thead><tbody><tr><td valign="top"><code>depth</code></td><td valign="top"><code>urn:ietf:params:restconf:capability:depth:1.0</code></td></tr><tr><td valign="top"><code>fields</code></td><td valign="top"><code>urn:ietf:params:restconf:capability:fields:1.0</code></td></tr><tr><td valign="top"><code>filter</code></td><td valign="top"><code>urn:ietf:params:restconf:capability:filter:1.0</code></td></tr><tr><td valign="top"><code>replay</code></td><td valign="top"><code>urn:ietf:params:restconf:capability:replay:1.0</code></td></tr><tr><td valign="top"><code>with.defaults</code></td><td valign="top"><code>urn:ietf:params:restconf:capability:with.defaults:1.0</code></td></tr></tbody></table>

For a description of the query parameter functionality, see [Query Parameters](#ncs.northbound.restconf.query_params).

## Query Parameters <a href="#ncs.northbound.restconf.query_params" id="ncs.northbound.restconf.query_params"></a>

Each RESTCONF operation allows zero or more query parameters to be present in the request URI. Query parameters can be given in any order, but can appear at most once. Supplying query parameters when invoking RPCs and actions is not supported, if supplied the response will be 400 (Bad Request) and the `error-app-tag` will be set to `invalid-value`. However, the query parameter `unhide` is exempted from this rule and supported for RPC and action invocation. The defined query parameters and in what type of HTTP request they can be used are shown in the table below (Query parameters).

<table><thead><tr><th width="185" valign="top">Name</th><th width="154" valign="top">Method</th><th valign="top">Description</th></tr></thead><tbody><tr><td valign="top"><code>content</code></td><td valign="top"><code>GET</code>,<code>HEAD</code></td><td valign="top">Select config and/or non-config data resources.</td></tr><tr><td valign="top"><code>depth</code></td><td valign="top"><code>GET</code>,<code>HEAD</code></td><td valign="top">Request limited subtree depth in the reply content.</td></tr><tr><td valign="top"><code>fields</code></td><td valign="top"><code>GET</code>,<code>HEAD</code></td><td valign="top">Request a subset of the target resource contents.</td></tr><tr><td valign="top"><code>exclude</code></td><td valign="top"><code>GET</code>,<code>HEAD</code></td><td valign="top">Exclude a subset of the target resource contents.</td></tr><tr><td valign="top"><code>filter</code></td><td valign="top"><code>GET</code>,<code>HEAD</code></td><td valign="top">Boolean notification filter for event stream resources.</td></tr><tr><td valign="top"><code>insert</code></td><td valign="top"><code>POST</code>,<code>PUT</code></td><td valign="top">Insertion mode for <em>ordered-by user</em> data resources</td></tr><tr><td valign="top"><code>point</code></td><td valign="top"><code>POST</code>,<code>PUT</code></td><td valign="top">Insertion point for <em>ordered-by user</em> data resources</td></tr><tr><td valign="top"><code>start-time</code></td><td valign="top"><code>GET</code>,<code>HEAD</code></td><td valign="top">Replay buffer start time for event stream resources.</td></tr><tr><td valign="top"><code>stop-time</code></td><td valign="top"><code>GET</code>,<code>HEAD</code></td><td valign="top">Replay buffer stop time for event stream resources.</td></tr><tr><td valign="top"><code>with-defaults</code></td><td valign="top"><code>GET</code>,<code>HEAD</code></td><td valign="top">Control the retrieval of default values.</td></tr><tr><td valign="top"><code>with-origin</code></td><td valign="top"><code>GET</code></td><td valign="top">Include the "origin" metadata annotations, as detailed in the NMDA.</td></tr></tbody></table>

### The `content` Query Parameter

The `content` query parameter controls if configuration, non-configuration, or both types of data should be returned. The `content` query parameter values are listed below.

The allowed values are:

<table><thead><tr><th width="148" valign="top">Value</th><th valign="top">Description</th></tr></thead><tbody><tr><td valign="top"><code>config</code></td><td valign="top">Return only configuration descendant data nodes.</td></tr><tr><td valign="top"><code>nonconfig</code></td><td valign="top">Return only non-configuration descendant data nodes.</td></tr><tr><td valign="top"><code>all</code></td><td valign="top">Return all descendant data nodes.</td></tr></tbody></table>

### The `depth` Query Parameter

The `depth` query parameter is used to limit the depth of subtrees returned by the server. Data nodes with a value greater than the `depth` parameter are not returned in response to a GET request.

The value of the `depth` parameter is either an integer between 1 and 65535 or the string `unbounded`. The default value is: `unbounded`.

### The `fields` Query Parameter <a href="#d5e1600" id="d5e1600"></a>

The `fields` query parameter is used to optionally identify data nodes within the target resource to be retrieved in a GET method. The client can use this parameter to retrieve a subset of all nodes in a resource.

For a full definition of the `fields` value can be constructed, refer to the [RFC 8040, Section 4.8.3](https://tools.ietf.org/html/rfc8040#section-4.8.3).

Note that the `fields` query parameter cannot be used together with the `exclude` query parameter. This will result in an error.

{% code title="Example: Example of How to use the Fields Query Parameter" %}

```http
GET /restconf/data/dhcp:dhcp?fields=subnet/range(low;high)
Accept: application/yang-data+xml

HTTP/1.1 200 OK
<dhcp xmlns="http://yang-central.org/ns/example/dhcp" \
      xmlns:dhcp="http://yang-central.org/ns/example/dhcp">
  <subnet>
    <range>
      <low>10.254.239.10</low>
      <high>10.254.239.20</high>
    </range>
  </subnet>
  <subnet>
    <range>
      <low>10.254.244.10</low>
      <high>10.254.244.20</high>
    </range>
  </subnet>
</dhcp>
```

{% endcode %}

### The `exclude` Query Parameter

The `exclude` query parameter is used to optionally exclude data nodes within the target resource from being retrieved with a GET request. The client can use this parameter to exclude a subset of all nodes in a resource. Only nodes below the target resource can be excluded, not the target resource itself.

Note that the `exclude` query parameter cannot be used together with the `fields` query parameter. This will result in an error.

The `exclude` query parameter uses the same syntax and has the same restrictions as the `fields` query parameter, as defined in [RFC 8040, Section 4.8.3](https://tools.ietf.org/html/rfc8040#section-4.8.3).

Selecting multiple nodes to exclude can be done the same way as for the `fields` query parameter, as described in [RFC 8040, Section 4.8.3](https://tools.ietf.org/html/rfc8040#section-4.8.3).

`exclude` using wildcards (\*) will exclude all child nodes of the node. For lists and presence containers, the parent node will be visible in the output but not its children, i.e. it will be displayed as an empty node. For non-presence containers, the parent node will be excluded from the output as well.

`exclude` can be used together with the `depth` query parameter to limit the depth of the output. In contrast to `fields`, where `depth` is counted from the node selected by `fields`, for `exclude` the depth is counted from the target resource, and the nodes are excluded if `depth` is deep enough to encounter an excluded node.

When `exclude` is not used:

{% code title="Example: Example of how to use the Exclude Query Parameter" %}

```http
GET /restconf/data/dhcp:dhcp/subnet
Accept: application/yang-data+xml

HTTP/1.1 200 OK
<subnet xmlns="http://yang-central.org/ns/example/dhcp"
          xmlns:dhcp="http://yang-central.org/ns/example/dhcp">
  <net>10.254.239.0/27</net>
  <range>
    <dynamic-bootp/>
    <low>10.254.239.10</low>
    <high>10.254.239.20</high>
  </range>
  <dhcp-options>
    <router>rtr-239-0-1.example.org</router>
    <router>rtr-239-0-2.example.org</router>
  </dhcp-options>
  <max-lease-time>1200</max-lease-time>
</subnet>
```

{% endcode %}

Using `exclude` to exclude `low` and `high` from `range`, note that these are absent in the output:

```http
GET /restconf/data/dhcp:dhcp/subnet?exclude=range(low;high)
Accept: application/yang-data+xml

HTTP/1.1 200 OK
<subnet xmlns="http://yang-central.org/ns/example/dhcp"
          xmlns:dhcp="http://yang-central.org/ns/example/dhcp">
  <net>10.254.239.0/27</net>
  <range>
    <dynamic-bootp/>
  </range>
  <dhcp-options>
    <router>rtr-239-0-1.example.org</router>
    <router>rtr-239-0-2.example.org</router>
  </dhcp-options>
  <max-lease-time>1200</max-lease-time>
</subnet>
```

### The `filter`, `start-time`, and `stop-time` Query Parameters

These query parameters are only allowed on an event stream resource and are further described in [Streams](#ncs.northbound.restconf.streams).

### The `insert` Query Parameter <a href="#d5e1657" id="d5e1657"></a>

The `insert` query parameter is used to specify how a resource should be inserted within an `ordered-by user` list. The allowed values are shown in the table below (The `content` query parameter values).

<table><thead><tr><th width="142" valign="top">Value</th><th valign="top">Description</th></tr></thead><tbody><tr><td valign="top"><code>first</code></td><td valign="top">Insert the new data as the new first entry.</td></tr><tr><td valign="top"><code>last</code></td><td valign="top">Insert the new data as the new last entry. This is the default value.</td></tr><tr><td valign="top"><code>before</code></td><td valign="top">Insert the new data before the insertion point, as specified by the value of the <code>point</code> parameter.</td></tr><tr><td valign="top"><code>after</code></td><td valign="top">Insert the new data after the insertion point, as specified by the value of the <code>point</code> parameter.</td></tr></tbody></table>

This parameter is only valid if the target data represents a YANG list or leaf-list that is `ordered-by user`. In the example below, we will insert a new `router` value, first, in the `ordered-by user` leaf-list of `dhcp-options/router` values. Remember that the default behavior is for new entries to be inserted last in an `ordered-by user` leaf-list.

{% code title="Example: Insert first into a ordered-by user leaf-list " %}

```bash
# Note: we have to split the POST line in order to fit the page
POST /restconf/data/dhcp:dhcp/subnet=10.254.239.0%2F27/dhcp-options?\
     insert=first
Content-Type: application/yang-data+xml

<router>one.acme.org</router>

# If the resource is created, the server might respond as follows:

HTTP/1.1 201 Created
Location /restconf/data/dhcp:dhcp/subnet=10.254.239.0%2F27/dhcp-options/\
         router=one.acme.org
```

{% endcode %}

To verify that the `router` value really ended up first:

```http
GET /restconf/data/dhcp:dhcp/subnet=10.254.239.0%2F27/dhcp-options
Accept: application/yang-data+xml

HTTP/1.1 200 OK
<dhcp-options xmlns="http://yang-central.org/ns/example/dhcp"
              xmlns:dhcp="http://yang-central.org/ns/example/dhcp">
  <router>one.acme.org</router>
  <router>rtr-239-0-1.example.org</router>
  <router>rtr-239-0-2.example.org</router>
</dhcp-options>
```

### The `point` Query Parameter <a href="#d5e1703" id="d5e1703"></a>

The `point` query parameter is used to specify the insertion point for a data resource that is being created or moved within an `ordered-by user` list or leaf-list. In the example below, we will insert the new `router` value: `two.acme.org`, after the first value: `one.acme.org` in the `ordered-by user` leaf-list of `dhcp-options/router` values.

{% code title="Example: Insert first into a ordered-by user leaf-list " %}

```bash
# Note: we have to split the POST line in order to fit the page
POST /restconf/data/dhcp:dhcp/subnet=10.254.239.0%2F27/dhcp-options?\
     insert=after&\
     point=/dhcp:dhcp/subnet=10.254.239.0%2F27/dhcp-options/router=one.acme.org
Content-Type: application/yang-data+xml

<router>two.acme.org</router>

# If the resource is created, the server might respond as follows:

HTTP/1.1 201 Created
Location /restconf/data/dhcp:dhcp/subnet=10.254.239.0%2F27/dhcp-options/\
         router=one.acme.org
```

{% endcode %}

To verify that the `router` value really ended up after our insertion point:

```http
GET /restconf/data/dhcp:dhcp/subnet=10.254.239.0%2F27/dhcp-options
Accept: application/yang-data+xml

HTTP/1.1 200 OK
<dhcp-options xmlns="http://yang-central.org/ns/example/dhcp"
              xmlns:dhcp="http://yang-central.org/ns/example/dhcp">
  <router>one.acme.org</router>
  <router>two.acme.org</router>
  <router>rtr-239-0-1.example.org</router>
  <router>rtr-239-0-2.example.org</router>
</dhcp-options>
```

### Additional Query Parameters <a href="#d5e1722" id="d5e1722"></a>

There are additional NSO query parameters available for the RESTCONF API. Commit behavior should be supplied through the shared `tailf-ncs-commit-params` model described in [Commit Parameters](/guides/operation-and-usage/operations/lifecycle-operations#d5e5048).

For RESTCONF, the shared commit-parameter structure is transported in one of two ways:

* The `params` query parameter, whose value is URL-encoded base64 JSON representing the contents of `tailf-ncs-commit-params:commit-params`.
* The `X-Cisco-NSO-Commit-Params` HTTP header, which carries the same base64-encoded JSON value. This is an alternative to the `params` query parameter when the client prefers not to place the often long, URL-encoded base64 commit-parameter payload in the request URI, for example when other query parameters are also used.

In other words, NSO expects the structured commit-parameter payload to be serialized as JSON and then base64-encoded for transport. URL encoding is additionally required only when that base64 string is placed in the `params` query parameter.

Example payload:

```json
{
  "label": "restconf-demo",
  "dry-run": {
    "outformat": "cli-c"
  }
}
```

Example request using the `params` query parameter:

```http
PATCH /restconf/data/.../search?params=eyJsYWJlbCI6ICJyZXN0Y29uZi1kZW1vIiwgImRyeS1ydW4iOiB7Im91dGZvcm1hdCI6ICJjbGktYyJ9fQ==
```

Example request using the header:

```http
X-Cisco-NSO-Commit-Params: eyJsYWJlbCI6ICJyZXN0Y29uZi1kZW1vIiwgImRyeS1ydW4iOiB7Im91dGZvcm1hdCI6ICJjbGktYyJ9fQ==
```

See the [examples.ncs/northbound-interfaces/commit-parameters](https://github.com/NSO-developer/nso-examples/tree/6.7/northbound-interfaces/commit-parameters) example for end-to-end requests using both the query parameter and the header form, together with the equivalent CLI commit flags.

Other NSO-specific query parameters are listed below.

<table data-full-width="false"><thead><tr><th width="200" valign="top">Name</th><th width="161" valign="top">Methods</th><th valign="top">Description</th></tr></thead><tbody><tr><td valign="top"><code>params</code></td><td valign="top"><code>POST</code><br><code>PUT</code><br><code>PATCH</code><br><code>DELETE</code></td><td valign="top">URL-encoded base64 JSON for the shared commit-parameter structure defined in <code>tailf-ncs-commit-params.yang</code>.</td></tr><tr><td valign="top"><code>limit</code></td><td valign="top"><code>GET</code></td><td valign="top">Used by the client to specify a limited set of list entries to retrieve. The value is either an integer greater than or equal to <code>1</code>, or the string <code>unbounded</code>. The string <code>unbounded</code> is the default value. See <a href="#ncs.northbound.partial_response">Partial Responses</a> for an example.</td></tr><tr><td valign="top"><code>offset</code></td><td valign="top"><code>GET</code></td><td valign="top">Used by the client to specify the number of list elements to skip before returning the requested set of list entries. The value is an integer greater than or equal to <code>0</code>. The default value is <code>0</code>. See <a href="#ncs.northbound.partial_response">Partial Responses</a> for an example.</td></tr><tr><td valign="top"><code>rollback-id</code></td><td valign="top"><code>POST</code><br><code>PUT</code><br><code>PATCH</code><br><code>DELETE</code></td><td valign="top">Return the rollback ID in the response if a rollback file was created during this operation. This requires rollbacks to be enabled in NSO to take effect.</td></tr><tr><td valign="top"><code>with-service-meta-data</code></td><td valign="top"><code>GET</code></td><td valign="top">Include FASTMAP attributes such as backpointers and reference counters in the reply. These are typically internal to NSO and thus not shown by default.</td></tr></tbody></table>

## Edit Collision Prevention

Two edit collision detection and prevention mechanisms are provided in RESTCONF for the datastore resource: a timestamp and an entity tag. Any change to configuration data resources will update the timestamp and entity tag of the datastore resource. This makes it possible for a client to apply precondition HTTP headers to a request.

The NSO RESTCONF API honors the following HTTP response headers: `Etag` and `Last-Modified`, and the following request headers: `If-Match`, `If-None-Match`, `If-Modified-Since`, and `If-Unmodified-Since`.

### Response Headers <a href="#d5e1892" id="d5e1892"></a>

* `Etag`: This header will contain an entity tag which is an opaque string representing the latest transaction identifier in the NSO database. This header is only available for the running datastore and hence, only relates to configuration data (non-operational).
* `Last-Modified`: This header contains the timestamp for the last modification made to the NSO database. This timestamp can be used by a RESTCONF client in subsequent requests, within the `If-Modified-Since` and `If-Unmodified-Since` header fields. This header is only available for the running datastore and hence, only relates to configuration data (non-operational).

### Request Headers <a href="#d5e1907" id="d5e1907"></a>

* `If-None-Match`: This header evaluates to true if the supplied value does not match the latest `Etag` entity-tag value. If evaluated to false, an error response with status 304 (Not Modified) will be sent with no body. This header carries only meaning if the entity tag of the `Etag` response header has previously been acquired. The usage of this could for example be a HEAD operation to get information if the data has changed since the last retrieval.
* `If-Modified-Since`: This request-header field is used with an HTTP method to make it conditional, i.e if the requested resource has not been modified since the time specified in this field, the request will not be processed by the RESTCONF server; instead, a 304 (Not Modified) response will be returned without any message-body. Usage of this is for instance for a GET operation to retrieve the information if (and only if) the data has changed since the last retrieval. Thus, this header should use the value of a `Last-Modified` response header that has previously been acquired.
* `If-Match`: This header evaluates to true if the supplied value matches the latest `Etag` value. If evaluated to false, an error response with status 412 (Precondition Failed) will be sent with no body. This header carries only meaning if the entity tag of the `Etag` response header has previously been acquired. The usage of this can be in the case of a `PUT`, where `If-Match` can be used to prevent the lost update problem. It can check if the modification of a resource that the user wants to upload will not override another change that has been done since the original resource was fetched.
* `If-Unmodified-Since`: This header evaluates to true if the supplied value has not been last modified after the given date. If the resource has been modified after the given date, the response will be a 412 (Precondition Failed) error with no body. This header carries only meaning if the `Last-Modified` response header has previously been acquired. The usage of this can be the case of a `POST`, where editions are rejected if the stored resource has been modified since the original value was retrieved.

## Using Rollbacks <a href="#ug.restconf.using_rollbacks" id="ug.restconf.using_rollbacks"></a>

### Rolling Back Configuration Changes <a href="#d5e1936" id="d5e1936"></a>

If rollbacks have been enabled in the configuration using the `rollback-id` query parameter, the fixed ID of the rollback file created during an operation is returned in the results. The below examples show the creation of a new resource and the removal of that resource using the rollback created in the first step.

{% code title="Example: Create a New dhcp/subnet Resource" %}

```http
POST /restconf/data/dhcp:dhcp?rollback-id=true
Content-Type: application/yang-data+xml

<subnet xmlns="http://yang-central.org/ns/example/dhcp">
  <net>10.254.239.0/27</net>
</subnet>

HTTP/1.1 201 Created
Location: http://localhost:8008/restconf/data/dhcp:dhcp/subnet=10.254.239.0%2F27

<result xmlns="http://tail-f.com/ns/tailf-restconf">
<rollback>
  <id>10002</id>
</rollback>
</result>
```

{% endcode %}

Then using the fixed ID returned above as input to the `apply-rollback-file` action:

```http
POST /restconf/data/tailf-rollback:rollback-files/apply-rollback-file
Content-Type: application/yang-data+xml

<input xmlns="http://tail-f.com/ns/rollback">
  <fixed-number>10002</fixed-number>
</input>

HTTP/1.1 204 No Content
```

## Streams <a href="#ncs.northbound.restconf.streams" id="ncs.northbound.restconf.streams"></a>

### Introduction <a href="#d5e1949" id="d5e1949"></a>

The RESTCONF protocol supports YANG-defined event notifications. The solution preserves aspects of NETCONF event notifications \[RFC5277] while utilizing the Server-Sent Events, [W3C.REC-eventsource-20150203](https://www.w3.org/TR/2015/REC-eventsource-20150203), transport strategy.

RESTCONF event notification streams are described in Sections 6 and 9.2 of [RFC 8040](https://www.ietf.org/rfc/rfc8040.txt), where also notification examples can be found.

RESTCONF event notification is a way for RESTCONF clients to retrieve notifications for different event streams. Event streams configured in NSO can be subscribed to using different channels such as the RESTCONF or the NETCONF channel.

More information on how to define a new notification event using Yang is described in [RFC 6020](https://www.ietf.org/rfc/rfc6020.txt).

How to add and configure notifications support in NSO is described in the `ncs.conf(3)` man page.

The design of RESTCONF event notification is inspired by how NETCONF event notification is designed. More information on NETCONF event notification can be found in [RFC 5277](https://www.ietf.org/rfc/rfc5277.txt).

### Configuration <a href="#d5e1964" id="d5e1964"></a>

For this example, we will define a notification stream, named `interface` in the `ncs.conf` configuration file as shown below.

We also enable the built-in replay store which means that NSO automatically stores all notifications on disk, ready to be replayed should a RESTCONF event notification subscriber ask for logged notifications. The replay store uses a set of wrapping log files on a disk (of a certain number and size) to store the notifications.

{% code title="Example: Configure an Example Notification" %}

```xml
<notifications>
  <eventStreams>
    <stream>
      <name>interface</name>
      <description>Example notifications</description>
      <replaySupport>true</replaySupport>
      <builtinReplayStore>
        <dir>./</dir>
        <maxSize>S1M</maxSize>
        <maxFiles>5</maxFiles>
      </builtinReplayStore>
    </stream>
  </eventStreams>
</notifications>
```

{% endcode %}

To view the currently enabled event streams, use the `ietf-restconf-monitoring` YANG model. The streams are available under the `/restconf/data/ietf-restconf-monitoring:restconf-state/streams` container.

{% code title="Example: View the Example RESTCONF Stream" %}

```http
GET /restconf/data/ietf-restconf-monitoring:restconf-state/streams
Accept: application/yang-data+xml

HTTP/1.1 200 OK

<streams xmlns="urn:ietf:params:xml:ns:yang:ietf-restconf-monitoring"
         xmlns:rcmon="urn:ietf:params:xml:ns:yang:ietf-restconf-monitoring">

  ...other streams info removed here for brewity reason...

  <stream>
    <name>interface</name>
    <description>Example notifications</description>
    <replay-support>true</replay-support>
    <replay-log-creation-time>
      2020-05-04T13:45:31.033817+00:00
    </replay-log-creation-time>
    <access>
      <encoding>xml</encoding>
      <location>https://localhost:8888/restconf/streams/interface/xml</location>
    </access>
    <access>
      <encoding>json</encoding>
      <location>https://localhost:8888/restconf/streams/interface/json</location>
    </access>
  </stream>
</streams>
```

{% endcode %}

Note the URL value we get in the *location* element in the example above. This URL should be used when subscribing to the notification events as is shown in the next example.

### Subscribe to Notification Events <a href="#d5e1982" id="d5e1982"></a>

RESTCONF clients can determine the URL for the subscription resource (to receive notifications) by sending an HTTP GET request for the `location` leaf with the `stream` list entry. The value returned by the server can be used for the actual notification subscription.

The client will send an HTTP GET request for the (location) URL returned by the server with the `Accept` type `text/event-stream` as shown in the example below. Note that this request works like a long polling request which means that the request will not return. Instead, server-side notifications will be sent to the client where each line of the notification will be prepended with `data:`.

{% code title="Example: View the Example RESTCONF Stream" %}

```http
GET /restconf/streams/interface/xml
Accept: text/event-stream

   ...NOTE: we will be waiting here until a notification is generated...

HTTP/1.1 200 OK
Content-Type: text/event-stream

data: <notification xmlns='urn:ietf:params:xml:ns:netconf:notification:1.0'>
data:     <eventTime>2020-05-04T13:48:02.291816+00:00</eventTime>
data:     <link-up xmlns='http://tail-f.com/ns/test/notif'>
data:       <if-index>2</if-index>
data:       <link-property>
data:         <newly-added/>
data:         <flags>42</flags>
data:         <extensions>
data:           <name>1</name>
data:           <value>3</value>
data:         </extensions>
data:         <extensions>
data:           <name>2</name>
data:           <value>4668</value>
data:         </extensions>
data:       </link-property>
data:     </link-up>
data: </notification>

   ...NOTE: we will still be waiting here for more notifications to come...
```

{% endcode %}

Since we have enabled the replay store, we can ask the server to replay any notifications generated since the specific date we specify. After those notifications have been delivered, we will continue waiting for new notifications to be generated.

{% code title="Example: View the Example RESTCONF Stream" %}

```http
GET /restconf/streams/interface/xml?start-time=2007-07-28T15%3A23%3A36Z
Accept: text/event-stream

HTTP/1.1 200 OK
Content-Type: text/event-stream

data: ...any existing notification since given date will be delivered here...

   ...NOTE: when all notifications are delivered, we will be waiting here for more...
```

{% endcode %}

### Errors

Errors occurring during streaming of events will be reported as Server-Sent Events (SSE) comments as described in [W3C.REC-eventsource-20150203](https://www.w3.org/TR/2015/REC-eventsource-20150203) as shown in the example below.

{% code title="Example: NSO RESTCONF Errors During Streaming" %}

```
: error: notification stream NETCONF temporarily unavailable
```

{% endcode %}

## Dynamic Subscriptions

This section describes how Subscribed Notifications and YANG-Push are implemented for RESTCONF. Dynamic subscriptions for RESTCONF are described in [RFC 8650](https://www.ietf.org/rfc/rfc8650.txt), YANG-Push is described in [RFC 8641](https://www.ietf.org/rfc/rfc8641.txt), and Subscribed Notifications are described in [RFC 8639](https://www.ietf.org/rfc/rfc8639.txt).

Subscribed notifications and YANG-Push in RESTCONF use the same underlying mechanism as NETCONF and therefore take the same input when establishing, modifying, deleting, killing, or re-syncing a subscription, as well as give the same notification messages for the same scenarios. The main difference is in how the subscription is started. This is more similar to how subscriptions to notification events are done for RESTCONF event streams. To start a subscription, one must first send a POST request to the establish-subscription RPC. This will respond with an ID for the subscription, as well as a URI to which a subsequent GET request can be made. This GET request will start a session for the subscription that will be used to receive notifications. The URI includes the ID for the subscription. The `Accept` header will be `text/event-stream` as shown in the example below. This process is described in more detail in [RFC 8650](https://www.ietf.org/rfc/rfc8650.txt). Just as with RESTCONF Event Streams, the GET request works like along polling request and will not return, instead waiting for notifications to arrive. Each line of the notification will have the prefix `data:` .

{% code title="Example: An Establish-subscription Request" %}

```http
POST /restconf/operations/ietf-subscribed-notifications:establish-subscription
Content-Type: application/yang-data+xml

<output xmlns='urn:ietf:params:xml:ns:yang:ietf-subscribed-notifications'>
  <id>1</id>
  <uri xmlns='urn:ietf:params:xml:ns:yang:ietf-restconf-subscribed-notifications'>\
              http://localhost:8080/restconf/subscriptions/1</uri>
</output>
```

{% endcode %}

{% code title="Example: Subscribed Notification" %}

```http
GET /restconf/subscriptions/1
Accept: text/event-stream

HTTP/1.1 200 OK
Content-Type: text/event-stream

  ...NOTE: we will be waiting here until a notification is generated...

data: <notification xmlns='urn:ietf:params:xml:ns:netconf:notification:1.0'>
data:     <eventTime>2020-05-04T13:48:02.291816+00:00</eventTime>
data:     <test xmlns='urn:test'>
data:       <name>notif1</name>
data:     </test>
data: </notification>

  ...NOTE: we will still be waiting here for more notifications to come...
```

{% endcode %}

{% code title="Example: A Subscribed Notifications Payload" %}

```xml
<input xmlns="urn:ietf:params:xml:ns:yang:ietf-subscribed-notifications">
  <stream-xpath-filter
    xmlns="urn:ietf:params:xml:ns:yang:ietf-subscribed-notifications"
    xmlns:test="urn:test">
      /test:test/name
  </stream-xpath-filter>
  <stream>interface</stream>
  <replay-start-time>2018-10-04T14:10:02.133651392+02:00</replay-start-time>
  <stop-time>2030-03-27T20:03:02.133651392+02:00</stop-time>
  <encoding>encode-xml</encoding>
</input>
```

{% endcode %}

To modify, delete, kill, or resync a subscription, a POST request is done to the modify-subscription, delete-subscription, kill-subscription, or resync-subscription RPC respectively.

Another way that RESTCONF dynamic subscriptions differ from NETCONF is when deleting a subscription. In NETCONF, when a subscription is deleted the session is not terminated, since it is possible to do other operations in the open session. In RESTCONF, however, a GET request session only receives notifications for a subscription, so when the subscription is deleted there is no reason to keep the session open. Therefore, a "subscription-terminated" notification will be sent when deleting a subscription, followed by the session closing.

Note that RFC 8650 states: "There cannot be two or more simultaneous GET requests on a subscription URI: any GET request received while there is a current GET request on the same URI MUST be rejected with HTTP error code 409." Therefore, if a GET request is already active for a subscription, no new GET request will be allowed to the same URI. If a user wants to be able to end a GET request session and then start a new one to the same subscription, they would have to set the ncs.conf setting `/ncs-config/webui/transport/tcp/keepalive` to true, as well as set the `/ncs-config/webui/transport/tcp/keepalive-timeout` to a desired value. This is needed to determine if a GET request session has been closed, so that a new one can be opened. Each `keepalive-timeout`, an SSE comment will be sent to the socket, which allows the process to notice if the socket has been closed.

### Limitations <a href="#d5e2039" id="d5e2039"></a>

RESTCONF Subscribed Notifications and YANG-Push have the same limitation as NETCONF\
Subscribed Notifications and YANG-Push. In addition, in RESTCONF subtree filtering is not supported.\
Furthermore, the JSON format is not supported, which deviates from RFC 8650. For details see section\
`Protocol YANG Modules`.

## Schema Resource

RFC 8040, Section 3.7 describes the retrieval of YANG modules used by the server via the RPC operation `get-schema`. The YANG source is made available by NSO in two ways: compiled into the `fxs` file or put in the loadPath. See [Monitoring of the NETCONF Server](#ug.netconf_agent.monitoring).

The example below shows how to list the available Yang modules. Since we are interested in the `dhcp` module, we only show that part of the output:

{% code title="Example: List the Available Yang Modules" %}

```http
GET /restconf/data/ietf-yang-library:modules-state
Accept: application/yang-data+xml

HTTP/1.1 200 OK
<modules-state xmlns="urn:ietf:params:xml:ns:yang:ietf-yang-library"
               xmlns:yanglib="urn:ietf:params:xml:ns:yang:ietf-yang-library">
  <module-set-id>f4709e88d3250bd84f2378185c2833c2</module-set-id>
  <module>
    <name>dhcp</name>
    <revision>2019-02-14</revision>
    <schema>http://localhost:8080/restconf/tailf/modules/dhcp/2019-02-14</schema>
    <namespace>http://yang-central.org/ns/example/dhcp</namespace>
    <conformance-type>implement</conformance-type>
  </module>

  ...rest of the output removed here...

</modules-state>
```

{% endcode %}

We can now retrieve the `dhcp` Yang module via the URL we got in the `schema` leaf of the reply. Note that the actual URL may point anywhere. The URL is configured by the `schemaServerUrl` setting in the `ncs.conf` file.

```http
GET /restconf/tailf/modules/dhcp/2019-02-14

HTTP/1.1 200 OK
module dhcp {
  namespace "http://yang-central.org/ns/example/dhcp";
  prefix dhcp;

  import ietf-yang-types {

  ...the rest of the Yang module removed here...
```

## YANG Patch Media Type <a href="#d5e2028" id="d5e2028"></a>

The NSO RESTCONF API also supports the YANG Patch Media Type, as defined in [RFC 8072](https://www.ietf.org/rfc/rfc8072.txt).

A YANG `Patch` is an ordered list of edits that are applied to the target datastore by the RESTCONF server. A YANG Patch request is sent as an HTTP PATCH request containing a body describing the edit operations to be performed. The format of the body is defined in the [RFC 8072](https://www.ietf.org/rfc/rfc8072.txt).

Referring to the example above (DHCP Yang model) in the [Getting Started](#ncs.northbound.restconf.getting_started) section; we will show how to use YANG Patch to achieve the same result but with fewer amount of requests.

### Create Two New Resources with the YANG Patch <a href="#d5e2039" id="d5e2039"></a>

To create the resources, we send an HTTP PATCH request where the `Content-Type` indicates that the body in the request consists of a `Yang-Patch` message. Our `Yang-Patch` request will initiate two edit operations where each operation will create a new subnet. In contrast, compare this with using plain RESTCONF where we would have needed two `POST` requests to achieve the same result.

{% code title="Example: Create a Two New dhcp/subnet Resources" %}

```http
PATCH /restconf/data/dhcp:dhcp
Accept: application/yang-data+xml
Content-Type: application/yang-patch+xml

<yang-patch xmlns="urn:ietf:params:xml:ns:yang:ietf-yang-patch">
  <patch-id>add-subnets</patch-id>
  <edit>
    <edit-id>add-subnet-239</edit-id>
    <operation>create</operation>
    <target>/subnet=10.254.239.0%2F27</target>
    <value>
      <subnet xmlns="http://yang-central.org/ns/example/dhcp" \
              xmlns:dhcp="http://yang-central.org/ns/example/dhcp">
        <net>10.254.239.0/27</net>
          ...content removed here for brevity...
        <max-lease-time>1200</max-lease-time>
      </subnet>
    </value>
  </edit>
  <edit>
    <edit-id>add-subnet-244</edit-id>
    <operation>create</operation>
    <target>/subnet=10.254.244.0%2F27</target>
    <value>
      <subnet xmlns="http://yang-central.org/ns/example/dhcp" \
              xmlns:dhcp="http://yang-central.org/ns/example/dhcp">
        <net>10.254.244.0/27</net>
          ...content removed here for brevity...
        <max-lease-time>1200</max-lease-time>
      </subnet>
    </value>
  </edit>
</yang-patch>

# If the YANG Patch request was successful,
# the server might respond as follows:

HTTP/1.1 200 OK
<yang-patch-status xmlns="urn:ietf:params:xml:ns:yang:ietf-yang-patch">
  <patch-id>add-subnets</patch-id>
  <ok/>
</yang-patch-status>
```

{% endcode %}

### Modify and Delete in the Same Yang-Patch Request

Let us modify the `max-lease-time` of one subnet and delete the `max-lease-time` value of the second subnet. Note that the delete will cause the default value of `max-lease-time` to take effect, which we will verify using a RESTCONF GET request.

{% code title="Example: Modify and Delete in the Same Yang-Patch Request" %}

```http
PATCH /restconf/data/dhcp:dhcp
Accept: application/yang-data+xml
Content-Type: application/yang-patch+xml

<yang-patch xmlns="urn:ietf:params:xml:ns:yang:ietf-yang-patch">
  <patch-id>modify-and-delete</patch-id>
  <edit>
    <edit-id>modify-max-lease-time-239</edit-id>
    <operation>merge</operation>
    <target>/dhcp:subnet=10.254.239.0%2F27</target>
    <value>
      <subnet xmlns="http://yang-central.org/ns/example/dhcp" \
              xmlns:dhcp="http://yang-central.org/ns/example/dhcp">
        <net>10.254.239.0/27</net>
        <max-lease-time>1234</max-lease-time>
      </subnet>
    </value>
  </edit>
  <edit>
    <edit-id>delete-max-lease-time-244</edit-id>
    <operation>delete</operation>
    <target>/dhcp:subnet=10.254.244.0%2F27/max-lease-time</target>
  </edit>
</yang-patch>

# If the YANG Patch request was successful,
# the server might respond as follows:

HTTP/1.1 200 OK
<yang-patch-status xmlns="urn:ietf:params:xml:ns:yang:ietf-yang-patch">
  <patch-id>modify-and-delete</patch-id>
  <ok/>
</yang-patch-status>
```

{% endcode %}

To verify that our modify and delete operations took place we make use of two RESTCONF `GET` requests as shown below.

{% code title="Example: Verify the Modified " %}

```http
GET /restconf/data/dhcp:dhcp/subnet=10.254.239.0%2F27/max-lease-time
Accept: application/yang-data+xml

HTTP/1.1 200 OK
<max-lease-time xmlns="http://yang-central.org/ns/example/dhcp"
                xmlns:dhcp="http://yang-central.org/ns/example/dhcp">
                1234
</max-lease-time>
```

{% endcode %}

{% code title="Example: Verify the Default Values after Delete of the " %}

```http
GET /restconf/data/dhcp:dhcp/subnet=10.254.244.0%2F27/max-lease-time?\
      with-defaults=report-all-tagged
Accept: application/yang-data+xml

HTTP/1.1 200 OK
<max-lease-time wd:default="true"
                xmlns:wd="urn:ietf:params:restconf:capability:defaults:1.0"
                xmlns="http://yang-central.org/ns/example/dhcp"
                xmlns:dhcp="http://yang-central.org/ns/example/dhcp">
                7200
</max-lease-time>
```

{% endcode %}

Note how we in the last `GET` request make use of the `with-defaults` query parameter to request that a default value should be returned and also be tagged as such.

## NMDA <a href="#d5e2067" id="d5e2067"></a>

Network Management Datastore Architecture (NMDA), as defined in [RFC 8527](https://www.ietf.org/rfc/rfc8527.txt), extends the RESTCONF protocol. This enables RESTCONF clients to discover which datastores are supported by the RESTCONF server, determine which modules are supported in each datastore, and interact with all the datastores supported by the NMDA.

A RESTCONF client can test if a server supports the NMDA by using either the `HEAD` or `GET` methods on `/restconf/ds/ietf- datastores:operational`, as shown below:

{% code title="Example: Check if the RESTCONF Server Support NMDA" %}

```
HEAD /restconf/ds/ietf-datastores:operational

HTTP/1.1 200 OK
```

{% endcode %}

A RESTCONF client can discover which datastores and YANG modules the server supports by reading the YANG library information from the operational state datastore. Note in the example below that, since the result consists of three top nodes, it can't be represented in XML; hence we request the returned content to be in JSON format. See also [Collections](#ncs.northbound.restconf.extensions.collections).

{% code title="Example: Check Which Datastores the RESTCONF Server Supports" %}

```http
GET /restconf/ds/ietf-datastores:operational/datastore
Accept: application/yang-data+json

HTTP/1.1 200 OK
{
  "ietf-yang-library:datastore": [
    {
      "name": "ietf-datastores:running",
      "schema": "common"
    },
    {
      "name": "ietf-datastores:intended",
      "schema": "common"
    },
    {
      "name": "ietf-datastores:operational",
      "schema": "common"
    }
  ]
}
```

{% endcode %}

## Extensions

To avoid any potential future conflict with the RESTCONF standard, any extensions made to the NSO implementation of RESTCONF are located under the URL path: `/restconf/tailf`, or is controlled by means of a vendor-specific media type.

{% hint style="info" %}
There is no index of extensions under `/restconf/tailf`. To list extensions, access `/restconf/data/ietf-yang-library:modules-state` and follow published links for schemas.
{% endhint %}

## Collections <a href="#ncs.northbound.restconf.extensions.collections" id="ncs.northbound.restconf.extensions.collections"></a>

The RESTCONF specification states that a result containing multiple instances (e.g. a number of list entries) is not allowed if XML encoding is used. The reason for this is that an XML document can only have one root node.

This functionality is supported if the `http://tail-f.com/ns/restconf/collection/1.0` capability is presented. See also [How to View the Capabilities of the RESTCONF Server](#ncs.northbound.restconf.capabilities).

To remedy this, an HTTP GET request can make use of the `Accept:` media type: `application/vnd.yang.collection+xml` as shown in the following example. The result will then be wrapped within a `collection` element.

{% code title="Example: Use of Collections" %}

```http
GET /restconf/ds/ietf-datastores:operational/\
    ietf-yang-library:yang-library/datastore
Accept: application/vnd.yang.collection+xml

<collection xmlns="http://tail-f.com/ns/restconf/collection/1.0">
  <datastore xmlns="urn:ietf:params:xml:ns:yang:ietf-yang-library"
            xmlns:yanglib="urn:ietf:params:xml:ns:yang:ietf-yang-library">
    <name xmlns:ds="urn:ietf:params:xml:ns:yang:ietf-datastores">
       ds:running
    </name>
    <schema>common</schema>
  </datastore>
  <datastore xmlns="urn:ietf:params:xml:ns:yang:ietf-yang-library"
             xmlns:yanglib="urn:ietf:params:xml:ns:yang:ietf-yang-library">
    <name xmlns:ds="urn:ietf:params:xml:ns:yang:ietf-datastores">
      ds:intended
    </name>
    <schema>common</schema>
  </datastore>
  <datastore xmlns="urn:ietf:params:xml:ns:yang:ietf-yang-library
             xmlns:yanglib="urn:ietf:params:xml:ns:yang:ietf-yang-library">
    <name xmlns:ds="urn:ietf:params:xml:ns:yang:ietf-datastores">
      ds:operational
    </name>
    <schema>common</schema>
  </datastore>
</collection>
```

{% endcode %}

## The RESTCONF Query API

The NSO RESTCONF Query API consists of a number of operations to start a query which may live over several RESTCONF requests, where data can be fetched in suitable chunks. The data to be returned is produced by applying an XPath expression where the data also may be sorted.

The RESTCONF client can check if the NSO RESTCONF server supports this functionality by looking for the `http://tail-f.com/ns/restconf/query-api/1.0` capability. See also [How to View the Capabilities of the RESTCONF Server](#ncs.northbound.restconf.capabilities).

The `tailf-rest-query.yang` and the `tailf-common-query.yang` YANG models describe the structure of the RESTCONF Query API messages. By using the Schema Resource functionality, as described in [Schema Resource](#schema-resource), you can get hold of them.

### Request and Replies <a href="#d5e2116" id="d5e2116"></a>

The API consists of the following requests:

* `start-query`: Start a query and return a query handle.
* `fetch-query-result`: Use a query handle to repeatedly fetch chunks of the result.
* `immediate-query`: Start a query and return the entire result immediately.
* `reset-query`: (Re)set where the next fetched result will begin from.
* `stop-query`: Stop (and close) the query.

The API consists of the following replies:

* `start-query-result`: Reply to the start-query request.
* `query-result`: Reply to the fetch-query-result and immediate-query requests.

In the following examples, we'll use this data model:

{% code title="Example: Example.yang : model for the Query API Example " %}

```yang
container x {
  list host {
    key number;
    leaf number {
      type int32;
    }
    leaf enabled {
      type boolean;
    }
    leaf name {
      type string;
    }
    leaf address {
      type inet:ip-address;
    }
  }
}]
```

{% endcode %}

The actual format of the payload should be represented either in XML or JSON. Note how we indicate the type of content using the `Content-Type` HTTP header. For XML, it could look like this:

{% code title="Example: Example of a start-query Request" %}

```http
POST /restconf/tailf/query
Content-Type: application/yang-data+xml

<start-query xmlns="http://tail-f.com/ns/tailf-rest-query">
  <foreach>
    /x/host[enabled = 'true']
  </foreach>
  <select>
    <label>Host name</label>
    <expression>name</expression>
    <result-type>string</result-type>
  </select>
  <select>
    <expression>address</expression>
    <result-type>string</result-type>
  </select>
  <sort-by>name</sort-by>
  <limit>100</limit>
  <offset>1</offset>
  <timeout>600</timeout>
</start-query>]
```

{% endcode %}

The same request in JSON format would look like:

{% code title="Example: JSON example of a start-query Request" %}

```http
POST /restconf/tailf/query
Content-Type: application/yang-data+json

{
 "start-query": {
   "foreach": "/x/host[enabled = 'true']",
   "select": [
     {
       "label": "Host name",
       "expression": "name",
       "result-type": ["string"]
     },
     {
       "expression": "address",
       "result-type": ["string"]
     }
   ],
   "sort-by": ["name"],
   "limit": 100,
   "offset": 1,
   "timeout": 600
 }
}]
```

{% endcode %}

An informal interpretation of this query is:

For each `/x/host` where `enabled` is true, select its `name`, and `address`, and return the result sorted by `name`, in chunks of 100 result items at a time.

Let us discuss the various pieces of this request. To start with, when using XML, we need to specify the namespace as shown:

```xml
<start-query xmlns="http://tail-f.com/ns/tailf-rest-query">
```

The actual XPath query to run is specified by the `foreach` element. The example below will search for all `/x/host` nodes that have the `enabled` node set to `true`:

```xml
<foreach>
  /x/host[enabled = 'true']
</foreach>
```

{% hint style="info" %}
Note that the `foreach` element, specifying an XPath, expects nodes qualified with YANG module prefix, not YANG module name as is customary elsewhere in RESTCONF.
{% endhint %}

Now we need to define what we want to have returned from the node set by using one or more `select` sections. What to actually return is defined by the XPath `expression`.

Choose how the result should be represented. Basically, it can be the actual value or the path leading to the value. This is specified per select chunk. The possible result types are `string`, `path`, `leaf-value`*,* and `inline`.

The difference between `string` and `leaf-value` is somewhat subtle. In the case of `string`, the result will be processed by the XPath function: `string()` (which if the result is a node-set will concatenate all the values). The `leaf-value` will return the value of the first node in the result. As long as the result is a leaf node, `string` and `leaf-value` will return the same result. In the example above, the `string` is used as shown below. Note that at least one `result-type` must be specified.

The result-type `inline` makes it possible to return the full sub-tree of data, either in XML or in JSON format. The data will be enclosed with a tag: `data`.

It is possible to specify an optional `label` for a convenient way of labeling the returned data:

```xml
<select>
  <label>Host name</label>
  <expression>name</expression>
  <result-type>string</result-type>
</select>
<select>
  <expression>address</expression>
  <result-type>string</result-type>
</select>
```

The returned result can be sorted. This is expressed as an XPath expression, which in most cases is very simple and refers to the found node-set. In this example, we sort the result by the content of the `name` node:

```xml
<sort-by>name</sort-by>
```

With the `offset` element, we can specify at which node we should start to receive the result. The default is 1, i.e., the first node in the resulting node set.

```xml
<offset>1</offset>
```

It is possible to set a custom timeout when starting or resetting a query. Each time a function is called, the timeout timer resets. The default is 600 seconds, i.e. 10 minutes.

```xml
<timeout>600</timeout>
```

The reply to this request would look something like this:

```xml
<start-query-result>
  <query-handle>12345</query-handle>
</start-query-result>
```

The query handle (in this example '12345') must be used in all subsequent calls. To retrieve the result, we can now send:

```xml
<fetch-query-result xmlns="http://tail-f.com/ns/tailf-rest-query">
  <query-handle>12345</query-handle>
</fetch-query-result>
```

Which will result in something like the following:

```xml
<query-result xmlns="http://tail-f.com/ns/tailf-rest-query">
  <result>
    <select>
      <label>Host name</label>
      <value>One</value>
    </select>
    <select>
      <value>10.0.0.1</value>
    </select>
  </result>
  <result>
    <select>
      <label>Host name</label>
      <value>Three</value>
    </select>
    <select>
      <value>10.0.0.3</value>
    </select>
  </result>
</query-result>
```

If we try to get more data with the `fetch-query-result`, we might get more `result` entries in return until no more data exists and we get an empty query result back:

```xml
<query-result xmlns="http://tail-f.com/ns/tailf-rest-query">
</query-result>
```

Finally, when we are done we stop the query:

```xml
<stop-query xmlns="http://tail-f.com/ns/tailf-rest-query">
  <query-handle>12345</query-handle>
</stop-query>
```

### Reset a Query

If we want to go back into the stream of received data chunks and have them repeated, we can do that with the `reset-query` request. In the example below, we ask to get results from the 42nd result entry:

```xml
<reset-query xmlns="http://tail-f.com/ns/tailf-rest-query">
  <query-handle>12345</query-handle>
  <offset>42</offset>
</reset-query>
```

### Immediate Query <a href="#d5e2226" id="d5e2226"></a>

If we want to get the entire result sent back to us, using only one request, we can do this by using the `immediate-query`. This function takes similar arguments as `start-query` and returns the entire result analogous with the result from a `fetch-query-result` request. Note that it is not possible to paginate or set an offset start node for the result list; i.e. the options `limit` and `offset` are ignored.

## Partial Responses <a href="#ncs.northbound.partial_response" id="ncs.northbound.partial_response"></a>

This functionality is supported if the `http://tail-f.com/ns/restconf/partial-response/1.0` capability is presented. See also [How to View the Capabilities of the RESTCONF Server](#ncs.northbound.restconf.capabilities).

By default, the server sends back the full representation of a resource after processing a request. For better performance, the server can be instructed to send only the nodes the client really needs in a partial response.

To request a partial response for a set of list entries, use the `offset` and `limit` query parameters to specify a limited set of entries to be returned.

In the following example, we retrieve only two entries, skipping the first entry and then returning the next two entries:

{% code title="Example: Partial Response" %}

```http
GET /restconf/data/example-jukebox:jukebox/library/artist?offset=1&limit=2
Accept: application/yang-data+json

...in return we will get the second and third elements of the list...
```

{% endcode %}

## Hidden Nodes

This functionality is supported if the `http://tail-f.com/ns/restconf/unhide/1.0` capability is presented. See also [How to View the Capabilities of the RESTCONF Server](#ncs.northbound.restconf.capabilities).

By default, hidden nodes are not visible in the RESTCONF interface. To unhide hidden nodes for retrieval or editing, clients can use the query parameter `unhide` or set parameter `showHidden` to `true` under `/confdConfig/restconf` in `confd.conf` file. The query parameter `unhide` is supported for RPC and action invocation.

The format of the `unhide` parameter is a comma-separated list of

```xml
<groupname>[;<password>]
```

As an example:

```
unhide=extra,debug;secret
```

This example unhides the unprotected group *extra* and the password-protected group `debug` with the password `secret;`.

## Trace Context

This functionality is supported if the `urn:ietf:params:xml:ns:yang:traceparent:1.0` and `urn:ietf:params:xml:ns:yang:tracestate:1.0` capability is presented. See also [How to View the Capabilities of the RESTCONF Server](#ncs.northbound.restconf.capabilities).

RESTCONF supports the IETF standard draft [I-D.draft-ietf-netconf-restconf-trace-ctx-headers-00](https://www.ietf.org/archive/id/draft-ietf-netconf-restconf-trace-ctx-headers-00.html), that is an adaption of the [W3C Trace Context](https://www.w3.org/TR/2021/REC-trace-context-1-20211123/) standard. Trace Context standardizes the format of `trace-id`, `parent-id`, and key-value pairs to be sent between distributed entities. The `parent-id` will become the `parent-span-id` for the next generated `span-id` in NSO.

Trace Context consists of two HTTP headers `traceparent` and `tracestate`. Header `traceparent` must be of the format

```
traceparent = <version>-<trace-id>-<parent-id>-<flags>
```

where `version` = "00" and `flags` = "01". The support for the values of `version` and `flags` may change in the future depending on the extension of the standard or functionality.

An example of header `traceparent` in use is:

```
traceparent: 00-100456789abcde10123456789abcde10-001006789abcdef0-01
```

Header `tracestate` is a vendor-specific list of key-value pairs. An example of the header `tracestate` in use is:

```
tracestate: key1=value1,key2=value2
```

where a value may contain space characters but not end with a space.

If a request does not include the header `traceparent`, a `traceparent` will be generated internally in NSO. Trace Context is handled by the progress trace functionality, see also [Progress Trace](/guides/development/advanced-development/progress-trace) in Development.

## Configuration Metadata <a href="#d5e2268" id="d5e2268"></a>

It is possible to associate metadata with the configuration data. For RESTCONF, resources such as containers, lists as well as leafs and leaf-lists can have such meta-data. For XML, this meta-data is represented as attributes attached to the XML element in question. For JSON, there does not exist a natural way to represent this info. Hence a special special notation has been introduced, based on the [RFC 7952](https://www.ietf.org/rfc/rfc7952.txt), see the example below.

{% code title="Example: XML Representation of Metadata" %}

```xml
<x xmlns="urn:x" xmlns:x="urn:x">
  <id tags=" important ethernet " annotation="hello world">42</id>
  <person annotation="This is a person">
    <name>Bill</name>
    <person annotation="This is another person">grandma</person>
  </person>
</x>
```

{% endcode %}

{% code title="Example: JSON Representation of Metadata" %}

```json
{
  "x": {
    "foo": 42,
    "@foo": {"tailf_netconf:tags": ["tags","for","foo"],
             "tailf_netconf:annotation": "annotation for foo"},
    "y": {
      "@": {"tailf_netconf:annotation": "Annotation for parent y"},
      "y": 1,
      "@y": {"tailf_netconf:annotation": "Annotation for sibling y"}
    }
  }
}
```

{% endcode %}

The meta-data for an object is represented by another object constructed either of an "@" sign if the meta-data object refers to the parent object, or by the object name prefixed with an "@" sign if the meta-data object refers to a sibling object.

Note that the meta-data node types, e.g., tags and annotations, are prefixed by the module name of the YANG module where the meta-data object is defined. This representation conforms to [RFC 7952 Section 5.2](https://www.rfc-editor.org/rfc/rfc7952.html#section-5.2). The YANG module name prefixes for meta-data node types are listed below:

<table><thead><tr><th valign="top">Meta-data type</th><th valign="top">Prefix</th></tr></thead><tbody><tr><td valign="top"><code>origin</code></td><td valign="top"><code>ietf-origin</code></td></tr><tr><td valign="top"><code>inactive/active</code></td><td valign="top"><code>tailf-netconf-inactive</code></td></tr><tr><td valign="top"><code>default</code></td><td valign="top"><code>ietf-netconf-with-defaults</code></td></tr><tr><td valign="top"><code>All other</code></td><td valign="top"><code>tailf_netconf</code></td></tr></tbody></table>

It is also possible to set meta-data objects in JSON format, except for setting the `default` and `insert` meta-data types, which are not supported using JSON.

## Authentication Cache <a href="#d5e2282" id="d5e2282"></a>

The RESTCONF server maintains an authentication cache. When authenticating an incoming request for a particular `User:Password`, it is first checked if the user exists in the cache and if so, the request is processed. This makes it possible to avoid the, potentially time-consuming, login procedure that will take place in case of a cache miss.

Cache entries have a maximum Time-To-Live (TTL) and upon expiry, a cache entry is removed which will cause the next request for that User to perform the normal login procedure. The TTL value is configurable via the `auth-cache-ttl` parameter, as shown in the example. Note that, by setting the TTL value to `PT0S` (zero), the cache is effectively turned off.

It is also possible to combine the client's IP address with the user name as a key into the cache. This behavior is disabled by default. It can be enabled by setting the `enable-auth-cache-client-ip` parameter to `true`. With this enabled, only a client coming from the same IP address may get a hit in the authentication cache.

{% hint style="info" %}
For considerations specific to RESTCONF deployments that use AAA package authentication, including high-frequency request patterns, see [Package Authentication](https://nso-docs.cisco.com/guides/development/core-concepts/northbound-apis/pages/oUchYxSeOKkwhPOOefVU#ug.aaa.packageauth).
{% endhint %}

{% code title="Example: NSO Configuration of the Authentication Cache TTL" %}

```xml
  ...
  <aaa>
     ...
     <restconf>
        <!-- Set the TTL to 10 seconds! -->
        <auth-cache-ttl>PT10S</auth-cache-ttl>
        <!-- Use both "User" and "ClientIP" as key into the AuthCache -->
        <enable-auth-cache-client-ip>false</enable-auth-cache-client-ip>
     </restconf>
     ...
  </aaa>
  ...
```

{% endcode %}

## Client IP via Proxy

It is possible to configure the NSO RESTCONF server to pick up the client IP address via an HTTP header in the request. A list of HTTP headers to look for is configurable via the `proxy-headers` parameter as shown in the example.

To avoid misuse of this feature, only requests from trusted sources will be searched for such an HTTP header. The list of trusted sources is configured via the `allowed-proxy-ip-prefix` as shown in the example.

{% code title="Example: NSO Configuration of Client IP via Proxy" %}

```xml
  ...
  <webui>
     ...
    <use-forwarded-client-ip>
      <proxy-headers>X-Forwarded-For</proxy-headers>
      <proxy-headers>X-REAL-IP</proxy-headers>
      <allowed-proxy-ip-prefix>10.12.34.0/24</allowed-proxy-ip-prefix>
      <allowed-proxy-ip-prefix>2001:db8:1234::/48</allowed-proxy-ip-prefix>
    </use-forwarded-client-ip>
     ...
  </webui>
  ...
```

{% endcode %}

## External Token Authentication/Validation

The NSO RESTCONF server can be set up to pass a long, a token used for authentication and/or validation of the client. Note that this requires `external authentication/validation` to be set up properly. See [External Token Validation](https://nso-docs.cisco.com/guides/development/core-concepts/northbound-apis/pages/oUchYxSeOKkwhPOOefVU#ug.aaa.external_validation) and [External Authentication](https://nso-docs.cisco.com/guides/development/core-concepts/northbound-apis/pages/oUchYxSeOKkwhPOOefVU#ug.aaa.external_authentication) for details.

With token authentication, we mean that the client sends a `User:Password` to the RESTCONF server, which will invoke an external executable that performs the authentication and upon success produces a token that the RESTCONF server will return in the `X-Auth-Token` HTTP header of the reply.

With token validation, we mean that the RESTCONF server will pass along any token, provided in the `X-Auth-Token` HTTP header, to an external executable that performs the validation. This external program may produce a new token that the RESTCONF server will return in the `X-Auth-Token` HTTP header of the reply.

To make this work, the following need to be configured in the `ncs.conf` file:

{% code title="Example: Configure RESTCONF External Token Authentication/Validation" %}

```xml
  ...
  <restconf>
     ...
    <token-response>
      <x-auth-token>true</x-auth-token>
    </token-response>
     ...
  </restconf>
  ...
```

{% endcode %}

It is also possible to have the RESTCONF server to return a HTTP cookie containing the token.

An HTTP cookie (web cookie, browser cookie) is a small piece of data that a server sends to the user's web browser. The browser may store it and send it back with the next request to the same server. This can be convenient in certain solutions, where typically, it is used to tell if two requests came from the same browser, keeping a user logged in, for example.

To make this happen, the name of the cookie needs to be configured as well as a `directives` string which will be sent as part of the cookie.

{% code title="Example: Configure the RESTCONF Token Cookie" %}

```xml
  ...
  <restconf>
     ...
     <token-cookie>
       <name>X-JWT-ACCESS-TOKEN</name>
       <directives>path=/; Expires=Tue, 19 Jan 2038 03:14:07 GMT;</directives>
     </token-cookie>
     ...
  </restconf>
  ...
```

{% endcode %}

## Custom Response HTTP Headers

The RESTCONF server can be configured to reply with particular HTTP headers in the HTTP response. For example, to support Cross-Origin Resource Sharing (CORS, <https://www.w3.org/TR/cors/>) there is a need to add a couple of headers to the HTTP Response.

We add the extra configuration parameter in `ncs.conf`.

{% code title="Example: NSO RESTCONF Custom Header Configuration" %}

```xml
    <restconf>
      <enabled>true</enabled>
      <custom-headers>
        <header>
          <name>Access-Control-Allow-Origin</name>
          <value>*</value>
        </header>
      </custom-headers>
    </restconf>
```

{% endcode %}

A number of HTTP headers have been deemed so important by security reasons that they, with sensible default values, per default will be included in the RESTCONF reply. The values can be changed by configuration in the `ncs.conf` file. Note that a configured empty value will effectively turn off that particular header from being included in the RESTCONF reply. The headers and their default values are:

* `xFrameOptions`: `DENY`

  The default value indicates that the page cannot be displayed in a frame/iframe/embed/object regardless of the site attempting to do so.
* `xContentTypeOptions`: `nosniff`

  The default value indicates that the MIME types advertised in the Content-Type headers should not be changed and be followed. In particular, should requests for CSS or Javascript be blocked in case a proper MIME type is not used.
* `xXssProtection`: `1; mode=block`

  This header is a feature of Internet Explorer, Chrome and Safari that stops pages from loading when they detect reflected cross-site scripting (XSS) attacks. It enables XSS filtering and tells the browser to prevent rendering of the page if an attack is detected.
* `strictTransportSecurity`: `max-age=31536000; includeSubDomains`

  The default value tells browsers that the RESTCONF server should only be accessed using HTTPS, instead of using HTTP. It sets the time that the browser should remember this and states that this rule applies to all of the server's subdomains as well.
* `contentSecurityPolicy`: `default-src 'self'; block-all-mixed-content; base-uri 'self'; frame-ancestors 'none';`

  The default value means that: Resources like fonts, scripts, connections, images, and styles will all only load from the same origin as the protected resource. All mixed contents will be blocked and frame-ancestors like iframes and applets are prohibited.

## Generating Swagger for RESTCONF <a href="#d5e2380" id="d5e2380"></a>

Swagger is a documentation language used to describe RESTful APIs. The resulting specifications are used to both document APIs as well as generating clients in a variety of languages. For more information about the Swagger specification itself and the ecosystem of tools available for it, see [swagger.io](https://swagger.io/).

The RESTCONF API in NSO provides an HTTP-based interface for accessing data. The YANG modules loaded into the system define the schema for the data structures that can be manipulated using the RESTCONF protocol. The `yanger` tool provides options to generate Swagger specifications from YANG files. The tool currently supports generating specifications according to OpenAPI/Swagger 2.0 using JSON encoding. The tool supports the validation of JSON bodies in body parameters and response bodies, and XML content validation is not supported.

YANG and Swagger are two different languages serving slightly different purposes. YANG is a data modeling language used to model configuration data, state data, Remote Procedure Calls, and notifications for network management protocols such as NETCONF and RESTCONF. Swagger is an API definition language that documents API resource structure as well as HTTP body content validation for applicable HTTP request methods. Translation from YANG to Swagger is not perfect in the sense that there are certain constructs and features in YANG that is not possible to capture completely in Swagger. The design of the translation is designed such that the resulting Swagger definitions are *more* restrictive than what is expressed in the YANG definitions. This means that there are certain cases where a client can do more in the RESTCONF API than what the Swagger definition expresses. There is also a set of well-known resources defined in the [RESTCONF RFC 8040](https://tools.ietf.org/html/rfc8040) that are not part of the generated Swagger specification, notably resources related to event streams.

### Using Y**anger** to Generate Swagger <a href="#d5e2390" id="d5e2390"></a>

The `yanger` tool is a YANG parser and validator that provides options to convert YANG modules to a multitude of formats including Swagger. You use the `-f swagger` option to generate a Swagger definition from one or more YANG files. The following command generates a Swagger file named `example.json` from the `example.yang` YANG file:

```
yanger -t expand -f swagger example.yang -o example.json
```

It is only supported to generate Swagger from one YANG module at a time. It is possible however to augment this module by supplying additional modules. The following command generates a Swagger document from `base.yang` which is augmented by `base-ext-1.yang` and `base-ext-2.yang`:

```
yanger -t expand -f swagger base.yang base-ext-1.yang base-ext-2.yang -o base.json
```

Only supplying augmenting modules is not supported.

Use the `--help` option to the `yanger` command to see all available options:

```
yanger --help
```

The complete list of options related to Swagger generation is:

```
Swagger output specific options:
  --swagger-host                    Add host to the Swagger output
  --swagger-basepath                Add basePath to the Swagger output
  --swagger-version                 Add version url to the Swagger output.
                                    NOTE: this will override any revision
                                    in the yang file
  --swagger-tag-mode                Set tag mode to group resources. Valid
                                    values are: methods, resources, all
                                    [default: all]
  --swagger-terms                   Add termsOfService to the Swagger
                                    output
  --swagger-contact-name            Add contact name to the Swagger output
  --swagger-contact-url             Add contact url to the Swagger output
  --swagger-contact-email           Add contact email to the Swagger output
  --swagger-license-name            Add license name to the Swagger output
  --swagger-license-url             Add license url to the Swagger output
  --swagger-top-resource            Generate only swagger resources from
                                    this top resource. Valid values are:
                                    root, data, operations, all [default:
                                    all]
  --swagger-omit-query-params       Omit RESTCONF query parameters
                                    [default: false]
  --swagger-omit-body-params        Omit RESTCONF body parameters
                                    [default: false]
  --swagger-omit-form-params        Omit RESTCONF form parameters
                                    [default: false]
  --swagger-omit-header-params      Omit RESTCONF header parameters
                                    [default: false]
  --swagger-omit-path-params        Omit RESTCONF path parameters
                                    [default: false]
  --swagger-omit-standard-statuses  Omit standard HTTP response statuses.
                                    NOTE: at least one successful HTTP
                                    status will still be included
                                    [default: false]
  --swagger-methods                 HTTP methods to include. Example:
                                    --swagger-methods "get, post"
                                    [default: "get, post, put, patch,
                                    delete"]
  --swagger-path-filter             Filter out paths matching a path filter.
                                    Example: --swagger-path-filter
                                    "/data/example-jukebox/jukebox"
  --swagger-only-actions            Only emit Swagger output for Yang actions
                                    [default: false]
  --swagger-only-nso-services       Only emit Swagger output for NSO Services
                                    [default: false]
  --swagger-hide-nso-services-data  Hide imported NSO services data
                                    [default: false]
  --swagger-only-list-keys          Only emit Swagger output for the keys in
                                    lists [default: false]
  --swagger-max-depth               Only emit Swagger output until Max-Depth
                                    is reached [default: -1]
  --swagger-unhide                  Unhide specified groups,
                                    example: --swagger-unhide "foo,bar"
  --swagger-unhide-all              Unhide all hidden groups [default: false]
```

Using the `example-jukebox.yang` from the [RESTCONF RFC 8040](https://tools.ietf.org/html/rfc8040), the following example generates a comprehensive Swagger definition using a variety of Swagger-related options:

{% code title="Example: Comprehensive Swagger Generation Example" %}

```
yanger -p . -t expand -f swagger example-jukebox.yang \
       --swagger-host 127.0.0.1:8080 \
       --swagger-basepath /restconf \
       --swagger-version "My swagger version 1.0.0.1" \
       --swagger-tag-mode all \
       --swagger-terms "http://my-terms.example.com" \
       --swagger-contact-name "my contact name" \
       --swagger-contact-url "http://my-contact-url.example.com" \
       --swagger-contact-email "my-contact-email@example.com" \
       --swagger-license-name "my license name" \
       --swagger-license-url "http://my-license-url.example.com" \
       --swagger-top-resource all \
       --swagger-omit-query-params false \
       --swagger-omit-body-params false \
       --swagger-omit-form-params false \
       --swagger-omit-header-params false \
       --swagger-omit-path-params false \
       --swagger-omit-standard-statuses false \
       --swagger-methods "post, get, patch, put, delete, head, options"
```

{% endcode %}

For a large YANG model the generated Swagger JSON output also becomes very large; so in order to restrict the amount of JSON output, a number of switches can be used, for example: `--swagger-only-actions` or `--swagger-max-depth`, etc.

Note that, per default, any hidden YANG elements will not show up in the JSON output. This behavior can be modified by using the switches: `--swagger-unhide-all` and `--swagger-unhide`.


# NSO SNMP Agent

Description of SNMP agent.

The SNMP agent in NSO is used mainly for monitoring and notifications. This guide covers the SNMPv3 `auth-priv` setup used in the examples, with SHA authentication and AES privacy.

The following standard MIBs are supported by the SNMP agent:

* SNMPv2-MIB [RFC 3418](https://www.ietf.org/rfc/rfc3418.txt)
* SNMP-FRAMEWORK-MIB [RFC 3411](https://www.ietf.org/rfc/rfc3411.txt)
* SNMP-USER-BASED-SM-MIB [RFC 3414](https://www.ietf.org/rfc/rfc3414.txt)
* SNMP-VIEW-BASED-ACM-MIB [RFC 3415](https://www.ietf.org/rfc/rfc3415.txt)
* SNMP-COMMUNITY-MIB [RFC 3584](https://www.ietf.org/rfc/rfc3584.txt)
* SNMP-TARGET-MIB and SNMP-NOTIFICATION-MIB [RFC 3413](https://www.ietf.org/rfc/rfc3413.txt)
* SNMP-MPD-MIB [RFC 3412](https://www.ietf.org/rfc/rfc3412.txt)
* TRANSPORT-ADDRESS-MIB [RFC 3419](https://www.ietf.org/rfc/rfc3419.txt)
* SNMP-USM-AES-MIB [RFC 3826](https://www.ietf.org/rfc/rfc3826.txt)
* IPV6-TC [RFC 2465](https://www.ietf.org/rfc/rfc2465.txt)

{% hint style="info" %}
The usmHMACMD5AuthProtocol authentication protocol and the usmDESPrivProtocol privacy protocol specified in SNMP-USER-BASED-SM-MIB are not supported, since they are not considered secure. The usmHMACSHAAuthProtocol authentication protocol specified in SNMP-USER-BASED-SM-MIB and the usmAesCfb128Protocol privacy protocol specified in SNMP-USM-AES-MIB are supported.
{% endhint %}

## Configuring the SNMP Agent <a href="#d5e2459" id="d5e2459"></a>

The SNMP agent is configured through any of the normal NSO northbound interfaces. It is possible to control most aspects of the agent through for example the CLI.

The YANG models describing all configuration capabilities of the SNMP agent reside under `$NCS_DIR/src/ncs/snmp/snmp-agent-cfg/*.yang` in the NSO distribution.

An example session configuring the SNMP agent through the CLI may look like:

```bash
admin@ncs# config
Entering configuration mode terminal
admin@ncs(config)# snmp agent udp-port 3457
admin@ncs(config)# snmp target monitor usm user-name initial
admin@ncs(config-target-monitor-usm)# snmp target monitor usm sec-level auth-priv
admin@ncs(config-target-monitor-usm)# commit
Commit complete.
admin@ncs(config-target-monitor-usm)# top
admin@ncs(config)# show full-configuration snmp
snmp agent enabled
snmp agent ip    0.0.0.0
snmp agent udp-port 3457
snmp agent version v3
snmp agent engine-id enterprise-number 32473
snmp agent engine-id from-text testing
snmp agent max-message-size 50000
snmp system contact ""
snmp system name ""
snmp system location ""
snmp usm local user initial
 auth sha password GoTellMom
 priv aes password GoTellMom
!
snmp target monitor
 ip       127.0.0.1
 udp-port 162
 tag      [ monitor ]
 timeout  1500
 retries  3
 usm user-name initial
 usm sec-level auth-priv
!
snmp notify foo
 tag  monitor
 type trap
!
snmp vacm group initial
 member initial
  sec-model [ usm ]
 !
 access usm auth-priv
  read-view   internet
  notify-view internet
 !
!
snmp vacm view internet
 subtree 1.3.6.1
  included
 !
!
snmp vacm view restricted
 subtree 1.3.6.1.6.3.11.2.1
  included
 !
 subtree 1.3.6.1.6.3.15.1.1
  included
 !
!
```

The SNMP agent configuration data is stored in CDB as any other configuration data, but is handled as a transformation between the data shown above and the data stored in the standard MIBs.

If you want to have a default configuration of the SNMP agent, you must provide that in an XML file. The initialization data of the SNMP agent is stored in an XML file that has precisely the same format as CDB initialization XML files, but it is not loaded by CDB, rather it is loaded at first startup by the SNMP agent. The XML file must be called `snmp_init.xml` and it must reside in the load path of NSO. In the NSO distribution, there is such an initialization file in `$NCS_DIR/etc/ncs/snmp/snmp_init.xml`. It is strongly recommended that this file be customized with another engine ID and site-specific SNMPv3 users and passwords.

If no `snmp_init.xml` file is found in the load path a default configuration with the agent disabled is loaded. Thus, the easiest way to start NSO without the SNMP agent is to ensure that the directory `$NCS_DIR/etc/ncs/snmp/` is not part of the NSO load path.

Note, that this only relates to initialization the first time NSO is started. On subsequent starts, all the SNMP agent configuration data is stored in CDB and the `snmp_init.xml` is never used again.

## Alarm MIB <a href="#d5e2482" id="d5e2482"></a>

The NSO SNMP alarm MIB is designed for ease of use in alarm systems. It defines a table of alarms and SNMP alarm notifications corresponding to alarm state changes. Based on the alarm model in NSO (see [NSO Alarms](/guides/administration/management/system-management#nso-alarms)), the notifications as well as the alarm table contain the parameters that are required for alarm standards compliance (X.733 and 3GPP). The MIB files are located in `$NCS_DIR/src/ncs/snmp/mibs`.

* **TAILF-TOP-MIB.mib**\
  **T**he tail-f enterprise OID.
* **TAILF-TC-MIB.mib**\
  Textual conventions for the alarm mib.
* **TAILF-ALARM-MIB.mib**\
  **T**he actual alarm MIB.
* **IANA-ITU-ALARM-TC-MIB.mib**\
  Import of IETF mapping of X.733 parameters.
* **ITU-ALARM-TC-MIB.mib**\
  Import of IETF mapping of X.733 parameters.

<div data-with-frame="true"><figure><img src="/files/Tm6gKn7JTK6QDm5zSuN9" alt=""><figcaption><p>The NSO Alarm MIB</p></figcaption></figure></div>

The alarm table has the following columns:

* **tfAlarmIndex**\
  An imaginary index for the alarm row that is persistent between restarts.
* **tfAlarmType**\
  This provides an identification of the alarm type and together with tfAlarmSpecificProblem forms a unique identification of the alarm.
* **tfAlarmDevice**\
  The alarming network device - can be NSO itself.
* **tfAlarmObject**\
  The alarming object within the device.
* **tfAlarmObjectOID**\
  In case the original alarm notification was an SNMP notification this column identifies the alarming SNMP object.
* **tfAlarmObjectStr**\
  Name of alarm object based on any other naming.
* **tfAlarmSpecificProblem**\
  This object is used when the 'tfAlarmType' object cannot uniquely identify the alarm type.
* **tfAlarmEventType**\
  The event type according to X.733 and based on the mapping of the alarm type in the NSO alarm model.
* **tfAlarmProbableCause**\
  The probable cause to X.733 and based on the mapping of the alarm type in the NSO alarm model. Note that you can configure this to match the probable cause values in the receiving alarm system.
* **tfAlarmOrigTime**\
  The time for the first occurrence of this alarm.
* **tfAlarmTime**\
  The time for the last state change of this alarm.
* **tfAlarmSeverity**\
  The latest severity (non-clear) reported for this alarm.
* **tfAlarmCleared**\
  Boolean indicated if the latest state change reports a clear.
* **tfAlarmText**\
  The latest alarm text.
* **tfAlarmOperatorState**\
  The latest operator alarm state such as ack.
* **tfAlarmOperatorNote**\
  The latest operator note.

The MIB defines separate notifications for every severity level to support SNMP managers that only can map severity levels to individual notifications. Every notification contains the parameters of the alarm table.

### SNMP Object Identifiers <a href="#d5e2580" id="d5e2580"></a>

{% code title="Example: Object Identifiers" %}

```
 tfAlarmMIB             node         1.3.6.1.4.1.24961.2.103
 tfAlarmObjects         node         1.3.6.1.4.1.24961.2.103.1
 tfAlarms               node         1.3.6.1.4.1.24961.2.103.1.1
 tfAlarmNumber          scalar       1.3.6.1.4.1.24961.2.103.1.1.1
 tfAlarmLastChanged     scalar       1.3.6.1.4.1.24961.2.103.1.1.2
 tfAlarmTable           table        1.3.6.1.4.1.24961.2.103.1.1.5
 tfAlarmEntry           row          1.3.6.1.4.1.24961.2.103.1.1.5.1
 tfAlarmIndex           column       1.3.6.1.4.1.24961.2.103.1.1.5.1.1
 tfAlarmType            column       1.3.6.1.4.1.24961.2.103.1.1.5.1.2
 tfAlarmDevice          column       1.3.6.1.4.1.24961.2.103.1.1.5.1.3
 tfAlarmObject          column       1.3.6.1.4.1.24961.2.103.1.1.5.1.4
 tfAlarmObjectOID       column       1.3.6.1.4.1.24961.2.103.1.1.5.1.5
 tfAlarmObjectStr       column       1.3.6.1.4.1.24961.2.103.1.1.5.1.6
 tfAlarmSpecificProblem column       1.3.6.1.4.1.24961.2.103.1.1.5.1.7
 tfAlarmEventType       column       1.3.6.1.4.1.24961.2.103.1.1.5.1.8
 tfAlarmProbableCause   column       1.3.6.1.4.1.24961.2.103.1.1.5.1.9
 tfAlarmOrigTime        column       1.3.6.1.4.1.24961.2.103.1.1.5.1.10
 tfAlarmTime            column       1.3.6.1.4.1.24961.2.103.1.1.5.1.11
 tfAlarmSeverity        column       1.3.6.1.4.1.24961.2.103.1.1.5.1.12
 tfAlarmCleared         column       1.3.6.1.4.1.24961.2.103.1.1.5.1.13
 tfAlarmText            column       1.3.6.1.4.1.24961.2.103.1.1.5.1.14
 tfAlarmOperatorState   column       1.3.6.1.4.1.24961.2.103.1.1.5.1.15
 tfAlarmOperatorNote    column       1.3.6.1.4.1.24961.2.103.1.1.5.1.16
 tfAlarmNotifications   node         1.3.6.1.4.1.24961.2.103.2
 tfAlarmNotifsPrefix    node         1.3.6.1.4.1.24961.2.103.2.0
 tfAlarmNotifsObjects   node         1.3.6.1.4.1.24961.2.103.2.1
 tfAlarmStateChangeText scalar       1.3.6.1.4.1.24961.2.103.2.1.1
 tfAlarmIndeterminate   notification 1.3.6.1.4.1.24961.2.103.2.0.1
 tfAlarmWarning         notification 1.3.6.1.4.1.24961.2.103.2.0.2
 tfAlarmMinor           notification 1.3.6.1.4.1.24961.2.103.2.0.3
 tfAlarmMajor           notification 1.3.6.1.4.1.24961.2.103.2.0.4
 tfAlarmCritical        notification 1.3.6.1.4.1.24961.2.103.2.0.5
 tfAlarmClear           notification 1.3.6.1.4.1.24961.2.103.2.0.6
 tfAlarmConformance     node         1.3.6.1.4.1.24961.2.103.10
 tfAlarmCompliances     node         1.3.6.1.4.1.24961.2.103.10.1
 tfAlarmCompliance      compliance   1.3.6.1.4.1.24961.2.103.10.1.1
 tfAlarmGroups          node         1.3.6.1.4.1.24961.2.103.10.2
 tfAlarmNotifs          group        1.3.6.1.4.1.24961.2.103.10.2.1
 tfAlarmObjs            group        1.3.6.1.4.1.24961.2.103.10.2.2
```

{% endcode %}

### Using the SNMP Alarm MIB

Alarm Managers should subscribe to the notifications and read the alarm table to synchronize the alarm list. To do this you need an access view that matches the alarm MIB and creates an SNMP target. In the examples, the alarm target uses SNMPv3 `auth-priv` with the `initial` user and SHA/AES credentials. A target is set up in the following way, assuming the SNMP Alarm Manager has IP address `192.168.1.1` and is configured with the same SNMPv3 user:

{% code title="Example: Subscribing to SNMP Alarms" %}

```bash
$ ncs_cli -u admin -C
admin@ncs# config
Entering configuration mode terminal
admin@ncs(config)# snmp notify monitor type trap tag monitor
admin@ncs(config-notify-monitor)# snmp target alarm-system ip 192.168.1.1 udp-port 162 \
        tag monitor usm user-name initial sec-level auth-priv
admin@ncs(config-target-alarm-system)# commit
Commit complete.
admin@ncs(config-target-alarm-system)# show full-configuration snmp target
snmp target alarm-system
 ip       192.168.1.1
 udp-port 162
 tag      [ monitor ]
 timeout  1500
 retries  3
 usm user-name initial
 usm sec-level auth-priv
!
snmp target monitor
 ip       127.0.0.1
 udp-port 162
 tag      [ monitor ]
 timeout  1500
 retries  3
 usm user-name initial
 usm sec-level auth-priv
!
admin@ncs(config-target-alarm-system)#
```

{% endcode %}


# NSO MCP Server

Use the NSO MCP server to expose NSO capabilities to MCP-compatible AI clients.

The Cisco NSO Adaptive MCP Server is an NSO package that bridges the Model Context Protocol (MCP) to Cisco NSO functionality. It allows MCP-compatible AI assistants and clients to securely interact with an NSO deployment through a standard MCP interface while continuing to use NSO authentication, authorization, and policy controls.

The MCP server provides a standard way for MCP-compatible clients to:

* Read NSO data through MCP resources
* Invoke NSO operations through MCP tools
* Use guided workflows exposed as MCP prompts

The server is delivered as an NSO package and exposes an MCP endpoint inside NSO instead of requiring a separate MCP server deployment.

## Requirements

The NSO MCP server has the following requirements:

| Requirement     | Version        |
| --------------- | -------------- |
| **NSO version** | 6.7 or higher  |
| **Java**        | 21 or higher   |
| **Protocol**    | MCP 2025-03-26 |

## Role of the MCP Server

The MCP server acts as a bridge between an MCP-compatible client and an NSO deployment.

In a typical workflow:

1. A user works through an MCP-compatible AI assistant or client.
2. The client connects to the NSO MCP endpoint.
3. NSO authenticates the user by using the authentication mechanisms configured for the deployment.
4. The MCP server exposes the capabilities that the authenticated user is allowed to access.
5. The client reads NSO context and invokes supported NSO operations on behalf of that user.

This allows engineers to use natural language or MCP-aware tooling to inspect NSO data and perform supported tasks while NSO remains the source of truth and enforcement point.

{% hint style="info" %}
The MCP server provides the integration layer, not the AI assistant itself. Customers connect an MCP-compatible client or assistant of their choice to the NSO MCP endpoint.
{% endhint %}

## Architecture

The solution uses a hybrid architecture with an Erlang proxy layer and a Java MCP backend.

At a high level:

* NSO exposes the external MCP endpoint over the NSO WebUI HTTP(S) listeners
* An Erlang proxy layer handles authentication-related request processing and forwards requests internally
* A Java `ApplicationComponent` implements the MCP server logic
* The Java backend discovers tools, resources, resource templates, and prompts from the live NSO schema and package content
* Requests are executed with the authenticated NSO user context
* The Java backend is not directly exposed to external clients

## Endpoint and Transport

The NSO MCP server exposes MCP by using Streamable HTTP at the NSO `/mcp` endpoint.

Clients connect to:

* `http://<nso-host>:<http-port>/mcp`
* `https://<nso-host>:<https-port>/mcp`

The server accepts MCP requests over HTTP `POST`. Other HTTP methods are not supported. In particular, `GET /mcp` is not a valid MCP request and returns `405 Method Not Allowed`.

Clients that support HTTP-based MCP can connect directly. Clients that support only stdio transport require an external proxy bridge, such as `mcp-proxy`.

{% hint style="info" %}
When `/ncs-config/webui/match-host-name` is `true`, the HTTP `Host` header must match the configured WebUI `server-name` or `server-alias`. With the default NSO WebUI settings this typically means using `localhost` rather than `127.0.0.1` for local access unless the WebUI host-name settings have been changed.
{% endhint %}

## Authentication and Authorization

The MCP server relies on existing NSO security mechanisms.

### Authentication

The MCP server does not introduce a separate authentication system. Clients authenticate by using the NSO authentication mechanisms configured for the deployment.

The proxy supports the same authentication styles expected for NSO WebUI access, including:

* WebUI session cookies
* HTTP Basic authentication
* `X-Auth-Token` or `Bearer` token authentication
* package-authentication fallback

{% hint style="info" %}
If package-authentication is used, note that the MCP endpoint uses the AAA context `mcp`. Package-authentication scripts that only recognize the `rest` context must be updated to also accept `mcp`.
{% endhint %}

### Authorization and User-Based Execution

Access to the MCP endpoint and to exposed capabilities is controlled by NSO authorization mechanisms, including NACM.

Operations requested through the MCP server run with the authenticated user's identity. MCP clients cannot read data or execute operations beyond that user's normal permissions.

The capabilities visible to a client are the effective intersection of:

* NSO authentication
* NSO authorization and NACM
* MCP exposure policy configuration

## What the MCP Server Exposes

The MCP server exposes NSO functionality through four MCP constructs:

* tools
* resources
* resource templates
* prompts

### Tools

Tools represent executable operations that MCP clients can invoke.

Examples include:

* service-related operations
* device operations such as `sync-from`, `sync-to`, `check-sync`, and `compare-config`
* NSO actions discovered from schema
* read-oriented helper tools for schema, configuration, and operational data

### Resources and Resource Templates

Resources represent readable NSO data that MCP clients can fetch.

Resource templates represent parameterized readable data, for example a device-specific or service-specific path.

Examples may include, depending on the NSO deployment and loaded packages:

* services
* devices
* device configuration
* device operational data
* configuration data at a path
* operational data at a path
* service schema views
* service instance-specific resource templates

### Prompts

Prompts are guided workflows exposed to MCP clients.

Examples include workflows for:

* service-related tasks
* device onboarding
* troubleshooting and recovery
* repeated change procedures

The MCP server automatically discovers supported tools, resources, resource templates, and prompts from NSO schema and package content.

{% hint style="info" %}
Prompt visibility does not by itself guarantee that every referenced tool is exposed. Under restrictive MCP policy settings, a prompt may be visible while some of the tools used by that workflow remain hidden.
{% endhint %}

## Supported MCP Methods

The NSO MCP server supports the following MCP methods:

* `initialize`
* `tools/list`
* `tools/call`
* `resources/list`
* `resources/templates/list`
* `resources/read`
* `prompts/list`
* `prompts/get`
* `ping`
* `notifications/initialized`

The endpoint also accepts:

* `resources/subscribe`
* `resources/unsubscribe`

but no active resource subscription capability is currently provided.

## Configuration

The MCP server is configured through the `cisco-nso-mcp` YANG module under the `mcp-server` container.

Key configuration parameters include:

| Parameter                 | Type        | Default      | Description                                       |
| ------------------------- | ----------- | ------------ | ------------------------------------------------- |
| `enabled`                 | boolean     | `true`       | Enable or disable the MCP server                  |
| `logging/level`           | enumeration | `info`       | Log level                                         |
| `policies/default-action` | enumeration | `restricted` | Behavior when no rule matches                     |
| `policies/rule`           | list        | none         | Ordered permit or deny rules by path or namespace |

## Policy Rules

The `policies/rule` list defines ordered rules for controlling which MCP tools and resources are exposed.

Each rule can match on:

* YANG namespace
* schema path

The first matching rule decides whether access is permitted or denied. If no rule matches, `policies/default-action` controls the outcome.

`policies/default-action` supports:

* `permit`
  * expose unmatched capabilities
* `deny`
  * deny unmatched capabilities
* `restricted`
  * expose only built-in tools, device-operation tools, and read helpers for schema, config, and operational data

With the default value `restricted`:

* core read and device-operation capabilities remain available
* service CRUD tools, many discovered actions, and other non-core tools remain hidden unless explicitly permitted

This is an important behavior difference for first-time users. If a deployment appears to expose only a small set of MCP tools, check the MCP policy configuration before assuming discovery failed.

Example targeted policy configuration:

{% code title="Example" overflow="wrap" %}

```xml
<config xmlns="http://tail-f.com/ns/config/1.0">
  <mcp-server xmlns="http://cisco.com/pkg/cisco-nso-mcp">
    <policies>
      <default-action>restricted</default-action>
      <rule>
        <sequence>100</sequence>
        <action>permit</action>
        <match>
          <path>/my-service:my-service/*</path>
        </match>
        <description>Expose my-service capabilities</description>
      </rule>
    </policies>
  </mcp-server>
</config>
```

{% endcode %}

For broad exposure in a lab environment, `default-action permit` can be used. In production deployments, targeted permit rules are usually preferred.

## Deployment and Setup

The NSO MCP server is delivered as a Cisco-provided NSO package. Deploy and configure it in your NSO environment instead of building a separate MCP server implementation.

The package is:

* exposed externally through the NSO `/mcp` endpoint
* managed through standard NSO package workflows

A typical setup flow is:

1. Ensure the MCP package is present in the NSO installation or deployment package set.
2. Place the package in the NSO packages directory if required by your deployment workflow.
3. Build and load the package.
4. Restart NSO with the `--with-package-reload` option.
5. Verify that the package is operational.
6. Review and adjust `mcp-server` configuration as needed.
7. Connect an MCP-compatible client to the `/mcp` endpoint by using NSO authentication.

## Connecting an MCP Client

After the package is installed and configured, an MCP-compatible client connects to the NSO MCP endpoint at `/mcp`.

The client authenticates by using the NSO authentication method configured for the deployment. Clients that support HTTP-based MCP can connect directly. Clients that support only stdio transport require a proxy bridge.

Once authenticated, the client can:

* list available MCP tools
* list available MCP resources
* list available MCP resource templates
* list available MCP prompts
* read permitted NSO data
* invoke permitted NSO operations

The capabilities visible to the client depend on NSO authentication, NACM authorization, and MCP exposure policy configuration.

## Verifying the Setup

A typical verification flow includes:

1. Confirm that the MCP package is loaded and operational.
2. Confirm that the NSO `/mcp` endpoint is reachable.
3. Connect an MCP-compatible client with valid NSO credentials.
4. Verify that the client can list tools, resources, resource templates, and prompts.
5. Test a permitted read or tool invocation.

Example `initialize` request:

{% code title="Example" overflow="wrap" %}

```bash
curl -u <user>:<password> \
  -H "Content-Type: application/json" \
  -d '{
        "jsonrpc": "2.0",
        "id": 1,
        "method": "initialize",
        "params": {
          "protocolVersion": "2025-03-26",
          "capabilities": {},
          "clientInfo": {
            "name": "example-client",
            "version": "1.0"
          }
        }
      }' \
  http://localhost:8080/mcp
```

{% endcode %}

Example `tools/list` request:

{% code title="Example" overflow="wrap" %}

```bash
curl -u <user>:<password> \
  -H "Content-Type: application/json" \
  -d '{
        "jsonrpc": "2.0",
        "id": 2,
        "method": "tools/list",
        "params": {}
      }' \
  http://localhost:8080/mcp
```

{% endcode %}

The package ships with validation script `test-mcp.sh`, that script can be used as a smoke test after the package is loaded.

## Interaction through AI Assistants

The MCP server allows MCP-compatible AI assistants and clients to interact with NSO through a standard interface.

This means that an engineer can use an AI assistant to:

* inspect NSO state
* discover available operations
* carry out supported tasks
* use guided workflows for common NSO procedures

The MCP server does not replace NSO review, authentication, authorization, or operational controls. Instead, it provides a structured bridge between AI tooling and the NSO deployment.

## Troubleshooting

The following checks are useful when the MCP endpoint does not behave as expected:

<details>

<summary><code>400 Bad Request</code> on <code>/mcp</code></summary>

If the endpoint returns `400 Bad Request`, verify the WebUI host-name settings.

When `/ncs-config/webui/match-host-name` is `true`, the HTTP `Host` header must match the configured WebUI `server-name` or `server-alias`. For local testing with the default NSO WebUI settings, use `localhost` rather than `127.0.0.1`.

</details>

<details>

<summary><code>405 Method Not Allowed</code></summary>

The MCP endpoint supports HTTP `POST`. Do not use `GET` for MCP requests.

</details>

<details>

<summary>Fewer tools are visible than expected</summary>

If discovery appears to work but only a small set of tools is visible, check:

* the authenticated user's NSO permissions and NACM rules
* the MCP exposure policy under `mcp-server/policies`
* whether `policies/default-action` is still `restricted`

</details>

<details>

<summary>Package-authentication works for RESTCONF but not for MCP</summary>

If package-authentication is used, confirm that the authentication script accepts the AAA context `mcp` in addition to any existing `rest` handling.

</details>

<details>

<summary><code>500 Internal Server Error</code> or backend unavailable</summary>

Check the package logs and the Java VM logs.

Useful log files include:

* `logs/cisco-nso-mcp-server.log` for the Erlang proxy layer
* `logs/ncs-java-vm.log` for the Java backend and component lifecycle

On platforms with short Unix-domain socket path limits, such as macOS, very long NSO runtime paths can prevent the internal MCP Unix socket from being created. If the Java backend reports an error such as `Unix domain path too long`, shorten the NSO runtime path and restart the package or deployment.

</details>

## Operational Considerations

AI-based operations can move quickly from intent to execution. In a production NSO environment, this carries a risk of unintended changes. This section offers guidance for careful and deliberate use of the NSO MCP server.

* Treat the MCP endpoint as another NSO northbound interface. An MCP client acts on behalf of the authenticated NSO user, so normal operational practices still apply: use least-privilege users, review intended changes, keep audit context, and verify the result after every state-changing operation.
* Read-oriented operations are usually the safest way to start. They are useful for discovering devices, services, schema, operational state, and available actions before making a change. However, read access can still expose sensitive configuration or operational data, so it should still be controlled through NSO authentication, NACM, and MCP exposure policies.
* State-changing tools and actions require more care. Some operations change NSO CDB, some push changes to devices, and some do both. For example, a device `sync-from` reads from the device but immediately updates NSO's copy of the configuration. A `sync-to`, service `re-deploy`, service create/update/delete operation, or another action may change device state or service ownership depending on the underlying NSO operation.

### Recommended Workflow for State-Changing Operations

Before invoking a state-changing MCP tool:

1. Use read operations first to identify the exact service, device, path, or action target.
2. Ask the client to describe the intended change, affected objects, and expected result.
3. Prefer narrow targets over broad selectors such as all devices or all services.
4. When the operation supports it, run a `dry-run` first and review the resulting diff.
5. Use commit metadata such as `label` and `comment` when available, so rollback files, events, compliance reports, and commit queue entries can be tied back to the reason for the change.
6. Consider commit parameters such as `no-overwrite`, `confirm-network-state`, `no-networking`, or `commit-queue` only when they match the operational intent and are supported by the invoked operation.
7. After the operation completes, verify the result with read operations, `check-sync`, `compare-config`, service `check-sync`, or commit queue status as appropriate.

For production use, keep `policies/default-action` set to `restricted` and expose paths, namespaces, or actions that enable state-changing operations only through targeted permit rules. Broad exposure, such as `default-action permit`, is better suited to labs and short-lived evaluation environments.

### Dos and Don'ts

Consider the following when using MCP-based operations.

<details>

<summary>Dos</summary>

* **Access control**: Use NSO users and NACM groups with only the permissions needed for the intended MCP workflows.
* **MCP policies**: Start from `restricted` and add targeted permit rules for known service paths, namespaces, or actions.
* **Read first**: Use resources, resource templates, and read helper tools to inspect state before invoking tools that change state.
* **Dry run**: Use `dry-run` where the underlying NSO operation supports it, especially for service deployment, re-deploy, `sync-to`, or package-related operations.
* **Auditability**: Use `label` and `comment` when available to make MCP-originated changes easy to find in rollback files, events, and commit queue results.
* **Verification**: Check the resulting NSO state, affected device sync state, service sync state, and commit queue result after state-changing operations.
* **Environment**: In local-install or development environments, source the NSO environment, for example `source ncsrc`, before running NSO commands, package scripts, or validation scripts such as `test-mcp.sh`.
* **Troubleshooting**: Check `logs/cisco-nso-mcp-server.log` for the Erlang proxy layer and `logs/ncs-java-vm.log` for Java backend or component lifecycle issues.

</details>

<details>

<summary>Don'ts</summary>

* **Broad changes**: Do not ask an MCP client to make vague changes such as "fix this service" or "clean up the devices" without first reviewing the proposed target and operation.
* **Production exposure**: Do not use broad `default-action permit` policies in production unless the deployment is intentionally designed for that level of MCP exposure.
* **Credentials**: Do not store long-lived administrator credentials in shared client configuration files such as `.vscode/mcp.json`.
* **Blind changes**: Do not invoke state-changing tools before checking the current NSO state and, where relevant, device or service sync state.
* **Sync operations**: Do not treat `sync-from` as a harmless read-only operation. It reads from the device but updates NSO's stored configuration.
* **Commit flags**: Do not use flags such as `no-networking`, `no-out-of-sync-check`, or `no-overwrite` unless the operator understands their NSO behavior and the operation supports them.
* **Logs**: Do not troubleshoot only from the MCP client response. Also inspect NSO package status, the MCP package logs, Java VM logs, and commit queue or rollback state when a state-changing operation fails.

</details>

## Example Client Configuration

After the package is installed and operational, the MCP endpoint is exposed by NSO at:

```
https://<host-name-or-ip>:<port>/mcp
```

If the deployment uses the default WebUI host-name settings for local access, use `https://localhost:<port>/mcp`.

### Example VS Code MCP Client Configuration

Visual Studio Code supports remote MCP servers over HTTP. A minimal workspace configuration looks like this:

{% code title="Example" overflow="wrap" %}

```json
{
  "servers": {
    "nsoProduction": {
      "type": "http",
      "url": "https://nso.example.com:<port>/mcp",
      "headers": {
        "Authorization": "Basic <base64-user-colon-password>"
      }
    }
  }
}
```

{% endcode %}

Adjust the URL, port, and authentication settings to match the NSO deployment and save this in `.vscode/mcp.json`.

If the deployment requires explicit HTTP authentication headers, configure them according to the NSO authentication method used in that environment.


# Advanced Development

Advanced-level NSO development.


# Development Environment and Resources

Useful information to help you get started with NSO development.

This section describes some recipes, tools, and other resources that you may find useful throughout development. The topics are tailored to novice users and focus on making development with NSO a more enjoyable experience.

## Development NSO Instance <a href="#ch_devenv.local" id="ch_devenv.local"></a>

Many developers prefer their own, dedicated NSO instance to avoid their work clashing with other team members. You can use either a local or remote Linux machine (such as a VM) or a macOS computer for this purpose.

The advantage of running local Linux with a GUI or macOS is that it is easier to set up the Integrated Development Environment (IDE) and other tools when they run on the same system as NSO. However, many IDEs today also allow working remotely, such as through the SSH protocol, making the choice of local versus remote less of a concern.

For development, using the so-called Local Install of NSO has some distinct advantages:

* It does not require elevated privileges to install or run.
* It keeps all NSO files in the same place (user-defined).
* It allows you to quickly switch between projects and NSO versions.

If you work with multiple projects in parallel, local install also allows you to take advantage of Python virtual environments to separate Python packages per project; simply start the NSO instance in an environment you have activated.

The main downside of using a local install is that it differs slightly from a system (production) install, such as in the filesystem paths used and the out-of-the-box configuration.

See [Local Install](/guides/administration/installation-and-deployment/local-install) for installation instructions.

## Examples and Showcases <a href="#ch_devenv.examples" id="ch_devenv.examples"></a>

There are a number of examples and showcases in this guide. We encourage you to follow them through. They are also a great reference if you are experimenting with a new feature and have trouble getting it to work; you can inspect and compare with the implementation in the example.

To run the examples, you will need access to an NSO instance. A development instance described in this chapter is the perfect option for running locally. See [Running NSO Examples](/guides/administration/installation-and-deployment/post-install-actions/running-nso-examples).

{% hint style="success" %}
Cisco also provides an online sandbox and containerized environments, such as a [Learning Lab](https://developer.cisco.com/learning/labs/nso-examples) or [NSO Sandbox](https://developer.cisco.com/catalogs/sandbox/nso), designed for this purpose. Refer to the [NSO documentation](https://nso-docs.cisco.com/learn-nso/learning-paths) for additional resources.
{% endhint %}

## IDE <a href="#ch_devenv.ide" id="ch_devenv.ide"></a>

Modern IDEs offer many features on top of advanced file editing support, such as code highlighting, syntax checks, and integrated debugging. While the initial setup takes some effort, the benefits of using an IDE are immense.

[Visual Studio Code](https://code.visualstudio.com/) (VS Code) is a freely available and extensible IDE. You can add support for Java, Python, and YANG languages, as well as remote access through SSH via VS Code extensions. Consider installing the following extensions:

* **Python** by Microsoft: Adds Python support.
* **Language Support for Java™** by Red Hat: Adds Java support.
* **NSO Developer Studio** by Cisco: Adds NSO-specific features as described in [NSO Developer Studio](https://nso-docs.cisco.com/resources/platform-tools/nso-developer-studio).
* **Remote - SSH** by Microsoft: Adds support for remote development.

The Remote - SSH extension is especially useful when you must work with a system through an SSH session. Once you connect to the remote host by clicking the `><` button (typically found in the bottom-left corner of the VS Code window), you can open and edit remote files with ease. If you also want language support (syntax highlighting and alike), you may need to install VS Code extensions remotely. That is, install the extensions after you have connected to the remote host; otherwise, the extension installation screen might not show the option for installation on the connected host.

<div data-with-frame="true"><figure><img src="/files/It1dpsDKiws5FNln25CH" alt="" width="563"><figcaption><p>Using the Remote - SSH extension in VS Code</p></figcaption></figure></div>

You will also benefit greatly from setting up SSH certificate authentication if you are using an SSH session for your work.

## Automating Instance Setup <a href="#ch_devenv.automate" id="ch_devenv.automate"></a>

Once you get familiar with NSO development and gain some experience, a single NSO instance is likely to be insufficient, either because you need instances for unit testing, because you need one-off (throwaway) instances for an experiment, or for something else entirely.

NSO includes tooling to help you quickly set up new local instances when such a need arises.

The following recipe relies on the `ncs-setup` command, which is available in the local install variant and requires a correctly set up shell environment (e.g., running `source ncsrc`). See [Local Install](/guides/administration/installation-and-deployment/local-install) for details.

A new instance typically needs a few things to be useful:

* Packages
* Initial data
* Devices to manage

In its simplest form, the `ncs-setup` invocation requires only a destination directory. However, you can specify additional packages to use with the `--package` option. Use the option to add as many packages as you need.

Running `ncs-setup` creates the required filesystem structure for an NSO instance. If you wish to include initial configuration data, put the XML-encoded data in the `ncs-cdb` subdirectory, and NSO will load it at the first start, as described in [Initialization Files](/guides/development/introduction-to-automation/cdb-and-yang#d5e268).

NSO also needs to know about the managed devices. In case you are using `ncs-netsim` simulated devices (described in [Network Simulator](/guides/operation-and-usage/operations/network-simulator-netsim)), you can use the `--netsim-dir` option with `ncs-setup` to add them directly. Otherwise, you may need to create some initial XML files with the relevant device configuration data—much like how you would add a device to NSO manually.

Most of the time, you must also invoke a sync with the device so that it performs correctly with NSO. If you wish to push some initial configuration to the device, you may add the configuration in the form of initial XML data and perform a `sync-to`. Alternatively, you can simply do a `sync-from`. You can use the `ncs_cmd` command for this purpose.

Combining all of this together, consider the following example:

1. Start by creating a new directory to hold the files:

   ```bash
   $ mkdir nso-throwaway
   $ cd nso-throwaway
   ```
2. Create and start a few simulated devices with `ncs-netsim`, using `./netsim` as directory:

   ```bash
   $ ncs-netsim ncs-netsim create-network $NCS_DIR/packages/neds/cisco-ios-cli-3.8 3 c
   DEVICE c0 CREATED
   DEVICE c1 CREATED
   DEVICE c2 CREATED
   $ ncs-netsim start
   ```
3. Next, create the running directory with the NED package for the simulated devices and one more package. Also, add configuration data to NSO on how to connect to these simulated devices.

   ```bash
       $ ncs-setup --dest ncs-run --netsim-dir ./netsim \
           --package $NCS_DIR/packages/neds/cisco-ios-cli-3.8 \
           --package $NCS_DIR/packages/neds/cisco-iosxr-cli-3.0
   ```
4. Now you can add custom initial data as XML files to `ncs-run/ncs-cdb/`. Usually, you would use existing files, but you can also create them on the fly.

   ```bash
   $ cat >ncs-run/ncs-cdb/my_init.xml <<'EOF'
   <config xmlns="http://tail-f.com/ns/config/1.0">
     <session xmlns="http://tail-f.com/ns/aaa/1.1">
       <idle-timeout>0</idle-timeout>
     </session>
   </config>
   EOF
   ```
5. At this point, you are ready to start NSO:

   ```bash
   $ cd ncs-run
   $ ncs
   ```
6. Finally, request an initial `sync-from`:

   ```bash
   $ ncs_cmd -u admin -c 'maction /devices/sync-from'
   sync-result begin
     device c0
     result true
   sync-result end
   sync-result begin
     device c1
     result true
   sync-result end
   sync-result begin
     device c2
     result true
   sync-result end
   ```
7. The instance is now ready for work. Once you are finished, you can stop it with `ncs --stop`. Remember to also stop the simulated devices with `ncs-netsim stop` if you no longer need them. Then, delete the containing folder (`nso-throwaway`) to remove all the leftover files and data.




---

[Next Page](/llms-full.txt/1)

