Salta al contenuto principale
Lympha technologies

Success stories

How we did it: 600+ applications, a Regional Authority and business continuity that pays for itself

First chapter of a series of cases told from the field: the multi-year management, evolution and operational continuity of a large Regional Authority's infrastructures — three teams, a single open-source governance dashboard and a continuity site that wor

With this article we open "How we did it": a series in which we take a real project and tell it the way it went — without names, for confidentiality, but with the real numbers and choices. We start with one of the most demanding: the multi-year management, evolution and operational continuity of the technology infrastructures of a large Italian Regional Authority.

We took an active part in the design, implementation and operation of a complete IT Service Management model for a fast-evolving information system, with a multi-site data centre of considerable technological complexity. The engagement covered the management and development of the infrastructures supporting more than 650 applications spread across three technology stacks, serving thousands of internal and external users, with a dedicated team of specialised professionals.

The project introduced strongly innovative elements: an integrated Operational Governance System, a Business Continuity and Disaster Recovery solution based on a both-active model and a methodological approach built on ITIL best practice, ensuring full compliance with Italian Legislative Decree 235/2010 (Art. 50-bis of the Digital Administration Code) and the Digital Agenda directives.

600+

applications managed across three technology stacks

+15%

application growth over two years

−50%

risk of losing critical services

4 h

RTO on critical services, with a 1-hour RPO

The context and the challenges

The Authority was facing a deep transformation of its information system. From a mid-range IT infrastructure, designed for the internal needs of the administration, it had progressively moved to a data centre with a particularly complex technology infrastructure in terms of networking, storage and servers, supporting services delivered not only to the Authority itself but also to numerous local bodies, trade associations, private companies and citizens across the territory.

The regional devolution process had substantially changed, both qualitatively and quantitatively, the perimeter of the information system, driving significant growth in the volumes, systems and applications to be managed: the number of published applications had grown by 15% in a single two-year period, across the three technology stacks.

The critical issues identified

The AS-IS analysis had highlighted problems that anyone working in the public sector will recognise at a glance:

  • fragmented monitoring tools: many different tools (trouble ticketing, monitoring, CMDB, project management) with uncorrelated data and no integrated view of IT services;
  • a gap between development and operations: the missing "weld" between the software project life cycle and the service life cycle was the weak link in the chain of control;
  • no structured Business Continuity and Disaster Recovery plan, in breach of the regulatory obligations introduced by Legislative Decree 235/2010;
  • storage not classified by tier (performance, capacity, reliability), with the resulting operational inefficiencies;
  • single points of failure at the level of internet connectivity and perimeter protection;
  • tape backup with time windows insufficient for critical services;
  • no integrated dashboard for operational governance and service reporting.

The strategic goals

The Authority had set clear goals — and none of them was "buying technology":

  1. adopt a process-oriented IT service management methodology, with the quality perceived by users as the cornerstone of action;
  2. equip itself with a Business Continuity and Disaster Recovery plan compliant with the regulations in force;
  3. implement an integrated governance platform for end-to-end control of IT services;
  4. guarantee seamless operational continuity throughout the transition;
  5. ensure full measurability and transparency of the services delivered through objective SLAs and KPIs.

Organisational model and governance

We therefore designed a lean and flexible organisational model, with clear decision-making roles, structured on three governance levels:

  • Strategic level — IT Steering Group: an executive committee made up of representatives of the administration, tasked with ensuring alignment with strategic goals and assessing the investments needed and the evolution path of the IT infrastructure;
  • Tactical level — IT Services Control Committee: the operational body for planning and coordinating implementation actions, with the Technical Manager as the main interface for day-to-day control, management and design activities;
  • Operational level — three integrated teams, described in the table below.
Team Scope Main activities
Operations Management Team (TGO)Service Operation, Service Transition End-to-end monitoring, operation control, backup/scheduling, second-level help desk, deploy & configuration, facility management
Centralised Systems Support Team (TSSC)Service Operation, Transition, Design Specialised infrastructure management across 6 vertical areas (Infrastructure, Middleware, Databases, Cartography, Open Source, SAP), system integration, third-level help desk
Project Development Team (TSP)Service Design, Service Strategy Design and implementation of new architectures, third-level help desk for critical cases, take-over start-up

The Change Area: integrating the software and service life cycles

One of the most innovative elements of the solution was the creation of the Change Area, an organisational unit spanning the three teams and dedicated to "taking over" the software produced and then "transforming" it into a service to the user — the same principle we apply today across the whole service life cycle. The Change Area ensured:

  • integration between the application design and development phase (Service Design) and the service delivery phase (Service Operation);
  • strong attention to the service test and validation phases (Service Transition);
  • planning of the release activities for every new service/product;
  • identification and proposal of strategies and innovation paths in the service life cycle;
  • coordination of the resources dedicated to defining and controlling standards and guidelines.

The Audit and Improvement Area

To guarantee constant, objective assessment of service quality, we established the Audit and Improvement Area, with the following responsibilities:

  • continuous monitoring of the predefined quality parameters (KPIs and SLAs);
  • planning of improvement actions even where levels were already satisfactory;
  • support in drafting an ICT Service Charter to formalise the "service pact" between the IT department and the General Directorates using the services;
  • management of the customer satisfaction process;
  • collection of statistical indicators for periodic reporting.

The Operational Governance System: one dashboard, almost all open source

We answered the tool fragmentation by taking part in the implementation of a complete, integrated platform supporting IT functions and processes — the heart of what we call IT governance — structured in three main areas and built largely on open-source components.

Service governance

  • integrated management of incidents, problems, changes and configuration via CMDBuild;
  • a trouble-ticketing tool ( RT) to track service requests and incidents;
  • interaction with the load-balancing systems ( LBL LoadBalancer) for Capacity Management;
  • a knowledge base built on Alfresco for centralised document management.

Infrastructure governance

  • Storage Management: data management, backup and restore;
  • System Fault & Performance Management: real-time monitoring via Zabbix with a centralised console, event filtering and correlation;
  • Application Performance Management: service control from the user's point of view with sample transactions;
  • data historicisation in an open-source data warehouse, SpagoBI, for analysis and reporting.

Knowledge management

  • a centralised document repository built on Alfresco;
  • organisation, persistence and use of the knowledge produced while delivering the services;
  • support for training and skill-building activities.

The governance Monitor Dashboard

On top of everything, the Monitor Dashboard built on SpagoBI. For a public administration this is no cosmetic detail: it is the difference between claiming service quality and being able to prove it, number by number. The dashboard offered:

  • a "traffic-light" overview of the global state of the services;
  • navigation of the KPIs associated with each service delivered;
  • drill-downs across different dimensions and time periods (real-time, day, week, month);
  • reports, OLAP, data mining and QBE;
  • a Logbook: an overall view of all the events characterising service delivery, filterable by time, service, system, triggering event and type;
  • a charge-back module allocating costs to the consuming directorates;
  • management of technical staff attendance and shifts;
  • Site Activity to track "who does what";
  • project activation with progress tracking via Redmine.

The ITIL core processes

A complete set of Core Processes (PF) based on ITIL v3 was then implemented, tailored to the Authority's specific needs:

Process Scope
PF01 — Incident ManagementRestoring operations in the shortest possible time
PF02 — Problem ManagementIdentifying root causes
PF03 — Application Service ManagementTake-over, evolution and termination of services (a process created ad hoc)
PF04 — Change ManagementControlled management of changes
PF05 — Service Asset & Configuration ManagementManaging the CMDB
PF06 — Release & Deploy ManagementManaging releases
PF07 — Service Level ManagementSLA monitoring and reporting
PF08 — Capacity ManagementResource planning and allocation
PF09 — Knowledge ManagementKnowledge management

Change Management and IMaaS

For infrastructure changes we introduced the concept of "Infrastructure Model as a Service" (IMaaS): the Project Development Team delivers infrastructure models which, once approved, are turned into operational IT architectures. The change process unfolds as follows:

  1. Request For Change and feasibility analysis;
  2. Service Design according to ITIL's 4 Ps (People, Products, Processes, Partners);
  3. release of the Service Design Package (SDP) containing requirements, technical specifications, test plans, transition plan and operation plan;
  4. approval by the Authority;
  5. Transition: development of test and acceptance environments;
  6. go-live with the support of the Operations Management Team;
  7. CSI (Continual Service Improvement) across all phases.

A virtual lab (IaaS) was also proposed to simulate the transition phases, with the aim of simplifying change adoption and management, standardising transition activities, safeguarding the integrity of production infrastructure configurations and reducing problems and incidents in production.

Business Continuity and Disaster Recovery: continuity that pays for itself

The most innovative chapter is business continuity. The solution was designed in full compliance with:

  • Italian Legislative Decree no. 235 of 30 December 2010, Art. 50-bis (operational continuity);
  • the DigitPA directives (now AgID) for technical feasibility studies;
  • the ISO 22301 / BS 25999 standards for Business Continuity Management;
  • the ITIL framework — the IT Service Continuity Management (ITSCM) process;
  • the M_o_R methodology (Management of Risk) for risk analysis and risk management.

Multi-site architecture

The classic approach involves a secondary site that costs money, consumes power and sits idle waiting for the emergency. Here we did the opposite, structuring the architecture on several levels.

Business Continuity Site (BCS) — inside the Authority's campus:

  • a both-active model: the BCS is not a passive cost but takes an active part in delivering services in real time;
  • asynchronous storage-level replication (Tier 4) for critical services, via IBM SVC virtualisers and Global Mirror/FlashCopy systems;
  • agent/snapshot-based replication (Tier 3) for lower-criticality services, via VMware Site Recovery Manager and vSphere Replication;
  • a 10 Gbit ethernet backbone and an extended Storage Area Network with dual optical-fibre fabric;
  • autonomous internet connectivity and dedicated perimeter protection systems.

Disaster Recovery Site (DRS) — at an external provider more than 80 km away:

  • native architectural readiness for emergency activation;
  • interconnection via DWDM (Dense Wavelength Division Multiplexing) technologies for distances above 100 km;
  • a largely virtualised infrastructure with bare-metal restore capability.

A continuity site that sits idle is a cost waiting for a disaster. One that works every day is productive capacity — and it pays for itself.

LBL Surface Cluster: orchestration and split-brain prevention

The most insidious technical risk of a two-active-site architecture is split brain: both sites convinced they are the "real" one, with logical data corruption as a result. To manage failover/failback procedures and prevent this scenario, thanks to the collaboration with our partner Oplon we introduced the LBL Surface Cluster suite (Decision Engine and WorkFlow), with the following capabilities:

  • the "Split Brain Assassin" algorithm: preventing logical data corruption in case of parallel feeding of the databases;
  • Decision Engine: 3 clustered instances per site, constantly verifying that critical services are working and taking decisions;
  • WorkFlow: coordinated execution of activity sequences (scripts, restarts, shutdowns) via RWC remote commands (Remote Workflow Command);
  • self-documentation of failover and failback procedures;
  • an intuitive web console allowing services to be reactivated even by non-specialist staff — because in real emergencies the specialist may not be there.

Protection levels (Tiers)

Not all services are worth the same, and the Business Impact Analysis translated that into differentiated protection levels — the same principle we tell in "Backup is not enough": RTO and RPO are defined per service, not for the organisation as a whole.

Tier RPO RTO Technology Services
Tier 41 hour 4 hours Asynchronous storage replication (Global Mirror) Critical end-to-end services
Tier 34 hours 4 hours Software agents, snapshots, Site Recovery Manager Medium-criticality services
Tier 2Vaulting with TSM, snapshots, deduplicated VTL Consolidated backups
Tier 1TSM, snapshots Non-critical services

The plans: BCP, DRP, BIA and risk analysis

A complete body of documentation was then drafted, and maintained over time:

  • the Business Continuity Plan (BCP): identification of exposure to internal and external hazards, hardware/software assets and processes to prevent outages and restore services;
  • the Disaster Recovery Plan (DRP): step-by-step operational procedures, checklists and technical diagrams for the reactive Response and Recovery phases;
  • the Business Impact Analysis (BIA): quantifying the impact of IT service interruption on the business;
  • Risk Analysis and Risk Management: identifying threats, assessing their probability and defining countermeasures.

Managing the application service

The process — created ad hoc — that governs the entire life of an application in operation, from birth to decommissioning.

Take-over and start-up

The take-over process for a new application service involved:

  1. verifying the completeness and adequacy of the documentation;
  2. impact analysis and change management (CAB);
  3. setting up the production environment (base software, application, database);
  4. setting up the monitoring environment (probes and specific rules);
  5. setting up the management environment (parameterisation and handover);
  6. registering the Configuration Items in the CMDB;
  7. acceptance testing of the entire infrastructure;
  8. training the Operations Management Team.

Service evolution

For evolutions (new features, configuration changes), the process involved identifying the changes and analysing their impact, running the regression test before releasing to production, and updating the CMDB and the operational documentation.

Service termination

For decommissioning a service: verifying the impact on the other services, saving the archives, uninstalling the application component, updating the CMDB and reallocating resources through Capacity Management.

Training, knowledge management and know-how transfer

A multi-year engagement is also judged by how it ends: the knowledge, in the end, must stay with those who have to exercise it.

Staff training

A process of continuous knowledge sharing was implemented, with:

  • an adequate number of training days per year for each resource;
  • ITIL certification paths (Foundation, Service Strategy, Service Design, Service Operation, CSI) for unit leads;
  • courses at the in-house ICT Advanced Training School;
  • distance learning and in-person seminars;
  • self-training on product manuals and documentation resources.

Knowledge management

The knowledge management process was structured in two recursive phases: organising knowledge (identifying needs, researching, cataloguing in a repository classified by service process) and valorising knowledge (distributing, using and continuously updating the know-how collected).

Know-how transfer at the end of the engagement

The exit management plan involved:

  • Step 1 — Planning: preparing the documentation, preliminary meetings, setting up seminars, verifying the state of the systems through checklists;
  • Step 2 — Shadowing: face-to-face learning sessions, autonomous execution of specific tasks, assessment and self-assessment, on-the-job training;
  • two months of on-demand assistance after the end of the engagement (e-mail, phone, online communication).

Resources and skills

The Centralised Systems Support Team operated with 24/7 coverage across six vertical areas:

  1. Infrastructure (IBM/HP blade servers, networking, VMware/Citrix virtualisation);
  2. Middleware (WebSphere, JBoss, IIS, Apache, Tomcat);
  3. Databases (Oracle, SQL Server, PostgreSQL, MySQL, DB2);
  4. Cartography (GIS and territorial systems);
  5. Open Source (Linux, Red Hat, FOSS solutions);
  6. SAP (ERP, BW, CRM environments).

All of it underpinned by a system of certifications and partnerships:

ISO 9001ISO 14001ISO/IEC 27001CMMI Level 3ITIL v3PMPPRINCE2

Results and benefits

Indicator Result
Service availability SLAs met and improved compared with contractual requirements
Operational coverage 24/7 coverage guaranteed across 6 critical technology areas
Specialist support activation (blocking) Within the next business day
Specialist support activation (non-blocking) Within 5 working days
Risk of losing critical services Reduced by 50% thanks to the active BCS model
Business Continuity ROI Almost immediate return on investment
Applications managed +15% over two years
Transparency and control A single dashboard with real-time KPIs and per-directorate reporting
Regulatory compliance Full adherence to Legislative Decree 235/2010 and the AgID guidelines

Infrastructure benefits

  • reliability: a solution certified by the vendors and validated through comparable implementations;
  • elasticity: hardware and software components chosen to absorb infrastructure changes with minimal impact;
  • standardisation: constant benchmarking against ISO standards, the ITIL framework and sector regulations;
  • resilience: correct sizing of the infrastructures and careful technology choices;
  • control: global management of all sites through centralised tools.

Economic benefits

  • economies of scale: the active BCS increased the overall volume of services delivered;
  • virtualisation: a reduction in the infrastructure's total cost of ownership (TCO);
  • storage optimisation: rationalisation through multi-tier technology (EasyTier) with SSD disks for critical data and capacity disks for the rest;
  • reuse of existing assets: limiting new hardware purchases by redistributing resources the Authority already owned.

Organisational benefits

  • overcoming the development/operations conflict thanks to the Change Area;
  • continuous improvement institutionalised through the Audit Area and the CSI process;
  • transparency in the cooperation between the Authority and the supplier;
  • the ICT Service Charter as the formalisation of the service pact;
  • reporting and charge-back of costs to the consuming directorates.

The elements of innovation

The "both-active" model

The Business Continuity Site is not a passive cost: it takes an active part in service delivery, generating immediate ROI.

An integrated Change Area

Overcoming the structural conflict between development and operations teams, with a controlled, documented take-over of every new service.

A unified governance platform

Tool fragmentation overcome in a single integrated ecosystem: CMDBuild, RT, Zabbix, Alfresco, Redmine, SpagoBI, LBL.

LBL Surface Cluster as DR orchestrator

Automation and self-documentation of failover/failback procedures, with split-brain risk managed at whole-data-centre level.

The IMaaS approach

Infrastructure Model as a Service: industrialising the infrastructure design process, reusing models and components in an "industrial" logic.

A virtual lab for transition

Simulating acceptance phases in a virtual environment (IaaS) to reduce risks and incidents in production.

A traffic-light Monitor Dashboard

Intuitive KPI navigation with multidimensional drill-downs and automated reporting (OLAP, data mining, RSS).

Conclusions

The project stands as a case of excellence in the integrated management of technology infrastructures for the public sector. The combination of specialist skills, ITIL methodology, open-source governance tools and innovative Business Continuity solutions enabled the Authority to:

  • guarantee the operational continuity of its IT services in compliance with the regulations in force;
  • obtain an integrated, transparent view of the entire technology infrastructure;
  • reduce operational and disaster risks by 50%;
  • optimise costs through virtualisation, asset reuse and the active BCS model;
  • continuously improve the quality of the services delivered through structured auditing and CSI processes;
  • preserve the knowledge capital through a structured knowledge management and know-how transfer plan.

The solution proved scalable, sustainable and fully integrated with the Authority's organisational and technological context: a model that can be replicated by other public administrations of comparable size and complexity — we talk about it on our local government page, while our measure-analyse-improve approach is told in our method. If your organisation is facing a similar transition — governance to build, continuity to bring up to standard, services growing faster than the capacity to manage them — let's talk.

For confidentiality we do not name the Authority; figures and indicators refer to the period in which the service was delivered.

Lympha Editorial Team

The articles on this blog come from the field experience of our Business Units and Competence Centres: the people writing are the people who design, run and support the systems we write about, every day. Content is provided for information purposes and reflects the state of the art at the date of publication.

Share this article

LinkedIn X Email

You might also like