Evidence behind my technology judgment.
Enterprise software and platform engineering, architecture, AI and MLOps, founder/operator work, research, and academic contribution. Together they show how complex systems are understood - and how operational problems are translated into technology direction that serves a business outcome.
Business outcomes first · Technology in service of strategy
Platform engineering, architecture & AI systems
Hands-on work on platforms, infrastructure, monitoring, MLOps, authentication, and cloud/on-prem systems - how technology operates at scale.
Process, people, decisions, information & architecture
Map how work happens across the system, identify the constraint, then evaluate technology options against the business objective.
Decisions under real constraints
Product, technology, architecture, partnership, and resource decisions where outcomes are personally owned.
Enterprise Decision Intelligence
Emerging research on human judgment, AI, software systems, information, and organizational processes in complex decisions.
Selected complex problems & strategic work
Verified work framed around context, constraint, options, direction, contribution, and outcome. Includes completed work, work in progress, and proposals - without inventing results still underway.
IFS Intelligent Observability - Enterprise AI Assurance
- Context
- IFS can already observe whether AI services are running - availability, latency, tokens, and infrastructure health. What is not yet standardized centrally is whether AI models and agents are behaving correctly: grounded answers, correct retrieval, correct agent workflows, tool correctness, and verified business outcomes.
- Problem
- Should IFS invest in a narrow, enterprise-wide AI Assurance capability that reuses existing observability foundations - or continue with fragmented, team-local monitoring that answers only “is the AI service running?”
- System & constraint
- Authored a cross-team strategic proposal distinguishing technical health from intelligent health. Mapped reuse of proven foundations (OpenTelemetry, Elastic APM, Prometheus/Cortex/Grafana, Langfuse for POC) against the risk of duplicating platforms. Framed the gap as false success: an agent can return HTTP 200 and still perform the wrong business action.
- Options
- Continue fragmented, product-local AI monitoring without a shared assurance contract
- Build a parallel proprietary AI observability stack that duplicates existing platforms
- Additive enterprise AI Assurance - standardize a small assurance signal set into the central monitoring platform
- Direction
- Sponsor structured discovery plus a controlled 12-week POC: define an IFS AI Observability Contract v0.1, instrument one representative agentic workflow end-to-end, publish low-cardinality assurance metrics into the central Grafana/Cortex path, and defer productionization until detection value, overhead, and safe data handling are proven.
- My contribution
- Initiative author and proposer - framing the strategic problem, architecture direction, collaboration map, POC design, and decision requests for managers and cross-functional teams.
- System implications
- Provider-neutral, OTel-first architecture; reuse existing collectors and backends; add IFS-specific assurance semantics only where business verification requires them; keep prompts/content capture opt-in and governed.
- Trade-offs & consequences
- Positions IFS to move from system health to intelligent outcome assurance - including verification of enterprise business state - while avoiding unnecessary platform spend and protecting privacy/compliance boundaries.
- Outcome
- Produced a decision-ready proposal for cross-team discovery and POC sponsorship. Formal enterprise rollout is explicitly deferred until the POC proves operational value.
- Proposed · discovery + 12-week POC
Observability Cost & Ownership - Elasticsearch Worst Culprits
- Context
- Verbose and high-volume logging patterns across containers, APIs, and methods were driving Elasticsearch observability cost without clear ownership. Teams lacked evidence packs that connected cost drivers to responsible owners and fix actions.
- Problem
- How should IFS identify the highest-cost log sources, assign them to owning teams with evidence, and prove cost reduction - without guessing or creating unowned operational noise?
- System & constraint
- Driving an enterprise epic under the Observability Framework to identify “worst culprits,” map them to owners, and establish a reusable methodology with before/after cost proof. Work spans non-production and customer-size cost views, FinOps calibration, cost dashboards, and ML-assisted detection of verbose patterns.
- Options
- Continue absorbing log cost without owner accountability
- One-off manual cleanups without reusable evidence or monitoring
- Systematic worst-offender identification with evidence packs, owner handoff, and cost proof
- Direction
- Run a staged operating model: identify top verbose containers and patterns, produce owner evidence packs with fix guidance and monitoring links, deploy cost visibility, and document a reusable methodology for remaining containers.
- My contribution
- Epic owner / lead - driving the initiative, coordinating child workstreams, and connecting observability engineering with FinOps and team accountability.
- System implications
- Cost dashboards in production, FinOps rate calibration, evidence-pack workflows, and ML categorization/alerting for verbose log patterns - integrated with the existing observability platform.
- Trade-offs & consequences
- Turns observability from an unbounded platform cost into a managed investment: owners see evidence, fix instructions, and measurable cost impact.
- Outcome
- Initiative in active delivery (Observability Framework epic). Core identification, evidence-pack, FinOps, and dashboard workstreams underway; full cost-reduction proof documented as definition of done, not claimed as completed savings.
- Epic in progress · FinOps + ownership model
Enterprise Platform Capability Enablement
- Context
- A shared monitoring and logging platform needed consistent understanding across a large global R&D organization - spanning on-prem and cloud deployments, observability architecture, and production diagnostics.
- Problem
- How should platform knowledge, operating practice, and troubleshooting capability be scaled across hundreds of engineers without fragmenting standards?
- System & constraint
- Assessed gaps between platform architecture complexity and day-to-day engineering readiness. Identified that uneven understanding of logging, alerting, and diagnostics created operational risk and slowed incident resolution.
- Options
- Ad-hoc peer support and tribal knowledge
- Documentation-only enablement
- Structured technical enablement with training media and operating guidance
- Direction
- Invest in structured platform enablement - translating complex observability architecture into shared organizational capability through training, documentation, and scalable technical practice.
- My contribution
- Platform engineer responsible for technical enablement design, content, and delivery to global R&D audiences.
- System implications
- Standardized understanding of distributed logging, alerting architecture, Kubernetes observability, and production diagnostics across on-prem and cloud contexts.
- Trade-offs & consequences
- Reduced dependency on a small set of specialists; improved consistency of platform operations; supported faster, more reliable engineering practice at organizational scale.
- Outcome
- Enabled 600+ global R&D engineers to operate and troubleshoot a shared enterprise monitoring platform through training videos, documentation, and deep-technical sessions.
- 600+ engineers enabled
Architecture Translation for Customer-Facing Decisions
- Context
- Customer-facing teams needed a clear narrative of monitoring architecture that non-technical stakeholders could use in commercial and solution conversations.
- Problem
- How should complex platform architecture be translated into decision-ready narratives for pre-sales and customer engagement?
- System & constraint
- Technical depth alone was insufficient for customer-facing conversations. The gap was between architecture accuracy and decision clarity for commercial stakeholders.
- Direction
- Develop non-technical enablement that preserves architectural integrity while framing capability, trade-offs, and value in language suitable for customer-facing teams.
- My contribution
- Technical contributor translating monitoring architecture into decision-ready narratives for pre-sales enablement.
- System implications
- Preserved accurate representation of monitoring architecture while reducing cognitive load for non-specialist audiences.
- Trade-offs & consequences
- Improved the organization's ability to communicate platform capability in commercial contexts - a direct bridge from technical complexity to stakeholder understanding.
- Outcome
- Delivered pre-sales enablement that made enterprise monitoring architecture usable in customer-facing decision conversations.
Operational Risk Reduction through Monitoring Automation
- Context
- SRE on-call load from platform monitoring and alerting was consuming engineering capacity and increasing operational noise.
- Problem
- Where should automation be applied to reduce operational alert volume without weakening visibility into critical systems?
- System & constraint
- Worked with SRE teams to identify alert patterns amenable to end-to-end automation while preserving coverage for critical failure modes.
- Direction
- Automate monitoring and alerting workflows for high-noise, high-volume patterns to restore signal quality and reduce on-call burden.
- My contribution
- Software engineer collaborating with Site Reliability Engineering on monitoring and alerting automation.
- System implications
- End-to-end monitoring and alerting automation integrated with platform operations.
- Trade-offs & consequences
- Lower operational overhead and improved reliability posture - freeing specialist capacity for higher-value work.
- Outcome
- Reduced SRE on-call alerts by approximately 50% through monitoring and alerting automation; maintained 99% uptime for critical on-prem and cloud environments.
- ~50% reduction in on-call alerts
- 99% uptime for critical systems
Enterprise technology & systems experience
Platform reliability, observability, AI-enabled systems, and translating technical complexity into guidance engineering teams can use.
Enterprise platform complexity
Work on systems used across global R&D - shared platforms, ownership, and friction between local teams and common standards.
Reliability & observability
How operational data, alerts, architecture, and human response interact. ~50% reduction in on-call alerts; 99% uptime on critical environments.
MLOps / AI systems
How AI services run in production - monitoring, CI/CD, safety validation, and a proposed Intelligent Observability / AI Assurance direction.
Platform standardization
Shared CI/CD, deployment, and platform patterns for on-prem and cloud - practices that affect how systems scale.
Technical enablement
Developed technical enablement for the Monitoring platform used by 600+ R&D engineers, translating observability concepts into operational guidance. Also supported pre-sales architecture explanation.
Observability cost & ownership (in progress)
Driving the Elasticsearch Worst Culprits work - identifying high-cost log sources, assigning owners with evidence, and tying cost visibility into operations.
Founder & operator experience
Product, technology, architecture, partnership, and resource decisions under real constraints - where outcomes are personally owned.
nZO Innovations
Nov 2020 – PresentFounder & Director
Technology consulting & solution advisory
Technology consulting and solution advisory: understand the business need before defining the system. Problem discovery, process understanding, solution advisory, architecture, platform design, integration, AI adoption, automation, and technology direction.
- Understand the business need before defining the system
- Solution design, platform architecture, integration, and automation
- AI adoption and digital platform decisions under real constraints
- Build vs buy vs partner framing where a capability is needed
Entertain Passport
May 2026 – PresentCo-Founder & Director
Multi-sided access platform
A multi-stakeholder access system: customers, venues, organizers, artists, tickets, identity, NFC, gates, payments, data, partner workflows, and platform architecture - physical and digital together. Evidence of systems thinking under founder/operator constraints.
- Multi-stakeholder ecosystem and partner workflows
- Identity, access, and gate operations across physical and digital channels
- Product prioritization under market and resource constraints
- Architecture choices for a multi-sided platform
Entertain Passport is owned and operated by Entertain Passport Pvt Ltd. Platform development and technology solution design are led by nZO Innovations.
Visit Entertain Passport →Research - Enterprise Decision Intelligence
Emerging research exploring how human judgment, AI, software systems, information, and organizational processes can work together to improve complex enterprise decisions and operations. Presented as a developing direction - not an established proprietary discipline.
Enterprise Decision Intelligence
Themes under development
- Human judgment under uncertainty
- Human-AI collaboration in operations
- Enterprise systems and information flow
- Decision support and technology direction
- AI adoption as an operating-model problem
- Build vs buy vs integrate vs partner
- Cognitive load, trust, and automation bias
My research
- Towards Neuro-Inspired Enterprise Intelligence: Computational Models of Human Cognition for Next-Generation Decision Support Systems
- Towards Tiny Transformers (ongoing - BrAIN Labs Inc.)
Researcher
- Towards Tiny Transformers (ongoing research).
Selected research supervision
Research projects under supervision at IIT - listed separately so they are not mistaken for personal publications.
- Towards Secure and Adaptive Knowledge Evolution for Retrieval-Augmented Generation over Continuously Evolving Enterprise Knowledge
- An Intelligent Framework for Client-Specific Business Rule Validation in ERP Master Data Migration
- Towards Reliable and Context-Aware Emotion Intelligence for Sinhala Customer Support Conversations
- Intelligent Decision Support for Fair and Adaptive Revenue Allocation in Collaborative Travel Booking Platforms
- Enhancing MRI Image Segmentation Through Quantization and AI-Powered Algorithms for Clinical Efficiency
- CCTV-Based Enhanced Public Security Management System for Sri Lanka
Education & professional development
Completed qualifications only. Planned programmes are not listed as completed.
- MSc Software EngineeringUniversity of Westminster (UK)Completed
- BSc (Hons) Information TechnologySLIITSpecialized undergraduate degree - completed
Supports progression from technology toward systems, strategy, and decision intelligence. Further study will be listed here when completed - not before.
Selected teaching, speaking & academic service
Communication, mentoring, and academic contribution - secondary to the core practice.
Visiting Lecturer, Research Supervisor & Viva Panel Examiner
- Undergraduate and master's student supervision, examination, and research guidance.
- Guest lectures connecting industry practice with structured analysis.


Data Structures & Algorithms
Guest lecture to 300+ undergraduate students - structured communication of foundational concepts.
View institutional reference →
Data Structures & Algorithms
Guest lecture to 150+ students - connecting enterprise engineering practice with academic foundations.


Competitive Programming Workshop
Workshop speaker - mentoring structured problem-solving for software engineers.


Inter-University Hackathon Judge
Invited judge for inter-university projects - technical judgment under time constraints.
Where media is available it is shown here; entries are backed by institutional references or published coverage.
Selected insights
Notes on systems, AI and human work, and technology decisions.
- Why AI strategy must begin with business architectureAI + Human Systems
Professional history
Chronological roles for reference. The sections above explain how this experience informs systems and technology problem solving.
Senior Software Engineer, Platform Infrastructure
- Authored the IFS Intelligent Observability proposal for enterprise AI Assurance - strategy, architecture direction, and 12-week POC design for cross-team sponsorship.
- Driving the Observability Framework epic on Elasticsearch Worst Culprits - FinOps cost visibility, owner evidence packs, and reusable cost-reduction methodology.
- Enabled 600+ global R&D engineers on the Monitoring & Logging Platform through structured technical enablement.
- Member of the IFS.AI research guild, contributing to enterprise AI initiatives.
- Standardized platform CI/CD, test automation, and deployment practice for on-prem and cloud deployments.
- Contributed to MLOps architecture: monitoring, CI/CD, and safety validation for ML services.
- Contributed to platform architecture initiatives including model inheritance, time zone support, EBR, backward compatibility, and near-zero-downtime upgrades.
- Delivered technical workshops for R&D new joiners; conducted engineering interviews and mentorship.
Software Engineer, Platform Infrastructure
- Reduced SRE on-call alerts by ~50% through end-to-end monitoring and alerting automation.
- Created training materials and documentation for internal teams and external customers.
- Managed on-prem and cloud environments with 99% uptime for critical systems.
Associate Software Engineer
- Contributed to zFactory - manufacturing operations, software, real-time information, and platform systems for apparel production. Relevant to later advisory work on process + systems + data.
- Built platform provisioning automation for client-wise configuration, reducing provisioning/configuration time by ~80%.
Software Engineer Intern
- Built Java REST APIs with Elasticsearch and MongoDB microservices enabling real-time access across 100+ devices - early exposure to integrations supporting operational workflows.
- Developed hybrid mobile applications used by 1,000+ users collecting 1M+ data points per day.