Executive summary
I organized the work as four coordinated delivery streams, modernized an approximately 10,000-line shell-based upgrade framework, and connected it to centralized monitoring, reporting, and event detection. The automated rollout grew from one system to approximately 2,000 upgrades in a night while recovery controls contained localized failures.
The business problem
Walmart needed to retire Informix versions 9 and 10 before the aging estate created a more expensive support position. A uniform upgrade plan would have been unsafe: some systems needed decommissioning, some required an operating-system migration, and shared or business-critical environments needed extensive dependency testing.
Each system therefore required a decision to upgrade in place, migrate, decommission, or defer while the program prioritized the estate exposed to the support deadline.
My role and scope
I served formally as the Informix SME and project lead. Although I was not a people manager or designated technical lead, two contractors and one employee relied on me for technical direction and program coordination. I owned program scope and sequencing, the common test strategy, upgrade automation, central metrics and reporting, release readiness, rollback planning, and live monitoring during the distributed rollout.
Four coordinated workstreams
Stores and pharmacies
Highly repeatable environments made automation possible, but narrow maintenance windows and archive-based recovery demanded strict preflight checks, timing controls, validation, and automatic rollback.
Distribution centers
Approximately 500 servers supported the heaviest Informix workloads. Six months of testing preceded roughly six months of deployment using disaster-recovery pairs, A/B failover, application validation, and a downgrade or failback path.
Home-office shared databases
Hundreds of teams depended on shared servers with multiple schemas and tightly coupled applications. Systems were ranked by version, deadline exposure, dependency risk, and operating cost; some were upgraded, some migrated, and others decommissioned.
Automation and instrumentation
I updated an approximately 10,000-line shell framework to coordinate package and checksum validation, deployment, connection shutdown, database upgrade, integrity checks, service health, log review, connectivity and timing validation, application reconnection, alerting, and recovery.
The framework did not treat successful installation as successful modernization. Application connections remained closed until database integrity and technical health checks passed. Central metrics, SQL reporting, and event monitoring identified systems needing intervention and supported evidence-based rollout decisions.
Scaling a controlled rollout
The store rollout expanded through waves of 1, 5, 10, 25, 50, 100, 250, 500, 1,000, 1,500, and finally approximately 2,000 systems. Production changes were limited to about two nights per week, outside peak business days, leaving recovery space between waves.
Each stage produced evidence about duration, failure modes, operational load, and rollback performance. Once the safety mechanisms had proved reliable, I recommended moving to 2,000 systems in one night to reduce the growing human cost of repeatedly supervising smaller batches.
Learning from failures
Localized slowness, client-connectivity symptoms, and infrastructure issues triggered recovery without becoming escalated events. Early failures also showed that an upgrade could expose existing disk corruption and then be blamed as its cause, so disk validation became a preflight control.
Recovery matched each environment: archive restore for stores and pharmacies, repair or downgrade when appropriate, and A/B failback for distribution centers. Reversibility allowed the program to increase speed without relaxing risk controls.
Measured results
- Modernized an Informix estate of approximately 18,000 servers across four portfolios.
- Automated upgrades across approximately 15,000 store and pharmacy servers.
- Scaled from one pilot system to approximately 2,000 upgrades in one night.
- Retired Informix v9 and v10, completing v10 nearly one year ahead of its IBM support deadline.
- Coordinated approximately 13 store teams, four pharmacy teams, 13 distribution-center teams, and hundreds of home-office teams.
- Contained localized failures through archive restore, downgrade, or A/B failback with zero escalated events.
Leadership lessons
Enterprise rollout speed comes from evidence and reversibility. The four portfolios could not share one schedule or recovery mechanism, but they could share a decision framework, automation foundation, readiness model, monitoring system, and escalation discipline.
Continue the conversation
Review my LinkedIn profile or contact contact@bengesoftwarellc.com.