Virgin Blue's checkin/everything system is back up and almost 12 hours earlier than predicted. More info here.
To be honest I am a little surprised that they managed to pull it off, let alone so far ahead of schedule, I guess a pessimist set the time line, which is better than an optimist.
I suppose though, that there is a pretty good possibility of the whole thing falling over in a big heap again, but here is hoping that they have it all sorted out, after all I have to fly with them again at least once so that I can use my flight credits and to be honest I actually usually like their service etc at least for flights between Sydney and Melbourne.
Showing posts with label VirginBlue. Show all posts
Showing posts with label VirginBlue. Show all posts
Wednesday, October 6, 2010
What? They're cancelling flights again!
What a crock! Yup, in part of the continuing saga related to the complete failure of their Navitaire 'New Skies' booking system two weeks ago Virgin Blue is cancelling more flights.
I talked about the planned downtime in a post on Monday, This really is Virgin on the Ridiculous, which as you can read there seems to be a terrible display of system and software engineering that makes me a little ashamed to be in the industry. The need to cancel flights screams to me that Virgin Blue still hasn't got a good backup check-in and booking system organised. While apparently they've accommodated everyone
So as this saga continues we continue to see flaws in both Virgin Blue and Navitaire, no system switch over should take 32 hours but really Virgin Blue should be able to handle the normal flight load on their backup system with advanced notice to bring in extra staff etc.
I talked about the planned downtime in a post on Monday, This really is Virgin on the Ridiculous, which as you can read there seems to be a terrible display of system and software engineering that makes me a little ashamed to be in the industry. The need to cancel flights screams to me that Virgin Blue still hasn't got a good backup check-in and booking system organised. While apparently they've accommodated everyone
"...on other flights and they will depart within an hour of their original departure time."I am curious why they couldn't with a week or so notice fix the problems associated with speed of check-in in their manual backup process to the point where they didn't need to cancel any flights what so ever. Perhaps it is purely a cost saving exercise, there being no point to fly the extra flights, but given that they usually have these extra flights running and presumably full enough to justify, it seems that it is because they expect delays in the "manual" check-in system again.
So as this saga continues we continue to see flaws in both Virgin Blue and Navitaire, no system switch over should take 32 hours but really Virgin Blue should be able to handle the normal flight load on their backup system with advanced notice to bring in extra staff etc.
Monday, October 4, 2010
This really is Virgin on the Ridiculous
Virgin Blue Minor Incident Report
So, I'm sure by now if you're reading this blog that you have seen my thoughts on the Virgin Blue check-in/reservation/booking system fiasco (Here they are if you missed them, Virgin on the Ridiculous, More background on Virgins IT Systems and Naive Navitaire). It just got a little more ridiculous in my book, they are still using their backup system, which in itself is pretty poor form, and they are planning to switch back over to their main system tomorrow. Here is the statement.
That's right, they expect the switch over to mean that their system will be down for1 day and 8 hours or 32 hours altogether during which time they will be checking people in manually and the phone and online reservation systems will be down.
How crazy is that? So not only did it take Virgin Blue and Navitaire 21 hours to switch over to the backup system apparently it is going to take them 32 hours to switch back to the 'primary' system. What a crock! What an absolute shambles! Is the system that badly designed that they have to untangle a mass of inter related crap instead of changing a few key routers or something like that to switch between the two systems.
I can understand that it would be complicated but anything more than a few hours seems crazy, can't they get the primary up and running in parallel and then flick a switch?
Anyway, thought you might all appreciate an update on the tom foolery.
So, I'm sure by now if you're reading this blog that you have seen my thoughts on the Virgin Blue check-in/reservation/booking system fiasco (Here they are if you missed them, Virgin on the Ridiculous, More background on Virgins IT Systems and Naive Navitaire). It just got a little more ridiculous in my book, they are still using their backup system, which in itself is pretty poor form, and they are planning to switch back over to their main system tomorrow. Here is the statement.
Online and telephone reservations services will be unavailable from 9.00pm AEST on Tuesday 5 October;
- Check in for domestic flights will open two hours before scheduled departure time; three hours for international flights;
- Web check, Kiosk check in and Check-mate will close at 8pm AEST on Tuesday 5 October through to 5am on Thursday 7 October;
That's right, they expect the switch over to mean that their system will be down for1 day and 8 hours or 32 hours altogether during which time they will be checking people in manually and the phone and online reservation systems will be down.
How crazy is that? So not only did it take Virgin Blue and Navitaire 21 hours to switch over to the backup system apparently it is going to take them 32 hours to switch back to the 'primary' system. What a crock! What an absolute shambles! Is the system that badly designed that they have to untangle a mass of inter related crap instead of changing a few key routers or something like that to switch between the two systems.
I can understand that it would be complicated but anything more than a few hours seems crazy, can't they get the primary up and running in parallel and then flick a switch?
Anyway, thought you might all appreciate an update on the tom foolery.
Wednesday, September 29, 2010
Naive Navitaire: Virgin on the Ridiculous part two
Following my 30 or so hour delay on a 90 minute trip, I've been investigating this failure a little bit. I guess I am still unsure how a single server failure could take out a major Australian airline for a day. Then there are the questions about why they don't have better backup and manual systems in place.
There is a bit more information around since my earlier post, Virgin Blue updated its press release on Monday afternoon says:
So basically, the SAN array died, Navitaire guys thought that they could fix it that didn't work and it took them 21 hours to get the backup system working in its place (or to get the hardware replaced and the data restored I'm not sure which). I am flabbergasted that a service provider that at services every Australian Airline and another 70 or so airlines around the world could have such a terrible response to the failure.
From what I can gather around the net the New Skies System is based on .NET and I presume some sort of SQL back-end. This sort of setup lends itself very well to redundancy via data mirroring and load balancing across a group of servers. So why wasn't there a redundant data server sitting there ready? Early quotes (that I can't seem to put my hands on now) indicate that Virgin Blue have a "cheaper" back up solution than Qantas, that should have kicked in within three hours. Obviously 21 hours is a lot longer than 3.
So for that I say screw you Navitaire! Where is my compensation?
There is a bit more information around since my earlier post, Virgin Blue updated its press release on Monday afternoon says:
"We are advised by Navitaire that while they were able to isolate the point of failure to the device in question relatively quickly, an initial decision to seek to repair the device proved less than fruitful and also contributed to the delay in initiating a cutover to a contingency hardware platform."and Virgin Blue and The Register reports that the failure was in a Solid State storage array.
So basically, the SAN array died, Navitaire guys thought that they could fix it that didn't work and it took them 21 hours to get the backup system working in its place (or to get the hardware replaced and the data restored I'm not sure which). I am flabbergasted that a service provider that at services every Australian Airline and another 70 or so airlines around the world could have such a terrible response to the failure.
From what I can gather around the net the New Skies System is based on .NET and I presume some sort of SQL back-end. This sort of setup lends itself very well to redundancy via data mirroring and load balancing across a group of servers. So why wasn't there a redundant data server sitting there ready? Early quotes (that I can't seem to put my hands on now) indicate that Virgin Blue have a "cheaper" back up solution than Qantas, that should have kicked in within three hours. Obviously 21 hours is a lot longer than 3.
A Virgin Blue spokesperson told iTWire that Navitaire was supposed to have a parallel system in place and in case of disaster this would go live within three hours. However, it did not actually come into play until almost a full day after the first incident.But why? I guess we'll possibly never know. There is however one thing that I know, if Navitaire did have that backup system working within 3 hours then while I would have probably been delayed it certainly wouldn't have taken more than 30 hours for me to get home.
So for that I say screw you Navitaire! Where is my compensation?
Monday, September 27, 2010
More background on Virgins IT Systems
Check out this article from May this year discussing Virgin Blue's IT infrastructure.
Virgin on the Ridiculous
Yesterday I was supposed to fly from Sydney back to Melbourne but like around 100000 other jet setters on what is apparently the third busiest air corridor in the world I spend a leisurely six hours at the airport. Two to check into my flight and then 4 sampling the excellent gourmet fair of the airport food court and the lovely community atmosphere generated by several thousand grumpy people in a confined space. The $12 worth of vouchers that Virgin Blue (VB) kindly provided me with secured a pie and a coffee.
Anyway, this post isn't about venting (ok well maybe a little bit) but more an under informed investigation and discussion of the last 24 hours.
Navitaire: Virgin Blue's* Services Provider
I managed to discover last night that VBs service provider for its reservation system is Navitaire and after a little poking around it seems that they use the Open Skies system. A look around the Navitaire website reveals that all of the Australian airlines use the same system. As the outage affects all aspects of the VB operations that the Open Skies system would handle it seems of little doubt that the outage is related to the Navitaire system. It would also seem that each airline has its own servers etc running the Navitaire software other wise more of the airlines would be affected.
Redundancy?
As someone that used to work on a critical business infrastructure product, corporate telephone systems, I have experienced first hand at work the business impact that downtime on one of these systems can have. Our system was supposed to be five 9s reliable, which means that the system was supposed to be down for a maximum of 5.26 minutes per year or 6 seconds a week, including scheduled upgrades etc. If there are no further outages this year then VBs flight reservation system is currently running somewhere around 99.7% reliability. Through a combination of factors we once caused a customer outage that lasted around 7 minutes and where forced to compensate in the order of $50 million.
In light of the high costs associated with downtime our systems had redundancy and lots of it. We had high availability (HA) servers that ran in pairs, you could shoot one and the other one didn't skip a beat. Then if they both failed another set of servers could take over, if the network suffered a large scale outage, servers at each site could take over the calls for the local phones.
So one of the big questions that needs to be asked is "where is the redundancy?" Where are the backup servers located in another location? Why is there not replacement server on standby ready to jump in and run from a current back up of the data?
I hope that at some point in the future we get analysis of the failures so that we can work to avoid such incidences again in the future. While it does seem that Navitaire is the culprit in this instance I'm moving on to talk about VBs handling of the situation.
Information Flows and Transparency
Let me just start out this section by saying I don't blame this on the people at the front line, the customer service folks that managed to, for the most part, maintain the smiles. They are doing a brilliant job. What is not working is the information hierarchy, it was evident that the staff were quick to pass on any and all information that they received but they were evidently not getting enough information and what information they were getting could not be effectively communicated.
I think I am done for now, I am sure that more things will spring to mind. But in the mean time Virgin Blue needs to be told that it is not what crises you face that defines you but rather how you respond to them and their response currently leaves a little to be desired. Studies have shown time and time again that customers will come back to a company after a failure such as this as long as the companies response is well handled.
On that note I am off to attempt to work from my mothers house on a tiny screen!
* I wonder as I write this whether I should be using Virgin's Blue as the plural of Virgin Blue.
Anyway, this post isn't about venting (ok well maybe a little bit) but more an under informed investigation and discussion of the last 24 hours.
Navitaire: Virgin Blue's* Services Provider
I managed to discover last night that VBs service provider for its reservation system is Navitaire and after a little poking around it seems that they use the Open Skies system. A look around the Navitaire website reveals that all of the Australian airlines use the same system. As the outage affects all aspects of the VB operations that the Open Skies system would handle it seems of little doubt that the outage is related to the Navitaire system. It would also seem that each airline has its own servers etc running the Navitaire software other wise more of the airlines would be affected.
Redundancy?
As someone that used to work on a critical business infrastructure product, corporate telephone systems, I have experienced first hand at work the business impact that downtime on one of these systems can have. Our system was supposed to be five 9s reliable, which means that the system was supposed to be down for a maximum of 5.26 minutes per year or 6 seconds a week, including scheduled upgrades etc. If there are no further outages this year then VBs flight reservation system is currently running somewhere around 99.7% reliability. Through a combination of factors we once caused a customer outage that lasted around 7 minutes and where forced to compensate in the order of $50 million.
In light of the high costs associated with downtime our systems had redundancy and lots of it. We had high availability (HA) servers that ran in pairs, you could shoot one and the other one didn't skip a beat. Then if they both failed another set of servers could take over, if the network suffered a large scale outage, servers at each site could take over the calls for the local phones.
So one of the big questions that needs to be asked is "where is the redundancy?" Where are the backup servers located in another location? Why is there not replacement server on standby ready to jump in and run from a current back up of the data?
I hope that at some point in the future we get analysis of the failures so that we can work to avoid such incidences again in the future. While it does seem that Navitaire is the culprit in this instance I'm moving on to talk about VBs handling of the situation.
Information Flows and Transparency
Let me just start out this section by saying I don't blame this on the people at the front line, the customer service folks that managed to, for the most part, maintain the smiles. They are doing a brilliant job. What is not working is the information hierarchy, it was evident that the staff were quick to pass on any and all information that they received but they were evidently not getting enough information and what information they were getting could not be effectively communicated.
- None of the display boards at the airport where updated. They were still showing the flights that where supposed to be going etc as if everything was running according to plan. We all know that even on a good day planes are delayed, cancelled etc and they should be able to reflect such changes on the boards quickly and easily.
- The website updates where too slow. I left for the airport four hours after the system went down because there was nothing on the website to say that I shouldn't. As soon as I arrived at the airport people where told that all non-essential flights should be postponed and to get out of the airport.
- Dissemination of credit and refund policies, I'm still not sure what the polices are!
- Change the messages at the call centre, I didn't need to spend an hour on hold to be told the system was still down. How about having a message at the start telling me that nothing could be done?! Or having that play instead of the "Your call is important to us" malarkey!
- You have email and phone details for a lot of the customers, use them! The only way I am finding semi-current information is Twitter. Check out the search for the latest. This morning I did find out that you would call or SMS people to let them know their new flight details. How about email as well? That could be easily automated!
I think I am done for now, I am sure that more things will spring to mind. But in the mean time Virgin Blue needs to be told that it is not what crises you face that defines you but rather how you respond to them and their response currently leaves a little to be desired. Studies have shown time and time again that customers will come back to a company after a failure such as this as long as the companies response is well handled.
On that note I am off to attempt to work from my mothers house on a tiny screen!
* I wonder as I write this whether I should be using Virgin's Blue as the plural of Virgin Blue.
Subscribe to:
Posts (Atom)
