Showing posts with label on-site service. Show all posts
Showing posts with label on-site service. Show all posts

2/6/08

A trip to Lansing?

Today, Mitch (one of my best friends) called me up from his work place in Lansing. We were doing some basic checks of his server... He needed to open the case and check to see the motherboard model and if all the RAM slots were full. He was planning a RAM upgrade, and I needed details, especially since I was getting different information about his server from different sources. Opening the case would be required.

He said he would call me back once he got the information. I told him not to break the server -- I was headed off to a service call in East Owosso, and would talk with him afterward. The roads were bad, a good 3" (7 cm) of slushy snow. Driving back, it was a little worse, but not too bad. The snow was tapering off, but still, I made the decision that all other service calls were canceled, save for mission critical support.

Once I walked into the store, returning from my service call less than an hour after I left, Mitch was already on the phone with Joel, co-owner of CyberMedics. First words out of Joel's mouth: "Mitch is on the phone, apparently the southbridge heatsync fell off... the server has shut itself off."

This was critical in Mitch's work environment. They only had one server, and everything was loaded on to it. Exchange (mail), Data Server, Domain Controller, everything. It simplifies administration, but it also is one whopping single point of failure, and it was down.

I stopped Joel as soon as he finished speaking and instantly responded, "Transfer him to me." I needed to speak with him immediately to get the details. I was quickly regretting not polishing off my cup of coffee before heading into work.

Yes, it was that bad. Mitch opened the case and the heatsync just fell off... No apparent reason why, it just fell off. It took only seconds for the server to over heat and shut down. And now, the server was unable to boot. Thoughts rolled though my head: circuitry failure from over heating, power supply failure, motherboard failure. All of which are hardware related, none of which were a quick fix.
Status: Server is Down, support required, mission critical, productivity has been stopped.

I was to prepare for a trip to Lansing, a 35 mile (55 km) drive. There is now nearly 8" (17 cm) of slushy snow on the ground. Calls were going crazy, authorization needed to be made before I headed out to work with Mitch, and I needed to schedule things with CyberMedics before heading out.

Within half an hour, a strategy was made, and the trip was planned. I was to head out, stop by my apartment and grab a second set of clothes, a blanket, some food, and something to drink (all but the first were in case my car broke down or I got into an accident, as well as planning on staying the night on site, since the roads were to get worse until after 1 AM) -- if I could, Mitch asked me to pick up some Cheetos for him, if I passed a convenience store, we were planning for a serious all-nighter. The 2-mile drive home took me nearly 20 minutes. While I was there, I took a moment to leave a note for my roommate with emergency contact information, in case I was in an accident or didn't return from Lansing within a day or so; I told myself, "Just in case, ya know?"

Upon leaving with my supplies, I got a few blocks before I determined that the storm had become a lot worse, there was now around 9" (22 cm) of snow on the ground. The snow had picked up visibility was down to nearly nothing, and my little car, basically a Ford Escort without anti-lock breaks, was unable to safely navigate the roads in town... The 35 miles to Lansing seemed quite impossible without getting into an accident or losing control of my car. And I was really beginning to think that those emergency contacts I left my roommate would need to be used if I did go to Lansing this evening. My logic stepped in full force, I had serious doubts that I'd be able to drive though town, let alone to Lansing, without getting into an accident. I needed to cancel on Mitch.

I felt bad. All I could do was call Mitch and talk with him over the phone. I really wanted to be there and help him with his network in this critical state, but it was too dangerous to drive this evening... I told him that I would try to come out the following day if he wasn't able to get the server up and running that evening, if the roads cleared up. I told him to call me with updates as he worked this evening.

The roads were so bad that I even called into work and talked with Joel, I told him that it was unwise to drive my car presently and would not be returning to work. It took me another 10 minutes to drive the half mile back home, down a single road, sliding with each turn, sharing a near-miss with the Shiawassee river at one point.

I just hope Mitch makes it home safely. I'm going to call him later this evening just to be sure.

10/11/07

Guess What Broke Down Again

Early this morning, my MS rep called me, wanting mostly to close the ticket. I bet his manager is breathing down his neck to get the account closed... We've already spent nearly 8 hours on the phone together. I finally let him talk me into it.

About 15 minutes later, I found out that the backup from the night before failed. ~sigh. Nothing sucks more than realizing that you just paid Microsoft $259 to waste two days of work. I'm still troubleshooting it now myself, while I've been in contact with my MS rep (via email only, apparently he doesn't feel like calling me back anymore) trying to get the back ups running again.

Damn you Exchange, you suck.

BTW: I also found this guy who actually did hit the 16 GB limit. Thank God that didn't happen in my Exchange environment... But it seems the two of us are having "post maintenance" troubleshooting. I'm glad that at least my users are able to send and receive easily, but without having decent backups, we definitely having a ticking time-bomb that is known as Microsoft Exchange Server.

[update]
I just finished some basic management of Exchange, I ran eseuitl /k (checksum checks) which turned out well, as well as eseutil /d, which defraged the database just fine. I then ran NTBackup and was able to backup the database... However, Veritas backups still don't work. ~sigh. I might look into troubleshooting that, but I'm not so sure that I'm as worried as I was earlier.

10/10/07

Exchange Sucks

I have no clue why Microsoft doesn't really design a better concept for Exchange. It works, and it works really well, until something goes wrong and undetected for a few months.

The repairs I ran with my Microsoft Support guy did get the job done... We did eseutil /k on the priv1.edb and priv1.stm files. Priv1.edb was find. No errors. Priv1.stm: a few thousand errors. My tech literally said "Oh My God" twice as the errors scrolled up the screen.

He then informed me that deleting the STM file and recreating it would be the best option for this situation. Repairing the file just wasn't as likely to succeed.

Unfortunately, both the database and our server were in bad shape. The server has very little storage space available on it. It was so low in fact, that we had to do much of our database management on an external USB hard drive and to expedite our troubleshooting, we also used the tape drive to back up the MDBData directory.

Once we got everything done, which we started at 5 PM and ended at 9:30, we made a backup of the MSExchangeIS and proceeded on to basically conclude the ticket. People were able to access their storage, the database was able to be backed up again, and I was even able to do an offline defrag of the database. Things were looking good and we were done long before I thought we would be.

I left the building after speaking with my "on-site boss." Neither of us got home until well after 11pm. I then spoke with Don, my "boss/mentor" down at CMC. We talked on the phone until 2 AM about the service call and what needed to be covered a few other things that needed to be done on site to be sure that we had everything under control. I was finally able to get into bed at 3AM.

I woke up at 7:30 to turn on my cell phone, just incase something happened on site, I went right back to bed.

At 8:30 AM I received a call from my on-site boss. There were problems. Apparently, after the work done on Monday night, now ever one on site who connected to Exchange had lost all of their external emails. They had emails from the local domain, but anything outside of that, including emails that have been saved for years were lost. It caused quite a stir come Tuesday morning. Fortunately, the most affected people were the IT staff.

I spent some time on the phone with Microsoft technicians and eventually we decided that further work was needed on the database. I worked with a new MS Technician for over an hour, before I got a call from the tech that I worked with the night before. We ended up deciding to switch the contact from my new support technician back to my original and continue work. I talked with my support tech and we decided that we needed to back restore the database to an alternative location and then we needed to get a second server up and running and mount the database on that guy to retrieve the data if at all possible.

Getting a decent computer, installing Windows 2000 Server and MS Exchange 2000 SP3 w/ Roll-up was troublesome enough. It was worse when I got back on site and received a call from my support tech at 4:30 and was informed that "Oops, we can't actually do that. It won't work unless everything is really similar." Great. 4 hours wasted. At least I got half an hour of sleep while the server OS was installed by Joel and Mitchell.

So, my support tech and I decided that it would be best to take our database offline, mount the old one and use ExMerge to pull out the six mailboxes that we needed. The process should be relatively quick compared to Monday's work.

Assuming, that is, if it works... Which it didn't.

ExMerge got one mailbox out of six. It was the largest mailbox too, at nearly 2 GB. At least that was something.

We then went and loaded up client machines and pulled their mailboxes off right from outlook instead of trying to do it through exchange. A much slower process, but at least it is much more likely to get the correct data. It seemed that it had. We remounted the "good" exchange database and all we needed to do was import the data back into Exchange.

Right then and there, my tech seemed really interested in getting me off the phone. We had pulled the mailboxes off of the server to back them up and we had reconnected the good exchange database, but we had not restored anything to Exchange.

I managed to talk him into helping me restore a single mailbox and keeping the trouble ticket open for another 24 hours before he closes it. And he quickly ushered me off of the phone. I can understand, the two of us have had busy days, but it was a little crappy there at the end... And I was not exactly pleased about the whole "Make a server by 4pm!" thing either.

~sigh. This is why I cannot stand Exchange.

10/7/07

Finally, The Horrors of Exchange, part 1

Last week, Wednesday, September 26, 2007, I was assigned to an on-site contract for a local Owosso business. On day 1, I focused on getting oriented with the network there, mostly just familiarizing myself with the 5 on site servers and the one off site server.

The network is interesting, it is international, two exchange servers, and a whole kaboodle of other traits of the system that give it a personality all its own.

Standard IT policy for managing a network, be it from day one administration or having a new administrator come into a network, the first place we look is at the domain controllers and fire up the event viewer to find out what errors the system is having.

I noticed several errors labeled as "MSExchange..." This is one of my nightmares. Exchange is really sensitive, and can be easy to break or lose data if it isn't dealt with according to Microsoft policy. I'm not sure why Microsoft made Exchange in such a sensitive way, but they must have had a reason for it, seeing as how they've been doing it for over a decade now [Microsoft Exchange Server, Wikipedia]. One of the interesting design concepts of Exchange is how messages are stored on the server. Exchange drops messages into one of several databases: priv1.edb priv1.stm pub1.edb pub1.stm or into an ever-increasing amount of log files. But one of the annoying features of Exchange is that it requires everything to be 100% operational for most of it's own functions to work correctly.

For example, while on site, I found that the Exchange wasn't backing up correctly each night and it had been doing this for several months. After doing some tinkering, I also noticed that it was running out of space. The MS Exchange database --the database is priv1.edb + priv1.stm, was 15.9 GB, while Microsoft has set maximum capacity on the Exchange database to be 16 GB. Email doesn't exactly come in at an extreme speed for the network, so it wasn't slated as a major thing, so we scheduled Exchange Database management for the upcoming Saturday, which would give me plenty of time to check to see what commands I need to run on the database and what contingencies I should be prepared to deal with (cough-disaster recovery-cough).

One day went by without major events, but then on Friday, I arrived on site to find that the Exchange had crashed and was re-enabled. It had crashed, according to event viewer, because the hard drive had gone down to less than 10 MB of free space (thankfully, Exchange was hosted on "D drive" instead of on C). However, when I fired up My Computer to see how much space was presently left on the drive, and found it to be under 512 KB, I immediately reported to my on-site boss that the Exchange server was going to go down any minute, due to the drive being filled up again and I was going to have to shut Exchange down to prevent it. He quickly remote into the server to take a look at the data, and noticed that the drive had 2 MB free and was growing?? And then several people immediately stated that they had suddenly lost email... Yup, Exchange had crashed again.

Now, this isn't a huge issue, Exchange is designed to protect itself in situations like this, by shutting down and displaying error messages so we avoid data loss. However, now that the drive had less than 10 MB free, we were in a bit of a bind. Exchange needs to be defraged (that would be the command eseutil /d) --the defrag command deletes information in the database that has been deleted by users or mailboxes that have been deleted by administrators, which should free up a decent bit of space in the database. The recommended procedure by Microsoft is that we take a backup of Exchange and then run the defrag command. While it is very common for the defrag command to cause problems with the database, it is still a possibility, especially if we have undetected hard ware problems.

So started the second fiasco. Backing up Exchange. Now, we have known for a while that the Veritas Backup Exec backups have been failing on that server for a while now, but they were backing up the entire server, and while it was getting an error while backing up Exchange, it wasn't clear if the Exchange errors were causing the backups to fail, or if it was more related to a tape/tape drive issue. So, I decided to just do an NTBackup.exe backup and save it to an external drive. It was redundant enough for our situation.

The backup started, and 2.5 hours later it had not displayed an error message and switched to data verification. 2.5 hours after that (a total of 5 hours after starting the backup), it returned "Backup failed." I checked the logs, and it said "\Mailsore (%servername%)" failed. The database may be corrupt or inaccessible. The file will not restore correctly. This was a becoming a major stressor for me. I have to do a defrag on the database, which a backup prior-to doing the defrag is highly recommended by Microsoft, yet I am unable to do the backup due to a corrupt database.

To keep this story from getting horribly long, I eventually was able to just run Defrag on priv1.edb, pub1.edb, and pub1.stm, which freed up nearly 6 GB of space. However, I was not able to defrag priv1.stm, which seems to be corrupt. I spent a lot of time with other variations of the eseutil command, but wasn't able to get the database back. I could use eseutil /p [database] /i to rebuild the indexes between priv1.edb and priv1.stm, however, the P switch will not only repair the database, but it will also delete anything it determines to be corrupt. Which, frequently will cause other problems with the database. But, there is one nice thing, even if it does delete files that it shouldn't, we can use the *.log files to replay transactions in exchange to get it back up and running with the correct information.

Once again however, my hopes were dashed. To save space, some log files had been deleted. There is no guarantee that we'll be able to use the log files to get the database up and running again. Hopes dashed. And upon further research, I'm suck between a rock and a hard place. There is no other options available. The only options I've got are to break the RAID array and tinker with one half of the mirrored image to see what happens (no guarantee of success, and it is possible to damage the RAID array in the process), or I can call Microsoft and start a support ticket. My boss at CyberMedics suggests the latter, after talking with the guys on site, we'll make the final decision come Monday.

More details as the story progresses.

3/7/07

what a day at cmc


Once again, there's been some trouble at CyberMedics. Why is it that the business I'm so tied to has such an abundance of luck?

First, I had to rush to a service call in Chesaning, Mi today... About a 20 mile drive. But, first came the delays. Customers rushing in and bosses leaving the store required me to stay at CyberMedics up until 2:30 (I was scheduled at Chesaning at 2:30). I managed to get a call to the site before it got late... Our reason was two: 1) our primary wiring tech was out sick... He had even spent some time in the hospital too (but is getting better) and 2) we were out of Cat5e Ethernet (I didn't inform the customer of this, but it was a cause for us being late).

Since our tech was out sick, I had to quickly pick up the pieces and get someone that didn't have much of an issue with heights to come along with me... That person was Steve. Zack was also supposed to show up, but he didn't make it to the store before 3pm, so I went ahead without him.

As soon as we left, we head off to our local Ethernet vendor, and it was after 3:30 before we were able to get the correct wire! The gods of fortune just weren't with us today! Even when we went to pay, we weren't even able to get though the registers with any expediency... The system wasn't taking our $100 bill, so the cashier had to do some adjustments to get our order processed though, taking twice as long as it should have at the register.

Leaving from Owosso after 3:30, we finally arrived on site at 4pm. Our 10' ladder was about 2' too high to use on site (which was also about how much of the ladder was sticking out of the back of Steve's truck on the drive over here as well), so we ended up having to stand on chairs at the site instead to do all our work, which looked horrible, but at least allowed us to get our work done. I spent most of the time I was there trying to make Ethernet cables to go from the wall jack to the computers. I was having no luck, spending most of my time cutting the wire and re-cutting it, trying to get it to fit correctly in the Ethernet head. ~sigh. I never even finished that... I ended up helping Steve every few minutes anyway, which caused a lot of slow down in my progress.

Steve was spending most of his time there drilling holes for the cabling to go though, since it had to go from a panel in the wall, up the wall, across a dropped ceiling, and then down again to another wall drop about 30' away (total distance is approximately 50' or so). We were doing okay with our work until we go half way across the room, when we ran into a rather tough piece of wood that prevented us form being able to drill though. We called Joel (the boss back at the store) and asked him what we should do... He said that we should reschedule it for another day to finish it, and then we can bring along some more powerful drills.

So, we talked to the receptionist, and scheduled a follow-up visit on Friday to work on the cables. We cleaned up our mess and finally began to head home. Packing up our stuff took quite a while, since we had so much of it, and that damn ladder was sticking out about a foot or so from the end of the truck, but at least with it being bright orange no one was going to hit it on the drive home.

Since it had been a rough day, and especially since Steve hadn't had anything to eat all day, we thought it would be nice to stop by Burger King to get a snack. We both ordered the same thing, that chicken tender sandwich thing. Whatever it is. We were there for 15 to 20 minutes before heading back to Owosso, spending most of our time chatting about work and relationships (Steve found out I was gay the other day, so I feel much more at ease talking about such things around him now). Steve finished his meal before I did, but I tossed the last half of my fries and we headed out. We had to do a quick acceleration out of the BK drive way to avoid getting hit (the speed limit out there is 45 or so). And then again we had some issues at the first stop sign we came to (who ever heard of a stop sign at a busy 4 lane intersection anyway!??!?!). We turned from that and accelerated quickly again. That's when Steve noticed a problem. He had looked in the rear view mirror to ensure that our networking equipment and ladder were still snugly in position for the drive home.

This time, there wasn't a ladder in the back. Our 10' ladder, which is not able to fit completely into the back of Steve's truck was missing! I told him we should turn around, and that it probably fell out at the intersection or when we pulled out of Burger King. I watched the side of the road intently and I kept watching for traffic behaving oddly (such as swerving to miss a ladder in the middle of the road) but I noticed nothing. Our bright orange ladder wasn't on the road/side of the road! It wasn't in BK parking lot, it wasn't around back or off to one side as a practical joke, and the management had no clue (apparently there aren't cameras out there).

I called Joel, he suggested we just head home, the police wouldn't really respond to a $139 ladder that we bought on sale for $99. It would just be a little report and that would be the end of it. Chances would be slim that we would ever see that ladder again. However, the BK manager wanted us to file a police report, so she called them and they said they would send an officer out to take our report. The police station was just two blocks away, so it wouldn't take long for them to get there at all, and even if they sent the State Police (which have a station in Saginaw, just a few miles away) we should hear from them in less than half an hour!

Riiight. An hour later and the cops weren't even there. The manager suggest that she take our contact information so we could get back to Owosso before the store closed. It was almost 7pm before we left the restaurant anyway, but I was getting tired and had some in-store systems to work on after hours anyway. So, we left all of our names, including Don and Joel, just in case they wanted to talk to them. Along with the address for CMC, and then we headed off back to the site to see if we managed to drop the ladder there (which we didn't) and then right next door to that is the police station, where we stopped by to chat with an officer to give them the information, but no one came to the door when we rang the door bell as requested by the note posted on the door. So we left, once again scouring the sides of the road to see if a bright orange ladder was tossed to the side, and once again, we found nothing.

Once we arrived at the store, it was pretty much business as usual. Don was on the phone, Joel had left for the day, and Steve and I tried to get caught up on a few things before heading home for the night. Steve left before I did, leaving just Don and I in the store.

While Don was in the middle of a conversation and I was working on a computer, we heard a knocking at our back door. Both Don and I went to check it out, I figured it was one of Don's guests knocking on the back door, but this time it wasn't... There was a young woman in her late 20s saying that an older woman fell outside the store and needed some help. Both Don and I went to assist. Don wearing short-sleeved scrubs and me wearing my long sleeve tee-shirt. It was quite cold in our sub-zero Celsius mid-Michigan winter, but helping was more important. We helped her to her apartment and she thanked us. While I was helping her up, I recognized her as the dog lady who frequently walks a black Labrador as I'm leaving CMC every evening between 7 and 9pm. I had only talked to her once, when her dog really started barking at me once when I was leaving the store. But, at least she was okay. A scrape here and there, and a few bruises, but nothing major. I was glad we didn't have to call the ambulance or something! What and end to a day that would have been.

Now I'm at home, listening to The Postal Service [alt link] trying to relax from this trying day. ~sigh, I'll be glad when tomorrow comes.