Tuesday, June 5, 2012

NetApp Experience: Shelf ADD => Disk Fail => Failover

During a shelf add last week, I experienced as big of a system outage as I've ever encountered on NetApp equipment.  We started seeing a few of these errors, which are normally spurious:


ses.exceptionShelfLog:info]: Retrieving Exception SES Shelf Log information on channel 0h ESH module A disk shelf ID 4.
ses.exceptionShelfLog:info]: Retrieving Exception SES Shelf Log information on channel 6b ESH module B disk shelf ID 5.


The first connection went smoothly, but when I unplugged the second connection from the existing loop, I started seeing some scary results.  Here's the order of important messages:

NOTE: Currently 14 disks are unowned. Use 'disk show -n' for additional information.
fci.link.break:error]: Link break detected on Fibre Channel adapter 0h.

disk.senseError:error]: Disk 7b.32: op 0x2a:1bc91268:0100 sector 0 SCSI:aborted command -  (b 47 1 4e)
raid.disk.maint.start:notice]: Disk /aggr3_thin/plex0/rg0/7b.32 Shelf 2 Bay 0  will be tested.
diskown.errorReadingOwnership:warning]: error 46 (disk condition triggered maintenance testing) while reading ownership on disk 7b.32
disk.failmsg:error]: Disk 7b.32 (JXWGA8UM): sense information: SCSI:aborted command(0x0b), ASC(0x47), ASCQ(0x01), FRU(0x00).
raid.rg.recons.missing:notice]: RAID group /aggr3_thin/plex0/rg0 is missing 1 disk(s).
Spare disk 0b.32 will be used to reconstruct one missing disk in RAID group /aggr3_thin/plex0/rg0.
raid.rg.recons.start:notice]: /aggr3_thin/plex0/rg0: starting reconstruction, using disk 0b.32

[disk.senseError:error]: Disk 7b.41: op 0x2a:190ca400:0100 sector 0 SCSI:aborted command - (b 47 1 4e)


diskown.errorReadingOwnership:warning]: error 46 (disk condition triggered maintenance testing) while reading ownership on disk 7b.41
[raid.disk.maint.start:notice]: Disk /aggr3_thin/plex0/rg1/7b.41 Shelf 2 Bay 9  will be tested
[disk.senseError:error]: Disk 7b.37: op 0x2a:190ca500:0100 sector 0 SCSI:aborted command - (b 47 1 4e)


raid.config.filesystem.disk.failed:error]: File system Disk /aggr3_thin/plex0/rg1/7b.37 Shelf 2 Bay 5  failed.
[disk.senseError:error]: Disk 7b.40: op 0x2a:190ca500:0100 sector 0 SCSI:aborted command - (b 47 1 4e)
raid.config.filesystem.disk.failed:error]: File system Disk /aggr3_thin/plex0/rg1/7b.40 Shelf 2 Bay 8  failed.
raid.vol.failed:CRITICAL]: Aggregate aggr3_thin: Failed due to multi-disk error



disk.failmsg:error]: Disk 7b.37 (JXWG6MLM): sense information: SCSI:aborted command(0x0b), ASC(0x47), ASCQ(0x01), FRU(0x00).

disk.failmsg:error]: Disk 7b.40 (JXWEEB3M): sense information: SCSI:aborted command(0x0b), ASC(0x47), ASCQ(0x01), FRU(0x00).
raid.disk.unload.done:info]: Unload of Disk 7b.37 Shelf 2 Bay 5 has completed successfully
raid.disk.unload.done:info]: Unload of Disk 7b.40 Shelf 2 Bay 8  has completed successfully

Waiting to be taken over.  REBOOT in 17 seconds.

cf.fsm.takeover.mdp:ALERT]: Cluster monitor: takeover attempted after multi-disk failure on partner


Long story short, this system had caused numerous issues in the past, and we replaced both a dead disk and an ESH module.  After that, the system stabilized: "Since the ESH module replacement there were no new loop or link breaks noticed in subsequent ASUPs."

Wednesday, April 4, 2012

BJJ: Armbar from Guard

Pretty amazing, how you learn even things you thought you knew.

My armbar from guard has always been terrible: I've spent years faking that move just to get to something else.  I could never seem to get my hips shifted quickly enough: putting my foot on their hip was too big of a giveaway.  

 But yesterday, everything changed.  Watch this video by world-class black belt of black belts Pedro Sauer: ignore the foot on the hip.  Focus on the other leg.  With no foot on the hip, you can swing your other foot across the person's back in one smooth movement, shifting your hips and catching them by surprise.  So slick.

NetApp Experience: Mixed Ownership

When determining where to add a shelf to a production system, disk show -o is useful in determining which loops/stacks contain disks owned by the controller you're planning for.  When the system is properly set up, this works just fine.  When the system already has mixed ownership on a loop but single ownership on the other loops, you would obviously prefer to expand the single ownership loops.  

But disk show -o will not indicate mixed ownership, so it's a bit of a trap.  The additional step you can take would be to either a)check autosupports before for mixed ownership or b) use disk show -v to show all disks, and verify there are no disks on that loop owned by the other controller.

Thursday, March 29, 2012

NetApp Experience: Amber LED

Quick hit: If you have an amber LED on your filer that won't turn off, try this:
1.  Disable CF.
2.  Halt -f  both heads.
3.  On reboot, enable CF.

Cool trick!

Thursday, January 5, 2012

NetApp Insights: NDU Ifgrp


I'm shamelessly plagiarizing some of my colleagues because this data is just too good to not share.

Q).  Has anybody converted a standalone physical interface into a single-mode vif/ifgrp non-disruptively before?

It should just be a matter of tweaking the partner interfaces and rc files at the appropriate times and doing some takeover/givebacks, but if someone has actually done it for real rather than me working out what I “think” will work in theory then that would be nice.

I’ve got cutting and pasting or “source” a script file that downs all the vlan interfaces, etc then creates a singlemode ifgrp, then recreates all the vlans, aliases, etc as an option, but am also looking for other options that might sound a little bit more NDU to a cust that’s a bit hot at the moment rather than playing with live interfaces on an active controller.

A). As far as the VIF goes I would copy to original rc file, place the new rc in place and yes do takeover/givebacks so long as the networking is done correctly the filers should boot up cleanly using the new rc file.

Tuesday, January 3, 2012

Downloads

This could accurately be filed under the category "rant."

Hey software companies: how do you earn money?  When people use your products, right?  Yeah.  That's when you earn money.

Here's a question: why would you ever make it more difficult for people to TRY your product?  Never.  That would be stupid.

So what's the #1 behavior you want to encourage?  Customer interaction with your product.  Well, how do they interact with it?  Most of the time, they have to download it.  So I would imagine that creating a labyrinth of hyperlinks that would lose and confuse users would be the last thing you'd want to do.  But you freakin guys do it all the time.

Case in point: Spybot Search and Destroy.  You've got users who probably already have a virus.  They're already cranky, they're already frustrated, they already have no patience.  They're dumb end users who just want their computers to work.  So experiment with me: how many clicks does it take you to go from their home page to actually downloading a product?  I count 6.  And that's WHILE knowing where to go.

You know where your download link should be?  On your home page.  Front and center.  People who visit the page should be asked (via text on your homepage) to download your tool BEFORE they know what they're downloading.  What on earth is more important for a homepage than getting your product in your customer's hands?