Showing posts with label Clustered ONTAP. Show all posts
Showing posts with label Clustered ONTAP. Show all posts

Monday, July 13, 2015

CDOT Tip #10: CDOT 8.3.1 Features

8.3.1 has a big payload for a minor release and will be out early fall. Here are the Top 10 you care about!


  1. AFF Performance improvements: Reads are 600-900us faster over 8.3.1 from re-engineering the IO path for SSD’s.  
  2. All Flash FAS Configuration
    1. Inline compression always-on by default, with same performance as 8.3 without compression
    2. Zero block detection and dedupe by default
  3. SVM DR (Storage Virtual Machine Disaster Recovery)
    1. This replicates an entire vserver from one cluster to another, including exports, IPs, etc, allow for easy failover to DR site.  Think SnapMirror for vservers!
    2. Automated change management, automated setup and provisioning. 
    3. Retains CIFS shares, NFS exports, permissions, names, data, network config, certificates, QOS policies, and much more.
  4. Additional CDOT MetroCluster configurations, including AFF, FlashPool, and 200km stretch
  5. Cluster Peering Enhancements: This enables CDOT to snapmirror or SVM DR to multiple other clusters in multiple IPspaces.
  6. Encryption for Cloud ONTAP, a CDOT virtual instance available in AWS or Azure.
  7. Foreign LUN Import enhancement: reduces downtime required to migrate LUNs onto NetApp by mirroring dataset and redirecting traffic.
  8. Upgrades to built-in system manager, including AFF-specific changes, enabling ONTAP upgrades from the GUI, and easier network config.
  9. Usability enhancements:
    1. Enhancements for audit log management
    2. Support of LDAP and NIS user authentication for cluster access 
    3. Support for the banner and Message of the Day (MOTD)
    4. New FlashPool caching policies
    5. Automated Workload Analyzer (AWA) volume-level reporting to predict the impact of larger FlashPool/FlashCache.
  10. Protocol enhancements:
    1. Support for dynamic DNS
    2. Support for Windows NFSv3 clients
    3. Support for SMB encryption for data transfers over SMB
    4. Support for configuring a guest UNIX user
    5. Support for mapping the administrators group to root
    6. Enhancements to SQL Server and Hyper-V over SMB solutions


Read more at 8.3.1 release notes: https://library.netapp.com/ecm/ecm_get_file/ECMP12456155

Monday, June 1, 2015

CDOT Tip #8

In order to clarify CDOT’s networking architecture, here’s a basic explanation of each object involved.

1.       Node SVM: aggregates, disks, and ports belong to the node Storage Virtual Machine.
2.       Data SVM: volumes, qtrees, and data LIFs belong to the data SVM.
3.       Ports: You issue commands to the three types of ports the same way.  Those three types are:
a.       Physical Port: at the Physical Port level you can set the MTU or Flow Control. 
b.      Ifgrp (interface group): these are now named in the convention “a0a.”  Ifgrps are made up of ports and exist for redundancy and load balancing.  You can set all the port properties here (changes will override the member ports). Ifgrps have these properties: Role, MTU, Flow Control, Duplex, and load balancing policy.
c.       VLAN: You can assign a VLAN to a port or ifgroup, which creates a virtual port.  VLANS have mostly the same properties as Ifgrps.
4.       LIF (logical interface): a LIF has a Name, an IP address, netmask, Role and a Home Port.  A LIF belongs to a Failover Group and a Routing Group
5.       Failover Group: list of ports a LIF is allowed to be on.  You usually want one for each node management, one for cluster management, and one for 10GbE data.
6.       Routing Group:  these allow a SVM to have different gateways for different VLANs or networks.  A Routing Group has these properties:  a name, address/mask combo (in CIDR notation), role, and metric.   Name data routing groups starting with a d, intercluster routing groups with an i, and cluster network routing groups with a c.

One thing you’ll notice is that Role is now important.  There are several Roles you can assign: Management, Data, Cluster, and Intercluster.  You’ll need to make sure the ports, LIFs, Failover Groups, and Routing Groups are in harmony as to their Role setting.


Also, one last point: SnapVault and SnapMirror are performed using the node’s Intercluster LIFs.  Your data SVMs will not have any intercluster LIFs. 

CDOT Tip #7

Security is always a top priority, so here are some things to consider as CDOT gets rolled out at TR.
1.       Full disk encryption (aka NetApp Storage Encryption)
2.       Non-returnable disk (NRD) entitlement
3.       SafeNet: the replacement for the DataFort data encryption devices is SafeNet StorageSecure.  SafeNet can do file encryption, key management, logging and auditing, and DB/APP encryption.
4.       RBAC – CDOT implements a command specific control, meaning you can give a group or user access to a single command or command tree.  For example, you can give someone access to just “network interface” or even more restricted “network interface show.”
5.       Firewall!  You can set system level, vserver level, and per interface firewall policies. 
6.       Use SSH (disable telnet and rsh)
7.       Alter ssh encryption algorithms per SVM
a.       Aes256-ctr, Aes192-ctr, Aes128-ctr
b.      Diffie-Hellman group exchange sha256
c.       Command: Security ssh show/security ssh modify -vserver -key-exchange-algorithms  - ciphers
6.       Reduce the default Config cli session time-out
a.       Command: System timeout modify 10
7.       SSL/TLS
a.       FIPS mode federal information processing standards 
b.      TLS only!  Command:  System services web modify -sslv3-enabled false
8.       Lock down export/share policies
a.       According to subnet
b.      NFS/CIFS ACL's
9.       Implement Off-box Antivirus
10.   Fpolicy: file based event notification.
a.       Based on file type, share/export, volume. 
b.      Allows you to monitor blocked access attempts.

11.   Log events to external syslog server (event command set)

CDOT Tip #4

Cool trick – You can watch traffic to an individual file in CDOT 8.3 with QOS policy groups:

cdot::qos policy-group> create -policy-group file_iops -vserver tdnas1                

cdot::qos policy-group> show
Name             Vserver     Class        Wklds Throughput 
---------------- ----------- ------------ ----- ------------
file_iops        tdnas1      user-defined 0     0-INF

Now you set this policy group per file:

cdot::> file modify -vserver tdnas1 -volume datastore_volume -file windows.vmdk -qos-policy-group file_iops

Then, you can do a perf analysis on the qos policy group

cdot::qos statistics performance> show -policy-group file_iops
Policy Group             IOPS      Throughput    Latency
-------------------- -------- --------------- ----------
file_iops                   1           0KB/s        0ms
file_iops                   1        0.00KB/s     2.00ms


Note: you only see the file_iops policy-group when doing the statistics command while there is actual traffic on the file with that policy group. otherwise you’ll only see –total- which reflects the whole cluster. 

CDOT Tip #3

Let’s tackle some networking!  Networking is admittedly not my strong suit, so please pepper me with questions if you see something amiss.  Some of these are general recommendations that may not be applicable to you, but they’re good to at least have reference to.
  • Remember, each lif type needs a routing group (mgmt, data, etc)
  • If you create a temporary IP address it may create a temporary routing group.  Make sure you go back and clean it up.
  • Remember to create the ifgrp before your lifs.  It’s a pain to go back!
  • If the switch port is type access, our ifgrps can’t have vlans. We recommend using switch port type trunk even if there’s only 1 vlan to allow for future flexibility
    • switchport trunk encapsulation dot1q
  • Portfast on
  • Disable IP fastpath
    • ::> node run -node * -command "options nodescope.reenabledoptions ip.fastpath"
    • ::> node run -node * -command options ip.fastpath.enable off
  • Disable flow control on all non-Unified Target Adapter (UTA) network interfaces and their associated switch ports
    • ::> net port modify -node -port   -flowcontrol-admin none
  • Create per-network/VLAN failover groups and modify network interface failover-group setting accordingly
    • ::> failover-groups create -failover-group -node -port [-vlan_id]
    • ::> network interface modify -vserver -lif -failover-group
  • You can do a net int show and use –fields pick field names (like routing-group) that aren’t showed by default.  Very useful
  • You can choose which lif to use when pinging.  This is a fantastic testing tool!  net ping -lif-owner svm -lif smlif -destination gwaddress
  • You can’t create a 2-node cluster unless the cluster network is up.  So if you’re setting up a switchless cluster, make sure you connect from each 10Gb port directly to the other node’s IC port.
  • If you’re setting up a switchless cluster, you need to follow these instruction on both nodes.
    • ::> set advanced      (y) 
    • ::*> network options switchless-cluster modify -enabled true


Bonus tip!
If you need to halt one controller without impacting the HA partner, there is no longer a cf disable option.  Use halt with an inhibit takeover switch (there’s also a restart inhibit takeover switch).

Thursday, April 24, 2014

CDOT 8.2.1 Summary

Take a moment to familiarize yourself with this CDOT 8.2.1 documentation.  There is a TON new and improved in 8.2.1 over 8.2, as well as some cautions.  I’ve highlighted a few here.

https://library.netapp.com/ecmdocs/ECMP1368924/html/GUID-45F85A02-114C-4192-8F1B-A4F50996D307.html

Features:
  • Support for FAS8000 series
  • V-Series feature now called “FlexArray,” a non disruptive on-the-fly licensable feature.
  • Support for qtree exports
  • Storage Encryption support
  • Support for direct attach E-Series configurations
  • Non-Disruptive shelf removal support
  • Log and core dumps available via http:///spi/
  • SQL over SMB3 non-disruptive operations support
  • VMware over IPv6 support
  • Offbox antivirus support
  • Health monitoring of Cluster Switches
  • Increased Max aggr sizes
  • 32-64-bit aggr conversion enhancements
  • Automatic Workload Analyzer, which assesses how the system would benefit from SSDs (Flash Pool)
  • Support for “Microsoft Previous Versions” tab on files (8.2 and later)

Cautions:
  • Some Hitachi or HP XP array LUNs might not be visible.  “In the case of Data ONTAP systems that support array LUNs, if the FC initiator ports are zoned with Hitachi or HP XP array target ports before the storage array parameters are set and the LUNs are mapped to the host groups, you might not be able to see any LUNs presented to the Data ONTAP interface“
  • NFSv2 not supported. Windows over NFSv3 not supported.
  • Verify management software versions are compatible.
  • First VLAN configuration may temporarily disconnect the port.
  • LUN revision numbers change during upgrades.  Windows 2008, 2012 interpret these as new LUNs.  
  • Dedupe space considerations and clearing stale metadata for upgrades.
  • Cautions for proper cluster and vserver peering methods
  • Cautions for proper vol move methods

Monday, February 10, 2014

CDOT Raid Groups

A cool bit of trivia about expanding RAID Group sizes.  

For the command storage aggregate modify
For the command storage aggregate add-disks(both quotes from 8.2 command manual)

Which means if you need to balance an unusual number of disks today but will add more disks at some future date, if you want to have small RG’s now and expand them in the future you have to be careful.  For example, if you have 45 disks and want a RG size today of 15 but 20 later, you’d do this:

Storage aggr create –aggregate aggr1 -diskcount 15 –disktype SAS –maxraidsize –s 20
Storage aggr add-disks –aggregate aggr1  -diskcount 15 -disktype SAS –raidgroup –g new
Storage aggr add-disks –aggregate aggr1 -diskcount 15 -disktype SAS –raidgroup –g new

But then we found this:
(Clustered Data ONTAP 82 Physical Storage.pdf)

While this excerpt doesn’t explicitly state you can, I tested in the lab and confirmed: you can manually expand RG sizes of all the RG’s in the aggregate by modifying the max RG size then manually adding disks to each RAID Group.

Storage aggregate add-disks –aggregate aggr1 -diskcount 5 -disktype SAS –raidgroup rg0

Wednesday, January 22, 2014

Adventures with CDOT

Had an interesting situation after an option 4 on a new CDOT system (3250’s), this is a bit long but I wanted to get it all onto paper and into our tribal knowledge.  While one controller was still clearing, I started configuring the other one (set HA true, set to switchless, update from 82p3 to 82p5).  Then I updated the other controller.  When the other system tried to join the cluster, I saw this:

Error: Node "cluster-04" on ring "Management" is offline. Check the health of the cluster using the "cluster show" command. For further assistance, contact support personnel.

Well, ok.  A little checking:
cluster::> cluster show
Node                  Health  Eligibility
--------------------- ------- ------------
cluster-01        true    true
cluster-04        true    true
Warning: Cluster HA has not been configured. Cluster HA must be configured on a
         two-node cluster to ensure data access availability in the event of
         storage failover. Use the "cluster ha modify -configured true" command
         to configure cluster HA.
2 entries were displayed.

cluster::> node show
Node      Health Eligibility Uptime        Model       Owner    Location 
--------- ------ ----------- ------------- ----------- -------- ---------------
cluster-01
          false  true         00:26:23.001 FAS3250              Minneapolis DR Site
cluster-04
          false  true         00:09:39.043 FAS3250
Warning: Cluster HA has not been configured. Cluster HA must be configured on a
         two-node cluster to ensure data access availability in the event of
         storage failover. Use the "cluster ha modify -configured true" command
         to configure cluster HA.
2 entries were displayed.

Well, then let’s modify cluster HA.
cluster::> cluster ha modify -configured true

Warning: High Availability (HA) configuration for cluster services requires
         that both SFO storage failover and SFO auto-giveback be enabled. These
         actions will be performed if necessary.
Do you want to continue? {y|n}: y
Error: command failed: Not enough online nodes in the cluster:
       SL_REMOVE_EPSILON_OOQ_ERROR (code 129)
       There are too few healthy nodes in the cluster to allow join of
       additional nodes. Ensure that the nodes are operational and re-issue the
       command. Use the "cluster show" command on a node in the target cluster
       to view the state of the cluster.

Well sheesh.  So I reboot and what happens?  A takeover.
Jan 22 14:53:28 [msp-cluster-04:callhome.sfo.takeover:CRITICAL]: Call home for CONTROLLER TAKEOVER COMPLETE AUTOMATIC
Jan 22 14:53:28 [msp-cluster-04:callhome.reboot.takeover:error]: Call home for PARTNER REBOOT (CONTROLLER TAKEOVER)

But the system doesn’t think it was taken over, or that it’s in HA mode.
cluster::> storage failover show-giveback
               Partner
Node           Aggregate         Giveback Status
-------------- ----------------- ---------------------------------------------
Warning: Unable to list entries on node cluster-01. RPC: Port mapper
         failure - RPC: Timed out

cluster::cluster ha> modify  -configured true
Warning: High Availability (HA) configuration for cluster services requires
         that both SFO storage failover and SFO auto-giveback be enabled. These
         actions will be performed if necessary.
Do you want to continue? {y|n}: y
Error: command failed: Could not enable auto-sendhome on partner node: Failed
       to set option cf.giveback.auto.enable. Reason: 169.254.97.26 is not
       healthy.

After another reboot I turned off and on HA, then everything cleared up and TO/GB’s were working perfectly.
cluster::> cluster ha modify -configured true
Warning: High Availability (HA) configuration for cluster services requires
         that both SFO storage failover and SFO auto-giveback be enabled. These
         actions will be performed if necessary.
Do you want to continue? {y|n}: y
Notice: HA is configured in management.

So on to the next problem: one of the vol0’s isn’t being recognized (click for larger image).




After a lot of searching, I found this magical solution. 

Curiously,  the vol0 is referred to in the excerpt above as a “7-Mode volume.”   But both vol0’s are, and there’s no way to change it.  The word from other engineers is that this is correct.
cluster::> vol show -is-cluster-volume -is-cluster-volume false
  (volume show)
Vserver   Volume       Aggregate    State      Type       Size  Available Used%
--------- ------------ ------------ ---------- ---- ---------- ---------- -----
cluster-01           vol0         aggr0_msp_cluster_01                                     online     RW        330GB    310.8GB    5%
cluster-02           vol0         aggr0_msp_cluster_02                                     online     RW        330GB    310.8GB    5%
2 entries were displayed.

Lastly, I needed to move one of the vol0’s.  I used this link to move the vol0 over to a new aggregate:
https://kb.netapp.com/support/index?page=content&id=1013762&actp=search&viewlocale=en_US&searchid=1390428133178

Thanks for making it all the way to the end with me!

Tuesday, July 23, 2013

Clustered ONTAP Networking

Basic
  • ifgrps are made up of physical ports.  ifgrps are to be named "a," e.g. a0a, a0b.
    • All members of a ifgrp must have the same role set (data, mgmt)
    • Cluster ports cannot be in ifgrps
    • A port with its own failover designation cannot be added to a ifgrp
    • All ports in an ifgrpmust be on the same physical node
  • A lif (logical interface group) is basically the IP addresses assigned to an ifgrp.  That IP address entity is given other properties like "home port," VLAN, failover group, etc.  Each IP needs to be able to have these properties independent of other IP's.
  • Spanning tree should not be enabled on the switch for any ports connecting to the controllers except the management ports.
  • It is a best practice to set flowcontrol to none for all ports except UTA ports.
    • UTA ports set to full.
  • You can of course add VLAN tagging and multiple IP's per LIF, named "-," e.g. a0b-441
  • Each LIF gets a routing group that you may or may not want to alter.  
    • You'll need to create a default routing group (default per vserver) for LIF's to default to.
Once you have those concepts figured out, you can move on to the advanced.
  • For snapmirror/snapvault, each node needs its own "intercluster LIF" to identify which ports and IP to be used for replication.
  • You'll need inter-cluster routing groups on each node to set up replication.
  • You have to create a cluster peer to establish a relationship for replication.
  • Network Interface Failover Groups are simply the cluster-wide list of ifgrps whose LIFs can migrate over to each other.  For example, you don't want your 10Gb NFS LIF failing over to another controller's 1Gb port ifgrp.
    • Your node management LIF can't fail over to other nodes (obviously, it's the portal to your physical node).  So it should have its own failover group that contains only local ports.
    • You want to be cognizant of not including e0M on any failover groups with data ports.

Some examples of real life setups:
  • A vserver (virtual storage machine) that was strictly SAN would need only one IP address, assigned to a physical port for management.