Showing posts with label storage. Show all posts
Showing posts with label storage. Show all posts

December 6, 2009

Day 6 - Storage: ATA over Ethernet

This article was written by Bob Feldbauer.

Continuing from yesterday's storage discussion on DRBD, let's introduce another option in the storage space.

When your environment has grown beyond direct attached storage (internal drives, or an external drive array) and basic network attached storage (NAS), the next step is generally to consider implementing a storage area network (SAN); however, the cost and complexity of existing Fibre Channel and iSCSI SAN solutions can be daunting. Fortunately, ATA over Ethernet (AoE) can often be used as a simpler, lower-cost alternative.

Before diving into AoE, you should be aware of its limitations. The major idiosyncrasy of AoE is that it does not use TCP/IP; therefore, it is not routable. Also, compared to iSCSI, it lacks encryption and user-level access control. AoE truly shines when you simply need storage over a local network, and can limit access by control over physical ports on a switch, VLANs, etc.

AoE is supported on a wide variety of operating systems, including Linux, Windows, FreebSD, VMWare ESX, Solaris, Mac OS X, and OpenBSD. Like Fibre Channel, iSCSI, and other storage protocols, the AoE protocol implements an initiator-target architecture - the initiator sends commands, and the target receives them. On Linux, various pieces of software are used to provide initiator and target functionality. Four independently developed AoE targets exist for Linux: qaoed, ggaoed, kvblade, and vblade. Since vblade is part of aoetools, it generally seems to have the most active development community; therefore, we'll use that for our example setup. Depending on the version of AoE you use, it can be used with Linux kernels 2.4.x or 2.6.x; however, it is best used with kernel 2.6.14 or higher and newer versions of aoetools.

For our example configuration, we'll be setting up AoE on Debian Linux (Lenny), and using LVM2. The AoE module should be included if you're using a standard Debian kernel, but let's check using:

grep ATA_OVER /boot/config-`uname -r`
Which should return:
CONFIG_ATA_OVER_ETH=m

That means AoE is supported through a kernel module. If you aren't already using AoE, the module is probably not loaded yet. Let's load it, and add it to /etc/modules so it is automatically loaded at when the system is started in the future:

modprobe aoe
echo "aoe" >> /etc/modules
On Debian, we'll use apt to install the remaining necessary components:
apt-get update
apt-get install lvm2 vblade vblade-persist aoetools

Note that aoetools includes the following tools:

aoecfg manipulate AoE configuration strings
aoe-discover trigger discovery of AoE devices
aoe-flush flush down devices out of the AoE driver
aoe-interfaces restrict network interfaces used for AoE
aoe-mkdevs create character and block device files
aoe-mkshelf create block device files for one shelf address
aoeping simple userland communication with AoE devices
aoe-revalidate revalidate the disk size of an AoE device
aoe-stat print status information for AoE devices
aoe-version print AoE-related software version information
coraid-update upload an update file to a Coraid appliance

For our example, we'll use LVM2 and AoE to allocate space and make it available over the network, using two disks. For a recent project, I used 13 drives on a hardware RAID controller and divided them into two RAID6 arrays of 6 (4TB) and 7 (5TB) drives.

Assuming the disks are the second and third disks on a Linux system, configure LVM2 to recognize the physical volumes:

pvcreate /dev/sdb
pvcreate /dev/sdc

Create two LVM2 volume groups on the physical volumes - for our example, let's call them content and backups:

vgcreate content /dev/sdb
vgcreate backups /dev/sdc

Then create two 1TB LVM2 logical partitions (/dev/content/server1, /dev/backups/server1):

lvcreate -L 1TB -n server1 content
lvcreate -L 1TB -n server1 backups

Although there are many different types of filesystems available under Linux, for familiarity and to avoid unnecessary complexity in our example, we'll use ext3:

mkfs.ext3 /dev/content/server1
mkfs.ext3 /dev/backups/server1

Now that we have our drives configured with LVM2 and formatted with a usable filesystem, we can setup the AoE target using vblade-persist:

vblade-persist setup 0 1 eth0 /dev/content/server1
vblade-persist setup 0 2 eth0 /dev/backups/server1
vblade-persist start 0 1
vblade-persist start 0 2

To mount our newly created AoE devices on a remote server, run the following on a second server:

modprobe aoe
apt-get install aoetools
aoe-discover
aoe-stat    # should show the available AoE exports
mkdir /mountpoint
# replace e0.1 with the appropriate device from aoe-stat
mount /dev/etherd/e0.1 /mountpoint 

Your options in the storage are many. This introduction should give you the necessary tools to decide if AoE is the storage solution that meets your requirements.

December 5, 2009

Day 5 - Storage Availability: DRBD

This article was written by Jeff Hengesbach.

Storage and storage technologies have really moved into the limelight in recent years. Previously, advanced storage and storage concepts were hidden away in large corporate datacenters; they were complicated, proprietary, and costly. Most small and mid-sized organizations used or still use systems with local or direct attached hard drives. In the event of drive failure, RAID was relied upon as the sole savior to keep things running uninterrupted. In the event of array or system loss, restore from backups was the primary or only option. The drives were expensive and storage needs modest (arguably due to storage costs).

Move forward a few years to the now. Storage needs are, relatively speaking, monumental, and space is much less costly. Virtualization has come to the forefront, and you are now behind the proverbial technology curve if not taking advantage of its benefits. System densities have increased (blades, virtualization, etc) and now, more than ever, storage is a critical component of the infrastructure. Most of these density-increasing technologies that improve space, power consumption, cooling, management, disaster recovery, and availability, require carefully designed, reliable, and centralized storage. Reliable storage starts with good drives and controllers with appropriate RAID levels for the application as well as connectivity to that storage. What happens when a controller fails, or a multiple drive failure brings down a RAID set, or connectivity is lost? The proprietary solutions of course have their options and some very nice options indeed, but for some organizations that need storage availability, many of the proprietary solutions do not scale to the available budgets.

This is where DRBD comes to the table. DRBD is graciously maintained, supported, and provided as open source by the people over at Linbit. As the website states: "DRBD can be understood as network based raid-1". Technically speaking, DRBD is a block device driver you layer into the device chain just like LVM or MD. It can be put anywhere in the stack: on top of the disk device, on top of LVM or MD, at the top just under the file-system, wherever it makes sense. Its specialty is taking block level changes, keeping track of them, and sending them over a network to another system where they are duplicated. DRBD has a few specific traits to highlight. First off, it is smart (and dumb?). It works at the block level and knows which bits may be out of sync and will only send those bits across the wire. It knows nothing of file-systems, files, etc, so the application using it can be anything: samba, iSCSI, NFS, EXT3, XFS, or others. Secondly it can be non-destructively added to existing data volumes. It is not necessary to go through backup/install/restore process - only install, configure and replicate.

Linbit also offers a closed source product, DRBD Proxy, that is designed for long haul, high latency(200ms) connections as well as greater than 2-node replication situations. If you want to replicate outside of a LAN using DRBD, then you'll need it. Recently, Linbit became the steward of the Heartbeat cluster messaging codebase. Heartbeat is the application that handles the monitoring and fail-over activities between the systems. DRBD and Heartbeat are a tried and true tag team for storage high availability / fail-over situations.

So what does a simple, highly available, solution look like? Two nodes with appropriate storage use DRBD in their block device stacks to replicated the data over a gigabit or better connection; one is primary (the source), the other secondary (the destination). Heartbeat runs on both systems monitoring for system and service availability. If Heartbeat detects a failure it coordinates the startup of services on the backup node including all DRBD and service (could be nfs, samba, iscsi, etc) related commands. It is a beautiful thing to see in action.

I'm not one to rehash out the commands needed to install, configure and run these solutions. The DRBD website has an excellent user manual that would be absolutely futile to reproduce here, and Heartbeat is relatively simplistic to get established and working. Both products have a multitude of options not necessary for basic operations but are very handy for more fine grained control and tuning.

What are the keys parts to a good solution? DRBD relies on TCP/IP networking to move data around. Make sure gigabit or better is being used along with quality cables and switching. If at all possible keep DRBD traffic on its own network - or even direct cross-over connections. Don't skimp on your backup node. A backup node that can not keep up with replication will cause IO blocking on the primary. Design connectivity to handle loss of a switch. If in doubt - ask Linbit or the community for assistance. DRBD, DRBD+Heartbeat have been in use for a while and they are sound technologies.

--

About the author: Jeff Hengesbach is a System Administrator with a history of working in smaller organizations. He is an avid proponent of virtulization and linux (and Microsoft, where appropriate) technologies. Between work and family he occasionally has the opportunity to post articles on his blog at: http://jeffhengesbach.blogspot.com.

Further reading: