Showing posts with label best practice. Show all posts
Showing posts with label best practice. Show all posts

December 23, 2008

Day 23 - Change Management

This post was contributed by Matt Simmons. Thanks, Matt! :)

It's been said that change is the only thing that ever stays the same, and whoever said that probably worked in IT. Transitions are a part of life, but we administrators are burdened by what I would judge to be more than our fair share.

Too frequently, we find ourselves picking up the pieces from the last major system change we made, while at the same time designing the next iteration of the infrastructure that we'll be putting in place. How many times have you chosen an implementation that wasn't ideal now, because a bigger change was just around the corner, and you wanted to "future proof" your design? Bonus points for having to make that decision due to a previous change that was still being implemented. It doesn't seem to matter how precisely you've planned a major upgrade, snags and snafus are expected to rear their ugly heads.

Is this something that we just have to deal with? Are we at the mercy of Murphy, or are there ways we can induce these issues to work to our benefit? Sure, it would be easy if we had a crystal ball, but too often we don't even have a rough guess as to where our plans will encounter problems.

Change itself isn't the enemy. Change promotes progress, and from the 10,000ft view, our long-term goals should work towards this progress. Dealing with change is a natural and positive endeavor.

Instead of being thrown about by the winds of chance, lets put some sails on our boat, and see if we can make headway by trying to manage the change on our terms. If we know that problems are going to be encountered, and we face those facts before we edit the first configuration, then we've taken the first step towards real change management.

The enemies of successful change (and the resulting progress) are imprecise requirements and lack of project leadership. Unless you plan around these pitfalls, your project may very well go into ventricular fibrillation, flip-flopping back and forth, unable to decide between two unforeseen evils midway through the work flow. While it's possible to recover from this with an injection of leadership, it's much easier to inoculate against the problem in the beginning.

If you're going to be planning a big project, you will probably want to follow a methodology. There are just about as many methods of managing a change as there are people who want you to pay them to do it, but with IT projects, I've found what I consider to be the most efficient for me. Your mileage may vary, of course.

  1. Team and goal formation

    Assuming your change is moderate to large scale, you've (hopefully) got a team of people involved, and one of them has been appointed leader. This is the point where you want to decide on your goals. Determine what success will be defined as at the end of the project, and how best to get there.

    Many times we don't yet know what or how success will be defined, or even what the target should be. Because of this, it's natural to perform step 2 before your goals have been decided upon. In fact, I'd recommend it.

  2. Analysis (Research) & Information Organization

    Too often (or not often enough, depending on your view point) we're asked to do too much with too little. Frequently, we don't even know how to do it. This Analysis step is here to allow you to make informed decisions, and to acquire the skills and resources necessary to succeed in your task. Sometimes the resources are people, in the form of new employees or contractors, or both.

  3. Design

    By this time, you know what the task entails, but you don't have a road map of to how you're going to get there. This step makes you the cartographer, planning the route from where you are to the implementation of your project and beyond. Some details of the design may change during development, but it's important to have the major framework laid out in this step as you proceed.

  4. Development

    In a perfect world, you would take the design produced in step three and translate it straight into something usable. We all know that this rarely, if ever happens. Instead, you encounter the first set of really difficult problems in this stage. Issues spring up with the technology that you're using, or with kinks in the design that you thought were smoothed over, but weren't. Development appears to follow Hofstadter's Law: 'It always takes longer than you expect, even when you take into account Hofstadter's Law'. Thorough testing at the end of the development stage will prevent misery in the next step.

  5. Implementation

    Here we find the second repository of unforeseen bugs and strange glitches that counteract your carefully planned designs. The good thing about issues at this point is that, provided you've tested thoroughly enough in development, you won't find many show stoppers. On the other hand, sometimes these bugs can appear as niggling details and intermittent issues, hard to reproduce.

  6. Support

    If you're designing, developing, and implementing a product, support is just another part of the game. This is where you pay for how carefully you performed the preceding steps. Garbage In, Garbage Out, they say, but because you've designed and built a solid system, your support tasks will be light, possibly just educating the users and performing routine maintenance.

  7. Evaluation

    Remember that part in step 1, where you decided what success would be defined as? Dust it off and evaluate your project according to those requirements. Discuss with your team what you could have improved on, and don't forget to give credit where it is due. Hard work deserves appreciation.

This method is really a modified ADDIE design, so named because it consists of Analysis, Design, Development, Implementation, and Evaluation. We've added a couple of steps to help it flow better in the IT world we live in. There are certainly other methods to look at. The Instructional Systems Design (ISD) is another one which is well known.

However you decide to manage change, it's important to stay with your plan and follow through. Remember to work and communicate with your teammates, and don't stress because the project is too big. Just take it one step at a time, follow your plan, and you'll get the job done. s

December 16, 2008

Day 17 - Time Management

This post was contributed by Ben Rockwood. Thanks, Ben!

During the holiday season there seems to be a mad rush between Thanksgiving and Christmas. An odd force compels us to hurry things up and get organized or finish projects before the long Christmas break. The great and glorious pay off is a new January, fresh with possibility and reward. Out with the old, in with the new.

The new year offers something special and unique... perspective. Maybe you don't finish your projects prior to Christmas but things somehow feel different in the new year. New plans, new schedules, and a fresh perspective. So how can we get that perspective on a more regular basis?

Principles of Time Management

  1. Write everything down:

    Really, everything. Work or home, write it all down. If the thought cross your mind, it should be recorded. The thought may be "get mail", or "see new Bond film", or "implement new backup solution". If it's not written down you will think about it again, and again, and again. So get it out of your head.

    I encourage you to set aside about 30 minutes to an hour for this purpose. Set aside some time, go to some place relaxing, get a cup of coffee and just let your mind flow. Things you always wanted to do as a kid. A book, comic, or movie you never really understood or only caught part of. A hobby you've always wanted to try. A skill you think would be fun and useful. A new language you've wanted to learn (programming or spoken).

    In this alone you'll feel a great sense of relief and comfort.

  2. Keep it all in a central place:

    I'll offer some systems below, but which system you use isn't as important as simply consistantly using it. For SysAdmin's this may actually be a small collection of places, such as your company ticket system and a personal planner. If you simply use scraps of paper or a legal pad you'll innevitably loose it, so opt for a specific peice of software or organizer or note book. When your brain knows that all your projects and tasks are in a place it can easily reference it will leave your concious mind alone.

  3. Keep multiple lists

    You should not just have a single simple TODO list... life just isn't that linear. Rather, you should have a big "braindump" list of things to do, and then break that out into daily, weekly, monthly lists, or whatever granularlity you need.

    The sad fact is that in our minds, we tend to have a hodge-podge of TODO tasks, large and small alike. We constantly are steam-rolling this in our concious mind, and eventually it becomes overwhelming. When you lay everything out, and then look at that list and say to yourself "What can I accomplish this month?", then create that list, you start making life managable.

    This is key. When you have everything written down, you can create smaller, more approachable sublists to execute on it and have a greater sense of confidence that you might actually do it.

  4. Create daily TODO lists:

    Each day you should have a TODO list. This is the end-result of your other lists, which can actually be directly executed on. Get mail, walk dog, buy milk, install backup agents on systems 1-14, upgrade customerX's MySQL instance, close at least 4 tickets. These tasks need to be very granular... do-able.

    When you have daily lists you have two big benefits. Firstly, you have a historical record of what you were doing day-by-day. What were you trying to get done on March 23rd? Now you can look back and find out. Secondly, you can "push" tasks from one day to the next. Don't have time to finish something today? Push it onto tomorows list now and move on with other tasks you can complete.

    This is something paper planners do very well, but is difficult to accomplish in software task managers.

  5. 'Clean the Garage' is a Project:

    The key to using software management tools, such as 'Things' or 'OmniFocus', is to think in projects. A "project" is defined as any goal or objective that is not accomplished in a single task. As an example, replacing a lightbulb could be either one. If you have lightbulbs and you just need to swap it, its a task. However if your out of light bulbs and need to go to the store first, it is now a project consisting of the tasks "Buy Lightbulbs" and "Replace bulb in hallway".

    I make special emphasis of this because if you don't think like this the software tools will be very difficult to manage. You'll just have growing piles of TODO's that seem unrelated and you'll spend more time digging through lists of tasks than accomplishing them.

  6. Think out the steps:

    It's very important to not just think about the end result you want, such as "Backup Oracle Database", but rather to think about its individual tasks and then lay them out. Even tasks that may be fairly simplistic can seem overhwelmingly complex when you mind floods your concious with all the possible permutations and unclear decisions you will need to make. If you just break it down you can stay more relaxed, focused, and objective. You may even need to create sub-projects to evaluate your options, such as "Benchmark RMAN", "Evaluate BakBone Oracle Agent", etc.

Management Systems

  1. Franklin Covey Paper Planners

    Franklin Covey is a leading name in time management tools. If you haven't heard of them, think of the old "DayRunners" or other paper planners you've seen, but more flexable and customized.

    When you get started with Franklin Covey you will need to select and purchase a binder, starter kit, and verious types of filler pages. If you visit a store an employee will walk you through it, and if you visit their website there is a guide. Planners come in all sizes and styles to meet your needs directly. I personally prefer the smallest size, so that I can put it in a pocket.

    The advantages to a paper planner is that you have a perminant record of each days work and schedule to archive, its easy and fast to use, and you can take it everywhere you go. On the downside, its more expensive than software. Expect a nice setup to run you about $80.

    If you are on a budget, you can emulate the same thing in any cheap binder. Many people use Moleskin notebooks for this purpose.

  2. OmniFocus

    OmniFocus for the Mac is the most popular time management software around these days. Its very powerful but can take some time to learn. Thankfully there are some excellent tutorial videos and the online help is useful.

    The software lets you easily gather new tasks in an inbox and then later sort them into categories and projects. One of the most useful concepts in OmniFocus is that of "Context". Any particular place in which tasks can be done is a "context". For instance, you can only do laundry when your at home, and you can only check the tape robots in the office. When you assign context to tasks you can later look only at tasks that you can actually accomplish now. You might need to "Wash the dog" today, but you don't need to look at that task when your in the office working.

    The key to OmniFocus is to use it as it was intended. If you don't take time to actually learn how the software is intended to be used you'll find it nifty for about a week and then start getting irritated or frustrated.

    One added bonus of OmniFocus is that if you have an iPhone you can buy OmniFocus for iPhone and sync it with your desktop, bringing the "everywhere you want to be" advantage of paper to OmniFocus.

  3. Things

    Thinks is also for the Mac and currently free. It adopted many of the concepts of OmniFocus but is not nearly as strict and rigid. Rather than structure it relies on tagging tasks, which can allow you to organize more freely.

    I highly recommend that anyone considering the software route start with OmniFocus's free trial. Learn OmniFocus and then move to tools like Things. If you don't and just try Things you'll immediately find it too free flow and unhelpful.

    There is also an iPhone version of Things, however at last check it did not support sync'ing with the desktop.

Further reading:

December 15, 2008

Day 15 - Documentation

Documentation is like automation. Good documentation will save you and your cowokers time, effort, and mistakes. Bad documetation will frustrate, anger, and annoy. No documentation means you get more interruptions, and you spend time not making progress on other tasks.

There are many things to document: designs, changes, APIs, policies and procedures. Each document should focus on a specific audience. Sending a change notification to your own group could be terse: "I rebooted system3 because <reason>." Documenting an important procedure that your customers will be invoking should be well written and probably not terse.

Good documentation takes effort and thought. Documentation should be written for a known audience. Choose your audience before writing a document. Express your intended audience before you begin your document. If you're revising a document, make sure the intended audience will still benefit after your revisions.

Documentation is about content, audience and findability. If the content is wrong or out of date, then your documentation is hurting the situation. If you don't know your audience, you can't best help the people who most need your information. If your document can't be found, no one will know it exists.

The content of your document is very important. The assumptions of prior knowledge in your content must be molded around your chosen audience. For instance, don't use an acronym without defining it unless you are certain your audience will know what it is or how to find out what it is. Content isn't necessarily always text. Good documents include diagrams or links to other resources, where possible.

The medium (wiki, email, paper, etc) of your documentation should reflect the needs of that documentation. Don't document long-term information over email. A printed new-hire todo list is good to have on said new-hire's desk on the first day. A network design should be published (perhaps on your wiki) with details including the problem, the solution, and benefits.

If the medium is electronic, you need to consider findability. Findability means that any person needing a piece of information can find it. Findability is difficult to provide without good search facilities and/or a easily browsable structure. Books have indeces to enhance findability, and your documentation should, too. These days, a wiki is a reasonable choice for containing your documents as they help you provide search and structure which improves findability. If you make roll out a major change, update your documentation and send an email to your audience (interested parties) indicating the change and where to learn about it.

You probably have a whole bucket of things that merit documentation. Prioritize these by what will gain you the most first. If you get multiple interrupts a week from customers making a frequently asked question, documenting it and respond with "Look on the wiki for <foo>" is a good way keep an interrupt short. Helping customers help themselves helps you. Of course, by "customer" I mean the people you are, as a sysadmin, supporting. Even if they're other employees, they're still customers.

Documentation, like code, needs to be tested. Testing documentation means having someone from your intended audience read it and report their level of understanding. If they didn't understand the information you were expressing, then you need to revise and re-test. If a document is intended for your customers, don't have a fellow team member review it.

Further, documentation being like code, suffers bit rot if neglected. Unmaintained documentation means people reading it will be misinformed. Don't ignore your existing documentation. It's worth more to you updated than neglected.

Knowing that documentation is important means that you should prioritize the act of improving and creating documentation among your other duties and tasks. Such things take time and effort, so be sure to consider documentation when budgeting your own time on work.

Lastly, it's worth pointing out that some documentation can lead to automation. For example, a well-documented alerts (or failure scenario) playbook often looks like a flowchart detailing operations to perform to debug and fix a problem. This kind of detail often lends itself to being transformed into a script. Once you have the script, you could have your alert system run the new script instead of paging you, or even just use the script to automate collection of some diagnostic information to help you more quickly debug a problem.

December 11, 2008

Day 11 - Home away from home

Logging in to a machine that isn't your own workstation can be scary. You are subject to the decisions made in someone else's configuration files that don't always align with your own configurations: different shell, different default shell configuration, different default editor, different editor configuration, etc.

This is a scary and unproductive place. Suddenly 'ls' output is colored, or vim uses a different indenting configuration, or worse, the default editor is not your favorite editor (which doesn't have to be vim). Dedication to mastering your basic tools over time has helped you create the One True Configuration for each tool; deviation from this configuration means a loss of productivity. You need to bring your home (directory) with you.

When your home directory doesn't magically appear through the miracles of network filesystems, you may need to fix the problem another way. One potential solution is to make sure that you copy all your configuration files (.vimrc, .zshrc, whatever) to every host you're going to login to. This doesn't scale. Further, it means everyone else has to repeat the same process for their own files.

The fix is to create a system which automatically keeps your home directory, on every machine, populated with your configurations. You can do this minimally with revision control and a cron job, but I prefer to add rsync to this process.

Step one is make a place in your revision control system for people to create home directories. For example, declaring that the path /trunk/home in your repository is where you should dump your homedirectory contents. This means if my username were 'jordan' then I'd check my '.vimrc' in as '/trunk/home/jordan/.vimrc' and should expect it to show up on any system I have access to.

Step two is to pick a server that has access to both revision control and other servers. Set up a cron job here that will check out and keep-updated your entire /trunk/home path. Run an rsync daemon here that exports this /trunk/home for other servers to update with. Set the rsync module name to 'homedirs' for readability.

Step three is to deploy a cron job on every necessary server that copies down all the obvious files from someserver::homedirs with rsync. You do have automation that lets you install a cron job on all of your servers, right? ;)

Before you go and write the one line of rsync invocation that it would take to copy someserver::homedirs to /home, you should take care to note the potential security implications of doing this as root. If I have a file checked in called /home/jls/test/shadow, and on one of the servers I sneakily symlink /home/jls/test/ to /etc, and you run the rsync blindly as root, you just let me overwrite your /etc/shadow file (or something else evil). Malicious or accidental, doing a single rsync may not be the best solution.

The fix is to run the rsync as each user. You can get the list of users to copy down by running 'rsync someserver::homedirs' to get the list of directories, which should include your usernames. Check out the completed version of the sync home directories script.

You should now be able to modify your home directory files in revision control and have them automatically propogate without your assistance.

Further reading:

The 'run rsync as the user' security idea from Pete Fritchman.

December 10, 2008

Day 10 - Config Generation

A few days ago we covered using a yaml file to label machines based on desired configuration. Sometimes part of this desired configuration includes using a config file that needs modification based on attributes of the machine it is running on: labels, hostname, ip, etc.

Using the same idea presented in Day 7, what can we do about generating configuration files? Your 'mysql-slave' label could cause your my.cnf (mysql's config file) to include settings that enable slaving off of a master mysql server. You could also use this machine:labels mapping to automatically generate monitoring configurations for whatever tool you use; nagios, cacti, etc.

The older ways of doing config generation included using tools like sed, m4, and others, to modify a base configuration file inline or writing a script that had lots of print statements to generate your config. These are both bad with respect to present-day technology: templating systems. Most (all?) major language have templating systems: ruby, python, perl, C, etc. I'll limit today's coverage, for the sake of providing an example, to ruby and ERB.

ERB is a ruby templating tool that supports conditionals, in-line code, in-line variable expansion, and other things you'll find in other systems. It gets bonus points because it comes standard with ruby installations. That one bonus means that most people (using ruby) will use ERB as their templating tool (Ruby on Rails does, for example), and this manifests itself in the form of good documentation and examples.

Let's generate a sample nagios config using ruby, ERB and yaml. Before that, we'll need another yaml file to describe what checks are run for each label. After all, the 'frontend' label might include checks for process status, page fetch tests, etc, and we don't want a single 'check-frontend' check since mashing all checks into a single script can mask problems.

You can view the hostlabels.yaml and lablechecks.yaml to get an idea of the simple formatting. Using this data we can see that 'host2.prod.yourdomain' has the 'frontend' label and should be monitored using the 'check-apache' and 'check-frontend-page-test' checks.

The ruby code and ERB are about 70 lines total, perhaps too much to write here, so here are the files:

Running 'configgen.rb' with all the files above in the same directory produces this output. Here's a small piece of it:
define hostgroup {
  hostgroup_name frontend
  members host2.prod.yourdomain
}

define service {
  hostgroup_name frontend
  service_description frontend.check-http-getroot
  check_command check-http-getroot
}

define service {
  hostgroup_name frontend
  service_description frontend.check-https-certificate-age
  check_command check-https-certificate-age
}

define service {
  hostgroup_name frontend
  service_description frontend.check-https-getroot
  check_command check-https-getroot
}
I'm not totally certain this generates valid nagios configurations, but I did my best to make it close.

If you add a new 'frontend' server to hostlabels.yaml, you can regenerate the nagios config trivially and see that the 'frontend' hostgroup now contains a new host:

define hostgroup {
  hostgroup_name frontend
  members host3.prod.yourdomain, host2.prod.yourdomain
}
(There's also a new host {} block declaring the new host3.prod.yourdomain not shown in this post)

Automatically generating config files moves you into a whole new world of sysadmin zen. You can regenerate any configuration file if it is corrupt or lost. No domain knowledge is required to add a new host or label. Knowing the nagios (or other tools) config language is only required when modifying the config template, not the label or host definitions (a time/mistake saver). You could swap nagios out for another monitoring tool and still make sure the underlying concepts (frontend has http monitoring, etc) are consistent. Being able to automatically generate configs means that you probably have both the templates and the source data (our yaml files here) stored in revision control, which is a whole other best practice to focus on.

Further reading:

December 7, 2008

Day 7 - Host vs Service

An important distinction when talking about servers and services is to talk about them separately. Build automation in terms of configuration sets, not in terms of servers.

I tend to think of servers, machines, devices, whatever, as having labels or tags. Each label refers to a particular configuration set. Your automation tools should know what labels are on a host and only apply changes based on those labels. Modern administration tools such as Capistrano and Puppet are designed with this distinction in mind. Capistrano calls them 'roles' and puppet calls them 'classes,' but ultimately they're just some kind of name you apply to configuration or change.

Labels can be anything, but they should be meaningful. You might have "mysql-debug" and "mysql-production" service labels which both cause mysql to install but the debug version means you have heavier logging features enabled like full query logging, etc.

Configuring with labels instead of individual hosts helps you scale up. Managing configuration changes for a specific service lets you make one change to a service and have it deploy on any host having that service. Further, if you buy new server hardware, simply adding the appropriate labels to a host will let your automation system do the hard work of installation and configuration.

It helps you scale down, too. Here's a fictional example:

Quality control requested a production-like environment to test release candidates before pushing to production, but the budget will only allow you to use two server hosts for this. Production uses many more than this. If you automate based on labels instead of hosts, you could easily spread the required services across your two servers by simply labelling them, and automation would take care of the installation and configuration.

Assuming you have the development time or the tools available, you can use labels all over your automation:

  • Generate dns entries for all hosts with a specific label
  • Configure your monitoring system based on labels on a host
  • Configure firewall rules
  • Configure backup policy
  • etc...
A simple implementation of this would be a small yaml file with host:label mappings:
host1.prod.yourdomain:
- mysql-debug
host2.prod.yourdomain:
- memcache
- frontend
The deployment of these labels is up to you and the needs of your automation system. Keeping this in revision control gives you history with logs. Along with the other automation code and configuration you should be keeping in revision control, you might just be one step closer to being able to do more while working less.
With puppet
If you're using puppet, telling each host what it's labels (aka, puppet classes) are is easy, you need only write a script to help puppet know what classes to apply to a host (or node, in puppet's case). This document will show you how in puppet.
With capistrano
You'll want some piece of code that turns your yaml file of host:label entries into 'role <label>, <host1, host2, ...>. Something like this may do (ruby): (I called our yaml file 'hostlabels.yaml')
# roles.rb
require "yaml"
labelmap = Hash.new { |h,k| h[k] = [] } # default hash value is empty array
hosts = YAML::load(File.new("hostlabels.yaml"))
hosts.each { |host,labels|
  labels.each { |label| labelmap[label] << host }
}
labelmap.each { |label,hosts|
  role label, *hosts
}
And in your Capfile:
load "roles"   # use 'load' not 'require'

task :uptime, :roles => "frontend" do
  run "uptime"
end
And now 'cap uptime' will only hit servers listed in your yaml file as having the label 'frontend'. Cool.
I wanted to provide an example with cfengine, too, but I'm not familiar enough with the tool and my time ran out learning how to do it.

The yaml file example is not totally ideal, but it's a start if you have nothing. Evolutions beyond the simple host:services are the state configuration management tools where you store information about what is truth - such as for every machine that exists, mac addresses, IPs, service labels, hardware type, etc. It might include the class of "enterprise inventory management" suites by Oracle and others, too.