← All posts
Apr 20, 2015

How to dockerize existing ruby on rails applications

If you’re like us, having an existing ruby on rails project and want to somehow run it inside a docker container so we can get those great benefits from the Docker ecosystem, this article may help.

Ben Cao · 10 min read · 0 comments

This article will focus on dockerizing existing rails app because it’s common and challenging.
But as long as a new app will eventually become old, you may find some tips here still be helpful.

If you’re like us, having an existing ruby on rails project and want to somehow run it inside a docker container so we can get those great benefits from the Docker ecosystem, this article may help.

TL;DR

There’re quite a few things we need to do to make our RoR app acts as a better container in the Docker world, and we can divide the task into several phases:

  • Setup a base image with ruby environment
  • Install ruby gems
  • Compile static assets
  • Define container interface for the app
  • Start web server

Setup a base image with ruby environment

Both Ruby and Rails have official images available.
But it’s also not too hard to build it from scratch.

Use Rails official docker images

We don’t recommend Rails official images for production use because

  • only rails 4+ has official images
  • we can’t control ruby version, which brings trouble in case we want to upgrade ruby for vulnerability issues
  • rails could simply be installed by bundle install

Use Ruby official docker images

Ruby official images are helpful, they offer a wild range of options from down to ruby 1.9.3 and up to the latest or stable version.
If you’re lucky, your target ruby is one of the official versions, then it could be the best choice to start from an official ruby image.

Build base Ruby image from scratch

But if you’re running a pretty old legacy RoR application, like us, we have legacy rails apps running on ree-1.8.7, so there are no better solutions except building a base ruby image by ourselves.

We’ve tried both RVM and ruby-build, and we finally chose ruby-build for its simplicity.

Note ree-1.8.7 need some patches to make it work correctly, create a directory called patch, and put ree-1.8.7–2011.12 inside of it.

Here’s the Dockerfile demonstrating how we build it:

FROM centos:7

# install those basic tools we will use in debugging
RUN yum install -y \
  git \
  tar vim \
  gcc gcc-c++ make patch \
  hostname nmap-ncat readline-devel; \
  yum -y clean all

# install ruby build
RUN git clone https://github.com/sstephenson/ruby-build.git /root/ruby-build && /root/ruby-build/install.sh

# install ree, gcc44 is a must-have
RUN yum install -y compat-gcc-44; yum -y clean all \
    && CC=gcc44 ruby-build ree-1.8.7-2011.12 /ruby \
    && echo "export PATH=$PATH:/ruby/bin" > /etc/profile.d/ruby.sh

ENV PATH $PATH:/ruby/bin

# install bundler
RUN gem install -N bundler

Then build an image with

docker build -t my-company/ruby-base .

Install ruby gems

Install System Library Dependencies

Before installing gems, we need some of those system libraries so that some gems can compile to native extensions.

For example, we need to install mysql-devel package manually in Centos before we install mysql2 gem.

Bundle Install

Simply run a bundle install in Dockerfile can be a solution. But the performance can be a problem for an app that has a lot of gem dependencies. For example given this Dockerfile:

FROM my-company/ruby-base:latest

COPY . /my-app
WORKDIR /my-app

RUN bundle install --jobs 3 --retry 3
RUN bundle clean --force

Since we copy project files first, the COPY command will invalidate docker build cache, bundle install becomes an expensive operation, in our case, without any optimization, this first RUN step took more than 10 minutes to finish.

How could we make it faster?

We could also split the total time spending into 2 buckets:

  • Time spending on downloading gems from remote gem servers
  • Time spending on compiling gems with native extensions

Let’s deal with them one by one.

Reduce time for downloading gems

We’ve tried many means and we found the following solution to be both helpful and elegant.

# this is a command line context
# like a Jenkins Job or our local terminal
# current working directory is the same directory 
# where we placed the above Dockerfile
bundle package --all       
# the above command utilizes the cache from previous runs on jenkins workers!
docker build -t my-company/incredible-app

The key is bundle package --all, with this command, gem sources will be cached in the vendor/cache directory, if this folder is copied into an image by the COPY command, when running bundle install later in the docker build process, we will spend no time downloading it from the remote server.

bundle package --all could be slow for the first time, but afterward, it takes almost no time.

The beauty of the solution is that we have the same solution for both local and CI environment, and there are no special tricks needed for the Dockerfile.

Without this cache, docker build will work. With this cache, docker build just runs faster. Pretty much like the progressive enhancement idea in the front-end world.

Reduce time for compiling gems

It’s easy to see which extension takes the most time to compile from CI logs that has a timestamp associated with each log line, let’s mark those gems that ate up the time heavily.

Say we have an app, gems are mostly stable for typical cases, we could introduce a new image layer which installs some gems in advance, so the later bundle install can skip installation for those already installed gems.

Here’s an example of a Dockerfile building the intermediate image:

FROM my-company/ruby-base:latest

# the only purpose for this image is to speed up normal build

# saves 3 min
RUN gem install -N nokogiri -v 1.6.2.1

# saves 1min10s
RUN gem install -N curb -v 0.8.5

# saves around 30s
RUN gem install -N poltergeist -v 1.5.1

# saves more than 30s
RUN gem install -N unicorn -v 4.8.3

# saves 16s
RUN gem install -N therubyracer -v 0.12.1

# saves 10s
RUN gem install -N oj -v 2.12.0

# saves 10s
RUN gem install -N mysql2 -v 0.3.18

# saves 10s
RUN gem install -N ruby-prof -v 0.15.2

After building a new image with docker build -t my-company/ruby-base-with-cache ., we can slightly modify our app’s Dockerfile’s FROM statement to be

# point to an extended base image with additional gem cache
# thus will run faster for bundle install
FROM my-company/ruby-base-with-cache:latest

Compile static assets

The step is as simple as RUN rake assets:precompile.

Define container interface for the app

Our app is now packaged as an image and the same image can be run in different environments, no matter it is development, staging or production.

Different environment has significantly different configurations, like the host and port for the database, and other depended upstream services.

In order to support all those versatile use cases, our app (as an image) should be designed as highly configurable.

Generate Config Files

A typical rails app have at least one config file for database.

Mount the config file as a docker volume is an option, but if you agree with Twelve Factor you may prefer to use environment variable as config option if possible.

We’ll focus on how to generate config files given ENV variables.

Here’s the thinking process:

  1. The config file needs to be generated at runtime, which means we should utilize either CMD or ENTRYPOINT instruction in Dockerfile.
  2. CMD is likely to be overridden by offering an additional command to docker run command, while ENTRYPOINT is less likely to be overridden.
  3. So the question becomes whether we want to ensure the generated config file would always be there even if end user run docker run with different commands instead of the default command that starts the web server?

In our case, we want to generate config files no matter what, such as the rspec spec command is provided to run tests within a container.

# this is how we expect to run the container in WEB SERVER mode
docker run my-company/my-rails-app:latest

# this is how we expect to run the container in TEST mode
docker run my-company/my-rails-app:latest rspec spec

So we chose to utilize ENTRYPOINT to generate config files, with that, the Dockerfile of the rails app became something like this.

FROM my-company/ruby-base:latest

COPY . /my-app
WORKDIR /my-app

RUN bundle install --jobs 3 --retry 3
RUN bundle clean --force

ENTRYPOINT ["/my-app/docker-entrypoint.sh"]
CMD ["rails", "server"]

And for docker-entrypoint.sh, we can start with

#!/bin/bash

# stop execution if any commands fail
set -e

# generate database.yml
source /my-app/docker-initializers/generate_database_yml.sh > /my-app/config/database.yml

# run command from either CMD instruction or docker run
exec "$@"

docker-initializers/generate_database_yml.sh

#!/bin/bash

cat << EOF
defaults: &defaults
  adapter: mysql2
  reconnect: true
  encoding: utf8
  host: $MYSQL_MY_APP_HOST
  port: $MYSQL_MY_APP_PORT
  username: $MYSQL_MY_APP_USERNAME
  password: $MYSQL_MY_APP_PASSWORD

test:
  <<: *defaults

development:
  <<: *defaults

production:
  <<: *defaults

EOF

You can notice from the above samples we rely on some MYSQL_MY_APP_* environment variables, and given different values of those environment variables will generate different config file in runtime.

Wait for Depended Services

Now our database.yml will be generated automatically when a container starts.

But another problem arises, if MySQL is still in initializing while the rails app tries to connect to it, the rails app will throw an error complaining about database connection and crash the container.

The reason is that rails needs to load the metadata information of DB tables to allow ActiveRecord to work correctly.

How to deal with this problem? We can address this issue from the external orchestration layer, but add some basic protections to make our rails app a bit more robust is still harmless.

Here’s how we do it:

We’ve employed nmap-ncat package in ruby-base image, now it’s time to take it to work.

docker-initializers/wait_support.sh

#!/bin/bash

function wait_for() {
  service=$1
  host=$2
  port=$3

  echo "waiting for $service to be up on $host:$port..."

  if [ -n "$host" -a -n "$port" ]; then
    # nc command is the key for the TCP probe
    while ! nc -w 1 -c echo $host $port
    do
      echo -n .
      sleep 1
    done

    echo 'ok'
  else
    echo "[ERROR] invalid host=$host or port=$port for $service"
    exit 1
  fi
}

wait_for "database connection - $MYSQL_MY_APP_DBNAME" $MYSQL_MY_APP_HOST $MYSQL_MY_APP_PORT

Adding wait support, docker-entrypoint.sh now becomes:

#!/bin/bash

# stop execution if any commands fail
set -e

# wait for other service ports to be ready, this can be enabled by a environment variable
if [ "$WAIT_FOR_DEPENDED_SERVICES" = "true" ]; then
  source /my-app/docker-initializers/wait_support.sh
fi

# generate database.yml
source /my-app/docker-initializers/generate_database_yml.sh > /my-app/config/database.yml

# run command from either CMD instruction or docker run
exec "$@"

Define Environment Variable Default Values

In the above section, we added a switch for wait support.

Only if user set WAIT_FOR_DEPENDED_SERVICES to be true we will enjoy the benefit of wait support.

It will be a bit verbose to supply such a ENV var everytime we run the container, so it’s common to consider setting a default value for an environment variable if it’s not defined or simply to exit if some critical environment variables are missing.

With this we introduce another shell script for defining environment vars:

docker-initializers/check_env_vars.sh

#!/bin/bash

MISSING_REQUIRED_ENV_VARIABLES=false

# the pattern to throw if missing
function required()
{
  if [ -z "${!1}" ]; then
    echo "please set required env variable $1" >&2

    MISSING_REQUIRED_ENV_VARIABLES=true
  fi
}

# the pattern to give default values
function optional()
{
  if [ -z "${!1}" ]; then
    export ${1}="$2"
  fi
}

required MYSQL_MY_APP_HOST
optional MYSQL_MY_APP_PORT 3306
required MYSQL_MY_APP_USERNAME
required MYSQL_MY_APP_PASSWORD
optional MYSQL_MY_APP_DBNAME my_app_database

optional WAIT_FOR_DEPENDED_SERVICES true

if [ "$MISSING_REQUIRED_ENV_VARIABLES" == "true" ]; then
  exit 1
fi

Including some simple cleanups, now docker-entrypoint.sh eventually looks like this:

#!/bin/bash

# stop execution if any commands fail
set -e

# define env vars, this becomes the container interface definition file
source /my-app/docker-initializers/check_env_vars.sh

# wait for other service ports to be ready
if [ "$WAIT_FOR_DEPENDED_SERVICES" = "true" ]; then
  source /my-app/docker-initializers/wait_support.sh
fi

# generate database.yml
source /my-app/docker-initializers/generate_database_yml.sh > /my-app/config/database.yml

# prepare log and tmp directories
mkdir -p /my-app/log
mkdir -p /my-app/tmp
rm -rf /my-app/tmp/*

# run command from either CMD instruction or docker run
exec "$@"

Start web server

Choose a web server

Depending on different team’s situation, maybe you have some experts about puma, maybe you prefer event machine based implementations like thin.
It’s total freedom to choose whatever we’re comfortable to power your rails app.

From our case, we were using unicorn as our web server so we kept using it in the dockerized app. Then this is how our Dockerfile eventually looks alike:

FROM my-company/ruby-base-with-cache:latest

COPY . /my-app
WORKDIR /my-app

RUN bundle install --jobs 3 --retry 3
RUN bundle clean --force

EXPOSE 3000

ENTRYPOINT ["/my-app/docker-entrypoint.sh"]
CMD ["unicorn_rails", "-l", "3000", "-c", "/my-app/unicorn.conf.rb"]

Conclusion

Some of you may already figure out we actually have created a new problem, now it would be quite a bit typing for us to start the application:

docker run \
  -e MYSQL_MY_APP_HOST=1.2.3.4 \
  -e MYSQL_MY_APP_USERNAME=testuser \
  -e MYSQL_MY_APP_PASSWORD=password \
  my-company/my-rails-app:latest

This is a great problem since it opens the door for container orchestration, which is another fascinating area where tools like docker-compose and Kubernetes are trying to solve.

Maybe in the future I will post a blog post about it.

It’s a long one, sincerely thank you for reading till here!

Originally published at blog.bencao.it on April 20, 2015. Revised on March 2019.

No comments yet. Be the first.

Optional. Leave it blank to post as Anonymous.