Wednesday, November 20, 2013

Integrating SQLAlchemy into Django

In my previous post I describe some of the limitations that I have hit while using Django's ORM for our RESTful API here at HireVue.  In this post I will describe how I integrated SQLAlchemy into our Django app for read-only (GET) requests, to handle the queries that Django doesn't allow.

Let's say we are running with a single database instance, so our django settings look something like

DATABASES {
    'default' : {
        'ENGINE' : 'django.db.backends.postgresql_psycopg2',
        'NAME' : 'mydatabase',
        'USER' : 'mydatabaseuser',
        'PASSWORD' : 'mypassword',
        'HOST' : '127.0.0.1',
        'PORT' : '5432',
    }
}

We will continue to let Django handle all of the connection management, so when SQLAlchemy needs a database connection, we want to just use the current connection.  The method to get this connection looks like

# custom connection factory, so we can share with django
def get_conn():
    from django.db import connections
    conn = connections['default']
    return conn.connection

Now we want SQLAlchemy to call get_conn whenever it needs a new connection.  In addition, we need to keep SQLAlchemy from trying to pool the connection, clean up after itself, etc.  We essentially need it to do absolutely no connection handling.  To accomplish this, we create the SQLAlchemy engine with a custom connection pool that looks like

# custom connection pool that doesn't close connections, and uses our
# custom connection factory
class SharingPool(NullPool):
    def __init__(self, *args, **kwargs):
        NullPool.__init__(self, get_conn, reset_on_return=False,
                          *args, **kwargs)

    def status(self):
        return 'Sharing Pool'

    def _do_return_conn(self, conn):
        pass

    def _do_get(self):
        return self._create_connection()

    def _close_connection(self, connection):
        pass

    def recreate(self):
        return self.__class__(self._creator,
                              recycle=self._recycle,
                              echo=self.echo,
                              logging_name=self._orig_logging_name,
                              use_threadlocal=self._use_threadlocal,
                              reset_on_return=False,
                              _dispatch=self.dispatch,
                              _dialect=self._dialect)

    def dispose(self):
        pass

The magic is where we pass get_conn to the NullPool constructor, and then we override most of the other methods to do nothing.  This allows SQLAlchemy to borrow the connection that Django is managing, without interfering.

Now we can create an engine with this pool like so:

# create engine using our custom pool and connection creation logic
engine = create_engine(db_url, poolclass=SharingPool)

Since we want to keep DRY, we use Django's settings to create the db_url that is passed to the engine constructor:

db_url = 'postgresql+psycopg2://{0}:{1}@{2}:{3}/{4}'.format(
            settings.DATABASES['default']['USER'],
            settings.DATABASES['default']['PASSWORD'],
            settings.DATABASES['default']['HOST'],
            settings.DATABASES['default']['PORT'],
            settings.DATABASES['default']['NAME'])

We also don't want to maintain table definitions separate from our Django models, so we use SQLAlchemy reflection to build the metadata:

# inspect the db metadata to build tables
meta = MetaData(bind=engine)
meta.reflect()

This is basically it!  Now you can get the tables you would like to query from the meta object, and create and run selects.  They will run on the connection provided by Django.  Don't try to do any transaction handling or it may mess up Django.  

Also, you'll want to make sure you reflect() on startup, so that tables are ready to go when you need them.

Monday, October 28, 2013

Limitations of Django's ORM

I'm currently working at HireVue, where we have a public RESTful API that we consume for our website. I have done quite a bit of work on the API, and in the process, have run into some of the limitations of the Django ORM.

Difficulty controlling joins

Let's say I have an Interview table for storing information about candidate interviews, an Evaluation table for storing interviewer evaluations, and a User table for information on the interviewers/evaluators.  The Evaluation table is a many-to-many connecting Users and Interviews, with evaluation information stored in the table as well.


Now I want to write a query that returns interviews that a specific user has evaluated highly.  I try something like this:

Interview.objects.filter(evaluation__rating__gt=5, evaluation__user__username='specificuser')

I cross my fingers and hope that Django only joins to the evaluation table once, and uses that same join for both filter conditions.  It seems to work in the cases I have tried, but there is no way for me to tell Django for sure that is what I want.

On the other hand, let's say I want to write a query that returns interviews that a specific user has evaluated low, but others have evaluated high.  How do I tell Django I need two different joins?

Interview.objects.filter(evaluation__rating__lt=3, evaluation__user__username='specificuser') \
        .filter(evaluation__rating__gt=5).distinct()

Will this work?  I have had mixed luck, and find myself banging my head against the Django ORM black box, wishing I could tell it more explicitly what I want.  The Django docs are very reassuring on this matter:
 Conditions in subsequent filter() or exclude() calls that refer to the same relation may end up filtering on different linked objects.
I'm also limited if I am trying to do filtering in different parts of my code, on the same relationship.  Let's say I have some 'permissions' functions that apply a filter so that the request can only see interviews that the current user has evaluated, and then other functions to filter based on search criteria, like whether the evaluation was high or low:

query = Interview.objects.all() 
query = apply_permissions(query) # filter(evaluation__user__username='specificuser')
query = apply_search(query) # filter(evaluation__rating__gt=5)

I don't know of any way to tell Django that both of the filters, applied by different methods, should both use the same joined table.  This makes it difficult to break code out into logical components like this.

Inability to force outer joins

This is the bigger problem, in my opinion:  I am unaware of any way to tell Django to use an outer join.  Others seem to agree that this is not possible.

What if I would like to return all interviews, and if the current user has evaluated the interview, I'd like to include the rating.  This is easily done using a LEFT OUTER JOIN in sql.  Django, however, does not allow a query like this, leaving us with extra queries and joins in code.  That may not seem like too big of a pain in most cases, but if you are querying against a large data set, and want to sort by the ratings in that left outer joined table, things get even more difficult.

In my experience, these are my biggest complaints with Django's ORM.  It's frustrating to be held back by limitations in the ORM, when I know that the database can handle the problem easily.  I find myself jumping through hoops, or coding around these limitations way too often, and thought I'd share, for others considering using Django.

I'll follow up this post with another describing how I recently bolted SQLAlchemy onto our existing Django app, to handle the 'tough' queries.

Wednesday, March 9, 2011

Spring Nested Transactions and problems with "Session is closed" Exceptions

I was really pulling my hair out over some code that did programmatic transaction handling using Spring's PlatformTransactionManager on top of Hibernate. It is fairly complicated, with up to 4 different transactions running concurrently. At one point it needs to process 3 different result sets, like this:

TransactionStatus status1 = transactionManager.getTransaction(new DefaultTransactionDefinition(TransactionDefinition.PROPAGATION_REQUIRES_NEW));
PreparedStatement statement1 = sessionFactory.getCurrentSession().connection().prepareStatement(...);

TransactionStatus status2 = transactionManager.getTransaction(new DefaultTransactionDefinition(TransactionDefinition.PROPAGATION_REQUIRES_NEW));
PreparedStatement statement2 = sessionFactory.getCurrentSession().connection().prepareStatement(...);


TransactionStatus status3 = transactionManager.getTransaction(new DefaultTransactionDefinition(TransactionDefinition.PROPAGATION_REQUIRES_NEW));
PreparedStatement statement3 = sessionFactory.getCurrentSession().connection().prepareStatement(...);

... // set parameters on the statements

ResultSet rs1 = statement1.executeQuery();
ResultSet rs2 = statement1.executeQuery();
ResultSet rs3 = statement1.executeQuery();

... // process the result sets

... // close the result sets and statements

transactionManager.commit(status1);
transactionManager.commit(status2);
transactionManager.commit(status3);


This part of the code ran fine the first time, but the next time, it would throw an exception saying "Session is closed", when trying to get the first transaction. Do you see the problem? I didn't see it for way too long. I spent time removing sections of code, staring at logs, and pulling lots of hair out, before I finally noticed that in the spring logging, after finishing that section of code, it was falling back to trying to use the wrong session.

Finally it dawned on me that Spring must be using a stack to keep track of the current session. When a transaction is started with PROPAGATION_REQUIRES_NEW, the current session is pushed on the stack, and a new session is created and becomes the current session. When that transaction finishes, the current session is closed, and the previous session is popped off of the stack to resume as the now current session.

Looking at the documentation about transaction propagation, it does talk about "inner" and "outer" transactions, but I don't think it's quite explicit enough at explaining the nesting relationship. And my problem is that I wasn't really considering them to be nested, thinking of them more as simply independent transactions (in my defense, the docs do say that they are "completely independent" transactions). That is why it took me so long to realize my mistake.

My problem was that I was creating transaction 1, then transaction 2, then 3. But I was trying to commit transaction 1 first, then 2, then 3. Doing this messed up the stack, and in the end Spring was left with a session that had been closed as the current session. So the next attempt to use the session would cause an exception to be thrown telling me that the "Session is closed".

Rearranging the code to commit transaction 3 first, then 2, then 1, fixed the problem. I now have a better sense of how Spring works, working code, and less hair. I thought I'd write this up in case it helps someone else avoid this simple mistake.

Wednesday, September 15, 2010

CruiseControl to Hudson

At work we recently switched from using CruiseControl to Hudson for continuous builds.

Originally I was just trying to upgrade CruiseControl so that we could add in a plugin to support Mercurial (we also recently switched from Subversion to Mercurial for version control). We were on a fairly old version of CruiseControl. After upgrading, we were having lots of headaches with the web interface freezing up, and I had to write a Mercurial label incrementer plugin, and I didn't really like the new web interface anyway. I finally got frustrated enough to try something new.

I threw Hudson on the machine, and had builds up and running so much more quickly than in CruiseControl that I was sold almost immediately. Some of the differences that I really liked were:

* No editing XML. All configuration in Hudson can be done through the web interface. It's also easy to create build configs for new clones/branches. When you create a new Job, you can create it based on an existing Job, and then just change a few names and paths.
* Plugins are listed and installed through the web interface. This was probably the single best part of Hudson compared to CruiseControl. With CruiseControl, I was left searching for plugins to see what existed, and going through the hassle of researching and installing plugins to try to get things working. Hudson lets you see all the available plugins in one place. That was really handy.
* Hudson automatically detects if test cases fail, and will mark the build as "unstable". It's also really good at showing you the test case failures within the Hudson web interface.
* More features. I could probably get all of these things in CruiseControl, but it's so hard to find plugins, that I didn't try very hard. In Hudson it was easy to get it to host/expose my javadocs. It's easy to link to the most recent artifacts with static urls. It's also possible to have "slave" build machines so that long builds don't hold everything up.
* The email notifications are better. I only get emailed the first time a build breaks, and not for every subsequent failure. It also emails me when the build is fixed.

Now I'm guessing that I could probably get CruiseControl to do most or maybe all of those things, but when Hudson makes it so easy, why try to figure CruiseControl out?

The only trick I had to work out was getting Hudson to show the build revision. We use build revisions in our bug tracking system, so that developers can let testers know what revision a bug fix was made in. The solution I found was to use the "Hudson Description Setter Plugin". I installed the plugin in Hudson, then modified my build script to output the Mercurial revision. I created an ant task that I run as a dependency in my build task:


<target name="revision" description="Stores the latest revision number in revision property">
   <echo message="os: ${os.name}"/>
   <condition property="unix">
      <os family="unix"/>
   </condition>
   <echo message="is unix: ${unix}"/>
   <if>
      <equals arg1="${unix}" arg2="true"/>
      <then>
         <exec executable="sh" outputproperty="revision" errorproperty="revision-error" dir="${basedir}">
            <arg value="-c"/>
            <arg line="hg identify -n | tr -d '+'"/>
         </exec>
      </then>
      <else>
         <property name="revision" value="unknown"/>

      </else>
   </if>
   <echo message="rev: ${revision}"/>
</target>



Then I activated the plugin in my build configuration in Hudson, under "Post-build Actions" (the checkbox named "Set build description"), and set the regular expression to "rev: (.*)" and the Description to "[version] \1". All subsequent builds will have a description with the Mercurial revision in it (e.g. "[version] 12345").

Tuesday, October 27, 2009

Storing a tree structure in a database

NOTE: I fixed a bug in the node update trigger on 16 August 2011.

I think that one of the more tough problems in database design is how to store tree data (arbitrary depth parent-child relationships, where a child has at most one parent).

The two most common approaches are the Adjacency List model, and the Nested Set model. Both are explained and compared here. This forum post also has some good links to information on the two models.

In my opinion the major advantage and disadvantage of each are:

Adjacency Lists:
  • Major advantage:  Simple and easy to understand.
  • Major disadvantage:  It takes multiple queries to find all ancestors or all descendants of a node.
Nested Sets:
  • Major advantage:  A single query can find all ancestors or all descendants of a node.
  • Major disadvantage:  Modifying the tree structure affects half the nodes in the tree, on average.
What if you have a large tree (hundreds of millions of nodes), with fairly frequent changes to the tree structure, and you need to be able to run queries that access all ancestors or descendants of a node?  Since the nested set model would make it very difficult to make frequent modifications to the hierarchy, the other option is to do some denormalization of the adjacency list model so that we can query for ancestors and descendants of a node.

Let's say we have a simple node table:

CREATE TABLE node (
 node_id INTEGER PRIMARY KEY AUTO_INCREMENT,
 parent_id INTEGER,

 CONSTRAINT fk_node__parent FOREIGN KEY (parent_id) REFERENCES node (node_id) ON DELETE CASCADE
) ENGINE=InnoDB;


Each node has at most one parent.  Root nodes have a parent of null.  Now we create an ancestor list, which is our denormalization table:

CREATE TABLE node_ancestry_link (
 node_id INTEGER UNSIGNED NOT NULL,
 ancestor_id INTEGER UNSIGNED NOT NULL,

 PRIMARY KEY(node_id, ancestor_id),
 INDEX ix_node_anc__anc_node (ancestor_id, node_id),
 CONSTRAINT fk_node_anc__node FOREIGN KEY (node_id) REFERENCES node(node_id) ON DELETE CASCADE,
 CONSTRAINT fk_node_anc__anc FOREIGN KEY (ancestor_id) REFERENCES node(node_id) ON DELETE CASCADE
) ENGINE=InnoDB;


Each node has an entry in the ancestry table for each of its ancestors.  This table will grow much more quickly than the node table, especially for deep trees.  This denormalization allows us to write queries to get all ancestors of a node:

SELECT l.ancestor_id FROM node n
JOIN node_ancestry_link l on n.node_id = l.node_id
WHERE n.node_id = :nodeid;


and all descendants of a node:

SELECT l.ancestor_id FROM node n
JOIN node_ancestry_link l on n.node_id = l.ancestor_id
WHERE n.node_id = :nodeid;


Now, to help us keep the ancestry table in sync as changes are made in the node table, we define some triggers:

DELIMITER |

CREATE TRIGGER tr_node_ins AFTER INSERT ON node
FOR EACH ROW
BEGIN
  INSERT INTO node_ancestry_link (node_id, ancestor_id) VALUES (NEW.node_id, NEW.node_id);
  IF NEW.parent_id IS NOT NULL THEN
    INSERT INTO node_ancestry_link (node_id, ancestor_id) SELECT NEW.node_id, l.ancestor_id FROM node_ancestry_link l WHERE l.node_id = NEW.parent_id;
  END IF;
END
|

CREATE TRIGGER tr_node_upd AFTER UPDATE ON node
FOR EACH ROW
BEGIN
  IF NEW.parent_id <> OLD.parent_id OR ((NEW.parent_id IS NULL) <> (OLD.parent_id IS NULL)) THEN
    IF OLD.parent_id IS NOT NULL THEN
      DELETE FROM links USING node_ancestry_link links
        JOIN node_ancestry_link anclinks ON links.ancestor_id = anclinks.ancestor_id
        JOIN node_ancestry_link deslinks ON links.node_id = deslinks.node_id
        WHERE anclinks.node_id = OLD.parent_id
        AND deslinks.ancestor_id = NEW.node_id;
    END IF;
    IF NEW.parent_id IS NOT NULL THEN
      INSERT INTO node_ancestry_link (node_id, ancestor_id)
        SELECT desnodes.node_id, ancnodes.ancestor_id
        FROM node_ancestry_link ancnodes
        CROSS JOIN node_ancestry_link desnodes
        WHERE ancnodes.node_id = NEW.parent_id
        AND desnodes.ancestor_id = NEW.node_id;
    END IF;
  END IF;
END
|

DELIMITER ;


With the triggers in place, we can edit the node hierarchy without having to worry about the ancestry table.  We only need to worry about inserts and updates because of the CASCADE DELETEs on the ancestry table foreign keys.

If you need to know how many descendants a particular node has, you may want to track the descendant count in the node table, since the count queries will be expensive for large trees. 

You could augment the triggers to update the descendant counts when nodes are inserted, updated, or deleted.  This would require adding a delete trigger.  The one gotcha with doing this in MySQL is that you can't depend on cascade deletes when you delete nodes.  MySQL has a bug/feature that cascade deletes don't fire delete triggers.  For keeping descendant counts up-to-date this isn't a problem, as long as you take it into account, and when a node is deleted, subtract its descendant count from its ancestors.

Tuesday, September 1, 2009

re-dispatching events in Flash

I recently ran into a head-scratcher, and couldn't find any help online, so I thought I'd post my solution.

I had a custom flex component, ZoomControl:

<mx:VBox>
  <mx:Metadata>
    [Event(name="zoomChanged", type="ZoomEvent")]
  </mx:Metadata>
...
</mx:VBox>>

ZoomControl can throw a "zoomChanged" event, of a custom class ZoomEvent. I then created a Toolbar custom component that contained the ZoomControl:

<mx:VBox>
  <mx:Metadata>
    [Event(name="zoomChanged", type="ZoomEvent")]
  </mx:Metadata>
...
  <ZoomControl
    zoomChanged="dispatchEvent(event)"
  />
</mx:VBox>>

When the ZoomControl dispatches a ZoomEvent, I want my Toolbar to re-dispatch the event. Seems simple, right? But when the Toolbar calls dispatchEvent, it throws an exception:

TypeError: Error #1034: Type Coercion failed: cannot convert flash.events::Event@19ea2df1 to footnote.imageviewer.events.ZooomEvent.
at flash.events::EventDispatcher/dispatchEventFunction()
at flash.events::EventDispatcher/dispatchEvent()
at mx.core::UIComponent/dispatchEvent()[E:\dev\3.0.x\frameworks\projects\framework\src\mx\core\UIComponent.as:9051]
at Toolbar/__zoomControl_zoomChanged()...

After some wasted time trying to make sure that it really was a ZoomEvent being passed to dispatchEvent, I finally took time to read the documentation on UIComponent.dispatchEvent:

* If the event is being redispatched, a clone of the event is created automatically.
* After an event is dispatched, its target property cannot be changed,
* so you must create a new copy of the event for redispatching to work.

My problem was that I needed to override the clone method in my ZoomEvent. The dispatchEvent method was calling the base Event.clone(), which was returning an Event, of course. Overriding the clone method to return a ZoomEvent solved the problem.

Friday, June 12, 2009

SuperDuper "Smart Update" doesn't stay smart

I use SuperDuper for backups on my work machine (MacBook Pro), and I have it set up to backup daily to an external drive using "Smart Update", which is supposed to be fast and only copy things that have changed. I really have no idea how it works, nor do I care, as long as it's working.

The problem is that it has started to take longer and longer to run. It had reached the point where it was taking 2 hours or more to finish (my drive is 120G and I don't back all of it up). I found very little help by searching on Google, so I asked the Sysadmin, who also uses SuperDuper, if he had seen the same problem. He said "I don't know, I only do full backups once a week".

That made me think that maybe SuperDuper just doesn't handle backing up repeatedly using Smart Update only. So I revised my schedule to do a full backup once a week, and Smart Updates daily. Success! My daily backups are back down to 15 minutes or less! I thought I'd post this in case it helps someone else.

Friday, April 24, 2009

SELECT DISTINCT with ORDER BY

I recently wrote a query in MySQL that didn't seem to be returning the right results, and at first I couldn't figure out why. Here is a toy example, where we are tracking pages and page views (one-to-many relationship):


create table page (
    page_id integer unsigned primary key,
    name varchar(32) not null,
    created datetime not null
) engine=InnoDB;

create table page_view (
    page_view_id integer unsigned primary key,
    page_id integer unsigned not null,
    created datetime not null,
    
    foreign key (page_id) references page (page_id) on delete cascade
) engine=InnoDB;


What I want to get is the most recently viewed pages. Let's say I have the following data in my tables:


mysql> select * from page;
+---------+--------+---------------------+
| page_id | name   | created             |
+---------+--------+---------------------+
|       1 | page 1 | 2000-01-01 00:00:00 |
|       2 | page 2 | 2000-01-02 00:00:00 |
|       3 | page 3 | 2000-01-03 00:00:00 |
|       4 | page 4 | 2000-01-04 00:00:00 |
+---------+--------+---------------------+
4 rows in set (0.00 sec)

mysql> select * from page_view;
+--------------+---------+---------------------+
| page_view_id | page_id | created             |
+--------------+---------+---------------------+
|            1 |       3 | 2000-01-01 00:00:00 |
|            2 |       1 | 2000-01-02 00:00:00 |
|            3 |       1 | 2000-01-03 00:00:00 |
|            4 |       3 | 2000-01-04 00:00:00 |
|            5 |       2 | 2000-01-05 00:00:00 |
|            6 |       4 | 2000-01-06 00:00:00 |
|            7 |       2 | 2000-01-07 00:00:00 |
+--------------+---------+---------------------+
7 rows in set (0.00 sec)


What I want to get back is page 2 (most recently viewed), then page 4, then page 3, then page 1.

So I write my query:


mysql> select distinct p.page_id, p.name, p.created from page p join page_view pv on p.page_id = pv.page_id order by pv.created desc;
+---------+--------+---------------------+
| page_id | name   | created             |
+---------+--------+---------------------+
|       4 | page 4 | 2000-01-04 00:00:00 |
|       2 | page 2 | 2000-01-02 00:00:00 |
|       1 | page 1 | 2000-01-01 00:00:00 |
|       3 | page 3 | 2000-01-03 00:00:00 |
+---------+--------+---------------------+
4 rows in set (0.00 sec)


That's not right at all! What's going on?

The problem is that I'm using distinct just on the page table, but ordering by the page_view table. Since there is a many-to-one, what is the database supposed to do when a page has multiple views? which view should it use for the order by?

What I wanted the query to do is first join, then order, then apply the distinct. That's not what MySQL does, though. It first joins, then applies the distinct, then orders the results (or something like that). You can think of it like MySQL going sequentially through the page_view table, finding rows with distinct page ids. So it would pick rows 1,2,5,6:


+--------------+---------+---------------------+
| page_view_id | page_id | created             |
+--------------+---------+---------------------+
|            1 |       3 | 2000-01-01 00:00:00 |
|            2 |       1 | 2000-01-02 00:00:00 |
|            5 |       2 | 2000-01-05 00:00:00 |
|            6 |       4 | 2000-01-06 00:00:00 |
+--------------+---------+---------------------+
4 rows in set (0.00 sec)


You can see that if you order those by created, you get the page order that the (badly written) query returned (4,2,1,3).

We can force MySQL to do things in the order we want by changing the query to:


mysql> select distinct p.page_id, p.name, p.created from (select p.page_id, p.name, p.created from page p join page_view pv on p.page_id = pv.page_id order by pv.created desc) as p;
+---------+--------+---------------------+
| page_id | name   | created             |
+---------+--------+---------------------+
|       2 | page 2 | 2000-01-02 00:00:00 |
|       4 | page 4 | 2000-01-04 00:00:00 |
|       3 | page 3 | 2000-01-03 00:00:00 |
|       1 | page 1 | 2000-01-01 00:00:00 |
+---------+--------+---------------------+
4 rows in set (0.00 sec)


But I think that's kind of a hack, and depends on MySQL doing the distinct in a certain order (I don't think order by in a subquery is standard sql, and shouldn't necessarily constraint the order of the entire query). So what's the "right" way to write this type of query?

Before I tackled that, I thought, "What would a strict database like PostgreSQL do with this type of query?" My hope was that it would throw it out altogether. And it does. Here's what I get:


postgres=# select distinct t1.id, t1.name, t1.created from table1 t1 join table2 t2 on t1.id = t2.table1_id order by t2.created desc;
ERROR: for SELECT DISTINCT, ORDER BY expressions must appear in select list


That's much better, and the error message is very helpful, and makes sense. So here's the query I came up with that will give the correct results, and is correct SQL, in MySQL...:


mysql> select p.page_id, p.name, p.created from page p join (select page_id, max(created) as created from page_view group by page_id) v on p.page_id = v.page_id order by v.created desc;
+---------+--------+---------------------+
| page_id | name   | created             |
+---------+--------+---------------------+
|       2 | page 2 | 2000-01-02 00:00:00 |
|       4 | page 4 | 2000-01-04 00:00:00 |
|       3 | page 3 | 2000-01-03 00:00:00 |
|       1 | page 1 | 2000-01-01 00:00:00 |
+---------+--------+---------------------+
4 rows in set (0.00 sec)

...and in PostgreSQL:

postgres=# select p.page_id, p.name, p.created from page p join (select page_id, max(created) as created from page_view group by page_id) v on p.page_id = v.page_id order by v.created desc;
page_id |  name  |       created       
---------+--------+---------------------
       2 | page 2 | 2000-01-02 00:00:00
       4 | page 4 | 2000-01-04 00:00:00
       3 | page 3 | 2000-01-03 00:00:00
       1 | page 1 | 2000-01-01 00:00:00
(4 rows)


Is there a better performing query out there to do the same thing? I'd love to know, please leave a comment! :)

Monday, April 6, 2009

safely editing MySQL triggers in a production database

MySQL does not provide an atomic CREATE OR REPLACE TRIGGER, or an ALTER TRIGGER statement that will safely modify a trigger on a database while it is in use. The only way to update a TRIGGER is with a DROP and then a CREATE.

Why is that a big deal? Say, for example, you are using triggers to keep row counts up-to-date in a summary table. You may miss some inserts while you are issuing the DROP and then the CREATE. To verify this, I used mysqlslap. Here is my schema script:


drop table if exists triggertest.record_count;
create table triggertest.record_count
(
id INTEGER UNSIGNED PRIMARY KEY AUTO_INCREMENT,
count_name VARCHAR(64) NOT NULL,
count_value INTEGER UNSIGNED NOT NULL DEFAULT 1,
UNIQUE (count_name)
) ENGINE=InnoDB;

drop table if exists triggertest.record_table;
create table triggertest.record_table
(
id INTEGER UNSIGNED PRIMARY KEY AUTO_INCREMENT,
some_value VARCHAR(64) NOT NULL
) ENGINE=InnoDB;

DROP PROCEDURE IF EXISTS triggertest.sp_increment_record_count;

DELIMITER |

CREATE PROCEDURE triggertest.sp_increment_record_count(IN countname VARCHAR(64))
BEGIN
INSERT INTO triggertest.record_count(count_name, count_value) VALUES (countname,1) ON DUPLICATE KEY UPDATE count_value = count_value + 1;
END
|

DELIMITER ;

DROP TRIGGER IF EXISTS triggertest.tr_record_table_ins;

CREATE TRIGGER triggertest.tr_record_table_ins AFTER INSERT ON triggertest.record_table
FOR EACH ROW CALL triggertest.sp_increment_record_count('record_table');


I used mysqlslap to run a lot of inserts against the record_table, and while that was running, I re-created the trigger by running this script a bunch of times:


DROP TRIGGER IF EXISTS triggertest.tr_record_table_ins;
CREATE TRIGGER triggertest.tr_record_table_ins AFTER INSERT ON triggertest.record_table
FOR EACH ROW CALL triggertest.sp_increment_record_count('record_table');


I then verified that the count_value in record_count was smaller than the number of records in record_table:


mysql> select * from record_count;
+----+--------------+-------------+
| id | count_name | count_value |
+----+--------------+-------------+
| 1 | record_table | 9944 |
+----+--------------+-------------+
1 row in set (0.00 sec)

mysql> select count(*) from record_table;
+----------+
| count(*) |
+----------+
| 10000 |
+----------+
1 row in set (0.00 sec)

mysql>


At first, I was not sure there would be a solution to this. I realize that you can lock tables, but my first guess was that since ddl (DROP, CREATE, etc.) statements cause an implicit commit, that my locks would be released.

Fortunately, as the MySQL documentation explains, if you use LOCK TABLES, implicit commits don't release your locks. From the docs:

...statements that implicitly cause transactions to be committed do not release existing locks.

So the safe way to recreate my trigger is like this:

set autocommit=0;
lock tables triggertest.record_table write;
DROP TRIGGER IF EXISTS triggertest.tr_record_table_ins;
CREATE TRIGGER triggertest.tr_record_table_ins AFTER INSERT ON triggertest.record_table FOR EACH ROW CALL triggertest.sp_increment_record_count('record_table');
unlock tables;


I had some trouble testing this with mysqlslap, even with only one thread running inserts, because of some locking issues (the inserts would error out with a 'Lock wait timeout'), but I did get a few tests to make it through, so I could verify that the record_count matched the number of rows in the record_table. So the worst-case seems to be that some of the inserts on the production database may hit a lock wait timeout, but no inserts will miss firing the triggers!

UPDATE: Soon after writing this post, I came across this: http://code.openark.org/blog/mysql/why-of-the-week, which may explain why I was having so many problems with deadlocks when I tried to run against a database that was in use. I'm not talking about deadlocks where MySQL detects it and rolls back a transaction. Things would just lock up. No deadlock detected, no lock wait timeout, just locked up, until I killed a query. So while the solution above should work in theory, beware of MySQL locking bugs...

Friday, January 9, 2009

PostgreSQL transactions and error handling

I'm taking our web app that runs on MySQL and seeing what it would take to get it running on PostgreSQL. I just discovered one rather glaring difference between PostgreSQL and most other DBMSs.

Here is an example to demonstrate:

I have a table: my_table, with the following row:


----------------
| id  |  val   |
----------------
| 3   | row 3  |
----------------


Let's say I want to make sure I have rows for ids 1 through 5 in the table. This is a case where I would use MySQL's 'INSERT IGNORE' statement, which doesn't exist in PostgreSQL. I have at least two options: I can create a stored procedure to do it, or I can do it in code.

Let's say I create a procedure:


CREATE OR REPLACE FUNCTION my_proc() RETURNS void AS
$$
BEGIN
  FOR i IN 1..5 LOOP
  BEGIN
    INSERT INTO my_table(id, val) VALUES (i, 'row ' || i);
  EXCEPTION WHEN unique_violation THEN
  -- do nothing
  END;
  END LOOP;
END;
$$ LANGUAGE plpgsql;


All is well, this type of exception handling is standard in stored procedures. Here's the entire test:


postgres=# create database my_test;
CREATE DATABASE
postgres=# \c my_test;
You are now connected to database "my_test".
my_test=# create table my_table (id INTEGER PRIMARY KEY, val VARCHAR(64));
NOTICE:  CREATE TABLE / PRIMARY KEY will create implicit index "my_table_pkey" for table "my_table"
CREATE TABLE
my_test=# CREATE OR REPLACE FUNCTION my_proc() RETURNS void AS
my_test-# $$
my_test$# BEGIN
my_test$#   FOR i IN 1..5 LOOP
my_test$#   BEGIN
my_test$#     INSERT INTO my_table(id, val) VALUES (i, 'row ' || i);
my_test$#   EXCEPTION WHEN unique_violation THEN
my_test$#   -- do nothing
my_test$#   END;
my_test$#   END LOOP;
my_test$# END;
my_test$# $$ LANGUAGE plpgsql;
CREATE FUNCTION
my_test=# insert into my_table values (3,'row 3');
INSERT 0 1
my_test=# select * from my_table;
 id |  val  
----+-------
  3 | row 3
(1 row)

my_test=# select my_proc();
 my_proc 
---------
 
(1 row)

my_test=# select * from my_table;
 id |  val  
----+-------
  3 | row 3
  1 | row 1
  2 | row 2
  4 | row 4
  5 | row 5
(5 rows)

my_test=# 


Now here's the twist - you can't (really) do the same thing outside of a stored procedure (without using savepoints, as I'll get to later). Here's what happens:


my_test=# \set AUTOCOMMIT OFF
my_test=# delete from my_table where id <> 3;
DELETE 4
my_test=# commit;
COMMIT
my_test=# insert into my_table values (1, 'row 1');
INSERT 0 1
my_test=# insert into my_table values (2, 'row 2');
INSERT 0 1
my_test=# insert into my_table values (3, 'row 3');
ERROR:  duplicate key value violates unique constraint "my_table_pkey"
my_test=# insert into my_table values (4, 'row 4');
ERROR:  current transaction is aborted, commands ignored until end of transaction block
my_test=# insert into my_table values (5, 'row 5');
ERROR:  current transaction is aborted, commands ignored until end of transaction block
my_test=# commit;
ROLLBACK
my_test=# select * from my_table;
 id |  val  
----+-------
  3 | row 3
(1 row)


PostgreSQL forces you to rollback a transaction that hits any error. There is no exception handling! Granted, you can get around this by using savepoints:


my_test=# insert into my_table values (1, 'row 1');
INSERT 0 1
my_test=# insert into my_table values (2, 'row 2');
INSERT 0 1
my_test=# SAVEPOINT my_hack;
SAVEPOINT
my_test=# insert into my_table values (3, 'row 3');
ERROR:  duplicate key value violates unique constraint "my_table_pkey"
my_test=# ROLLBACK TO SAVEPOINT my_hack;
ROLLBACK
my_test=# insert into my_table values (4, 'row 4');
INSERT 0 1
my_test=# insert into my_table values (5, 'row 5');
INSERT 0 1
my_test=# commit;
COMMIT
my_test=# select * from my_table;
 id |  val  
----+-------
  3 | row 3
  1 | row 1
  2 | row 2
  4 | row 4
  5 | row 5
(5 rows)

my_test=# 


But I see this as more of a hack than a real solution.

I can see arguments for both sides of this issue. I really like this conversation about the issue: http://www.nabble.com/25P02,-current-transaction-is-aborted,-commands-ignored-until-end-of-transaction-block-td3710080.html . But what bugs me is that stored procedures can do exception handling, but no one else can. Especially in my case, where I'm using Java and JDBC, which throws exceptions for any errors that come back. I would rather do my own exception handling, or at least have the option of doing my own.

So here's my vote for changing PostgreSQL to behave more like MySQL in this case.

Tuesday, November 25, 2008

should you name foreign key constraints in MySQL?

In MySQL, if you don't name your foreign key constraints, the database generates a name for them automatically. Foreign key constraint names must be globally unique within the database, so unless you have a reason to name them, it is probably more hassle than it is worth.

I happen to have a reason to want to name them, and our setup is probably not that uncommon, so others may discover that it is advantageous to name them, too. We have three different databases where I work - one for development ('dev'), a staging database for testing each iteration, and patches ('staging'), and, of course, our live database that the site runs on ('live'). Additionally, each developer has a local database on their personal machines, to develop against.

For each iteration, we create a single schema migration script as we are developing. It will likely get run in pieces on the dev database, and there may be multiple revisions to a table as the iteration develops. Usually the script is very final by the time it is run against staging, but there is always the possibility of additional changes late in the game.

So where foreign key constraint names come into play is when you want an alter statement that can be run against all the different databases without changing. If you leave the naming up to the database, they are typically named in a sequential fashion (first foreign key constraint will probably have a '_1' at the end, the next will have '_2', etc.).

The problem is that if you create and drop foreign keys in different orders on different databases, the names won't match up. The foreign key named 'blah_1' on dev might be on a different column than the one with the same name on staging. You have to alter them by name, so there is no way to have a single script that will run correctly on all of the databases.

Wednesday, October 15, 2008

Spring AOP and @annotation pointcuts

I'm working on a Spring-AOP and annotation-based solution for caching web service requests. Here's the overview:

I have created an Annotation named "Cacheable", that you use like this:


@Cacheable(seconds=60)
public Data getData(int id)
{
  ...
}


Then using Spring's AOP functionality, I want to wrap every method that is annotated as @Cacheable in around advice that uses memcache to return cached results.

The problem I ran into was getting access to the seconds attribute of the annotation in the advice (for setting the cache timeout).

What I didn't understand, and wasn't clear to me from the Spring docs, was how to pass the annotation to the advice.

(Note: I'm using schema-based aop configuration)

Normally, if you are just trying to match methods that are annotated in a pointcut expression, you would make a pointcut definition with "@annotation(com.xyz.AnnotationName)". But what if you want to have access to the annotation in the advice? Your advice looks like this:


public Object aroundCacheable(ProceedingJoinPoint pjp, Cacheable cachable) throws Throwable
{
  int timeout = cacheable.seconds();
  ...
}


Your pointcut expression has to specify what to pass as the cacheable parameter in the advice. So you have to modify the pointcut definition and add arg-names like this:


<aop:around pointcut="@annotation(cacheable)" method="aroundCacheable" arg-names="cacheable"/>


My understanding is that this tell Spring to look for an argument named cacheable in the advice method, and figure out the type from that. This seems a little strange because the pointcut definition is dependent on the advice. It seems like a pointcut should be self-contained, and not depend on how it is used. But maybe I'm missing something. I'll have to look into it more later, but for now, I'm just glad I got it working.

Thursday, July 31, 2008

MySQL: triggers + replication = frustration

For the most part, I've been impressed with MySQL, but every once in a while I hit a problem that really surprises me. It seems like MySQL has this mentality that if there is a bug that is hard to fix, just document it, and then it's a feature and not a bug. Nice.

MySQL claims to have most of the power features of a robust RDBMS like triggers, foreign keys (although I consider that the most basic of features), stored procedures, etc. But it sure is frustrating to find out that most are incomplete.

Sure MySQL has foreign keys, and cascade deletes, but don't expect cascade deletes to fire triggers. That's a documented feature (bug).

Sure MySQL has triggers, but don't try using them if you are also using replication. We recently got bit by a feature (bug) where stored procedures or triggers that insert multiple records in tables with auto-increment don't work with replication. The auto-increment values on the replica will get off, and replication will break.

The problem is that before each insert statement in the binlog, there is a statement to set the auto-increment value. This is important to make sure that the values will always be the same between master and replica. On the master you may have two transactions that run in parallel, and use interleaved auto-increment values:

tx1: insert into table1 ... (uses auto-increment value 1)
tx2: insert into table1 ... (uses auto-increment value 2)
tx1: insert into table1 ... (uses auto-increment value 3)
tx1: commit;
tx2: commit;

In the binlogs, the transactions are serialized, so the auto-increment value has to be explicitly set:

tx1:
set auto-increment to 1;
insert into table1...
set auto-increment to 3;
insert into table1...

tx2:
set auto-increment to 2;
insert into table1...

But consider the case where the insert is on a table with a trigger that inserts a record into a second table (like an auditing table). The binlog only sets the auto-increment value for the actual insert statement. The trigger's insert will use whatever the replica's auto-increment value for the second table is set to. Since simultaneous transactions on the master are serialized in the binlogs, inserts on the second table may happen out of order, and auto-increment values will no longer match the master.

It seems that MySQL is not planning on fixing this bug in 5.0. My understanding is that in 5.1 the solution will be to use row-based replication. Statement based replication will still be broken, from what I can tell. I did see one bug report where someone said something about mixed mode replication, and switching to row-based temporarily for any statement or stored procedure call that will insert multiple records. Sounds like a can of worms to me.

Regardless of what happens in 5.1, there will be no solution for this in 5.0. And from my experience, I'll be nervous to move to 5.1 anytime soon. So it looks like I'll be stuck with this "feature" for a while to come. Nice.

Wednesday, April 2, 2008

stubbing out java.util.Random with EasyMock and Spring

Assuming you are familiar with Spring and EasyMock, here is a little tutorial on how to stub out randomness for your JUnit tests.

First I probably need to give a little context on how I have my test framework set up with Spring. I have two spring config files, one of them is just for testing and injects stubs/mock objects where appropriate. I have a base test class that extends AbstractDependencyInjectionSpringContextTests. It has members (and setters) for all of my classes that I want to test, and all of the stubs that get injected.

I make all of my classes to test scope="prototype" and my stubs are all singletons, so that I can easy get access to the same stub that will be injected into my class under test. Before running a test, I set up the stubs appropriately, and get them ready to replay. After the test, I call EasyMock.reset() to reset them for the next test (since they are singleton).

Ok, on to the example. Let's say you have a method that uses java.util.Random:


public class MyRandomClass
{
  private void Random rand;

  public int doSomethingRandom()
  {
    return rand.nextInt();
  }

  public void setRand(Random rand)
  {
     this.rand = rand;
  }
}


In your real spring config file you would have something like this:

<bean id="myRandomClass" class="MyRandomClass">
   <property name="rand" ref="rand"/>
</bean>

<bean id="rand" scope="prototype" class="java.util.Random"></bean>


And in your test spring config you would set up your 'rand' bean like this:

<bean id="rand" class="org.easymock.classextension.EasyMock" factory-method="createMock">
   <constructor-arg value="java.util.Random"/>
</bean>


Notice that I am using the EasyMock Class Extension because Random does not have an interface (Sun should add one, in my opinion). The regular EasyMock library can only mock interfaces, and the class extension adds the ability to mock classes themselves.

Now when setting up your test case, you can specify what 'random' values you want, like this:


EasyMock.expect(rand.nextInt()).andReturn(0);
EasyMock.expect(rand.nextInt()).andReturn(1);
...


And then you know what to expect:

assertEquals("returned wrong result", 0, myRandomClass.doSomethingRandom());
assertEquals("returned wrong result", 1, myRandomClass.doSomethingRandom());


That's it! (at least for this simple toy example)

One thing I noticed is that when using the EasyMock Class Extension, you have to remember to use the org.easymock.classextension.EasyMock createMock(), reset(), replay(), verify() etc. methods on any actual classes that you stub/mock. But I was able to use org.easymock.EasyMock.expect(), etc. methods when setting up the stubs.

Tuesday, April 1, 2008

using Spring in an Axis Web Service Impl

This is a re-post from another blog of mine:

I have an existing code base using Axis for web services, and am working on integrating Spring (2.0) into the system, for transaction management.

This is not straightforward, because Spring likes to be the one that creates your objects, so that it can inject dependencies. But with web services, Axis creates the service implementation class, not the Spring container. A 'hack' is necessary as an alternative to the standard Spring injection. And while it is not straightforward, it's not difficult, either. It just seemed that way to me because all the examples I found on the internet were confusing.

Here's the short version of how to get things working:

  1. If you are generating a Skeleton class with wsdl2java (-skeletonDeploy), stop doing it.
  2. Create a wrapper class for your service impl. For me this meant renaming my MyServiceSoapBindingImpl class to MyServiceImpl, and then regenerating MyServiceSoapBindingImpl. MyServiceSoapBindingImpl is now the wrapper class, and you have it contain a MyServiceImpl object and delegate all calls to that object.
  3. Change the wrapper class to extend ServletEndpointSupport, and make sure that your wsdd points to that wrapper class.
  4. Override ServletEndpointSupport.onInit(), and get the real impl (MyServiceImpl) as a bean using getWebApplicationContext().getBean("myServiceImplBean");

Now your impl class can use Spring for dependency injection just like any other class.

Here's the long version:

30,000 feet

What we want is to have a service interface (MyService), a service impl (MyServiceImpl) with the actual code, and be able to inject dependencies into MyServiceImpl (dao, etc.).

public interface MyService
{
public int doSomething();
}

public class MyServiceImpl implements MyService
{
private HelperBean myHelperObj;

public void setMyHelperObj(HelperBean myHelperObj)
{
this.myHelperObj = myHelperObj;
}

public int doSomething()
{
return myHelperObj.doTheSomething();
}
}

We want to be able to set this up in Spring:


<bean id="helperObj" class="HelperBean"></bean>

<bean id="myService" class="MyServiceImpl">
<property name="myHelperObj" ref="helperObj"/>
</bean>


And then set that service class up as a web service. This means that your wsdd would have something like this:


<service name="MyService" provider="java:RPC"></service>
...
<parameter name="className" value="MyServiceImpl"/>
...
</service>

The problem is that since MyServiceImpl is your web service class, Axis is in charge of it, not Spring. Since Spring does not create the class, it can't inject the HelperObj!

So let's see what the workaround is. First, I need to explain how I have things set up.

Axis/service class setup

I generate the wsdl from the service interface using the ant java2wsdl task, and then generate the stub, locator, etc. classes from the wsdl using the ant wsdl2java task. Because of this, my service impl class is actually named MyServiceSoapBindingImpl, just because that's what the ant task generates. The classes generated are:
  • MyServiceSoapBindingStub
  • MyServiceService
  • MyServiceServiceLocator
  • MyServiceSoapBindingImpl
Originally we were also generating a Skeleton class for the service (skeletonDeploy="yes"). This ended up causing me headaches, and I'm not really sure what the purpose of the Skeleton class is, anyway. Bottom line: don't do it.

To get around Spring not being able to inject dependencies into your service impl, it's necessary to make your service impl class (the one that your deploy.wsdd points to) a wrapper around the real impl class that you want to inject dependencies into. The wrapper class extends Spring's ServletEndpointSupport class, which provides an onInit method that you can override to get access to the spring application context:

public class MyServiceImpl implements MyService
{
private HelperBean myHelperObj;

public void setMyHelperObj(HelperBean myHelperObj)
{
this.myHelperObj = myHelperObj;
}

public int doSomething()
{
return myHelperObj.doTheSomething();
}
}

public class MyServiceSoapBindingImpl extends ServletEndpointSupport implements MyService
{
private MyService impl;

protected void onInit() throws ServiceException
{
impl = (MyService) getWebApplicationContext().getBean("myService");
}

public int doSomething()
{
return impl.doSomething();
}
}


So the 'hack' is that you have to use ServletEndpointSupport and access the Spring context directly, but from that point on, your real impl class can behave like any other Spring bean.

The part that messed me up was the Skeleton class. Since the wsdd pointed to the Skeleton class instead of the Impl, Spring for some reason didn't call my onInit() callback method. Hopefully this helps someone else having the same problem.

web.xml

Your web.xml doesn't need to have anything special. Just the normal configation. I got confused because all the examples I found on the web made it seem that you had to register a DispatcherServlet, which is not necessary just for what we are trying to do here.