Tuesday, September 23, 2008

String constructor considered useless turns out to be useful after all (film at 11)

When you were still a Java neophyte, chances are you wrote some code that looked like this:
String s = new String("test");
Brings back some embarrassing memories, doesn't it? You soon learned that instantiating Strings via the constructor was hardly ever done, and the String constructors seemed to be, well, utterly useless. When do you really ever need to do "new String(oldString)"? Come on, IntelliJ IDEA even flags all occurrences of this as "redundant"!

It turns out that this constructor can actually be useful in at least one circumstance. If you've ever peeked at the String source code, you'll have seen that it doesn't just have fields for the char array value and the count of characters, but also for the offset to the beginning of the String. This is so that Strings can share the char array value with other Strings, usually results from calling one of the substring() methods. Java was famously chastised for this in jwz'  Java rant from years back:
The only reason for this overhead is so that String.substring() can return strings which share the same value array. Doing this at the cost of adding 8 bytes to each and every String object is not a net savings...
Byte savings aside, if you have some code like this:
// imagine a multi-megabyte string here
String s = "0123456789012345678901234567890123456789";
String s2 = s.substring(0, 1);
s = null;
You'll now have a String s2 which, although it seems to be a one-character string, holds a reference to the gigantic char array created in the String s. This means the array won't be garbage collected, even though we've explicitly nulled out the String s!

The fix for this is to use our previously mentioned "useless" String constructor like this:
String s2 = new String(s.substring(0, 1));
It's not well-known that this constructor actually copies that old contents to a new array if the old array is larger than the count of characters in the string. This means the old String contents will be garbage collected as intended. Happy happy joy joy.

Sometimes, seemingly useless constructs reveal hidden gems of usefulness. Be sure to check out the source for String to find how this works!

Friday, August 22, 2008

Multiline grep

I recently needed to do some multiline grepping through some log files. The problem is, the default grep and egrep on Linux systems don't seem to support regex patterns that extend over several lines.

Luckily, at least RHEL come with pcregrep installed, which does Perl-compatible regex matching. And by adding a -M switch you get multiline matching!

So you can just go
pcregrep -M 'a\nb' files...
Easy peasy!

Thursday, May 15, 2008

EBCDIC trick

You can use good old /bin/dd to convert EBCDIC files to their ASCII equivalent.

Just do
dd if=infile.txt of=outfile.txt conv=ascii
Which does the conversion automatically.

Monday, April 7, 2008

Maven 2 help plugin

(I always forget this bit, so I'm putting it here for posterity.)

The help plugin for Maven 2 can print the possible goals and other info for a plugin:
mvn help:describe -Dplugin=eclipse -Dmedium=true
-Dfull=true gives some more detailed output.

For more info see the help plugin pages.

Thursday, October 18, 2007

On clean code and refactoring

Bob Martin (also known as "Uncle Bob") visited us at BEKK today and gave us a short pep talk on "clean code". Using his argument parser written in Ruby as an example, he showed some refactoring tricks and some nice Ruby idioms (passing blocks as arguments, etc.). His main point was that you should never check in bad code and leave the fixing for later, cause "later" never comes. If you make a mess, you should clean it up, and clean it up now.

I've thought along these lines before, and one of the issues I've wondered about is: Should you refactor just after you write the code or should you leave it until the next time you need to change it? This is really just another way of asking "How do you know where your refactoring should end up?"

On one hand, I agree that you should clean up after yourself if you've duplicated something or in some other way made the code worse than it was. But on the other hand, you don't know now what the best structure is since you don't know what changes you want to make next time. You'll have to guess in what form you'll need the code to be the next time you change it. This might lead you to a "local optimum" of code quality, but when you revisit the code you might need to take in a totally different direction, forcing you to abandon the local optimum for a solution that is better in the big picture.

If you leave the code in its current (messy) state and wait until the next time you have to change it, it could be a lot easier to see in which direction you should go. With concrete requirements for how you want to change the code, you can probably see that "hey, this would be a lot easier to change if the code was structured this way...". You would then have to make the necessary refactorings, before making any changes.

Of course, there's always the chance that you'll be in a hurry the next time, and so you won't really have time for refactoring. Your changes will make the design worse, and all of a sudden you have a big ball of mud. Also, working this way feels a bit like always working uphill. For every change you want to make, you have to refactor something first. It's a bit like having to clean up the kitchen, when you really just want to make dinner.

Either way, I guess I prefer cleaning up after I've written some (messy) code to leaving it for later. You can probably make educated guesses in most cases as to where you should take the design. But, just as you should be ready to throw away some messy code, I think you should be prepared to back out some refactorings if they conflict with the way you feel the code should be going.

Anyway, just some food for thought.

Thursday, October 4, 2007

Spring dynamic proxy problem

I encountered a strange problem at work today: The following test code failed:

ServiceImpl service = applicationContext.getBean("serviceImpl");

The application context looks like this:

<beans>
<bean id="serviceImpl" class="ServiceImpl"/>
</beans>

ServiceImpl implements the Service interface.

Note that the reference type in the test code is ServiceImpl. I sometimes do this when I need to test drive methods that are not in the implemented interface.

The error was a ClassCastException, which made no sense. Looking at the error, it said that it was unable to cast Proxy-yadayada to ServiceImpl.

It turns out that someone added some AOP stuff to the service classes. Spring then proxies the service implementation in one of two ways: if the class doesn't implement an interface it uses cglib to create a proxy.

But if the class does implement an interface Spring will interject a dynamic proxy which implements that interface between the original interface and implementation, and the above code will fail.

So what's the solution? The only immediate solution I see is to instantiate the service implementation yourself if you need to test drive methods that are not in the interface. Hardly ideal, but easy enough.

Tuesday, September 18, 2007

Is the DAO pattern still valuable in the time of ORM?

The DAO discussion is surfacing again (previously), and there's one point I think is lacking from the current discussion: Testing.

It's a lot easier to mock or stub out PersonDao.getPersonByFirstName() method than EntityManager.createNamedQuery() blah blah blah. Don't mock infrastructure. Avoiding the database access is crucial to making your tests run as fast as possible.

The DAO tests themselves should access the database, of course, as well as the end-to-end functional tests. But running every test in the system against the database is simply uneconomical.