Here are my notes from PyData with links for more details. This isn't a complete list, and in some cases my notes don't really do justice to the actual talks, but I hope that these will be helpful to anyone who's feeling PyData FOMO until the videos are released.
This is the first post in a multi-part series wherein I will explain the details surrounding the language prediction model I presented in my Pycon 2014 talk. If you make it all the way through, you will learn how to create and deploy a language prediction model of your own.
A little over a year ago I was frustrated with the lack of data meetups in the Philadelphia area, so I started DataPhilly. I quickly learned that when you start a tech meetup you're going to have to do some public speaking to get the ball rolling.
After my successful talks at PyData Boston in July, I decided to submit one of my talks to Pycon. I'm happy to say my talk was accepted! This will be my first Pycon and I'm really excited!
For some light vacation reading, I started reading Hadoop Beginner's Guide. I made it through about half of the book, and I wanted to share some random facts that I found particularly enlightening.
When it comes to tokenization, email content presents some unique challenges. Some messages have a plain text version, some have a HTML version, and some have both.