Tuesday, 23 October 2012

Coursera: Data Programming in R, post II

I've finished my course work for Computing for Data Analysis, on Coursera, so I thought I'd take the time do do a quick review.

Overall:
I'm glad I took the course. A structured learning timeline with specific targets and an active discussion board is very valuable. It's far better than learning in isolation through random web tutorials and linked resources. 

The course is not for everyone:
 If you haven't done any programming, R is not a good first language. This course will not be a friendly introduction to programming. In particular, R is quirky and the command syntax is difficult to read. Many commands have similar names, but subtly different behaviors.  The help files are opaque and the examples frequently esoteric. The R programming environment lacks some very basic coding tools such as code completion, although these are probably available in other environments such as R-Studio or ESS.  If you want to learn basic programming, take a course in Python.

R will be a lot easier to digest if you are comfortable with statistics or matrix algebra, can read mathematical notation without difficulty, and have done a bit of programming. If you want flexible data analysis and publication quality output from free software, you'll be very happy. Most of what you want to do can be done with the core functions of R. It's probably best to learn what they can do before reaching for a package. This may save a lot of time later on when the package gets superseded by another one. The core of R will still be there, unchanged. That said, R appears to have a pretty good package management system, so incorporating packages that rely on packages seems to work very well.

If you are coming from an object oriented language such as Java, Ruby or C++, the scoping rules are a leetle different. This seems to be very powerful when used well, but it's mind bending. I haven't really managed to bend my mind around this one enough yet.

Level of the course: 
It would definitely help to have some programming and problem solving background going in. Students needed good problem solving to do the exercises. The basics of the language were taught in Powerpoint slides. The information in the videos was enough to get through the quizzes, but the exercises required more: R-help, R-bloggers, stackoverflow were very useful. 
There was very little discussion of speed optimisation or the tradeoffs in using different programming approaches in R. The final exercise could be solved with for loops, and judging by the forums, many students resorted to them. There was no penalty for this in the grading. The exercises, however were thoughtfully put together, and did provide a good platform for learning to leverage the language, with a little creativity and perseverance. 

There was an introduction to the differences between S3 and S4 classes, but no discussion of more advanced technologies such as refactoring, unit testing, version control, or documentation. I saw one reference to software carpentry on the forums, but there was no reference to how to incorporate these methods in R specifically. Function prototypes were provided for the exercises, and these included useful comments, setting a good standard. However, there was no mention of Runit (testing, TDD), Git (version control), or .Rd files or Roxygen (creating documentation). So if you want to learn how to incorporate these into your work with R, don't look here.  

Time Commitment: 
The course website suggests 2-4 hours / week. I was able to fit the video viewing and exercises into that time frame, on the outside. I spent an extra couple of hours reading and commenting on the forums and looking further / honing more satisfying solutions to the exercises.

What next?
R is quirky. I won't remember much of what I learned beyond a month or two at most. There is clearly a steep learning curve here, and thus a big difference between introduction, competence, and mastery.  At this point, I've had an introduction. In order to progress, I'll need some projects to work on.  

Resources for the future:
... which is just a tiny tip of the iceberg. Let me know about more in the comments.


Tuesday, 2 October 2012

Spirograph mania!

According to RetroWow, among other sources, Spirograph was invented by Denys Fisher in 1965. It was first intended as a drafting tool, but was marketed as a toy. There were multiple versions, and there is a modern remake from Hasbro.

 We've tried out three versions. The new version didn't stay in the house. It was too difficult to use because the gears kept slipping underneath the outer template. The pocket version shown below works OK, but the designs are somewhat limited (no epitrochoids). The antique Kenner version is our favorite. This is getting a lot of use right now. It's great at 8 or 9 years old and up, but your six year old would have to be very adept with a pen to enjoy it for long.
Mom! Can I get out the Spirograph?

The old pens were dry, so I did get a wonderful set of Stabilos, which are working perfectly. The pens need to be narrow enough to fit through the holes in the gears. Felt tips are a bit softer than roller balls, so they don't seem to make holes in the paper as easily, although the ink sometimes runs a bit.

Once you do a few patterns, it's nice to be able to predict what the wheels will do. There is a handy chart on the inside of the box lid, but the math is rather fun, too. We're not quite up to common denominators and gear ratios yet, but predicting the number of points on a spirograph pattern will be a good tool when we get there:
  • Outer wheel: 96 teeth
  • Inner wheel: 60 teeth
  • Step 1: Factor the number of teeth
    • 96 = 32 x 3 = 2 x 2 x 2 x 2 x 2 x 3
    • 60 = 20 x 3 = 5 x 2 x 3
  • Step 2: Compare the numbers. Find factors that are the same.
    • 96 = 32 x 3 = 2 x 2 x 2 x 2 x 2 x 3
    • 60 = 10 x 6 = 5 x 2 x 2 x 3
  • Step 3: Calculate the largest common denominator:
    • = 2 x 2 x 3 = 12
  • Step 4: Divide the number of teeth by the LCD to get the gear ratio:
    • 96 / 12 = 8, 
    • 60 / 12 = 5 
    • for a gear ratio of 8/5
    So it will take 5 round trips inside of the larger wheel to create a pattern with 8 points.
    When the circle is rotated around the inside of the fixed circle, as in this pattern, the result is called a hypotrochoid. 'Hypo' is a commonly used Greek root for 'under' or 'inner' as in hypoglycaemia for low blood sugar and hypoxia for lack of oxygen. If the inner circle is fixed and the outer one is rotated, the pattern is an epitrochoid. I think of 'epi' as a Greek root meaning 'surface' or 'outer' as in the medical name for the outer skin - epidermis.

    Of course these geometric forms have equations which can be used to describe them. And these equations have been implemented as interactive demos on the web. I don't find these as much fun as the physical drawing. Partly this is because the process of drawing a hypotrochoid is rather pleasant loopy-loop feeling. Also, though, the restriction to an integer number of teeth on the gears makes for a restriction on the variety of patterns. As in flowers, we don't really notice that number of petals is restricted to a multiple of 2, 3, or 5. We just find it pleasing.

    Sunday, 30 September 2012

    Learning Some R

    I'm following the free course: Programming for Data analysis, taught by Roger D. Peng of Johns Hopkins Bloomberg School of Public Health, and the Simply Statistics blog. It's offered through Coursera.org.

    I should be over-prepared for this course. It's supposed to be completely introductory, with only limited programming or statistics background required, but I'm interested in learning R. So this provides an interesting introduction where I can find out about MOOC's and how they work. My latest app update is 'waiting for review' on iTunes, so I should be able to find a couple hours a week to put into it. 

    The first week's lectures covered downloading and installing R, how to get help, some basic data types and selecting data from vectors, lists, etc. Data input and output. Although this is a programming class, the lectures were presented as slide-decks with voice-over. The slides could be downloaded as pdf, and translations were available as subtitles. Each slide had 2-5 sentences with some information about an R command, possibly including an example. I found this easy to follow, but low on information density. I listened to the lectures while making cupcakes for a bakesale. Then I increased the speed to 1.5x. I'll get more from referring to the pdf slides while doing the exercises.

    There were no suggestions for outside materials, although some resources were referred to in the 'getting help' lecture video. 

    Some 30,000+ students have apparently signed up for the course. The discussion board has several pages worth of questions. The introduce - yourself thread has the most views with nearly 5000. The other most popular threads have 1100 views or so at this point, so there is quite a bit of activity. Several of the "students" are already quite accomplished with R, and they are posting visualizations and code. This is a great learning resource. The rest of us are trying to share resources we find on the web:

    Links to free R resources


    It's not clear how many students will make it through the 1st quiz, much less the 1st assignment, but complaints on the discussion forum are significant, and are mostly coming from people without much programming background. Since the 1st assignment is not due until the end of week 2, the 1st week's lectures did not cover all the relevant material, many students feeling lost. The 'Not sure where to start…" thread has 2100 views. It would be nice if the 1st lectures were designed to allow you to get started writing code. That's not really necessary to get started using R, but it is necessary to do the course. Fortunately, the boards are monitored, and the lectures for the 2nd week were released a bit early to help with this. 

    I do wish the lectures had a more theoretical founding. This is one of my pet peeves about unix world in general, although I don't think I'm alone. Nothing seems to ever be related to anything else. Although there is a full lecture on the history of R, it goes through where R was developed, who developed it, but not what it's theoretical bases were or why it was made the way it was. With every language there are underlying assumptions about what forms of data are important and how it should be saved and treated. Understanding these can make learning and using the language much easier. No such luck here. Is read.table fundamental? How does it relate to read in unix or C or other common languages of the early R days? This seems like arcane knowledge, but it is the kind of thing that forms an actual education. If programming languages aren't taught as isolated functions and control characters, but as a historical web of intellectual developments, students have a framework for learning. I'm not sure Professor Peng, as a statistician, has this framework himself (it isn't encouraged outside of liberal arts schools), so I'm probably just shouting into the wind. 

    If you are used to learning in a classroom, the MOOC may seem very unusual. It is somewhere between independent learning and actually going to class. So far, the materials presented in the lecture have not been enough to even get full marks on the 1st quiz. (There's a gotcha about vector recycling in R which will not be apparent unless you either know it already, get lucky, or try it at the R command line.) The best way to learn is probably to have the R command line open and try things out during lecture. Even using the command line is not made easy, however, because the lectures are not structured as problem, discussion, resolution, but rather as a list of things you might find useful. You have to come up with the examples yourself, and most people will need outside help. 

    So, many students are pooling expertise on the discussion forum to figure it out. This includes posts that benchmark possible solutions, and a lot of hints on how to get started. The great majority of answers are helpful and supportive. Perusing the forum will make the assignments doable. 

    So overall, I'm not particularly impressed with the lecture style or structure of the teaching and information. On the other hand, I'm impressed with the breadth of the emerging student community. The assignments are challenging, and the outcome should be a reasonably good understanding. This learning won't come from the lectures, though. In other words, follow the excellent suggestions on how to succeed in a MOOC. 

    And if you think that my complaints must be unique to this course, try looking at this blog post. 






    Saturday, 3 September 2011

    Backyard Archeology

    We fell in love with BBC's Time Team. The idea that you can dig pretty much anywhere in Britain and find some really interesting archeology. It's fabulous. Wow. Now I do know that lots of people have been watching the show for ages, but, coming from California, it's a real eye opener. You mean the Romans and the Middle Ages and the Neanderthals all left stuff behind in people's back yards?

    Yes, but not mine...

    Or at least not in this hole.

    We did, however, find:
    Broken pottery-- mostly modern.
    (A few pieces may date back to the 1800's)
    Dog skeleton
    Broken window frames
    Bicycle -- with hand brakes, and balloon tires, so not ancient
    Meerschaum pipe pieces -- Unknown ages



    The house was built in the 1870's, and this looks mostly like the detritus of people living in it and remodeling it since then. We didn't find any indication of remains from the Mortlake tapestry works, the brewery or the pottery works that were nearby. Those could have dated back to the 1400's, easily. There also didn't seem to be anything related to St. Mary's Church. We're rank amateurs and didn't do a thorough job, so there's no guarantee we didn't miss something, but it seems reasonable to say that this patch was market gardens until Victorian times.

    We hit Thames River sand about 4 ft down. It all looked geological from then on down. The hole reached 6ft in the center, in the end. We dug some of the sand out for the sand box and re-filled the hole. We added manure from a local stable and voila - a new vegetable patch!

    On to the next project...

    Thursday, 21 January 2010

    A little sunshine on vitamin D vs Sunscreen

    I was a bit shocked when a friend of mine said she puts sun screen on every morning as a face cream. It would be a fine thing to do someplace where there is lots of sunshine, or at least lots of UV light, but she lives in Seattle; it’s so far north that there is no way of getting a sunburn during the winter months. There just isn’t enough light.

    It got me thinking about how sunscreen has been sold as this necessary thing. We have to block the harmful UV rays. They might cause cancer (which can usually be easily removed) or wrinkles (which we get anyway) or skin damage. Of course, we also know that vitamin D is necessary for human health, and we’re currently learning that it’s not just for kids and growing bodies. It’s necessary for a healthy body - and lack of vitamin D has now been linked to heart disease, cancers, auto-immune diseases and what most of us would call general health. It’s been in the news quite a bit recently, and there’s a selection of the information here.

    So how do we get vitamin D? Well, either we eat LOTS of fatty fish (i.e. live like an Inuit), we take vitamin supplements, or we make sure to spend some time in the sun. Not weak sun, not filtered winter sun, not even cloudy day spring sun, but real, honest to goodness bright sunshine.

    Now, I do know that UV light causes skin cancer, particularly in pale skinned Europeans who spend time outside in sunny climes. Caucasians living in Australia really should take care. Even in northern climes, I do not mean that you should never use sunscreen. I do mean that everyone needs to figure out their own balance based on where they are, who they are, and their lifestyle. Even if your skin is black as night, you need to wear some sunscreen, sometime -- like on a sailboat near the equator, for instance.

    So how do you know when to put on sunscreen? Well, first make sure you’re getting your vitamin D. If there isn’t enough UVB to bring up your vitamin D levels, then there isn’t enough to cause skin damage, either. Some dermatologists will say that any UV is bad because it can cause changes in DNA. UV light does change DNA, and I don’t dispute the molecular effects. I just think that we are highly developed organisms, and our bodies can repair slight damage. As scientists, we don’t yet fully understand our own biology, so, for the moment, I’d rather get my vitamin D the way my ancestors. A daily daily regimen of supplements and sunscreen is a relatively new development. These products don’t have extensive epidemiological studies over decades to back them up. Frankly, given the way that sunscreen ingredients are tested, I would not be surprised if they are eventually shown to be as bad for your skin as a *little* UV. After all, skin cancer doesn’t typically show up until over 40 years of age. No sunscreen has been tested for that long, at least not without changing its formulation!

    To figure out how much time in the sun you need to get enough vitamin D, have a look at the calculator at the Norwegian Institute of Air Research. They use 30 ug/L in the blood as a healthy level, but it's been shown that your body will find a healthy equilibrium. Even doubling the exposure time needed for that amount of vitamin D won’t cause a sunburn. You need about 4x the vitamin D exposure to get skin damage. The calculation depends on your skin tone (how easily you burn), how much UVB is coming in from the sun, and how much UVB is reflected from your surroundings. To figure out how much UVB is coming from the sun, it needs to know your latitude, the time of year, and the weather conditions. It assumes that you have 25% of your skin exposed to the sun, which is a bit much for me, since I never go sleeveless.

    Since I live in London, time of year, time of day, and weather are the deciding factors in whether or not to wear sunscreen:

    At noon on a sunny day in June or July, it takes about 6 minutes of sunshine to get enough vitamin D, but in January and December, there isn’t enough UVB available to make the vitamin at all, even in full sun at noon. There’s also variation during the day: If I get my sun during the school run at 9 am and 3 pm, I cannot make nearly as much vitamin D as if I go out briefly at lunchtime. Even at the height of summer, though, if I'm wearing a long sleeved shirt, it will take 12 minutes to get my vitamin D at noon. If the skies aren't crystal clear, it will take longer. So I think I can sit in the garden with my sandwich for a few minutes without reaching for the sunscreen first. It's probably good for me! On the other hand, if I'm at the beach, at high altitude, or on snow I'll reach for the sunscreen. Reflected light adds up, and I've been burned before under those conditions.

    A lot of websites try to provide ‘one-size-fits-all’ advice about sunlight, but it really isn’t appropriate. I’m fair skinned, so I need less sunlight to make vitamin D than someone with darker skin. Moreover, if I add up the periods of sun exposure during a typical school day, it looks like the fair-skinned children would benefit from sunscreen on some days where the darker-skinned children would be getting a barely sufficient amount of vitamin D. Since vitamin D deficiency is well documented, particularly in children with dark skin, and since this group is not at risk of developing skin cancer while living in England, they should not be told to put sunscreen on every morning. Unfortunately, the current dogma is UV=bad, and it will probably stay that way for a while.

    Monday, 7 December 2009

    Schroedinger's chat

    is when you invite someone to a skype chat and they haven't replied yet....

    Tuesday, 1 December 2009

    electricity footprints

    About those fuels...

    I was trying to calculate my carbon footprint with various calculators on the web. One of the main differences between the calculations is due to how electricity use is counted. In the UK, where most electricity is produced by coal burning power plants, we produce 0.537 kg CO2  for every kWh of electricity. In California, PG&E estimates 0.238 CO2 / kWh. That's more than a factor of 2! Why the big difference?

    You don't have to go very far to find a lot of statements like 'natural gas produces less CO2 than coal'. This bothered me because it seemed like both were carbon compounds from more or less the same source (plants and animals of the carboniferous, mostly), and combining them with oxygen shouldn't be that different. But it is.

    It turns out that the main difference is due to breaking carbon-hydrogen (C-H) bonds and forming hydrogen-oxygen (H-O) bonds.

    When these fuels burn, the carbon and hydrogen in them combine with oxygen to form CO2 and water.

        fuel + O2 ---> CO2 + H2O

    The different grades of coal have different ratios of carbon to other atoms, but commercial coal used in power plants (bituminous coal) is at least 70% carbon. The rest is mostly water. Wikipedia lists sub-bituminous coal as 6% or less hydrogen. Moreover, if this hydrogen is already tied up in water, it isn't available to be burned.

    In contrast, methane, the main component of gas delivered to my water heater, has a formula CH4. Each mole contains 12 g of carbon and 4 g of hydrogen, so it is 25% hydrogen by weight. Also, that hydrogen is all bonded to carbon, so it is available for burning.

    The result is that a power plant burning coal produces something like 1.2 kg CO2 per kWh of electricity, while a power plant burning natural gas only produces 0.7 kg CO2 for the same energy output.

    Now add in a few wind farms, solar projects, and a hydroelectric dam and the differences between the UK and California electricity footprints make a bit more sense.

    Too bad my computer doesn't run on natural gas.