Showing posts with label Demonstrations. Show all posts
Showing posts with label Demonstrations. Show all posts

Monday, March 18, 2013

R: Reproducible Research - Watermarks in Plots

Research is an iterative process: data is constantly being acquired, cleaned, and finalized while the analysis is being conducted. Since we work in a constantly evolving data and program landscape, version control is extremely important when programming, making research decisions, drafting reports, and writing manuscripts. Thankfully, reproducible research techniques are making this easier, which can mean less version control headaches, and more productivity.  Watermarks are a quick way to improve version control and better communicate the scope of results.

The problem: Images can be abstracted and embedded, obscuring their context or data provenance. This can lead to preliminary, dummy, or simulation results being confused for a final result or privileged/confidential results being accidentally disseminated.

The solution: Image watermarks can’t be easily removed by cropping and embedding by the lay user, don’t obscure results, and provide information about data provenance and information sensitivity. These can be added in your software package of choice, so that no post-processing of images is necessary, and can be easily suppressed in final results by flag variables.

The tools:
  • mtext, text: add text to margins (mtext) or within the plot (text)
  • grconvertX, grconvertY: find coordinates in a device, independent of axes
  • rgb: create translucent text from RGB values
 Let's save our results in a date stamped.png format: first we open a new device, use the generic R plot function, and use text and mtext to add watermarks in the image.

run.date <- format(Sys.Date(), "%m-%d-%Y")
    png(paste0("Watermark 1 - ", run.date, ".png"))
    plot(rnorm(100))
    text(x = grconvertX(0.5, from = "npc"),  # align to center of plot X axis
       y = grconvertY(0.5, from = "npc"), # align to center of plot Y axis
        labels = "CONFIDENTIAL", # our watermark
        cex = 3, font = 2, # large, bold font - hard to miss
        col = rgb(1, 0, 0, .2), # translucent (0.2 = 20%) red color
        srt = 45) # srt = angle of text: 45 degree angle to X axis
    # Add another watermark in lower (side = 1) right (adj = 1) corner
    watermark <- paste("Data embargo until", run.date)
    mtext(watermark, side = 1, line = -1, adj = 1, col = rgb(1, 0, 0, .2), cex = 1.2)
dev.off() # close device, create image


The result:
 

Watermarks in ggplot2 and lattice:

Monday, March 4, 2013

What's Ailing Introductory Statistics?



Introductory courses are the most important in any academic department. They are often a student's first and only exposure to a discipline, and reach the widest audience of any course in the department. A bad first impression can not only cheats people out of a potential career option, but can leave them with a lasting aversion to an entire field. Introductory statistics has the added importance of being the basis of every scientific field – no pressure there.

So why are there so many introductory statistics horror stories? Introductory statistics has a high conceptual overhead, but very little computational demands – you can get very far with basic arithmetic, and even further with a little calculus. Without a firm grasp of the conceptual basis, introductory statistics can easily become a vacant exercise in arithmetical procedures.
How do we keep introductory statistics from becoming a rote arithmetic exercise, devoid of all its utility? In my experience, there are a couple concepts that people don’t seem to be grasping and retaining:
  • What is randomness? Where does it come from? How does it fit into research?
  • What’s the difference between probability and statistics?
  • What does a null hypothesis mean? Why can we only collect evidence against it?
  • What are Type I Error and Power? What’s the intuition behind test statistics and null distributions?
There are a lot of challenges in constructing a good introductory statistics curriculum. Lots of introductory statistics books are garbage, and developing good lecture material takes a lot of time, which is often in very short supply. We have a wide audience to reach, from the next generation of statisticians to those who just care about degree requirements and letter grades – our job is to convert as much of the latter as possible into the former. This is a formidable task, but not an impossible one.

Fortunately, the tools available to teachers are evolving. I think R has the potential to make a huge impact in statistics education at all levels, for a number of reasons:
·       Free – students can use it outside of computer labs and after they graduate.
·       Open source – makes tinkering easier, and that’s the best way to learn a great deal in a limited time.
·       Community – Users groups, bloggers, free open courses, and more, all lending support.
·       R Markdown – weave R into blogs, labs, presentations, and such. When people are copying from a blackboard or slides, they’re not spending as much effort listening.
·       Shiny – This seems like an incredibly powerful tool for creating and disseminating interactive didactic tools which previously were only accomplished using JS.
R isn’t the only game in town, but I do think it’s the best way forward for students.