Thursday, November 14, 2013

Fernando Perez: An ambitious experiment in Data Science takes off:...

Fernando Perez: An ambitious experiment in Data Science takes off:...: Today, during a White House OSTP event combining government, academia and industry, the Gordon and Betty Moore Foundation and the Alfred P...

Wednesday, September 4, 2013

Those funny offset numbers in matplotlib

If you are a heavy matplotlib user, you are bound to have seen the funny offset numbers in the top left of the plot window:

They are obviously there to help the viewer focus on the level where the numbers are really changing, removing the area where there's no change happening.

But I am claiming that due to pattern recognition, there are quite a few cases where this confuses more than it helps. In this example I (and the people in my team) are used to see 5-digit numbers and it takes quite some time to figure out here, that these are indeed 5-digit numbers.

Therefore I researched how to switch this behavior off.


First, one imports the ScalarFormatter class from the matplotlib.ticker module:

from matplotlib.ticker import ScalarFormatter
Then, one creates a formatter object with the use of offset numbers switched off:

y_formatter = ScalarFormatter(useOffset=False)

Finally, you apply it to an axis object that you either receive via the fig.subplot() command, via plt.gca() (acronym for Get Current Axis) or you catch it when it is being returned after a plot command:

ax.yaxis.set_major_formatter(y_formatter)
There you go, hope this helps someone.

Here is the stackoverflow issue that helped me to find the solution.

Update (2013-10-20) :

An easier way is to catch the axis object from the plot command and apply the following command:
ax.ticklabel_format(useOffset=False)
I initiated a github issue to have this included in matplotlib, which has been responded already with a solution, so this will be configurable in the future, yay!

Update 2, same day:
Weird, I thought I had the above shortcut working at some time, now it doesn't. If anyone knows the circumstance under this can work and can not, please comment.

Wednesday, June 5, 2013

polyfit

A follow-up to the previous post.

Polynomial fitting is also very easy with the numpy packages polyfit and poly1d.


In [196]: x = range(100)
In [197]: y = randn(100)
In [198]: plot(x,y)
Out[198]: []
Here I am asking polyfit to fit me a 2nd degree polynomial.
In [199]: polyfit(x,y,2)
Out[199]: array([-0.00018313,  0.01669275, -0.09621319])
The polyfit function returns the polynomial coefficients in a list.
If I want to use them directly as a fit function, just embed them in a new polynomial object:
In [200]: fitfunc = poly1d(polyfit(x,y,2))
In [201]: plot(x,fitfunc(x))
Out[201]: []
Saving the plot like this
In [202]: savefig('/Users/maye/Desktop/blog_polyfit.png')
and looks like this:


Friday, May 31, 2013

Polynomials with Python

Seriously, can it be any easier? ;)

If you are not in a pylab session, import the module like this:
In [148]: from numpy import poly1d
 Otherwise, just "import poly1d" should work.
Now let's get a polynomial for the coefficients of [3,2,1] (always in decreasing order!):
In [149]: p = poly1d([3,2,1])
Printing it provides a semi-analytical printout:
In [150]: print p
   2
3 x + 2 x + 1
Applying new x values to it is easy, because the poly1d object is a function:
In [152]: newx = linspace(0,10,10)
In [153]: p(newx)
Out[153]:
array([   1.        ,    6.92592593,   20.25925926,   41.        ,
         69.14814815,  104.7037037 ,  147.66666667,  198.03703704,
        255.81481481,  321.        ])
Lots of other things are possible with this object. IPython's object inspection makes it easy to discover them:
In [154]: p.
p.coeffs    p.deriv     p.integ     p.order     p.variable

In [155]: p.deriv()
Out[155]: poly1d([6, 2])
In [156]: pderiv = p.deriv()
In [157]: print pderiv
6 x + 2
Roots for this polynomial can be either determined by the roots function that is imported in a pylab session (or importable like from numpy import roots)
In [158]: roots(p)
Out[158]: array([-0.33333333+0.47140452j, -0.33333333-0.47140452j])
In [159]: p.r
Out[159]: array([-0.33333333+0.47140452j, -0.33333333-0.47140452j])




PS: One of these days I really have to find out how to do code high-lighting in Blogger, or, preferably, go all the way and do IPython notebook posts.

Friday, February 25, 2011

A scientific Python starter

As requested by some, here is a list of important websites, doc-sites and modules for a successful start in Python.

1. Getting Python
If there's one problem with Python, it is how some of the available modules depend on other modules to be the right version. Therefore I recommend wholeheartedly to install big packages that include all the modules you need as one big chunk.
I personally made good experiences with http://www.enthought.com., they provide academic licenses for free and also offer 64-bit version free on personal email request (they did for me at least).
There are other packages like PythonXY, but I think Enthought is the only one, that creates a package for Mac, Windows AND Linux.

2. First steps
The tutorial on http://docs.python.org is really excellent! When I worked through it years ago, I continued the next day with Python GUI programming tutorials and was coding my own graphical user interfaces with Python within 1 day! The clarity of Python enables this, I believe. Haven't learned a language before that is that easy to learn (and I used BASIC, C, C++, FORTRAN77, Java and IDL).

Read the tutorial until inclusive section 5 (Data Structures, up to here is a MUST!) and then you can go on for now to the scientific tutorials, but if you feel puzzled sometime later you really should continue this tutorial at least until section 9 inclusive to see what classes are all about and, important, how to read files in section 7).

3. A warning for IDL switchers:
Before we go on to Sciency stuff in Python, a warning:

If you do this in IDL:
a = [1,2,3]
b = a

then you have a new array b, copied from a, and you can do what you want with it without influence on a.
This has advantages for ease of use, but makes your IDL very fast very slow, because you carry your data around multiple times after a while.
As Python is a real programming language, it tries to be memory efficient and it does so, by avoiding copies if not explicitly asked for.
So in Python when you do:

a = [1,2,3]
b = a

then you don't get a copy but a link to the a-array. So if you do b[0]=4, 'a' has changed as well! (Try it out!)
So what do you do, if you really want a copy without changing the original?
Well, you ask for a copy:
b = a.copy()

4. Numpy and Scipy
Now on to the most import Python modules for scientists called 'numpy' and 'scipy'.
Scipy uses Numpy so they are closely linked.

The documentation for both can be found here:
http://docs.scipy.org/doc/ or just start some browsing around at http://www.scipy.org it's very interesting.
I recommend to start with the last linked document, the Scipy Reference guide on the docs page.
Why? Because it has a tutorial for Numpy, Scipy and also introduces you to the plotting module 'matplotlib' at the same time! So really worth reading.

Before I leave you alone in your python adventures (i put all links together again at the end), one more import comment for beginners confusion:
An import difference between numpy arrays (which look a lot like lists) and the original Python lists.
Python lists can take anything, so a list like this is possible:

myList = ['aString', 3.1415, (atuple, atuple)]

but the Python compiler needs some time to be able to deal with all this different things, so if one wants efficient arrays that only deal with the same type of elements at a time, then one needs numpy arrays.
So one important difference is the type of elements (many in Python lists, only one for numpy arrays), the other is how they react on mathematical operations.

For Python lists it can be quite handy, that it is possible to do:
[3,4]*3
to get
[3,4,3,4,3,4]

For writing text files in certain formats this is quite useful.
But for scientific calculations that doesn't make any sense of course, that's why numpy arrays do exactly what you expect for this:

import numpy as np
a = np.array([3,4])
print a*3

and you get
array([9, 12])

What you saw here as well, that the np.array function can transform normal Python lists to numpy arrays without problems (that also works the other way around in case you need it).

Ok, have fun in Python and don't shy to ask questions, the Python community is very helpful, I think because everybody is so happy about it. ;)

Here are the promised links of all important Python websites for scientists:

http://www.enthought.com
(to get a full scientific Python environment, that runs exactly the same on Win, Mac and Linux)
http://docs.python.org
(Overview of only the core Python stuff itself, home of the Tutorial !)
http://www.scipy.org
(the home of Scientific Python, very nice to browse through...)
http://docs.scipy.org/doc/
(the docs page of scipy)
http://matplotlib.sourceforge.net/
(The home of the matplotlib Python module, a very powerful plotting library. I always look at the gallery to find what I need and one get's the example code for each graph! Very helpful)

Maybe one last tip: The enthought environment installs an Examples folder as well, so you get a lot of example code installed on your computer, for the times when you are not online!

Now, enjoy!
I was getting a lot of MDS related errors in my system logfile (visible via the Console.app in Mac).
Like these:

Feb 25 12:42:00 paradigm /System/Library/Frameworks/SecurityFoundation.framework/Versions/A/dotmacfx.app/Contents/MacOS/dotmacfx[89682]: MDS Error: unable to create user DBs in /var/folders/PO/PO-GUydDF4mpVHTQxsRatE+++TI/-Caches-//mds
Feb 25 12:42:02 paradigm /System/Library/Frameworks/SecurityFoundation.framework/Versions/A/dotmacfx.app/Contents/MacOS/dotmacfx[89699]: MDS Error: unable to create user DBs in /var/folders/PO/PO-GUydDF4mpVHTQxsRatE+++TI/-Caches-//mds
Feb 25 12:42:34 paradigm /System/Library/PrivateFrameworks/GoogleContactSync.framework/Versions/A/Resources/gconsync[89761]: MDS Error: unable to create user DBs in /var/folders/PO/PO-GUydDF4mpVHTQxsRatE+++TI/-Caches-//mds
Feb 25 12:55:34 paradigm /System/Library/Frameworks/PubSub.framework/Versions/A/Resources/PubSubAgent.app/Contents/MacOS/PubSubAgent[91445]: MDS Error: unable to create user DBs in /var/folders/PO/PO-GUydDF4mpVHTQxsRatE+++TI/-Caches-//mds
Feb 25 12:59:30 paradigm /Applications/Safari.app/Contents/Safari Webpage Preview Fetcher[91954]: MDS Error: unable to create user DBs in /var/folders/PO/PO-GUydDF4mpVHTQxsRatE+++TI/-Caches-//mds

and many more.
I thought that can't be good for performance so I digged around a bit and found that the caches of the kernel modules can become a bit problematic after security updates from Apple.
But there is a very simple remedy: Boot in safe mode, because then the caches are automatically deleted and renewed at the next boot.
To do that you keep pressing the SHIFT key while rebooting. Once you are fully booted into safe mode, just restart again and you will see (if you are as sensitive as me to the snappiness of a GUI) that things move a tad snappier now, like it was when the Mac was still fresh. ;)

Good luck!

Find out on which computer you are 'python-ing'

Often, data is stored at different locations, depending on which computer the code is running.
To find out one can use the sys module like this:


import sys
if sys.platform == 'darwin':
    base = '/Users/aye/Data/hirise/'
else:
    base = '/processed_data/'

Thursday, February 24, 2011

Great Python resource!

This site diggs through many useful stuff of the Python standard library.

http://www.doughellmann.com/PyMOTW/contents.html

Python lists as function parameters

I often stumble about something like this:
A function needs 2 parameters like
myfunction(a,b)
and i have them, but inside a list 'myList'=[a,b].
Sure I could write
myfunction(myList[0],myList[1]),
but where's the famous Python beauty in this, right?
Until I realised that this is what the *args thingie is for that I never used before.
Damn, all those typing hours lost when I typed the explicit unpacking of lists!!!
Now I can just do
myfunction(*myList)
and all is beautiful again! ;)
Thank you, Python

Tuesday, February 15, 2011

Deutsche Service Wueste, jetzt auch in der Schweiz!

Schade, dass die deutsche Service-Wueste schon in die Schweiz migriert ist:
Kreditkarte bei der PostFinance am Schalter bestellt.
Karte kommt, Name falsch, wir zurueck zum Schalter um es zu aendern:
"Nein, das muessen sie jetzt mit der Kreditkartenfirma ausmachen, da haben wir ueberhaupt keinen Einfluss darauf."
Warum bezahl ich dann fuer eine PostFinance Kreditkarte?
Darauf wurde beim Vertragsabschluss natuerlich nicht hingewiesen, das man eventuelle Probleme dann woanders loesen muss. Echt traurig sowas...

Tuesday, June 1, 2010

ExWi Unite! ;)

The cafeteria bosses want to kick out Margrite (?, the red haired friendly older one) 6 months before her pension due to 'economic constraints' even so they wrote elsewhere, that they made good earnings in the last year!! WTF? Obviously a case of 'kick u out, get somebody cheaper=younger'. I wondered, if anybody would join a signature list to the cafeteria bosses, to stop this bullshit! Also, we could use this opportunity to complain about some 'worseprovements' they introduced after the change of ownership, like canceling the speaker phone etc.
Who is with me and would sign something like that?

Sunday, January 17, 2010

Intuitive division in Python

When one is hot-coding some idea, it is all about writing down an idea fast, not pretty (even so it's quite hard to write ugly code in Python due to its indental paradigm), not clean or precise, just trying out an idea.
It is then that it could happen to overlook that a division of 2 integer numbers is forcing an integer solution of this calculation (the underlying C-paradigm for divisions in Python is enforcing this), so 3/4 is 0, not 0.75, because one forgot to write something like 3/4.0 to enforce the floating point division.
Many errors in un-counted Python (and C) classes were caused by this intricacy, maybe even some good ideas have been discredited because the code produced weird results. And because of the caused frustrations, the Python inventors have decided to get rid of this, beginning with the 3.x versions of Python, so 3/4 is actually 0.75.
And because this was such a great idea, it was decided to make it available as an option in Python 2.x (before you judge this as a minor change, admit to yourself how often you have been caught by this bug!?).
So, to enable this 'safe' division in Python 2.x code, you have to write the following line before all other imports:
from __future__ import division
(thats 2 underscores in front AND behind the future)

Enjoy your hot-coding now! ;)

Thursday, December 31, 2009

Macbook Memory and Swap files tip

If you are like me, meaning (in this special case):
  • don't like to reboot your Macbook and have it running for weeks,
  • from time to time have some heavy computing sessions with dozens of progs running,
  • realising, that OS X Snow Leopard still does not manage its memory efficiently, by for example releasing the many and huge swap files created for these heavy computing sessions after these programs have been closed, even so your memory monitor shows you, that you have plenty of free RAM left
  • and you still don't want to reboot to remedy this problem (= OSX still using the swap files, even so it has free RAM).
you could try the following in a Terminal (found here: Link, but I save you from reading all that.. ;) )
  • sudo purge
  • sudo killall -HUP dynamic_pager
From the manpage of purge:
Purge can be used to approximate initial boot conditions with a cold disk buffer
cache for performance analysis. It does not affect anonymous memory that has been
allocated through malloc, vm_allocate, etc.

dynamic_pager is the prog responsible for the swap files on the Mac. This command is what you could call a 'friendly' kill, because it asks to 'HangUP' on the program, not just to terminate it.
I have tried it and so far it looks very efficient, and saved me already 1 or 2 reboots after only 1 day (of heavy computing.. ;)

Enjoy and Happy New Year to all my 2.3 readers.. ;)

Wednesday, December 9, 2009

Mac data analysis tip of the day

So, you have a lot of plots/images in a folder, but you now only want to look at some of them.
Fortunately, in a Terminal, you could easily 'grab' the right ones by shell globbing (=a form of pattern matching for you non-geeks out there), meaning:

Let's say you have plots/images ending in *.profile.png and *.histo.png, so you easily 'grab' only the profiles by ls *.profile.png, nothing easier than that, but how do you get Mac's Preview now to show this choice to you?

Well, as so often, Mac OS just does the right thing:
You type: 'open *.profile.png' and you get all plots/images nicely put into one Preview window.

A slight setback, but logical:
So far they are treated as single documents, so if you just hit Cmd-P to get a print of all your analyis work, you only get offered to print the currently-selected one.
Even a Cmd-A to select all does NOT change this completely, but is on the right track: After selecting all images you can use the File->Print Selected Images menu entry (Or Alt-Cmd-P) to do just that.
And of course with that you can now print again into a new PDF file to have your plots nicely saved all together.

So, I would call this 'efficiency-with-a-hickup', but nevertheless a much more workable and faster solution to produce a pdf out of a shell-globbed (you don't wanna mouse select every 2nd or 3rd image in your folder now, would you?) selection than possible under pure *nix or *doze..
(I timed it, took me rougly 20 seconds to get the pdf, and only because I couldn't decide nor type the resulting filename ;)

PS.: The way to combine/edit PDFs has slightly changed in SL, get the details here.

Wednesday, November 4, 2009

Using the space station: where does the US go from here?

With the International Space Station nearly completely assembled, attention now turns to how to best utilize it. Taylor Dinerman explains how that will depend on how much access scientists will have to it once the shuttle is retired.

.
.
.
.
A lot!

Monday, June 22, 2009

A2DP on iPhone with P590

Just tried to connect my Plantronics P590 to my iPhone 3G that I updated last week to 3.0.
I am running this for 20 minutes now, and so far I love it in total.
Some points I noticed:
  • Quite some bass coming from the P590.
  • There are some problems, some songs show skipping like an old CD player, and even more enerving is a pitch-down effect (not good for musical ears).I found in some forums that there are indications that this only happens for self-imported tracks and not for iTunes tracks bought on iTunes? I would find this a weird selection effect and would propose that this is caused by differences between mp3 (less efficient) and ac3 (more efficient)?
  • Volume control is after pairing with headset blocked on iPhone and solely done via the headset
  • With the P590 next and previous track selection does NOT work.
  • Even high data rate songs in m4a format (lossless) play nicely via the headset. They /might/ sound a bit better, but that might be wishful thinking as the bluetooth encoder might reduce that quality all away
So overall I am very pleased that this is finally in the iPhone OS, pity that next track does not work with the P590, but I am looking anyway for a new Stereo Bluetooth headset, want to try in-ear plugs this time.