Sunday, July 10, 2016

Keras pre-trained models how-to

In keras Functional API there are few examples for  combining models in a different way, but if you want to mix and match parts of models , like combine new classifier-head with a headless pre-trained model, you might encounter some issues. Let's explain how it works behind the scenes.

A saved model is a combination of three things:
  1. Layer definitions like "Dense" layer with it's regularizes ,output_dim and name
  2. Layer graph: which layer is connected to which other, and from which direction (in/out)
  3. Layer weights: for layers which do have them (dropout do not , for example)


When you call: model.to_json() , you get the first two. When you call model.save_weights() you get the third, where the key is the layer-name from (1)
If you want to save and load the exact same model, just call m=model_from_json(json) and then m.load_weights(file).  Easy.
If you want to create a head-less model, which does not contain the last few layers, you will have to re-create the model by code, create the connectivity and then use your own function to read the weights file(.h5).  The h5 API is quite clean and it's an easy-enough task.

If you want to create a new model from a headless-one and a new head you created, again you will have to define (1) and (2) in code, and then manually load-weights per layer.

Code for connectivity between layers
# assuming this was called already: graph = Graph() 
# graph.add_input(name='input1', ndim=2)

new_layer = (Dense(32, 4)  #define layer
graph.add_node(new_layer , name='dense1'         ,input='input1') #connect layer to input
#note that if you want two inputs, just use instead  ,inputs=['input1,'input2']

# and in the end .... graph.add_output(name='output', input='...')

The functional API hides the Graph API and allows you to create a layer and connect it to a node in one line
#assuming input- Input(shape=784,)
Dense(32,4)(input)  

It's important to remember that the node connectivity is actually added to the layer instance itself, this sadly means you can't re-use the layer in multiple graphs with totally different connectivity.
You will have to re-instantiate (redefine) a new identical, unconnected, layer.

FAQ:
"UserWarning: Model inputs must come from a Keras Input layer, they cannot be the output of a previous non-Input layer"
You can't instantiate a Model which does not start with Input layer, in other words, you can't just cut the few last layers of an existing model and call Model (input='other_input', output='...same')



Wednesday, July 6, 2016

Features regression/localization

There are many approaches, we shall start with the simple ones (which perform badly) and continue to those performing a bit better.

Loss via Linear regression

Combine all the features  into one loss function, and calculate it (diff of mouth + diff of nose + diff of eyes, etc).  The results are usually quite rough, so different cascade are suggested to fine tune the windows on which the CNN is running.

  • Facial recognition with no cascade: Tutorial with full code Lasgne 
  • Joint location with cascade: DeepPose - Run one DNN to get rough estimates of joint locations. cut a small window around each joint location and run the same network architecture (but with different parameters) on it to get finer estimation. Do it again one more time to get best results.
  • Deep Convolutional Network Cascade for Facial Point Detection
    Modification 1: abs on tan, instead of  ReLU, on some layers.
    Modification 2: locally-shared weights instead of globally shared weights in the Convolution layer.
    Modification 3: lots of networks structure:Train 3 networks with the same architecture (high1) to detect eyes region, nose region and mouth region.
    Then pass it to multiple shallow architecture (shallow1), again each trained by itself, and it only gets a small region. This one can only slighly modify output location, as we assume it is more accurate, but 'does not see the big picture.
    Then again pass it to multiple shallow arch...  with even smaller region.

Classification as Loss function

Cascade: This is usually used for bounding-box calculations on classification tasks.  As you already have a good classifier model, you want to re-use it.  This is done by cutting the image into multiple very small images, and checking each one of the sub-images for the existance of a "nose"/"eye". if so mark what there is there. If you apply it to the whole image, you can a rough estimate of where the nose is.

Detection
region proposal networks - faster rcnn
yolo - 

Heatmap as Loss function

The output of the network is a heatmap  (WxH pixels with intensity levels from 0 to 1), where few pixels around the feature are highlighred. This apears to provide better results, as it is a more "natual" calculation for a CNN.  Do not that regular classification models are shaped like a cone, with smaller and smaller layers till the FC result.  This architecture is not adequate here.
In theory, one may not need a cascade here.

Localization as side-effect

The CAM technique (Class activation mapping) can generate it automatically , as a side-effect of the attention model.


Appendix, Honorable mentions: Linear combination of the previous layers


 this is good for the whole face pixels, not as good for location of occulded features in the face which is fully not facing-the-camera.


Datasets and competion results for Object Segmentaion
Coco - Huge one.


Datasets and compettion results for Faces

HPDatabase 
Youtube face DB
TAU article TCNN for Facial Landmark Detection with Tweaked Convolutional Neural Networks. implementation in caffe


Datasets and competition results for human joints 

2016 - Coco key point challenge - new (July 2016) and probably the best one to use.  90K person instances labeled with keypoints (the majority of people in COCO at medium and large scales) and over 1 million total labeled keypoints.

The Human Pose Recovery and Behavior Analysis HuPBA 8k+ dataset , from cha learn

FLIC: Frames Labeled In Cinema contains 4000 training and 1000 test images obtained from popular Hollywood movies. The images contain people in diverse poses and especially diverse clothing. For each labeled human, 10 upper body joints are labeled.
Note: FLIC_Full contains more challanging(occulded) scenes. too hard to train from.
FLIC-motion-dataset
 includes short clips. in this case, maybe the motion will help estimate better.

Leeds Sports Dataset [12] and its extension [13], which we will jointly denote by LSP. Combined they contain 11000 training and 1000 testing images. These are images from sports activities and as such are quite challenging in terms of appearance and especially articulations. In addition, the majority of people have 150 pixel height which makes the pose estimation even more challenging. In this dataset, for each person the full body is labeled with total 14 joints.

The dataset includes around 25K images containing over 40K people with annotated body joints. The images were systematically collected using an established taxonomy of every day human activities. Overall the dataset covers 410 human activities and each image is provided with an activity label
Best results for MPII

2016 Convolutional Pose Machines in caffe/matlab ; Stack hourglass in torch
2015 DeepCut: Joint Subset Partition and Labeling for Multi Person Pose Estimation






Features regression/localization

There are many approaches, we shall start with the simple ones (which perform badly) and continue to those performing a bit better.

Loss via Linear regression

Combine all the features  into one loss function, and calculate it (diff of mouth + diff of nose + diff of eyes, etc).  The results are usually quite rough, so different cascade are suggested to fine tune the windows on which the CNN is running.

  • Facial recognition with no cascade: Tutorial with full code Lasgne 
  • Joint location with cascade: DeepPose - Run one DNN to get rough estimates of joint locations. cut a small window around each joint location and run the same network architecture (but with different parameters) on it to get finer estimation. Do it again one more time to get best results.
  • Deep Convolutional Network Cascade for Facial Point Detection
    Modification 1: abs on tan, instead of  ReLU, on some layers.
    Modification 2: locally-shared weights instead of globally shared weights in the Convolution layer.
    Modification 3: lots of networks structure:Train 3 networks with the same architecture (high1) to detect eyes region, nose region and mouth region.
    Then pass it to multiple shallow architecture (shallow1), again each trained by itself, and it only gets a small region. This one can only slighly modify output location, as we assume it is more accurate, but 'does not see the big picture.
    Then again pass it to multiple shallow arch...  with even smaller region.

Classification as Loss function

Cascade: This is usually used for bounding-box calculations on classification tasks.  As you already have a good classifier model, you want to re-use it.  This is done by cutting the image into multiple very small images, and checking each one of the sub-images for the existance of a "nose"/"eye". if so mark what there is there. If you apply it to the whole image, you can a rough estimate of where the nose is.

Heatmap as Loss function

The output of the network is a heatmap  (WxH pixels with intensity levels from 0 to 1), where few pixels around the feature are highlighred. This apears to provide better results, as it is a more "natual" calculation for a CNN.  Do not that regular classification models are shaped like a cone, with smaller and smaller layers till the FC result.  This architecture is not adequate here.
In theory, one may not need a cascade here.

Localization as side-effect

The CAM technique (Class activation mapping) can generate it automatically , as a side-effect of the attention model.


Appendix, Honorable mentions: Linear combination of the previous layers

 this is good for the whole face pixels, not as good for location of occulded features in the face which is fully not facing-the-camera.



Datasets and competition results for human joints 

FLIC: Frames Labeled In Cinema contains 4000 training and 1000 test images obtained from popular Hollywood movies. The images contain people in diverse poses and especially diverse clothing. For each labeled human, 10 upper body joints are labeled.
Note: FLIC_Full contains more challanging(occulded) scenes. too hard to train from.
FLIC-motion-dataset
 includes short clips. in this case, maybe the motion will help estimate better.

Leeds Sports Dataset [12] and its extension [13], which we will jointly denote by LSP. Combined they contain 11000 training and 1000 testing images. These are images from sports activities and as such are quite challenging in terms of appearance and especially articulations. In addition, the majority of people have 150 pixel height which makes the pose estimation even more challenging. In this dataset, for each person the full body is labeled with total 14 joints.

The dataset includes around 25K images containing over 40K people with annotated body joints. The images were systematically collected using an established taxonomy of every day human activities. Overall the dataset covers 410 human activities and each image is provided with an activity label


Best results:

2016 Convolutional Pose Machines
2015 DeepCut: Joint Subset Partition and Labeling for Multi Person Pose Estimation






Keras Installation on Windows - CPU performance


1. Install the basics.  start with this stackoverflow:

  • Install TDM GCC x64.  (need to add to begin of path? probably not)
  • Install Anaconda x64.
  • Open the Anaconda prompt
  • Run conda update conda (You might need to open-command-prompt with admin privalage)
  • Run conda update --all
  • Run conda install mingw libpython
  • Install the latest version of Theano, pip install git+git://github.com/Theano/Theano.git
  • Run pip install git+git://github.com/fchollet/keras.git
  • Update the cxx option in the env variables: set THEANO_FLAGS=floatX=float32,device=cpu,cxx=E:\\TDM-GCC-64\\bin\\g++.exe
  • Check that a simple Keras Hello-World is working .
2. Now let's improve performance. Here I use CPU and not GPU. If you have a good nvidia GPU, check keras cuda install guide.

  • Change your BLAS library : OpenBlas/MKL can give you 400% speedup. Download a good BLAS library and add it (the \bin folder) to the system path. I used openblas 2.0.14 and got a huge boost  update the THEANO_FLAGS env variable with it:
    set THEANO_FLAGS=floatX=float32,device=cpu,cxx=E:\\TDM-GCC-64\\bin\\g++.exe,blas.ldflags=-LE:\\code\\openblas\\bin -lopenblas

    I tried to use intel mkl library, which should be faster, but could not successfully configure it for keras. (If you did, please leave a comment....)
  • OMP: usually 10-20% boost
  • set OMP_NUM_THREADS=2  (benchmark X on your-system,  maybe 4 is better?)
    update THEANO_FLAGS:
    set THEANO_FLAGS=openmp=True,floatX=float32,device=cpu,cxx=E:\\TDM-GCC-
    64\\bin\\g++.exe,blas.ldflags=-LE:\\code\\openblas\\bin -lopenblas

    If you get this error :
    UserWarning: Your g++ compiler fails to compile OpenMP code. We know this happen with some version... then make
    Make sure your cxx is configured properly, but one some machine-configuration (my very-old desktop for example), I could not solve the issue.

FAQ:
InvalidValueError: InvalidValueError  ...
        type(variable) = TensorType(float32, (True, True))
        variable       = TensorConstant{(1L, 1L) of inf}
        context        = ...
  TensorConstant{(1L, 1L) of inf} [id A]

Check THEANO_FLAGS , mode should not be debug_mode.  if it does not help do "conda update -all"

Sunday, May 1, 2016

Python after C#/Java - Part 2 - Classes


Classes

class Account(object) :  # inherits object directly
     
     instanceCount= 0 # like static variable in java
     #this is similiar to java constructor, called automatically during creation : Account("assaf")                 def __init__(self, balance, someBool=true):
       self.Balance= balance#note:  no need to define them before, nice!!!
       self.SomeBool= someBool   #public style
       self._XProtected = 5 # one '_' prefix is like java-public, but the programmer ask politly.
       self.__Yprivate = 8  # two '_' prefix is like java-private. you 'll get exception when accessing.
       Account.instanceCount +=1

    #destructor (usually unused, like in Java)  called when no reference exists any more .
    #for example by using x=None,   or explicit calling del(x)
    def __del__(self):
         Account.instanceCount -=1 #static variable: note the usage of Account. and not self.
   
   #like the ToString method. optional, of course
    def __str__(self):
         return "balance: %f".format(self.Balance)
 
   #regular method
   def  deposit(self, x):  #note, when calling it "self" is not needed
        self.Balance += x
     

Usage:
account = Account("assaf")
account._XProtected = 666       #works, but '_' it means the programmer asked you not to do it
print( account._XProtected)    --> 666
print( account.__Yprivate)  --> AttributeError
print (account._Account__YPrivate) --> 8    #showing you there is no real way in python to defend against this. if someone want's, he can access it anyway



inheritance and static variables


class Counter:
    instanceCount = 0
    def __init__(self):
           type(self).instanceCount +=1 # and not Counter.instanceCount, cause it will be all sons.
    def __del__(self):
           type(self).instanceCount -=1

class Account(Counter):
      def __init__(self, x,y,z):
               Counter.__init__(self)  #explicit call is needed!

class MultipleInherit(Counter, Shouter):
        def __init__(self, x,y,z):
                 Counter.__init__(self)
                  Shouter_init__(self, y)

storage optimization (__slots__)

each instance has a built-in __dict__ hashmap, which contains all the dynamic memebers.  It means that you can always add members to a class, but it's heavier in storage on the RAM.
account = Account("assaf")
account.dynamicNewMember = 8  #works just fine

If you instantiate millions of these instances, instead of using the built-in __dict__, you can use a tighter static structure. Note: Don't optimize this way unless you have millions of instances.
Use __slots__ and define them before hand , by name , like:
def AccountWithLessStorage:
  __slots__ = ['Balance', 'someBool', '_XProtected', '__YPrivate']
   #everything else in the class is exactly the same, including usage in __init__
   

MetaClass , annotations and Reflection

java reflection-like operations is much easier in python,
instance.__dict__   # is a dictionary of the members (variables and methods), so calling the method append on a list can be done "in-reflection" very easily.
lst = [ 1 , 2  , 3 ]
non-reflection:   list.append(4)
reflection:           list.__dict__["append"](lst, 4) # the 1st parameter is for instance method is "self"     
Use decorators to wrap a method
Use metaclass for decorators like count call/timer/logging.



Monday, April 25, 2016

TensorFlow Installation


Tensor flow installation 

$sudo apt-get install oracle-java8-installer sudo apt-get install pkg-config zip g++ zlib1g-dev unzip wget https://github
Tried to install on GCE using two methods.
The first is the default installation instructions, when running the minst demo, it appears a bit slow (650-700ms)
Also tried this script.
Also tried to install from source: see original page

git clone --recurse-submodules https://github.com/tensorflow/tensorflow -b r0.7
$ sudo apt-get install python-numpy swig python-dev
$ sudo add-apt-repository ppa:webupd8team/java
$ sudo apt-get update
$ sudo apt-get install oracle-java8-installer sudo apt-get install pkg-config zip g++ zlib1g-dev unzip wget https://github.com/bazelbuild/bazel/releases/download/0.2.1/bazel_0.2.1-linux-x86_64.deb
 then google "how to install .deb"
./configure    --> say you don't have GPU  (my case. if it's not, you read yourself...)
bazel build -c opt --copt=-mavx //tensorflow/cc:tutorials_example_trainer


Jupiter notebook
look for their installation (pip..., including the dev too)
I also installed plots:
sudo apt-get install libfreetype6-dev libxft-dev
pip install matplotlib 
open port in GCE console:  gcloud compute firewall-rules create tcp8888 --allow=tcp:8888
run on the linux shell:  (ip 0.0.0.0 is a must, otherwise only local host will work)
jupyter notebook --ip=0.0.0.0 --port=8888 --no-browser
plotting:  if not visible, add this to the cell:  %matplotlib inline.   see a permanent solution here



installed via:
sudo pip install --upgrade https://storage.googleapis.com/tensorflow/linux/cpu/tensorflow-0.7.1-cp27-none-linux_x86_64.whl

$ python -m tensorflow.models.image.mnist.convolutionalSuccessfully downloaded train-images-idx3-ubyte.gz 9912422 bytes.Successfully downloaded train-labels-idx1-ubyte.gz 28881 bytes.Successfully downloaded t10k-images-idx3-ubyte.gz 1648877 bytes.Successfully downloaded t10k-labels-idx1-ubyte.gz 4542 bytes.Extracting data/train-images-idx3-ubyte.gzExtracting data/train-labels-idx1-ubyte.gzExtracting data/t10k-images-idx3-ubyte.gzExtracting data/t10k-labels-idx1-ubyte.gzInitialized!Step 0 (epoch 0.00), 7.2 msMinibatch loss: 12.053, learning rate: 0.010000Minibatch error: 90.6%Validation error: 84.6%Step 100 (epoch 0.12), 698.7 msMinibatch loss: 3.279, learning rate: 0.010000Minibatch error: 6.2%Validation error: 7.1%Step 200 (epoch 0.23), 670.1 msMinibatch loss: 3.503, learning rate: 0.010000Minibatch error: 12.5%Validation error: 3.6%Step 300 (epoch 0.35), 666.0 msMinibatch loss: 3.199, learning rate: 0.010000Minibatch error: 7.8%Validation error: 3.4%Step 400 (epoch 0.47), 658.3 msMinibatch loss: 3.239, learning rate: 0.010000Minibatch error: 10.9%Validation error: 2.6%Step 500 (epoch 0.58), 657.4 msMinibatch loss: 3.283, learning rate: 0.010000Minibatch error: 9.4%Validation error: 2.6%

Wednesday, April 20, 2016

Python over C# / Java

Most people say the python is elegnet over Java. They are correct :)

Let's go over the basics, and I believe you will agree in the end...


Syntax


Curly brackets for scopes are gone.  we use the indentation instead.
if (x>10) :
  print("somewhat big")
  if x>50 :
     print "big"

Strings can be anything between 'x'  or "x"  or """x""" , in the later case new line in the editor is translated to new line (\n) in the string, so what you see is what you get.


and, or and not  are the actual operators(!) not &&,||,!.
"in" replaces the contains operator: if  'a' in ('a','b','c')
[not so great?]  instead of max=(a>b)?a:b  ,  use max = a if (a>b) else b
very easy way to wrap a function with a decorator on function like @args-check




loops

while and for loops have an optional "else" clause which will be triggered only when the loop exists normally on the condition failure (and not in "break").  This is very useful and save ugly code in end of java loops where you are unsure what caused you to pass the loop.

for loops are always using "in"
for odd in range(1,10,2) : print odd  #   from 1 , as long as <10 ,jump of 2
for item in list : print item
for key in dictionary: print key
for value in dic.itervalues() : print value

built in data structures

no arrays!  python asks us to use a higher level structure, like list.
list = []  is like Java ArrayList of Objectstuple=()  is immutalbe ArrayList of Objects

some syntactic sugar for all sequences data structures (list,tuple and string too!):

list[0]
list[0,3] returns a slice with elements 0,1,2
list += [ "hello" , "world"]   #adds all elements of the second list
if  "hello" in list : print "world"

dict = { "key1" : "value1" , "key2" : value2 }
len(dict)  # return 2
dict.update( { "key1": "updated1" , "key100":"value100"}) #update by another dict
if "key100" in dict:  print ("this was expected")
dict["key100"] = "updated100"
del(dict["key100"])  #both key and value are deleted
dict.items = list of key,value tuples   [ ("key1","value1", "key2,value2")]
dict(dict.items)  # transform a list with tuples into a dictionary

set is similiar to hashset
mySet = { "key1" , "key2" , "key3")
mySet = set(["key1","key2","key3" )


Better than the ugly :  if (list!=null && list.length>0)

if  list :   #nil=false and empty list = false too
  #do something with the list
else
  print("empty list or nil list, do exception path!")


Super cool documentation and unit-testing feature

the first declaration inside a method can be a string, with code examples.
def add(x,y)
  """ adds x to y and return the sum
  >>> add(1,1)
  2
 >>> add(-1,101)
 100
"""
  return x+y

the documentation is accessed using add.__doc__
doctest can run all the code samples in the documentation.

import doctest
if __name__ == "__main__" ;  #optional "if" ,means run this code only if you don't import this
  doctest.testmod()
it will search for all the mehod documentation, and run each example.  if there is a problem, it will output the expected and the actual value.


printing


simple, in this case we indent to the left the name, for column size of 20, then price as float.
print("Name={name:<20s} Price:{price:8.2f}. format(a="golden-crown", b=2000.01) )
print("Name={name:<20s} Price:{price:8.2f}. format("golden-crown", 2000.01) )  #same


in many case, you have a varialbes instead, you can do this trick and use the general method locals() which pass the dictionary of current scope values:
print("Name={name:<20s} Price:{price:8.2f}. format(**locals() ))
As a sidenote: You can use the same trick with your own dictionary.
  def foo(a,b,c,d) :  print (a,b,c,d) #method with four arguments
  myDic = { "c":0.4 , d:"description" , "a": 55 , "b": true , }  #dictionay somewhere in the code
  foo(**myDic)  #will pass the right values from the dictionary to the right arguments
  myList = [ 55, true, 0.4 , "description"]
  foo(*myList) #does the same, but here order is important

print itlsef is quite strong
print( a,b, sep="\n")  will seperate with newline, default is one space.
fh = open ("data.txt", "w")
print("lets write to file. why should it be difficult or different then regular print?", file=fh)
fh.close()



exceptions (see the else part)

try:
   f = open("file.txt")
except IOError as (errnum, errstr):
  print("IO exception number {0} : {1}".format(errnum,errstr)
except (ValueError, InventedError)
  print("other known error happened")
except:  #other unknown error
   print("unknown error, for this, we propogate up", sys.exc_info()[0]))
   raise
else:
    #read the file
    print( f.readlines())
    f.close()

finally:
   print("just like in java, very optional"

assert is a synatactic sugar to if  not <condition>: raise AssertionError("msg")
assert  x>0 , "x must be above zero, it wasn't! fix this for assertion to pass"