Tuesday, July 12, 2016

StateFarm experiment 1

Let's start with a simple and quick to run model.

150x150 (32,3,3) (32,3,3) (64,3,3) -> Dense( 3x200) -> dropout0.5-> dense10


each epoc is  train: 5*1024 . validate= 1*1024. batch-32

model_chapter3
epoc 0 699s - loss: 18.1189 - acc: 0.2369 - val_loss: 2.1892 - val_acc: 0.3574
epoc 1 765s - loss: 7.5443 - acc: 0.4570 - val_loss: 1.5257 - val_acc: 0.4697
epoc 2 689s - loss: 3.5896 - acc: 0.6488 - val_loss: 1.9699 - val_acc: 0.3590
epoc 3 697s - loss: 1.8959 - acc: 0.7616 - val_loss: 1.8912 - val_acc: 0.3887
epoc 4 707s - loss: 1.2178 - acc: 0.7992 - val_loss: 1.5978 - val_acc: 0.4756
epoc 5 710s - loss: 0.9396 - acc: 0.8277 - val_loss: 1.6677 - val_acc: 0.5829
epoc 6 702s - loss: 0.8008 - acc: 0.8520 - val_loss: 1.9146 - val_acc: 0.5781
epoc 7 707s - loss: 0.6810 - acc: 0.8798 - val_loss: 1.3611 - val_acc: 0.5752
epoc 8 707s - loss: 0.6647 - acc: 0.8748 - val_loss: 1.8251 - val_acc: 0.5314
epoc 9 706s - loss: 0.6234 - acc: 0.8936 - val_loss: 1.5517 - val_acc: 0.5908
epoc 10 709s - loss: 0.5812 - acc: 0.9054 - val_loss: 1.8407 - val_acc: 0.5225

 Usually we will plot loss, but here I plot the accuracy graph (training converges to 95% while validation does not pass the 58%). 


continuing till epoc 30 reduce the training loss a bit, and the accuracy, but the validation does not improve. 
epoc 30 - loss: 0.3641 - acc: 0.9482 - val_loss: 1.3679 - val_acc: 0.6270



Notes on this run:
After epoc 5 (in this case epoc is sample of 1/4 of the images), we start to overfit.   Further epocs do not help  (validation stays the same while training loss reduced to be extremely small)

There could be two main reasons:
1. Model is too strong and not regularized enough -  Not the case here... it's small , heavy-regularzation and dropout.
2. Model is too strong compared to the data. I think this is the case.

The data
The number of training images is small (20k), further more, they are taken from ~20 videos of 20 actors, cut by frames, while the test set is from different video of different actors.
20 actors is not enough to regularize on all the people in the world.

What can be done?
  • More data is the obvious solution, but there is none.
  • Pretrained models are allowed in the competition, if they are pulic and can be used commercialy. Great imporevments were achieved using VGG-16  (10 times better) which can't be  commericaly used. What does the pretrained network give us?
    • Better visual filters on the lower filters.
    • Cellphone detection on the higher filters.
    • Probably good human detection, but not clear if good hand localization detection
  • Or use a cascade of a 2 pretrained-models creating features, combine them into an image/new-channel and provide this to a small model.
    • A good one for humans exist, but runs in 17s x20,000 images =  340K second / 86,400 = 3.93 days

Further experiment with similiar architectures



experiment 3
Dense 3x200. l2(0.01). BN on all layers exect the 1st dense. adam optimizer
711s - loss: 0.4788 - acc: 0.9227 - val_loss: 2.0435 - val_acc: 0.5019
Saved model to disk model_chapter3_17epoc
#Validation : SCORE of model_chapter3_17epoc 0.290623311932 accuracy 0.434080421885
#  Leader-board score = 1.64778



experiment 4
experiment 4 ran with: dense: 200-100-50 . Full BN. Pre-relu  SGD(lr=0.001, decay=1e-7, momentum=.9) optimizer. 


experiment 5


expeiment 5 ran with: dense 256-124-64. BN on allbut the 1st dense. regular Relu. Adam optimizer

5120/5120 - 1012s - loss: 0.4410 - acc: 0.9084 - val_loss: 1.0536 - val_acc: 0.6631
Saved model to disk model_chapter5_18epoc



Sunday, July 10, 2016

Keras pre-trained models how-to

In keras Functional API there are few examples for  combining models in a different way, but if you want to mix and match parts of models , like combine new classifier-head with a headless pre-trained model, you might encounter some issues. Let's explain how it works behind the scenes.

A saved model is a combination of three things:
  1. Layer definitions like "Dense" layer with it's regularizes ,output_dim and name
  2. Layer graph: which layer is connected to which other, and from which direction (in/out)
  3. Layer weights: for layers which do have them (dropout do not , for example)


When you call: model.to_json() , you get the first two. When you call model.save_weights() you get the third, where the key is the layer-name from (1)
If you want to save and load the exact same model, just call m=model_from_json(json) and then m.load_weights(file).  Easy.
If you want to create a head-less model, which does not contain the last few layers, you will have to re-create the model by code, create the connectivity and then use your own function to read the weights file(.h5).  The h5 API is quite clean and it's an easy-enough task.

If you want to create a new model from a headless-one and a new head you created, again you will have to define (1) and (2) in code, and then manually load-weights per layer.

Code for connectivity between layers
# assuming this was called already: graph = Graph() 
# graph.add_input(name='input1', ndim=2)

new_layer = (Dense(32, 4)  #define layer
graph.add_node(new_layer , name='dense1'         ,input='input1') #connect layer to input
#note that if you want two inputs, just use instead  ,inputs=['input1,'input2']

# and in the end .... graph.add_output(name='output', input='...')

The functional API hides the Graph API and allows you to create a layer and connect it to a node in one line
#assuming input- Input(shape=784,)
Dense(32,4)(input)  

It's important to remember that the node connectivity is actually added to the layer instance itself, this sadly means you can't re-use the layer in multiple graphs with totally different connectivity.
You will have to re-instantiate (redefine) a new identical, unconnected, layer.

FAQ:
"UserWarning: Model inputs must come from a Keras Input layer, they cannot be the output of a previous non-Input layer"
You can't instantiate a Model which does not start with Input layer, in other words, you can't just cut the few last layers of an existing model and call Model (input='other_input', output='...same')



Wednesday, July 6, 2016

Features regression/localization

There are many approaches, we shall start with the simple ones (which perform badly) and continue to those performing a bit better.

Loss via Linear regression

Combine all the features  into one loss function, and calculate it (diff of mouth + diff of nose + diff of eyes, etc).  The results are usually quite rough, so different cascade are suggested to fine tune the windows on which the CNN is running.

  • Facial recognition with no cascade: Tutorial with full code Lasgne 
  • Joint location with cascade: DeepPose - Run one DNN to get rough estimates of joint locations. cut a small window around each joint location and run the same network architecture (but with different parameters) on it to get finer estimation. Do it again one more time to get best results.
  • Deep Convolutional Network Cascade for Facial Point Detection
    Modification 1: abs on tan, instead of  ReLU, on some layers.
    Modification 2: locally-shared weights instead of globally shared weights in the Convolution layer.
    Modification 3: lots of networks structure:Train 3 networks with the same architecture (high1) to detect eyes region, nose region and mouth region.
    Then pass it to multiple shallow architecture (shallow1), again each trained by itself, and it only gets a small region. This one can only slighly modify output location, as we assume it is more accurate, but 'does not see the big picture.
    Then again pass it to multiple shallow arch...  with even smaller region.

Classification as Loss function

Cascade: This is usually used for bounding-box calculations on classification tasks.  As you already have a good classifier model, you want to re-use it.  This is done by cutting the image into multiple very small images, and checking each one of the sub-images for the existance of a "nose"/"eye". if so mark what there is there. If you apply it to the whole image, you can a rough estimate of where the nose is.

Detection
region proposal networks - faster rcnn
yolo - 

Heatmap as Loss function

The output of the network is a heatmap  (WxH pixels with intensity levels from 0 to 1), where few pixels around the feature are highlighred. This apears to provide better results, as it is a more "natual" calculation for a CNN.  Do not that regular classification models are shaped like a cone, with smaller and smaller layers till the FC result.  This architecture is not adequate here.
In theory, one may not need a cascade here.

Localization as side-effect

The CAM technique (Class activation mapping) can generate it automatically , as a side-effect of the attention model.


Appendix, Honorable mentions: Linear combination of the previous layers


 this is good for the whole face pixels, not as good for location of occulded features in the face which is fully not facing-the-camera.


Datasets and competion results for Object Segmentaion
Coco - Huge one.


Datasets and compettion results for Faces

HPDatabase 
Youtube face DB
TAU article TCNN for Facial Landmark Detection with Tweaked Convolutional Neural Networks. implementation in caffe


Datasets and competition results for human joints 

2016 - Coco key point challenge - new (July 2016) and probably the best one to use.  90K person instances labeled with keypoints (the majority of people in COCO at medium and large scales) and over 1 million total labeled keypoints.

The Human Pose Recovery and Behavior Analysis HuPBA 8k+ dataset , from cha learn

FLIC: Frames Labeled In Cinema contains 4000 training and 1000 test images obtained from popular Hollywood movies. The images contain people in diverse poses and especially diverse clothing. For each labeled human, 10 upper body joints are labeled.
Note: FLIC_Full contains more challanging(occulded) scenes. too hard to train from.
FLIC-motion-dataset
 includes short clips. in this case, maybe the motion will help estimate better.

Leeds Sports Dataset [12] and its extension [13], which we will jointly denote by LSP. Combined they contain 11000 training and 1000 testing images. These are images from sports activities and as such are quite challenging in terms of appearance and especially articulations. In addition, the majority of people have 150 pixel height which makes the pose estimation even more challenging. In this dataset, for each person the full body is labeled with total 14 joints.

The dataset includes around 25K images containing over 40K people with annotated body joints. The images were systematically collected using an established taxonomy of every day human activities. Overall the dataset covers 410 human activities and each image is provided with an activity label
Best results for MPII

2016 Convolutional Pose Machines in caffe/matlab ; Stack hourglass in torch
2015 DeepCut: Joint Subset Partition and Labeling for Multi Person Pose Estimation






Features regression/localization

There are many approaches, we shall start with the simple ones (which perform badly) and continue to those performing a bit better.

Loss via Linear regression

Combine all the features  into one loss function, and calculate it (diff of mouth + diff of nose + diff of eyes, etc).  The results are usually quite rough, so different cascade are suggested to fine tune the windows on which the CNN is running.

  • Facial recognition with no cascade: Tutorial with full code Lasgne 
  • Joint location with cascade: DeepPose - Run one DNN to get rough estimates of joint locations. cut a small window around each joint location and run the same network architecture (but with different parameters) on it to get finer estimation. Do it again one more time to get best results.
  • Deep Convolutional Network Cascade for Facial Point Detection
    Modification 1: abs on tan, instead of  ReLU, on some layers.
    Modification 2: locally-shared weights instead of globally shared weights in the Convolution layer.
    Modification 3: lots of networks structure:Train 3 networks with the same architecture (high1) to detect eyes region, nose region and mouth region.
    Then pass it to multiple shallow architecture (shallow1), again each trained by itself, and it only gets a small region. This one can only slighly modify output location, as we assume it is more accurate, but 'does not see the big picture.
    Then again pass it to multiple shallow arch...  with even smaller region.

Classification as Loss function

Cascade: This is usually used for bounding-box calculations on classification tasks.  As you already have a good classifier model, you want to re-use it.  This is done by cutting the image into multiple very small images, and checking each one of the sub-images for the existance of a "nose"/"eye". if so mark what there is there. If you apply it to the whole image, you can a rough estimate of where the nose is.

Heatmap as Loss function

The output of the network is a heatmap  (WxH pixels with intensity levels from 0 to 1), where few pixels around the feature are highlighred. This apears to provide better results, as it is a more "natual" calculation for a CNN.  Do not that regular classification models are shaped like a cone, with smaller and smaller layers till the FC result.  This architecture is not adequate here.
In theory, one may not need a cascade here.

Localization as side-effect

The CAM technique (Class activation mapping) can generate it automatically , as a side-effect of the attention model.


Appendix, Honorable mentions: Linear combination of the previous layers

 this is good for the whole face pixels, not as good for location of occulded features in the face which is fully not facing-the-camera.



Datasets and competition results for human joints 

FLIC: Frames Labeled In Cinema contains 4000 training and 1000 test images obtained from popular Hollywood movies. The images contain people in diverse poses and especially diverse clothing. For each labeled human, 10 upper body joints are labeled.
Note: FLIC_Full contains more challanging(occulded) scenes. too hard to train from.
FLIC-motion-dataset
 includes short clips. in this case, maybe the motion will help estimate better.

Leeds Sports Dataset [12] and its extension [13], which we will jointly denote by LSP. Combined they contain 11000 training and 1000 testing images. These are images from sports activities and as such are quite challenging in terms of appearance and especially articulations. In addition, the majority of people have 150 pixel height which makes the pose estimation even more challenging. In this dataset, for each person the full body is labeled with total 14 joints.

The dataset includes around 25K images containing over 40K people with annotated body joints. The images were systematically collected using an established taxonomy of every day human activities. Overall the dataset covers 410 human activities and each image is provided with an activity label


Best results:

2016 Convolutional Pose Machines
2015 DeepCut: Joint Subset Partition and Labeling for Multi Person Pose Estimation






Keras Installation on Windows - CPU performance


1. Install the basics.  start with this stackoverflow:

  • Install TDM GCC x64.  (need to add to begin of path? probably not)
  • Install Anaconda x64.
  • Open the Anaconda prompt
  • Run conda update conda (You might need to open-command-prompt with admin privalage)
  • Run conda update --all
  • Run conda install mingw libpython
  • Install the latest version of Theano, pip install git+git://github.com/Theano/Theano.git
  • Run pip install git+git://github.com/fchollet/keras.git
  • Update the cxx option in the env variables: set THEANO_FLAGS=floatX=float32,device=cpu,cxx=E:\\TDM-GCC-64\\bin\\g++.exe
  • Check that a simple Keras Hello-World is working .
2. Now let's improve performance. Here I use CPU and not GPU. If you have a good nvidia GPU, check keras cuda install guide.

  • Change your BLAS library : OpenBlas/MKL can give you 400% speedup. Download a good BLAS library and add it (the \bin folder) to the system path. I used openblas 2.0.14 and got a huge boost  update the THEANO_FLAGS env variable with it:
    set THEANO_FLAGS=floatX=float32,device=cpu,cxx=E:\\TDM-GCC-64\\bin\\g++.exe,blas.ldflags=-LE:\\code\\openblas\\bin -lopenblas

    I tried to use intel mkl library, which should be faster, but could not successfully configure it for keras. (If you did, please leave a comment....)
  • OMP: usually 10-20% boost
  • set OMP_NUM_THREADS=2  (benchmark X on your-system,  maybe 4 is better?)
    update THEANO_FLAGS:
    set THEANO_FLAGS=openmp=True,floatX=float32,device=cpu,cxx=E:\\TDM-GCC-
    64\\bin\\g++.exe,blas.ldflags=-LE:\\code\\openblas\\bin -lopenblas

    If you get this error :
    UserWarning: Your g++ compiler fails to compile OpenMP code. We know this happen with some version... then make
    Make sure your cxx is configured properly, but one some machine-configuration (my very-old desktop for example), I could not solve the issue.

FAQ:
InvalidValueError: InvalidValueError  ...
        type(variable) = TensorType(float32, (True, True))
        variable       = TensorConstant{(1L, 1L) of inf}
        context        = ...
  TensorConstant{(1L, 1L) of inf} [id A]

Check THEANO_FLAGS , mode should not be debug_mode.  if it does not help do "conda update -all"

Sunday, May 1, 2016

Python after C#/Java - Part 2 - Classes


Classes

class Account(object) :  # inherits object directly
     
     instanceCount= 0 # like static variable in java
     #this is similiar to java constructor, called automatically during creation : Account("assaf")                 def __init__(self, balance, someBool=true):
       self.Balance= balance#note:  no need to define them before, nice!!!
       self.SomeBool= someBool   #public style
       self._XProtected = 5 # one '_' prefix is like java-public, but the programmer ask politly.
       self.__Yprivate = 8  # two '_' prefix is like java-private. you 'll get exception when accessing.
       Account.instanceCount +=1

    #destructor (usually unused, like in Java)  called when no reference exists any more .
    #for example by using x=None,   or explicit calling del(x)
    def __del__(self):
         Account.instanceCount -=1 #static variable: note the usage of Account. and not self.
   
   #like the ToString method. optional, of course
    def __str__(self):
         return "balance: %f".format(self.Balance)
 
   #regular method
   def  deposit(self, x):  #note, when calling it "self" is not needed
        self.Balance += x
     

Usage:
account = Account("assaf")
account._XProtected = 666       #works, but '_' it means the programmer asked you not to do it
print( account._XProtected)    --> 666
print( account.__Yprivate)  --> AttributeError
print (account._Account__YPrivate) --> 8    #showing you there is no real way in python to defend against this. if someone want's, he can access it anyway



inheritance and static variables


class Counter:
    instanceCount = 0
    def __init__(self):
           type(self).instanceCount +=1 # and not Counter.instanceCount, cause it will be all sons.
    def __del__(self):
           type(self).instanceCount -=1

class Account(Counter):
      def __init__(self, x,y,z):
               Counter.__init__(self)  #explicit call is needed!

class MultipleInherit(Counter, Shouter):
        def __init__(self, x,y,z):
                 Counter.__init__(self)
                  Shouter_init__(self, y)

storage optimization (__slots__)

each instance has a built-in __dict__ hashmap, which contains all the dynamic memebers.  It means that you can always add members to a class, but it's heavier in storage on the RAM.
account = Account("assaf")
account.dynamicNewMember = 8  #works just fine

If you instantiate millions of these instances, instead of using the built-in __dict__, you can use a tighter static structure. Note: Don't optimize this way unless you have millions of instances.
Use __slots__ and define them before hand , by name , like:
def AccountWithLessStorage:
  __slots__ = ['Balance', 'someBool', '_XProtected', '__YPrivate']
   #everything else in the class is exactly the same, including usage in __init__
   

MetaClass , annotations and Reflection

java reflection-like operations is much easier in python,
instance.__dict__   # is a dictionary of the members (variables and methods), so calling the method append on a list can be done "in-reflection" very easily.
lst = [ 1 , 2  , 3 ]
non-reflection:   list.append(4)
reflection:           list.__dict__["append"](lst, 4) # the 1st parameter is for instance method is "self"     
Use decorators to wrap a method
Use metaclass for decorators like count call/timer/logging.