{"id":538,"date":"2018-12-28T20:43:58","date_gmt":"2018-12-29T00:43:58","guid":{"rendered":"http:\/\/dominionsw.com\/wordpress\/?p=538"},"modified":"2019-11-25T07:51:31","modified_gmt":"2019-11-25T11:51:31","slug":"classifying-tree-seedlings-with-machine-learning-part-4","status":"publish","type":"post","link":"https:\/\/www.dominionsw.com\/?p=538","title":{"rendered":"Classifying tree seedlings with Machine Learning, Part 4."},"content":{"rendered":"<p><a href=\"http:\/\/dominionsw.com\/wordpress\/wp-content\/uploads\/2018\/12\/GSBG-1.png\"><img loading=\"lazy\" decoding=\"async\" class=\"alignnone wp-image-302 size-full\" src=\"https:\/\/www.dominionsw.com\/wp-content\/uploads\/2018\/12\/GSBG-1.png\" alt=\"\" width=\"1932\" height=\"644\" srcset=\"https:\/\/www.dominionsw.com\/wp-content\/uploads\/2018\/12\/GSBG-1.png 1932w, https:\/\/www.dominionsw.com\/wp-content\/uploads\/2018\/12\/GSBG-1-300x100.png 300w, https:\/\/www.dominionsw.com\/wp-content\/uploads\/2018\/12\/GSBG-1-768x256.png 768w, https:\/\/www.dominionsw.com\/wp-content\/uploads\/2018\/12\/GSBG-1-1024x341.png 1024w, https:\/\/www.dominionsw.com\/wp-content\/uploads\/2018\/12\/GSBG-1-624x208.png 624w\" sizes=\"auto, (max-width: 1932px) 100vw, 1932px\" \/><\/a><\/p>\n<h1>Image Recognition with the Matlab Deep Learning Toolbox<\/h1>\n<p>After some time in the Artificial Intelligence Doghouse, Neural Networks regained status after Alex Krizhevsky, Ilya Sustkever, and Goeffrey Hinton won the\u00a0Large Scale Visual Recognition Challenge in 2012 using their Neural Network, now known as AlexNet.<\/p>\n<p>Details can be found in their <a href=\"https:\/\/papers.nips.cc\/paper\/4824-imagenet-classification-with-deep-convolutional-neural-networks.pdf\">paper<\/a>.<\/p>\n<p>The Matlab Deep Learning Toolbox makes it really easy to develop our own Deep Learning network to use for our tree seedling classification task. And, as it turns out, we can take advantage of the work done by the winners above by using their trained network to bootstrap our own network using a technique called <em>Transfer Learning.<\/em><\/p>\n<h2>Overview of Deep Learning Neural Networks<\/h2>\n<p>A Neural Network is designed to mimic the way a biological neuron &#8220;Activates&#8221; on a given stimulus. How close it comes to recreating an actual biological neuron I will leave to the Neuroscientists to decide. But that is the concept. These neurons are then connected into &#8220;Layers&#8221;. The &#8220;Deep&#8221; part of &#8220;Deep Learning&#8221; refers to the fact that there are many layers of neurons in a particular network, therefore creating a &#8220;Deep&#8221; Neural Network.<\/p>\n<p>Let&#8217;s take a look at the typical diagram for a section of a Neural Network:<\/p>\n<p><a href=\"http:\/\/dominionsw.com\/wordpress\/wp-content\/uploads\/2018\/12\/Capture-6.png\"><img loading=\"lazy\" decoding=\"async\" class=\"alignnone wp-image-546 size-full\" src=\"https:\/\/www.dominionsw.com\/wp-content\/uploads\/2018\/12\/Capture-6.png\" alt=\"\" width=\"1868\" height=\"1180\" srcset=\"https:\/\/www.dominionsw.com\/wp-content\/uploads\/2018\/12\/Capture-6.png 1868w, https:\/\/www.dominionsw.com\/wp-content\/uploads\/2018\/12\/Capture-6-300x190.png 300w, https:\/\/www.dominionsw.com\/wp-content\/uploads\/2018\/12\/Capture-6-768x485.png 768w, https:\/\/www.dominionsw.com\/wp-content\/uploads\/2018\/12\/Capture-6-1024x647.png 1024w, https:\/\/www.dominionsw.com\/wp-content\/uploads\/2018\/12\/Capture-6-624x394.png 624w\" sizes=\"auto, (max-width: 1868px) 100vw, 1868px\" \/><\/a><\/p>\n<p>The circles on the left are the activations. The lines connecting to the circle on the right are the weights. The circle on the right is the output which is the weighted sum of the activations plus a bias, b, plugged in as a parameter to\u00a0\u03c3(), which is a function known as a Rectifier which simply sets everything &lt; 0 to 0. We&#8217;ll skip over the bias and the sigma function, as they are important but not essential to the basic concept.<\/p>\n<p>These diagrams certainly didn&#8217;t mean all that much to me at first; as a programmer, I immediately thought &#8220;what does this look like in code&#8221;?<\/p>\n<p>Well, it turns out that if you know a bit about image processing filters, it is easy to understand the basic concepts of convolutional neural networks.<\/p>\n<p>It turns out that for the image specific tasks, like classifying photos, the activations are the input pixels. The weights are the values in an image filter. So, for instance, a filter could be a 3 by 3 edge detection filter,<\/p>\n\n<table id=\"tablepress-5\" class=\"tablepress tablepress-id-5\">\n<tbody>\n<tr class=\"row-1\">\n\t<td class=\"column-1\">-1<\/td><td class=\"column-2\">0<\/td><td class=\"column-3\">1<\/td>\n<\/tr>\n<tr class=\"row-2\">\n\t<td class=\"column-1\">-1<\/td><td class=\"column-2\">0<\/td><td class=\"column-3\">1<\/td>\n<\/tr>\n<tr class=\"row-3\">\n\t<td class=\"column-1\">-1<\/td><td class=\"column-2\">0<\/td><td class=\"column-3\">1<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<!-- #tablepress-5 from cache -->\n<p>or any other type of image filter. The activations are the pixels of the input image, and so one is left with the &#8220;classic&#8221; image processing task of sliding a filter across an image, taking the dot product and setting a new pixel &#8211; potentially a new activation in the next layer.<\/p>\n<p><a href=\"http:\/\/dominionsw.com\/wordpress\/wp-content\/uploads\/2018\/12\/ConvolutionFilter.png\"><img loading=\"lazy\" decoding=\"async\" class=\"alignnone wp-image-564 size-full\" src=\"https:\/\/www.dominionsw.com\/wp-content\/uploads\/2018\/12\/ConvolutionFilter.png\" alt=\"\" width=\"1000\" height=\"1000\" srcset=\"https:\/\/www.dominionsw.com\/wp-content\/uploads\/2018\/12\/ConvolutionFilter.png 1000w, https:\/\/www.dominionsw.com\/wp-content\/uploads\/2018\/12\/ConvolutionFilter-150x150.png 150w, https:\/\/www.dominionsw.com\/wp-content\/uploads\/2018\/12\/ConvolutionFilter-300x300.png 300w, https:\/\/www.dominionsw.com\/wp-content\/uploads\/2018\/12\/ConvolutionFilter-768x768.png 768w, https:\/\/www.dominionsw.com\/wp-content\/uploads\/2018\/12\/ConvolutionFilter-624x624.png 624w\" sizes=\"auto, (max-width: 1000px) 100vw, 1000px\" \/><\/a><\/p>\n<p>Now, here&#8217;s where the &#8220;Learning&#8221; part of Deep Learning comes in. Instead of first specifying what filters to use, we tell the network what result we want, and it keeps adjusting the filters until it gets that result, or close to it. This feedback is known as back-propagation. So, let&#8217;s say we want to train the network to recognize an image with the label &#8220;Cat&#8221;. We keep feeding the images we want to be recognized as a Cat, and the algorithm keeps updating the filters until they agree that the images can be classified as a Cat. This is a bit of a simplification, but it is the basic essence of what is happening during the training process.<\/p>\n<p>Once the network is trained properly, the filters will then be able to process an unseen before image as a &#8220;Cat&#8221;.\u00a0<a href=\"http:\/\/dominionsw.com\/wordpress\/wp-content\/uploads\/2018\/12\/Capture-7.png\"><img loading=\"lazy\" decoding=\"async\" class=\"alignnone wp-image-551 size-full\" src=\"https:\/\/www.dominionsw.com\/wp-content\/uploads\/2018\/12\/Capture-7.png\" alt=\"\" width=\"1884\" height=\"1329\" srcset=\"https:\/\/www.dominionsw.com\/wp-content\/uploads\/2018\/12\/Capture-7.png 1884w, https:\/\/www.dominionsw.com\/wp-content\/uploads\/2018\/12\/Capture-7-300x212.png 300w, https:\/\/www.dominionsw.com\/wp-content\/uploads\/2018\/12\/Capture-7-768x542.png 768w, https:\/\/www.dominionsw.com\/wp-content\/uploads\/2018\/12\/Capture-7-1024x722.png 1024w, https:\/\/www.dominionsw.com\/wp-content\/uploads\/2018\/12\/Capture-7-624x440.png 624w\" sizes=\"auto, (max-width: 1884px) 100vw, 1884px\" \/><\/a><\/p>\n<p>With Matlab, we can actually look at the filters in the first layer of AlexNet.<\/p>\n<p>Load AlexNet<\/p>\n<pre>net = alexnet;<\/pre>\n<p>Let&#8217;s look at the Layers:<\/p>\n<p>net.Layers<\/p>\n<pre>ans = \r\n  25x1 <a href=\"matlab:helpPopup nnet.cnn.layer.Layer\">Layer<\/a> array with layers:\r\n\r\n     1   'data'     Image Input                   227x227x3 images with 'zerocenter' normalization\r\n     2   'conv1'    Convolution                   96 11x11x3 convolutions with stride [4  4] and padding [0  0  0  0]\r\n     3   'relu1'    ReLU                          ReLU\r\n     4   'norm1'    Cross Channel Normalization   cross channel normalization with 5 channels per element\r\n     5   'pool1'    Max Pooling                   3x3 max pooling with stride [2  2] and padding [0  0  0  0]\r\n     6   'conv2'    Convolution                   256 5x5x48 convolutions with stride [1  1] and padding [2  2  2  2]\r\n     7   'relu2'    ReLU                          ReLU\r\n     8   'norm2'    Cross Channel Normalization   cross channel normalization with 5 channels per element\r\n     9   'pool2'    Max Pooling                   3x3 max pooling with stride [2  2] and padding [0  0  0  0]\r\n    10   'conv3'    Convolution                   384 3x3x256 convolutions with stride [1  1] and padding [1  1  1  1]\r\n    11   'relu3'    ReLU                          ReLU\r\n    12   'conv4'    Convolution                   384 3x3x192 convolutions with stride [1  1] and padding [1  1  1  1]\r\n    13   'relu4'    ReLU                          ReLU\r\n    14   'conv5'    Convolution                   256 3x3x192 convolutions with stride [1  1] and padding [1  1  1  1]\r\n    15   'relu5'    ReLU                          ReLU\r\n    16   'pool5'    Max Pooling                   3x3 max pooling with stride [2  2] and padding [0  0  0  0]\r\n    17   'fc6'      Fully Connected               4096 fully connected layer\r\n    18   'relu6'    ReLU                          ReLU\r\n    19   'drop6'    Dropout                       50% dropout\r\n    20   'fc7'      Fully Connected               4096 fully connected layer\r\n    21   'relu7'    ReLU                          ReLU\r\n    22   'drop7'    Dropout                       50% dropout\r\n    23   'fc8'      Fully Connected               1000 fully connected layer\r\n    24   'prob'     Softmax                       softmax\r\n    25   'output'   Classification Output         crossentropyex with 'tench' and 999 other classes<\/pre>\n<p>We can take a look at the 2nd layer and visualize the image filters. There are 96 filters, 11 pixels by 11 pixels by 3 pixels (one set of pixels each for R, G, B colors).<\/p>\n<p>We offset the pixels all into the positive range so we can see them more easily, then write them as a tiff file.<\/p>\n<pre>layer = net.Layers(2);\r\nimg = imtile(layer.Weights);\r\nmn = min(min(min(img)));\r\nimg = img - mn;\r\n\r\nimgu = uint8(img.*255);\r\n\r\nimwrite(imgu,'Tiles.tif','tif');<\/pre>\n<p><a href=\"http:\/\/dominionsw.com\/wordpress\/wp-content\/uploads\/2018\/12\/Capture-8.png\"><img loading=\"lazy\" decoding=\"async\" class=\"alignnone wp-image-554 size-full\" src=\"https:\/\/www.dominionsw.com\/wp-content\/uploads\/2018\/12\/Capture-8.png\" alt=\"\" width=\"1323\" height=\"1324\" srcset=\"https:\/\/www.dominionsw.com\/wp-content\/uploads\/2018\/12\/Capture-8.png 1323w, https:\/\/www.dominionsw.com\/wp-content\/uploads\/2018\/12\/Capture-8-150x150.png 150w, https:\/\/www.dominionsw.com\/wp-content\/uploads\/2018\/12\/Capture-8-300x300.png 300w, https:\/\/www.dominionsw.com\/wp-content\/uploads\/2018\/12\/Capture-8-768x769.png 768w, https:\/\/www.dominionsw.com\/wp-content\/uploads\/2018\/12\/Capture-8-1024x1024.png 1024w, https:\/\/www.dominionsw.com\/wp-content\/uploads\/2018\/12\/Capture-8-624x624.png 624w\" sizes=\"auto, (max-width: 1323px) 100vw, 1323px\" \/><\/a><\/p>\n<p>This is pretty interesting. 96 image filters that were &#8220;learned&#8221; from the 1.2 million images with 1000 different labels. I also notice the cyan and green colors, which remind me of the Lab color space from a previous post.<\/p>\n<p>There a lot of details that one can go into here, of course. We&#8217;ll skip ahead to how we can use this network to recognize our own images.<\/p>\n<h2>Transfer Learning<\/h2>\n<p>It turns out that we can retrain an existing network to classify our seedlings using what is known as &#8220;Transfer Learning&#8221;. Using the Matlab Deep Learning Toolbox, we can do this by replacing two of the layers in the original AlexNet with our own.<\/p>\n<p>First, we load the network, and set up a path to our Labeled image folders.<\/p>\n<pre>net = alexnet;\r\nimagePath = 'C:\\Users\\rickf\\Google Drive\\Greenstand\\TestingAndTraining';<\/pre>\n<p>&nbsp;<\/p>\n<p>We can set up our datastore, and use a Matlab helper function to pick 80 percent of the images for training. In addition, the\u00a0augmentedImageDatastore() function let&#8217;s us scale the images to the required 227&#215;227 inputs that AlexNet requires.<\/p>\n<pre>imds = imageDatastore(imagePath,'IncludeSubfolders',true,'LabelSource','foldernames');\r\n[trainIm, testIm] = splitEachLabel(imds, 0.8,'randomized')\r\n\r\ntrainData = augmentedImageDatastore([227 227], trainIm)\r\ntestData = augmentedImageDatastore([227 227], testIm)<\/pre>\n<p>If we look above we can see that the final layers in the AlexNet network are set up for 1000 different labels.<\/p>\n<pre>23 'fc8' Fully Connected 1000 fully connected layer \r\n24 'prob' Softmax softmax \r\n25 'output' Classification Output crossentropyex with 'tench' and 999 other classes<\/pre>\n<p>What we need to do is to replace layer 23 with a fully connected 6 input layer (for our 6 labels),<br \/>\nand replace 25 with a new classifier that will accept those 6 inputs from the softmax layer.<\/p>\n<pre>layer(23) = fullyConnectedLayer(6,'Name','fcfinal');\r\nlayer(end) = classificationLayer('Name','treeClassification');<\/pre>\n<div>Set up training options, and train the network.<\/div>\n<div><\/div>\n<pre>opts = trainingOptions('adam','InitialLearnRate',0.0001,'Plots','training-progress','MaxEpochs',10);\r\n[net,info] = trainNetwork(trainData,layer,opts);<\/pre>\n<p><a href=\"https:\/\/www.dominionsw.com\/wp-content\/uploads\/2018\/12\/Capture-9.png\"><img loading=\"lazy\" decoding=\"async\" class=\"alignnone wp-image-556 size-full\" src=\"https:\/\/www.dominionsw.com\/wp-content\/uploads\/2018\/12\/Capture-9.png\" alt=\"\" width=\"2168\" height=\"1493\" srcset=\"https:\/\/www.dominionsw.com\/wp-content\/uploads\/2018\/12\/Capture-9.png 2168w, https:\/\/www.dominionsw.com\/wp-content\/uploads\/2018\/12\/Capture-9-300x207.png 300w, https:\/\/www.dominionsw.com\/wp-content\/uploads\/2018\/12\/Capture-9-768x529.png 768w, https:\/\/www.dominionsw.com\/wp-content\/uploads\/2018\/12\/Capture-9-1024x705.png 1024w, https:\/\/www.dominionsw.com\/wp-content\/uploads\/2018\/12\/Capture-9-624x430.png 624w\" sizes=\"auto, (max-width: 2168px) 100vw, 2168px\" \/><\/a><\/p>\n<p>We specified the option to plot progress, which is shown above.<\/p>\n<p>The final results can be seen in our confusion matrix.<\/p>\n<p><a href=\"https:\/\/www.dominionsw.com\/wp-content\/uploads\/2018\/12\/Capture-10.png\"><img loading=\"lazy\" decoding=\"async\" class=\"alignnone wp-image-557 size-full\" src=\"https:\/\/www.dominionsw.com\/wp-content\/uploads\/2018\/12\/Capture-10.png\" alt=\"\" width=\"833\" height=\"636\" srcset=\"https:\/\/www.dominionsw.com\/wp-content\/uploads\/2018\/12\/Capture-10.png 833w, https:\/\/www.dominionsw.com\/wp-content\/uploads\/2018\/12\/Capture-10-300x229.png 300w, https:\/\/www.dominionsw.com\/wp-content\/uploads\/2018\/12\/Capture-10-768x586.png 768w, https:\/\/www.dominionsw.com\/wp-content\/uploads\/2018\/12\/Capture-10-624x476.png 624w\" sizes=\"auto, (max-width: 833px) 100vw, 833px\" \/><\/a><\/p>\n<p>Much better results that the Support Vector Machine, and no color segmentation was needed!<\/p>\n<p>I\u00a0 would suspect with more training data we can get the Broadleaf and Fernlike classifications to improve. We&#8217;ll try that, and then put the network to a test of classifying unseen new images in the next post.<\/p>\n<p>&nbsp;<\/p>\n","protected":false},"excerpt":{"rendered":"<p>Image Recognition with the Matlab Deep Learning Toolbox After some time in the Artificial Intelligence Doghouse, Neural Networks regained status after Alex Krizhevsky, Ilya Sustkever, and Goeffrey Hinton won the\u00a0Large Scale Visual Recognition Challenge in 2012 using their Neural Network, now known as AlexNet. Details can be found in their paper. The Matlab Deep Learning [&hellip;]<\/p>\n","protected":false},"author":1,"featured_media":0,"comment_status":"closed","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[1],"tags":[],"class_list":["post-538","post","type-post","status-publish","format-standard","hentry","category-uncategorized"],"_links":{"self":[{"href":"https:\/\/www.dominionsw.com\/index.php?rest_route=\/wp\/v2\/posts\/538","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.dominionsw.com\/index.php?rest_route=\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.dominionsw.com\/index.php?rest_route=\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.dominionsw.com\/index.php?rest_route=\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/www.dominionsw.com\/index.php?rest_route=%2Fwp%2Fv2%2Fcomments&post=538"}],"version-history":[{"count":18,"href":"https:\/\/www.dominionsw.com\/index.php?rest_route=\/wp\/v2\/posts\/538\/revisions"}],"predecessor-version":[{"id":609,"href":"https:\/\/www.dominionsw.com\/index.php?rest_route=\/wp\/v2\/posts\/538\/revisions\/609"}],"wp:attachment":[{"href":"https:\/\/www.dominionsw.com\/index.php?rest_route=%2Fwp%2Fv2%2Fmedia&parent=538"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.dominionsw.com\/index.php?rest_route=%2Fwp%2Fv2%2Fcategories&post=538"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.dominionsw.com\/index.php?rest_route=%2Fwp%2Fv2%2Ftags&post=538"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}