Assorted Figurative Idioms Collection

‘Watch Your Tongue’ by Mark Abley on what your everyday sayings and idioms literally and figuratively mean was a great entertainer. I’ve captured some sayings and idioms that I liked which I assume you may also enjoy

English has shaped a kind of body language –one made up of gestures and postures but of idioms. These turns of phrase rely on physical selves to evoke ideas and emotions.

Body

Saying/Idiom Comment
Keep your body and soul together survive
stiff upper lip uniquely British form of stoicism
A sight for sore eyes feeling of happiness or relief
all ears listen intently
all mouth and no trousers British expression of a braggart who can’t backup his threats
saving face and loosing face direct translations from Chinese
heads-up as in Let me give you a heads-up update you of the happenings
keep your chin up don’t lose hope
To play it by ear from the world of music, when you don’t have printed score, you play it by ear
ear candy & eye candy appealing with no substance
An eye for an eye age-old, revenge
keep abreast of stay informed…akin to keeping close to my chest where chest is the seat of affections and emotions
gut feelings and gut instincts sometimes seat of affections point to lower parts – guts, in American English guts is synonym for courage
‘a shot in the arm’ may lead to ‘kick up your heels’ receive one and get excited
Don’t ‘get off on the wrong foot’ but ‘put your best foot forward’ or even ‘foot in the door’ avoid the pit and proceed to higher ground or enter/have a start
toe the party line speak or vote the way your party says
sleight of the hand to Shakespeare and et-all ‘sleight’ meant craft, strategy, deceit or cleverness
cold shoulder opposite of warm embrace
limbs akimbo akimbo – something spread or bent
grid up your loins, fruit of loins prepare for battle, child, With mother – it becomes fruit of the womb
pea in the pod, bin in the oven, up the duff, knocked up/laded up, expecting a happy event pregnancy explained with – food, bakery, pudding, hitting based-objectionable, sentimental

Feelings

Saying/Idiom Comment
have a crush teenage infatuation, pleasant enough while it lasts but nothing too serious
falling in love suggests the joyous, agonizing and scary loss of balance that love provokes.. Later it sweeps you ‘off your feet’ and your world turns upside down and you’re heads over heels in love
lock lips, make eyes at him, love at first sight, hit it off together, burns of desire followed by fires of love kiss a girl, she reacts, then follows love between two, if all goes well – make it
wins your heart, stolen your heart, an affair of heart symbolizes heart for love
peeping Tom voyeur who takes sexual pleasure in furtively watching others and there’s an untrue story behind Godiva in the English city of Coventry, Tired of appealing to his husband for nixing townspeople taxes, he challenged her to ride naked in horse which she did asking no one to see but the Tailor Thomas saw her but was killed later
drop-dead gorgeous, beauty is skin deep, hunk & Mr.Right stunning beauty, beauty is not all, male beauties
freinds with benefits/significant others living together
loverbirds when they’re whispering sweet nothings how love blossoms
All’s equal in love and war, at each other’s throats, on the rocks when it goes south
kiss and makeup, patch things up and tie the knot – when ready, so you pop the question and get hitched steps to mend and marry and when you marry – you ‘walk down the aisle’ as in a church
clandestine marriage, arranged marriage, mail-order bride, picture bride, lavender marriage, green card marriage types and means of marriages and finally marriage of convenience
bodice ripper bodice – girdle or whalebone corset, now usually a vest or top part of woman’s dress
self-abuse/self-consolation, previously: beat the meat or spank the rooster male masturbation
buttering the corn, flicking the bean, buffin’ the muffin female masturbation
putting your wand in the chamber of secrets Harry Potter way euphemism to sex
vasocongestion accompanied by testicular pain, lover’s nuts, blue balls sex from scientific to harrowing affliction
get laid, sleep with, sleep together, having relations with alternatives to doing that
go all the way is to hit home run, first base (kissing), second base (petting) and even third base (orals) from boys baseball speak for dealing with girls
horizontal refreshment, carnal knowledge, doing the nasty, making a baby, planting the parsnip, pounding the duck, dipping the wick, storming the cotton gin, jelly roll, happy happy, jig jig, in-and-out how not to say it

Tree Therapy

Excerpts from an excellent book on Trees, ‘The Secret Therapy of Trees’ by Marco Mencagli & Marco Nieri.

Biological effects of Ions

Are negative ions really so important for our health? The scientific literature on the topic leaves no room for doubt. Dozens of studies completed from 1975 to today from all over the world have not only demonstrated the beneficial effects of negative ionization in the air, such as cleaning the atmosphere, but also thoroughly investigated its effects on psychophysical well-being.
Asthma was one of the first diseases to be the subject of clinical experimentation in environments rich in negative ions. A few dozen minutes of exposure each day to intense negative ionization was enough to improve the entire array of patients’ symptoms. In the 1980s and ’90s, researchers had already reported a reduction in asthmatic conditions with a significant reduction in the administration and dosage of medications such as theophylline and cortisone. Other studies focused on elements connected to stress, mood disorders, and physical and mental fatigue. A predominance of positive ions in the air brought with it an increase in these disturbances, while a prevalence of negative ions diminished them.

The agitation or even discomfort felt by some meteoropathic people before a storm or when certain winds are blowing (such as the Föhn in Alpine regions, the Sirocco in parts of Italy, the Santa Anas in the United States, or the Khamsin n the Middle East) is a result of the greater quantity of positive ions generated in the lower layers of the atmosphere. The aforementioned winds are shown to be disagreeable even to animals, and not only domesticated ones—if dogs and cats behave worse than their owners during these times, they are probably not doing it out of empathy with human beings but because of the increased concentration of positive ions.

The phenomenon has been extensively examined in several nations, and not only because of increased absences from work due to various symptoms such as respiratory stress, migraines, fatigue, and various other maladies. Indeed, the agitated and irritable states clearly caused by these environmental conditions have led some judges to extend greater leniency in sentences for crimes of passion. In some hospitals, elective surgery is even postponed or ionizers are employed in the operating rooms. Even some well-informed teachers know to expect less-brilliant results from their pupils during these days. In addition to consulting her horoscope, a university student, thinking about exams, might also want to check the weather forecast!

On the other hand, environments with a predominance of negative ions are shown to be effective in reducing states of stress, depression, and psychophysical maladies connected to stressful events. Various studies and experiments have noted the antidepressant effects of high concentrations of anions in the air. A published in the United States in 2013 presents a meta analysis of five different clinical studies that all confirmed a direct correlation between high negative ionization of environments and improvement in depressive states. Studies conducted in Japan and Romania showed that a high concentration of negative ions in the environment (greater than 10,000 ions per cubic centimeter) enables quicker recovery from intense physical exertion, with a normalization of blood pressure in less time (moreover, evidence of the latter phenomenon for subjects with hypertension has also been shown in a number of other studies). Studies also show that cognitive performance, especially memory, is improved in both children and adults. Even sleep disturbances can be significantly reduced when the environment is enriched with negative ions, while an excess in positive ions, as well as a scarcity in the total amount of ions in the air, produces the opposite effect. In individuals who suffer from chronic stress, positive ionization of the air does nothing but accentuate their symptoms, as well as weaken their immune defenses.

Various studies have ascertained that when positive ions in the air are absorbed by the human body, the majority transform into free radicals that can give rise to oxidation processes, increasing acidity in the body and accelerating aging and launched lab experiments to discover to what extent treatment with anions can reduce the risk of developing cancer. Even taking into account the need for further investigations and in depth studies, the preliminary results appear very favorable. Protecting ourselves from the excess of positive ions in the environment is thus very important, but it is not always easy given that 80 to 85 percent of ionic particles enter the body through the skin and the remaining 15 to 20 percent through respiration. Balancing both the excess positive ionization and the low level of ionization in our environments, the best solution is still improving air quality by increasing small anions, namely, those that are most biologically active. But it would also be wise to change our lifestyles, with more time spent in natural environments, which show greater potential negative ionization.

Moreover, there is no toxicity threshold regarding the absorption of negative ions: even with anionic concentrations greater than 100,000 units per cubic centimeter, no organ malfunctions or symptoms of illness were recorded for subjects participating in clinical experiments.

Advice on buying an Air-Ionizer

Now, if negative ions are scarce in the environments where we pass most of our time, is it possible to take care of this in some way and “produce” them? Air ionizers have been recommended (especially by their manufacturers) for purifying domestic and work environments of foul odors and pollutants. The idea that negative ions produced by a home appliance can reduce unpleasant odors in a closed space is not quite correct: doing this “dirty job” of ionization also generates a byproduct: ozone, Indeed, many devices labeled as ionizers or air purifiers have this “small” problem —that together with negative ions they emit ozone, which is widely known to be a greenhouse gas and which, more important, is an irritant for the respiratory system. Producing it at home is obviously not ideal, especially if we have small children living with us. Ozone has a characteristic strong odor, while ions have neither odor nor flavor. After all, these come from the ionization of certain molecules in the air, which is by definition an odorless, tasteless mix of gases.

Fortunately, our nose has a fairly low threshold for detecting ozone, ranging from 0.005 to 0.02 parts per million in the air, while irritation of the mucous membranes of the nose, eyes, and throat, together with headache and labored breath, arise at a concentration at least 5 to 10 times greater than this. Not all ionizers produce ozone; some models are built so as to minimize its production. When purchasing an environmental ionizer it is a good idea to make sure the manufacturer states the level of ozone emissions: look carefully, as many companies do not.

One more tip: a good ionization device should not have air filters, because it is likely most of the ions will be absorbed by these filters, and you’ll be left with a simple dust filter. In the absence of an ionizer, one can always resort to natural ventilation, opening the windows and letting the air circulate for a few minutes, especially after a storm or intense rain when the atmosphere is rich in negative ions.

The presence of plants inside our living spaces or closed workplaces can contribute significantly to improving air quality”, including ionization and reduction of pollutants. In calm, sunny weather conditions, if it cannot take a through a park rich in tall trees or a forest area. we can at east allow ourselves a long, lovely shower each day, taking advantage in small part of what physicists call the Lenard effect (ionization due to the collision of water particles). It’s not much, but it’s better than passing long periods of time in an ion-poor environment without attempting to mitigate its effects. So natural environments are able to furnish our body with the supply of ions it needs?

A Natural Ion Shower

Whenever there are bodies of moving water there is always negative ionization, created by the Lenard or “waterfall” effect: the greater the kinetic energy with which a body of water crashes into a solid object or disperses into the air, the more notable and effective the production of ions.

The separation of electrons from water molecules are electrically neutral, produces as many positive ions as negative. But while the majority of negative ions remain suspended in the air until their neutralization, positively charged water particles rapidly fall to the ground (or the body of water) after impact, and are in turn absorbed and deactivated, For this reason, the atmosphere near bodies of colliding water can contain as many as tens of thousands of small negative ions per cubic centimeter, these ions being much more active from a biological perspective as well.

The spectacular waterfalls formed by rivers large and small are an important source of negative ionization, with a high therapeutic quality. The shore of the sea or ocean can likewise be an effective source, with its power correlated to the water’s degree of agitation: it is clear that a wave colliding forcefully against a cliff can produce more negative ions than a placid lake lapping a sandy shore.

Of course you don’t have to take a cruise to Niagara Falls or watch the waves crash at Thunder Hole in New England’s Acadia National Park to get an effective ion shower. If well planned, even small waterfalls or fountains with a flow of water comparable to a stream can fill the air in their immediate vicinity with several thousand negative ions per cubic centimeter. The important thing is that the water’s movement is sufficiently high and forceful to produce a significant Lenard effect. The body of water’s beneficial activities can be further strengthened by lush greenery, with trees and shrubbery capable of holding the aerosols in the area benefiting from negative ionization. Even decorative elements such as fountains that reproduce the effect of small waterfalls can help in creating effective green spaces aimed at psychophysiological well-being. For several years the study of air ionization has also been extended to hydrothermal areas. Hot springs have biological and therapeutic properties due especially to their chemical composition and in part to their temperature. A thermal or hyperthermal complex in which waters circulate and bubble so as to release a high amount of aerosols into the air is definitely a significant source of small anions.

If you aren’t near an active body of water. other sources negative ions accessible to are natural environments characterized by sufficiently dense. lush and a prevalent, well-developed tree The ideal model would be a rainforest the lower density of a temperate forest. It isn’t difficult to find this model: many deciduous and coniferous forests can these characteristics. as long as there is enough water to guarantee good levels of humidity in the air. A good amount of natural light is thanks to the generation of negative ions during. photosynthesis. We can also infer that many forest environments suitable for the practice of forest bathing also have a sufficient level of air ionization, often with a predominance of negative ions positive ones.

If the forests are located in a mountain environment, it is likely that the amount of small anions is even more significant. There is also a natural tendency for negative ions to distance themselves from the earth’s surface, so those produced in adjacent foothill areas also “transit” through these higher altitudes.

Furthermore, there is another phenomenon found more frequently in mountainous areas that leads to air ionization. This is the “corona effect,” characterized by objects with a pointed shape such as certain mountain peaks or especially sharp rocks. Background radiation can bring greater potential energy to the point of an object exposed to open air, creating a spontaneous release of electrons.

These negatively charged atmospheric nitrogen and oxygen molecules initiate a chain reaction that lasts as long as the potential energy present at the pointed tip remains high. These different elements explain why a greater concentration of ions is recorded high in the mountains as opposed to, for example, an area with open fields.

Waterfalls, seashores, large bodies of moving water, hot springs, forests, woods, mountains: nature gives us a wide array of therapeutic opportunities. The forest bath and ion shower are two examples of the purifying, regenerative power of natural environments. Now it’s time to explore a new resource provided exclusively by contact with plant energy.

Spotting Narcissists and Psychopaths in Workplace

Why do so many incompetent men/women become leaders? )_how to fix them…yes this indeed the title of the book (Tomas Chamorro) that took my attention and the gist of the fixing is how to spot narcissism and psychopathy in workplace from this book, some excerpts below:

Why Bad Guys Win?

“He is a dreadful manager,” said a worker. “I have found  it impossible to work for him . Very often, when told a  idea, he will immediately attack it and say it is worth-  less or even stupid, and tell you that it was a waste of tune  to work on it. This alone is bad management, but if the  idea is a good one he will soon be telling people about it as  though it was his own.”  Few people would like the idea of working for such a  boss. And even fewer would expect a boss like that to be  held up as one of the best business leaders of all time. But  remarkably, the quote describes none other than Steve  Jobs, the founder of one of the most successful companies  in history. (The quote comes from Jef Raskin, who led  the design of the original Mac computer.) Apple has just  become the first trillion-dollar company in US history, even though it has not released a blockbuster product since Jobs’s death in 2011.

The Jobs paradox kept many puzzled, in part because it fits with a familiar archetype: exacting, visionary perfectionist who is turbocharged by unstoppable force of a gargantuan ego. unveilings, his stark uniform of black turtlenecks, and his megalomaniac mission, Jobs seemed to present a model for ambitious leaders to follow. It has even been said that he was capable of creating a cultish reality distortion when he talked about Apple products, convincing employees, investors, and suppliers that anything was possible. As we do with many tormented artists, we tend to see Jobs’s personality quirks as inseparable from his genius.

In reality, few leaders succeed when they are as difficult and badly behaved as Jobs was. A self-made billionaire with a flawed personality succeeds despite his or her character defects, not because of them. What makes the Jobs story a true exception is not only that he was hired back as Apple CEO—after being fired from his own company—but also that he achieved such extraordinary levels of success. As much as his fans would like to attribute Jobs’s unrivaled success to his eccentric and uncompromising personality, many narcissistic leaders have no problem distorting reality or coming up with colossal ideas or megalomaniac visions for the future. Their main problem is that they are not Steve Jobs, and without his talents, their delusions of grandeur will never become the next Apple. We have, alas, a tendency to generalize from unrepresentative examples, mostly because they are so memorable. Einstein’s lack of brilliance in his early years at school does not imply that bad grades will help you win a Nobel prize. Likewise, John Coltrane’s musical genius did not result from his heroin addiction—his talent somehow managed to survive the heroin. The only advantage of a difficult personality is that it may make a person unfit for traditional employment and can consequently propel them to launch their own business out of sheer necessity, if not revenge. But there is a big gap between being a mega-successful entrepreneur and being unemployable, and that gap is a function of talent rather than personality.

Many obnoxious leaders manage not only to remain employed but also to attain impressive levels of personal career success, despite their toxic personalities. To this end, this chapter explores the relationship between leadership and the two best-known examples of such toxic traits: narcissism and psychopathy. Looking at these two character traits will allow us to examine problematic leadership in more depth than we could by just talking about difficult bosses in general.

Spotting Narcissism

What do we mean when we say that someone is narcissistic? Primarily, narcissism involves an unrealistic sense of grandiosity and superiority, manifested in the form of vanity, self-admiration, and delusions of talent. Yet underlying this apparent superiority complex is often an unstable self- concept: because narcissists’ self-esteem is high but fragile, they often crave validation and recognition from others.

This craving is hardly surprising: if you are constantly showing off, you are probably desperate for others’ admiration. Such inner insecurity is rarely found in naturally humble people. Second, narcissists tend to be self-centered. They are less interested in others and have deficits in empathy, the ability to feel what others are feeling. For this reason, narcissists are rarely found displaying any genuine consideration for people other than themselves. A third defining feature of narcissism concerns high levels of entitlement. Narcissists commonly behave as if they deserve certain privileges or enjoy higher status than their peers enjoy. Examples abound: “Do I really need to apply for a promotion?” “Why didn’t I get a bigger bonus?” “Do I have to wait in line?” Such entitlements may justify narcissists’ exploitative behaviors at work and elsewhere. When you think you are better than others, you perceive unfairness where there is none and behave in demeaning and condescending ways toward people.

For decades, psychologists have devised and tested different tools for detecting narcissism. The most common method is self-report questionnaires, which simply provide respondents with a list of statements relating to their personal habits, preferences, or dispositions. Examples of these statements are “I am a natural leader” and “I am more talented than most of the people I know.” And if you think that this method is too transparent to work, you are wrong. A recent study led by Sara Konrath of Indiana University showed that you can spot whether someone is a narcissist “To what extent do you agree with with a single question, this statement: ‘I am a narcissist.’ Note: The word ‘narcissist’ means egotistical, self-focused, and vain.

Participants then answered the question on a scale of 1 (not very true of me) to 7 (very true of me). To the researchers’ surprise, narcissistic individuals were quite happy to con- fess to being narcissistic, and the single question captured people’s narcissism with an accuracy comparable to longer, well-established tests, which the researchers demonstrated in eleven studies. Narcissism was easily detected by the single question because narcissists are not only aware of their extraordinary self-love, but also proud of it, for they truly love loving themselves, unashamedly.

Nonetheless, various less transparent methods are also available to detect narcissism. For example, executives’ narcissism can be inferred from the size and attractiveness of their corporate profile picture, the number of times they are mentioned in their organizations’ brochures and press releases, and the frequency with which they use the word I and other self-referential pronouns.5 For CEOs, their narcissism can also be inferred from their compensation: the bigger their egos, the bigger the gap between their salaries and those of everyone else in their organizations!6 More recently, several studies have shown that you can detect narcissism from a person’s digital footprint. For example, sexier, more attractive, self- promoting Facebook pictures and, of course, an excess of selfies, all suggest narcissism.

Spotting Psychopathy

Let us now turn our attention to the other major dark-side trait. Psychopathy is often discussed in connection with leadership, particularly when it comes to famous political and business leaders. Unlike narcissism, which is wide- spread, psychopathy is rare. And yet few toxic character traits have attracted as much public fascination and media attention as psychopathy has—even though only about 1 percent of the general population is thought to have psychopathic tendencies. Perhaps part of our obsession with psychopaths stems from the disproportionate rate at which they seem to succeed. Professor Robert Hare, a pioneer in the field of criminal psychology and coauthor of the influential book Snakes in Suits, famously noted that “not all psychopaths are in prison; some are in the boardroom.” According to estimates he reports in a subsequent study, there are three times as many psychopaths in management roles than in the overall population. More recently, a much higher figure of around 20 percent (one in five) has been reported for another US corporate sample. This large range in variability reflects how people measure psychopathy, but psychopathy levels do increase with levels of career success.

So, what makes someone psychopathic? The first salient feature is a lack of moral inhibition, which at an extreme is manifested in the form of strong antisocial tendencies and an intense desire to break the rules, even just for the sake of doing so. And when psychopaths do break the rules, they feel no guilt or remorse to avoid a repeat of events. people with psychopathic tendencies are also more prone to making reckless behavioral choices. For instance, psychopaths are more likely to drink, smoke, take drugs, and have promiscuous sexual relations and extramarital affairs. To be clear, not all adrenaline junkies are psychopathic, but the vast majority of psychopathic individuals are thrill seekers, and their reduced concern for danger will put them and others at risk. A third defining feature of psychopaths is their lack of empathy. They don’t care about what others are thinking or feeling, despite being able to understand those feelings.24 As a result, psychopaths are known for their cold dispositions. The absence of empathy is probably a major cause for their lack of moral constraints; it’s obviously much harder to behave in prosocial ways when you don’t care about people.

Unsurprisingly for a trait once described as “the mask of sanity,” psychopathy is not easily detected by laypeople.39 For this reason, you want to be alert to the potential risks of basing hiring decisions on short-term interactions with candidates. In fact, given their deceptive nature, fearless attitude, short-term likability, and skilled impression management, you can expect psychopaths to perform quite well during job interviews.40 Yet, just as you wouldn’t marry someone after only a first date, you should not select some- one for a leadership job solely because of the person’s inter- view performance—which is exactly that, a performance. Psychopaths are hard to detect, but you can simultaneously evaluate a leader’s psychopathy and predict its effects on his or her subordinates. How? You can simply ask the leader’s subordinates to rate their boss on critical indicators of psychopathy In one study, for example, employees were asked to rate their bosses on various personality aspects, such as “can make a joke out of anyone,” ‘enjoys being disruptive,” and “is not sincere.

Scientist have developed concise measures of psychopathy, such as the Short Dark Triad assessment.42 With just fifteen self-report statements, you can get a good sense of an individual’s psychopathy level. Here are some of those statements:

I am a thrill seeker.
I like to get revenge on authorities.
I never feel guilty
People who mess with me always regret it.

Of course, test takers can certainly fake their answers by portraying themselves in a less psychopathic way and by presenting a much more prosocial and conformist aspect of their personality. But such misrepresentation doesn’t happen enough to invalidate this test. Rather, people with psychopathic tendencies seem proud to answer honestly or at least are too defiant to hide their views, perhaps because they have little guilt about their personality or care little about what others make of them. their psychopathic inclinations. For example, a study found that the number of selfies people post on social media reliably indicates their psychopathy level. Psychopathy can also be detected in language, as psychopathic individuals speak and write in a more dominant and coercive way and express more aggression and irritability. For instance, the tendency to swear is a consistent indicator of higher psychopathy. Another linguistic feature associated with psychopathy is the proclivity to talk about power, money, sex, and physical needs, whereas lower-psychopathy individuals tend to talk more about family, friends, and spirituality. In short, we have numerous intuitive signals to detect people with psychopathic tendencies.

Tesla’s “World System” in 1900 mirrors our Wireless Obsession in 2020

Tesla’s “World-System” of wireless transmission which he  undertook to commercialize on his return to New York  in 1900. This system kind of presages what he envisaged a world to become and in our days, we could reflect upon this great stalwart’s vison of the future and rest on his laurels of inventions. He could clearly see where the world would be headed in centuries to come. An impressive list indeed to achieve. And in his own words: as below (Source: My Inventions Nikola Tesla):

—————

As to the immediate purposes of my enterprise,  they were clearly outlined in a technical statement of  that period from which I quote, “The ‘World-System’  has resulted from a combination of several original  discoveries made by the inventor in the course of long  continued research and experimentation. It makes  possible not only the instantaneous and precise wireless  transmission of any kind of signals, messages or characters, to all parts of the world, but also the  inter-connection of the existing telegraph, telephone,  and other signal stations without any change in their present equipment. By its means, for instance, a tele-  phone subscriber here may call up and talk to any other  subscriber on the globe. An inexpensive receiver, not  bigger than a watch, will enable him to listen anywhere,  on land or sea, to a speech delivered or music played  in some other place, however distant. These examples  are cited merely to give an idea of the possibilities of  this great scientific advance, which annihilates distance  and makes that perfect natural conductor, the Earth,  available for all the innumerable purposes which human  ingenuity has found for a line-wire. One far-reaching  result of this is that any device capable of being operated  through one or more wires (at a distance obviously  restricted) can likewise be actuated, without artificial  conductors and with the same facility and accuracy, at  distances to which there are no limits other than those  imposed by the physical dimensions of the globe. Thus,  not only will entirely new fields for commercial exploitation be opened up by this ideal method of transmission  but the old ones vastly extended. The ‘World-System’  is based on the application of the following important  inventions and discoveries: 

1. The ‘Tesla Transformer.’ This apparatus is in  the production of electrical vibrations as revolutionary as gunpowder was in warfare.  Currents many times stronger than any ever  generated in the usual ways, and sparks over  100 feet long, have been produced by the  inventor with an instrument of this kind.

2. The ‘Magnifying Transmitter.’ This is Tesla’s best invention, a peculiar transformer specially  adapted to excite the Earth, which is in the  transmission of electrical energy what the telescope is in astronomical observation. By the  use of this marvelous device he has already set  up electrical movements of greater intensity  than those of lightning and passed a current,  sufficient to light more than two hundred  incandescent lamps, around the globe. 

3. The ‘Tesla Wireless System.’ This system  comprises a number of improvements and is  the only means known for transmitting  economically electrical energy to a distance  without wires. Careful tests and measurements  in connection with an experimental station of  great activity, erected by the inventor in  Colorado, have demonstrated that power in  any desired amount can be conveyed, clear  across the globe if necessary, with a loss not  exceeding a few percent. 

4. The ‘Art of Individualization.’ This invention  of Tesla’s is to primitive ‘tuning’ what refined  language is to unarticulated expression. It  makes possible the transmission of signals or  messages absolutely secret and exclusive both  in the active and passive aspect, that is, non-interfering as well as secure. Each signal is like  an individual of unmistakable identity and there  is virtually no limit to the number of stations or instruments which can be simultaneously operated without the slightest mutual disturbance.

5. ‘The Terrestrial Stationary Waves.’ This wonderful discovery, popularly explained, means that the Earth is responsive to electrical vibrations of definite pitch just as a tuning fork to certain waves of sound. These particular electrical vibrations, capable of powerfully exciting the globe, lend themselves to innumerable uses of great importance commercially and in many other respects.

The first ‘World-System’ power plant can be put in  operation in nine months. With this power plant it will  be practicable to attain electrical activities up to ten  million horsepower and it is designed to serve for as  many technical achievements as are possible without  due expense. Among these the following may be  mentioned: 

1. The inter-connection of the existing telegraph  exchanges or offices all over the world; 
2. The establishment of a secret and secure  government telegraph service; 
3. The inter-connection of all the present telephone exchanges or offices on the globe; 
4. The universal distribution of general news, by  telegraph or telephone, in connection with the  press;
5.The establishment of such a ‘World-System’ of intelligence transmission for exclusive  private use; 
6. The inter-connection and operation of all  stock tickers of the world; 
7. The establishment of a ‘World-System’ of  musical distribution, etc.; 
8. The universal registration of time by cheap  clocks indicating the hour with astronomical  precision and requiring no attention whatever; 
9. The world transmission of typed or hand-  written characters, letters, checks, etc.; 
10. The establishment of a universal marine  service enabling the navigators of all ships to  steer perfectly without compass, to determine  the exact location, hour and speed, to prevent  collisions and disasters, etc.; 
11. The inauguration of a system of world-printing  on land and sea; 
12. The world reproduction of photographic  pictures and all kinds of drawings or records.” 

I also proposed to make demonstrations in the wire-  less transmission of power on a small scale but  sufficient to carry conviction. Besides these I referred  to other and incomparably more important applications of my discoveries which will be disclosed at some  future date.

Tesla’s Inventor Instincts in School Days

Had an opportunity to delve into Nikola Tesla’s “My Inventions” a book co-written by him around 1900. He’s a prolific inventor of many electrical devices and apparatuses – equal in ambition, inventiveness and precocity to that of Edison.  Perhaps this from his own book provides an insight of what his formula and how he used his visions in due course to come with great many inventions, notably alternating current (AC) generator and transmission – quintessential to our days’ ubiquitous power generation and transmission needs topping that with wireless power transfer as well. Excerpts from his book:

His visions and how he took care of his health:
I all dwell briefly on these extraordinary experiences, on account of their possible interest to students of psychology and physiology and also because this period of agony was of the greatest consequence on my mental development and subsequent labors. But it is indispensable to first relate the circumstances and conditions which preceded them and in which might be found their partial explanation.

From childhood I was compelled to concentrate attention upon myself. This caused me much suffering but, to my present view, it was a blessing in disguise for it has taught me to appreciate the inestimable value of introspection in the preservation of life, as well as a means of achievement.

The pressure of occupation and the incessant stream of impressions pouring into our consciousness through all the gateways of knowledge make modern existence hazardous in many ways. Most persons are so absorbed in the contemplation of the outside world that they are wholly oblivious to what is passing on within themselves.

The premature death of millions is primarily traceable to this cause. Even among those who exercise care it is mistake to avoid imaginary, and ignore the a common real dangers. And what is true of an also applies, more or less, to a people as a whole.
Witness, in illustration, the prohibition movement.A drastic, if not unconstitutional, measure is now being put through in this country to prevent the consumption of alcohol and yet it is a positive fact that coffee, tea,tobacco, chewing gum and other stimulants, which are freely indulged in even at the tender age, are vastly more injurious to the national body, judging from the number of those who succumb. So, for instance, during my student years I gathered from the published necro-logues in Vienna, the home of coffee drinkers, that deaths from heart trouble sometimes reached 67% of the total. Similar observations might probably be made in cities where the consumption of tea is excessive. These delicious beverages super-excite and gradually exhaust the fine fibers of the brain. They also interfere seriously with arterial circulation and should be enjoyed all the more sparingly as their deleterious effects are slow and imperceptible. Tobacco, on the other hand, is conducive to easy and pleasant thinking and detracts from the intensity and concentration necessary to all original and vigorous effort of the intellect. Chewing gum is helpful for a short while but soon drains the glandular system and inflicts irreparable damage, not to speak of the revulsion it creates. Alcohol in small quantities is an excellent tonic, but is toxic in its action when absorbed in larger amounts, quite immaterial as to whether it is taken in as whiskey or produced in the stomach from sugar. But it should not be overlooked that all these are great eliminators assisting Nature, as they do, in upholding her stern but just law of the survival of the fittest. Eager reformers should also be mindful of the eternal perversity of mankind which makes the indifferent “laissez-faire” by far preferable to enforced restraint.

The truth about this is that we need stimulants to do our best work under present living conditions, and that we must exercise moderation and control our appetites and inclinations in every direction. That is what I have been doing for many years, in this way maintaining myself young in body and mind. Abstinence was not always to my liking but I find ample reward in the agreeable experiences I am now making. Just in the hope of converting some to my precepts and convictions I will recall one or two.

Just the One Here:
I fell into a worse predicament once. There was a large flour mill with a dam across the river near the city where I was studying at that time. As a rule the height of the water was only two or three inches above the dam and to swim out to it was a sport not very dangerous in which I often indulged. One day I went alone to the river to enjoy myself as usual. When I was a short distance from the masonry, however, I was horrified to observe that the water had risen and was carrying me along swiftly. I tried to get away but it was too late. Luckily, though, I saved myself from being swept over by taking hold of the wall with both hands. The pressure against my chest was great and I was barely able to keep my head above the surface. Not a soul was in sight and my voice was lost in the roar of the fall. Slowly and gradually I became exhausted and unable to withstand the strain longer. Just as I was about to let go, to be dashed against the rocks below, I saw in a flash of light a familiar diagram illustrating the hydraulic principle that the pressure of a fluid in motion is proportionate to the area exposed, and automatically I turned on my left side. As if by magic the pressure was reduced and I found it comparatively easy in that position to resist the force of the stream. But the danger still confronted me. I knew that sooner or later I would be carried down, as it was not possible for any help to reach me in time, even if I attracted attention. I am ambidextrous now  but then I was left-handed and had comparatively little  strength in my right arm. For this reason I did not dare  to turn on the other side to rest and nothing remained  but to slowly push my body along the dam. I had to  get away from the mill towards which my face was  turned as the current there was much swifter and deeper.  It was a long and painful ordeal and I came near to  failing at its very end for I was confronted with a depression in the masonry. I managed to get over with the last  ounce of my force and fell in a swoon when I reached  the bank, where I was found. I had torn virtually all the  skin from my left side and it took several weeks before  the fever subsided and I was well. These are only two  of many instances but they may be sufficient to show  that had it not been for the inventor’s instinct I would  not have lived to tell this tale.

Tolstoy’s Love & Literature

I had intimate interactions with Tolstoy by reading his biography and scrapes of non-fiction. There used to be saying that in his time, there were two tsars in Russia, one was him and the other, the real tsar, such was his fame from literary works. His early works figured human experiences like love, hope, marriage, death, etc. with fiction and in his late forties as he was grapping with human frailties, idiosyncrasies and ironies and addressing them was his pre-occupation which led to essays and works that attempted to answer them in his style and substance. Countless people have been inspired by his non-fiction writing including Gandhi, whose understanding and practice of non-violence, civil dis-obedience, self-industriousness, simplicity stems from Leo and they had had an active letter communication. Tolstoy’s reading, I suppose should have been varied and perhaps might have had a penchant for eastern thoughts as his later year’s non-fiction emphasized brotherhood, humanism and love. Recently had a chance to read ‘Leo Tolstoy – A Very Short Introduction’ by Liza Knapp, a perfect companion for those who want to delve deeper into Tolstoy’s works and to understand him more. Some short excerpts that I want to capture to highlight his  thoughts on Love and his literary devices from this very short introduction.

Love – through his experiences and writings:

Tolstoy’s brothers took him to a brothel soon after his fourteenth birthday. After he ‘committed the act’, he ‘stood by the woman’s bedside and wept’. From this loss of virginity, through love affairs and marriage, through the birth of over a dozen children, through advocacy of chastity later in life (before the birth of his youngest children), to acrimony between him and his wife that ended in him leaving home right before his death, Tolstoy’s love life has been extensively documented and debated. Tolstoy himself addressed all this directly in letters, in diaries, and in frank conversation with memoirists; others involved, including his wife, also left their own accounts.

Tolstoy’s major fiction, known for its autobiographical and ‘autopsychological’ elements, roughly follows the trajectory of Tolstoy’s life and loves. Tolstoy begins with the yearnings of the motherless child for love and the myths of happy Russian gentry childhood in Childhood, Boyhood, Youth. He celebrates Russian marriage and family life (and tames sexual desire) in War and Peace. He then explores adulterous passion, the tender joys of conjugal love, and family unhappiness in Anna Karenina. Ile excoriates sex, marriage, and family life in ‘The Kreutzer Sonata’. In his late love fiction, Tolstoy favors quests to expiate sexual guilt or else parables of love and death among simple folk.

In his depictions of loving families and of romantic passion, conjugal or adulterous, Tolstoy includes reminders of the call to love God and neighbor, as he, along his characters, asks whether it’s possible somehow to combine, balance, or reconcile these loves. Or do certain forms of love, by their very nature, exclude others? Tolstoy’s heroes often have insights that look ahead to Sigmund Freud’s questions about loving your neighbor: doesn’t it conflict with one’s duty to ‘one’s own people’? Tolstoy himself came to see both sex and family life as service of the self rather than of God. Tolstoyan heroes are often haunted by the Ant Brothers’ dream of love and happiness for all. It may be an insurmountable obstacle on their course toward family happiness or sexual bliss.

As Tolstoy put it in his Confession, marriage and family life seemed to bring him happiness and give meaning to his life, until he began to regard them, like all worldly activities, as a diversion from what really mattered, which was faith. He needed to find ‘meaning in life that inevitable death awaiting [him] would not destroy’ (5, 24). Leaving children or Anna Karenina behind after death was not a consolation to him. Nor were they enough to divert him any longer. But even fiction written before this realization reveals Tolstoy’s anxieties about loving and being loved in the face of death.

Literature – devices employed

Realism

As Dmitry Merezhkovsky saw it, ‘there is simply no writer equal to Tolstoy in depicting the human body’. Tolstoy is a ‘seer of the flesh’ while Dostoevsky is a ‘seer of the spirit’. When compared to Victorian novelists writing at the same time, Tolstoy does much more with the flesh. Tolstoy is often praised for his childbirth scenes, notably the one that tracks Levin as Kitty gives birth in Anna Karenina.

Tolstoy has also been accused of excessive ‘naturalism’, that is, for his excessive interest in human bodies. For example, the epilogue of War and Peace dwells on the transformation of Natasha, now married and the mother of four, into nothing but a maternal body, whose face is bereft of its former animation and whose soul is said to be not even ‘visible’ (Epilogue 1:10, 1242). Tolstoy’s view of motherhood as the end-all and be-all of women’s existence has disappointed and infuriated many readers. Still, his respect for motherhood was genuine. In War and Peace, he presents it as much more meaningful than the struggle for glory on the battlefield or for power in the empire.

Defamiliarization, or ‘looking at things afresh’

In his 1917 essay ‘Art as Technique’ (sometimes translated ‘Art as Device’), the Russian critic Viktor Shklovsky heralded Tolstoy a.s a master, although not the sole practitioner, of an artistic technique that he called defamiliarization. The technique consists of de-automatizing perception by presenting what is already familiar to the reader in a new way that makes it seem unfamiliar and strange. Rather than naming something, the artist will describe it.

Before it was christened ‘defamiliarization’, Tolstoy’s technique had already attention. In Tolstoy and his Message (1904), Ernest Crosby, an American follower of Tolstoy, took special note of Tolstoy’s ‘habit of looking at things afresh as if no one had ever considered them before’. The Russian literary critic Vinokenty Veresaev, writing in 1911, described what he saw as ‘an extremely unique at work in Tolstoy’s writings: ‘it is as if an attentive, all-noticing child looked at a phenomenon and described it, not to convention, but simply the way it is, such that all the habitual, hypnotizing conventions fall away from the phenomenon so that it appears in all its bare absurdity.’

Another example: : ‘And then, one after another, living men are pushed off the benches which are from under their feet, and by their own weight suddenly tighten the nooses round their necks and are painfully strangled. Men, alive a minute before, become corpses dangling from a rope, at first swinging slowly and then resting motionless’ (395-6). This is a classic example of Tolstoy’s technique of defamiliarization. here, Tolstoy uses it to denounce and to expose, as as to ‘prick at our conscience.’

This technique of defamiliarizing lends itself quite naturally to the critique of culture and convention that is embedded in Tolstoy’s fiction and made overt in his non-fiction. Through defamiliarization, Tolstoy attacks assumptions taken as axiomatic. He subjects all to analysis and takes nothing on faith.

Depicting inner life

When Tolstoy first broke into print with Childhood and the Sevastopol tales, Russian reviewers noticed something special in the way Tolstoy depicted the inner life of his characters. He even caught, in the words of Dmitry Pisarev, ‘the mysterious, unclear movements of the soul that have not reached consciousness and are not completely understood even by the person who experiences them’. As his fiction became more widely known outside of Russia in the early 20th century, Tolstoy’s powers for rendering consciousness continued to astound. Virginia Woolf wrote, ‘Tolstoy seems able to read the minds of different people as certainly as we count the buttons on their coats.’ Although Tolstoy gravitates to the interiority of some characters more than others, he gives us at least fleeting access to a large number of different form of consciousnesses. For Tolstoy, a glimpse is often enough to reveal a whole soul.

Throughout his career, Tolstoy developed his techniques for presenting consciousness, moving from rendering what goes on in the mind of a young officer facing his first bombardment at Sevastopol, to what is often hailed as the prototype of ‘stream of consciousness’ (a technique dear to Woolf and fellow Modernists) in Anna Karenina’s final moments. Tolstoy moves in and out of Anna’s consciousness in the last four chapters of Part 7, capturing those mysterious movements of her soul as she becomes more and more unstrung. In this respect, they recall her disjointed ramblings out loud as she lay dying after giving birth in Part 4. Those ravings, which contained more meaning and more truth than her normal speech, brought Vronsky, Karenin, and herself together in love and forgiveness in the face of death.

Time & Plot

Tolstoy consciously and willfully deviated from conventions of genre and expectations about unity, form, and plot. In a draft of a foreword to the first part of 11tzr and Peace, Tolstoy declared that he was not writing the kind of conventional novel culminating in ‘a happy or unhappy denouement’—like marriage or death—that would ‘destroy the interest of the narrative’. The interest of his narratives transcends these conventional endings. Tolstoy seemed to foresee the accusations of artlessness that would later be made against him. But Tolstoy was determined to ‘be true to [his] own practices and [his] own powers’ instead of following the norms of the novel. Tolstoy’s desire to capture human experience in its natural rhythms is reflected in his early habit of naming his works after units of time, such as mornings and years. Thus, when the first part of War and Peace was conceived and published, it wms called simply The Year 1805′. Tolstoy was following what was already a pattern for him. For example, his first attempt at fiction was called ‘A History of Yesterday’ and his first attempt at a novel was the trilogy Childhood, Boyhood, Youth. What was to be the final part published in aborted form as ‘Landowner’s Morning’. And the titles of his early war stories, ‘Sevastopol in December’, ‘Sevastopol in May, and ‘Sevastopol in August, 1855’, all favor time over plot. Instead a chronicle of a given time period, Tolstoy describes selected moments that capture its essence. Thus, in Childhood, Boyhood, Youth, Tolstoy avoids the more continuous Sequential plot of Dickens’s David Copperfield (1850), a novel that in so many respects was one of his key inspirations. Dickens begins with a chapter called ‘I Was Born’ and carries on to the novel’s denouement in (re)marriage.

Simile and the power of comparison

George Steiner has suggested that whereas novels recall tragedies, Tolstoy’s recall epics. One device Tolstoy borrows from the epic is the simile, a comparison, often long, that (in English) usually begins with ‘like’ or ‘as’ This device marks a given point in the action, then carries the focus elsewhere to the realm of the comparison, and brings that elsewhere to bear on the point in the action that prompted the simile. A simile involves an associative or intuitive leap. Homer used similes to draw the different realms of his epic universe together and to include all aspects of life in it. Tolstoy would do the same.

Tolstoy also often compared humans to animals or to other aspects of the natural world. A simile that recalls one that Virgil used for Carthage in the Aeneid likens Moscow, emptied out as Napoleon arrives, to a queenless beehive—and develops the analogy for a whole chapter (3.3:20, 938—40). At other points, the novel compares the troops marching to the forces at play in the natural world, in order to prepare the reader for Tolstoy’s theorizing about human agency and God’s role in the universe. The simile invites us to wonder: do the same forces govern both the natural and the human realms?

In his Lectures on Russian Literature, Vladimir Nabokov remarked on a whole class of Tolstoyan similes that follows the formula: someone felt ‘like a person who.. .’. Thus, in Anna Karenina, Karenin, on discovering the emotional distance that o has arisen between him and Anna as a result of her passion for Vronsky, experiences ‘a sensation’ ‘like that which might be experienced by someone who has returned home and found his house locked’ (2:9, 148). These similes make the particular experience of a character accessible and suggest that this experience reveals something about human experience.

Similes are a special case of Tolstoy’s general fascination with comparisons. Within the action of his works, Tolstoy shows his characters engaging in the mental act of making comparisons. For example, Anna Karenina depicts Kitty comparing Levin and Vronsky, her two suitors, imagining each on his own and then ‘both together’ (1:13; 49). Hadji Murat describes Tsar Nicholas thinking about the frightened look of the young woman he has just seduced, and ‘now of the full, powerful shoulders of his established mistress, Nelidova’, and then ‘comparing the two’ (15, 421).

As a young man, Tolstoy had already recognized the making comparisons as to the of the In a notebook he kept during the spring of 1847 (as he was dropping out of Kazan University), as Tolstoy set forth a program for self-improvement on various fronts, he identified ‘five main mental faculties.’ They are: ‘the faculty of conceptualization, the faculty of memory, the faculty of comparison, the faculty of drawing conclusions from these comparisons. and finally the faculty of putting these conclusions in order (46:271). From the rest of the diary entry, it’s clear that Tolstoy was especially interested in comparisons. Mental operations for him were not just on a linear axis, with one thing connecting directly to another. He was aware that the human mind works by leaps and bounds and he muted to harness these his art.

Hidden symmetries

The young Tolstoy also composed a ‘philosophical treatise’ on Symmetry, a foundational concept in aesthetics that hinges On similarity and opposition. The manuscript didn’t survive.However, in Boyhood. the hero Nikolai is seen musing, with chalk in hand, about symmetry, wondering whether it is an inborn feeling, and wanting to know if human life is governed by it. When Tolstoy turned to writing. he made use of mental operations, aesthetic principles, and artistic based in comparisons, opposition, and symmetry —the concepts that he identified and described in his philosophical and aesthetic enquiries of younger years.

Flaunted or Flouted Admissions

Recently was reading a wired article which was criticizing the ingenuity in forgery committed to get rich kids to colleges they never deserve by forging the persona and their achievements, essentially flout the sport quota to gain entry into prestigious institutions by backdoor. The article’s way of disparaging the rich’s laziness being compensated by their money is worth reading and have given the excerpt here which caught my attention….well said — you need to earn by sweat & dedication and not by indolence & ostentation!!

Unlike actual education and sweaty sports, brands live in two dimensions, now almost always on screens. They can be conjured and rebuilt in short order by Singer types adept in damage control—scrubbing Google references, working the media, and creating an illusion of trustworthiness.

“The creativity was perversely impressive,” said Allen Koh in a video interview with The Wall Street Journal. Koh, another opulently compensated college consultant, once considered Singer a competitor. “I’ve heard of résumé exaggeration; I’ve never heard of wholesale fabrication.”

But the real Key to the story of the admissions cheats is that they left the fabrication—the dirty work—to someone else. Fabrication comes from the Latin for forge; Singer did the forgery so his clients didn’t have to. Thorstein Veblen’s master­piece, The Theory of the Leisure Class, argues that the rich leave “production” of any kind to the middle and working classes, and then flaunt gloriously unproductive antiwork, largely to mark their distance from their subordinates. This wanton squandering of resources is what Veblen called “conspicuous consumption.”

The twist of our time is that the rich are no longer content to fritter away millions on pieds-à-terre in Paris and kennels of Tibetan mastiffs. Oh, no: What’s insidious about the college admissions scandal is the suggestion that at least some of the leisure class got tired of coming across as jackasses in garish watches and decided to spare no expense to seem like striving bourgeois warriors.

But the new pose of the rich as hard workers is indeed another version of consuming conspicuously. What more spiteful thing for the entitled to do than rob others of even their most earned, try-hard moments … without lifting a finger? Everyone in the scam is set up to protect the leisure, languor, ignorance, and ego of the cosseted student—the coaches, the parents, the bribed officials, the friends who know she doesn’t play sports. Even as they say they only half-knew what was going on, the flagrancy of their deceptions suggest they didn’t mind if it was an open secret. Like the bronzed glamour girls of Biarritz in the 1920s, the admission-­scandal families cherish their laziness, as well as the sight of money being vaporized. How else to demonstrate to the less fortunate how big their surplus is?

If the parents had really wanted their kids to be handsomely educated, they would have gone the traditional route. True, it kinda sucks and doesn’t guarantee USC, but maybe perseverance—for one or two of these kids—could have been an actual plan of action, and not just a brand made in a lab. Here’s how it goes, and it doesn’t require Photoshop: You practice horrid, abdomen-­splitting butterfly strokes, study polynomials deep into the night, and go early to goddamn class, sit in the front, read all the Kant, and earn actual As.

The Wicked Wit of Winston Churchill

Some wit from those compiled, edited and introduced by Dominique Enright.

Just before the wit some wisdom – notable works of WSC to refer and get a bit of wisdom:

  1. The World Crisis
  2. Marlborough
  3. The Second World War
  4. The Histories of English Speaking Peoples
  5. Besides numerous volumes of speeches, broadcasts, volumes of autobiography, a biography of his father and one rather poor novel Savrola

Be Killed by Many Times

  1. The world today is rules by harassed politicians absorbed in getting into office or turning out the other man so that not much room is left for debating great issues on their merits
  2. On the qualities required by a politician: ‘The ability to foretell what is going to happen tomorrow, next week, next month, and next year. And to have the ability afterwards to explain why it didn’t happen.’
  3. ‘No one pretends that democracy is perfect or all wise. Indeed, it has been said that democracy is the worst form of Government except all those other forms that have been tried from time to time.’
  4. Of WSC’s (then) fellow Conservatives: ‘They are a class of right honorable gentlemen — all good men, all honest men — who are ready to make great sacrifices for their opinions, but they have no opinions. They are ready to die for the truth, if only they knew what the truth was.’
  5. ‘Some men change their party for the sake of their principles; others change their principles for the sake of their party.’
  6. ‘To improve is to change; to be perfect is to change often.’

Terminological Diversions: Words

  1. ‘Men will forgive a man anything except bad prose.’
  2. ‘We must have a better word than “prefabricated”. Why not “ready-made”?’
  3. ‘Short words are best and the old words when short are best of all’

Pigs treats US Equals

  1. On being advised his fly buttons were undone: ‘Dead birds don’t fall out of their nests.’
  2. ‘An appeaser is one who feeds a crocodile hoping that it will eat him last.’

Casting pearls: Speeches

  1. ‘I’m going to make a long speech because I’ve not had the time to prepare a short one.
  2. ‘ On verbosity: ‘It is sheer laziness not compressing thought into a reasonable space.’
  3. On MP Lord Charles Beresford: ‘He is one of those orators of whom it was well said, “Before they get up, they do not know what they are going to say; when they are speaking, they do not know what they are saying; and when they have sat down, they do not know what they have said.” ‘ ‘I can well understand the Honorable Member’s wishing to speak on. He needs the practice badly.’
  4. ‘Let us not shrink from using the short expressive phrase even if it is conversational.’

Friends Like These

  1. On Stanley Baldwin: ‘The Government cannot make up their minds, or they cannot get the Prime Minister to make up his mind. So they go on, in strange paradox, decided only to be undecided, resolved to be irresolute, adamant for drift, solid for fluidity, all powerful to be impotent.’
  2. ‘It is a fine thing to be honest, but it is also very important to be right.’
  3. On Field Marshal Sir Bernard ( 1st Viscount Montgomery of Alamein, one of Britain’s most successful Military leaders): ‘In defeat, unbeatable; in victory, unbearable.’
  4. On Clement Attlee: ‘A sheep in sheep’s clothing.’ ‘If any grub is fed on Royal Jelly it turns into a Queen Bee.’ ‘He is a modest man who has a good deal to be modest about.’ ‘An empty taxi arrived at 10 Downing Street, and when the door was opened Attlee got out.’
  5. On Herbert Morrison (Labor statesman, and deputy PM in Attlee’s administration, 1945—51): ‘A curious mixture of geniality and venom.’
  6. On Stafford Cripps (Labor statesman, Chancellor the 1947-50): •He has all the virtues I dislike and none of the vices I admire.’
  7. On George Bernard Shaw: ‘Few people practice what they preach and none less so than George Bernard Shaw Saint, sage and clown; venerable, profound and irresistible.’

Of Course I’m an Egoist

  1. ‘Eating my words has never given indigestion.’
  2. On his friend Leo Amery. • I shall suck to you with all the loyalty o fa leech.’
  3. ‘l always manage somehow to adjust to any new level of luxury Without whimper or complaint. It is one of my more winning traits.’
  4. On his seventy-fifth birthday: ‘I am ready to meet my Maker. Whether my Maker is ready for the ordeal of meeting me is another matter.’
  5. On attending a dinner for the Prince of Wales, later Edward 1711, in 1896: ‘I realized that I must be on my best behavior punctual, subdued, reserved — in short, display all the qualities with which I am the least endowed.’

What kind of People

  1. ‘The English never draw a line without blurring it.’

An ineradicable habit: Drink

  1. ‘A single glass of champagne imparts a feeling of exhilaration. The nerves are braced, the imagination is agreeably stirred, the wits become more nimble. A bottle produces the contrary effect.’
  2. ‘Good cognac is like a woman. Do not assault it. Coddle and warn) it in your hands before you sip it.’

‘Our maxims will remain’: Epigrams

  1. Among his most famous words is the epigraph: In war: resolution In defeat: defiance In victory: magnanimity In peace: goodwill.
  2. ‘A fanatic is one who can’t change his mind and won’t change the subject.’
  3. ‘It is a fine thing to be honest, but it is also very important to be right.’
  4. ‘Youth is for freedom and reform, maturity for judicious compromise, and old age for stability and repose.’
  5. ‘Diplomacy is the art of telling plain truths without giving offence.’
  6. ‘Never stand so high upon a principle that you cannot lower it to suit the circumstances.’
  7. ‘It is always wise to look ahead, but difficult to look farther than you can see.’
  8. You will never get to the end of journey if you stop to shy a stone at every dog that barks.’
  9. ‘Virtuous motives, trammeled by inertia and timidity, are no match for armed and resolute wickedness.’
  10. ‘The worst quarrels only arise when both sides are equally in the right and in the wrong.’

Fathoming the Deep in Deep Learning – A Practical Approach

Deep in ‘Deep Learning’ is elusive yet approachable with a bit of mathematics. This beckons a practical question: Is elementary calculus sufficient to unravel deep learning? The answer is yes indeed. Armed with an unbound curiosity to learn and re-learn new and old alike and possibly if you can methodically follow below sections, I reckon you’ll cross the chasm to intuitively understand and apply every concepts including calculus in their glory to de-clutter all intricacies of deep learning.

Learning ‘deep learning’ is fun yet fraught. Mathematics is fundamental and can help understand clearly and etch the concepts indelibly. When I started on Michael Nielsen’s book ‘Neural Networks and Deep Learning’, I stumbled to understand thoroughly at first go with lots of basic concepts involved requiring a refresher and I thought if I could share my journey with others to simplify their learning, then I think I’ve done my little part in facilitating Deep Learning.

In this post, I’m covering the steps I took and what I researched, read and understood – being captured to reveal each concept as intuitively as possible and additional topics that piques your interest. Each step detailed under a heading, to understand the whole edifice of Deep Learning, the cornerstone of AI which is still evolving. Feel free to skip sections familiar and dive deep that interests much.

Steps to fathom the depth:

  1. The Beginnings – Modelling Decisions with Perceptrons
  2. Workhorses inside Nodes – Activation Functions
  3. A Gentle Detour on Basics – Differential Calculus
  4. The Underpinnings – Essential Statistics and Loss Reduction
  5. The Grand Optimization – Gradient Descent
  6. Intuitive Examples to the Rescue – Descent Demystified
  7. Ensemble directed Back & Forth – Feed Forward & Back Propagation
  8. Inner Workings of Bare NeuralNet – Matrices matched to Code
  9. Learning Curve Retraced – References & Acknowledgements

The Beginnings – Modelling Decisions with Perceptrons

Top

Biological neuron models explain the underlying mechanism by which nervous system functions to facilitate perception be it sensory or its allied activities. From this perception came ‘perceptron’ an allegorical mathematical unit or model. Simply put, a perceptron (1) takes multiple inputs, (2) weighs (multiplies by a set weight), (3) sums and adds up to a constant called bias and finally (4) sends to an activation function to (5) provide an output.

Typically you could visualize a perceptron as a mathematical model to decide whether I can go for a movie today given a set of inputs and a threshold value. Lets say the inputs are cloud cover, holiday occurrence, time of the day – all as a scaled real number, which are each modified by multiplying to their corresponding preset weighage and added to a threshold value (bias factor) and then this combined number is stepped or smoothed out wherein the resultant of a stepper function is 1 or 0 whereas a smoother function the value ranges between 1 or 0 – which can be interpreted as a probability of the result happening.

To make it clear, below ia an illustration of the above example:

mathematically, if we have j number of perceptron inputs, the output is stated as:

output=\left\{\begin{array}{c}0\;\;if\sum_jw_jx_j\;\leq\;threshold\\1\;\;if\sum_jw_jx_j\;>\;threshold\end{array}\right.

To simplify further, we can think of the above equation as dot product or scalar product which is an algebraic operation that takes two equal-length sequences of numbers (usually coordinate vectors) and returns a single number. The dot product of two vectors w = [w1, w2, …, wn] and x = [x1, x2, …, xn] is defined as: \mathrm w.\mathrm x\;=\;\overset{\mathrm n}{\underset{\mathrm                         i=1}{\sum\;w_{\mathrm i}}}x_{\mathrm i\;\;=}\;w_1x_1+\;w_2x_2+w_3x_3+...+\;w_{\mathrm                         n}x_{\mathrm n}

Expressing the above example in this way, a 1 × 3 matrix (row vector) is multiplied by a 3 × 1 matrix (column vector) to get a 1 × 1 matrix that is identified with its unique entry:

\;\begin{bmatrix}1\\-8\\2\end{bmatrix}\;\ast\;\begin{bmatrix}1&-2&-3\end{bmatrix}\;=\;-9

We can simply write the perceptron equation more equitably, by taking the threshold value as bias and introducing dot product to sum of multiples of inputs and weights. First we write \textstyle\underset j{\sum{\text{w}}_j{\text{x}}_j\text{ }} as a dot product w \cdot x                     \equiv \sum_j w_j x_j where w and x are weights and inputs and secondly we move the threshold to the right side of the inequality and use bias instead where b \equiv -\mbox{threshold}
output=\left\{\begin{array}{c}0\;\;\;if\;w.x\;+\;b\;\leq\;0\\1\;\;\;if\;w.x\;+\;b\;>\;0\end{array}\right.\;\;\;\;\;\;\;\;(2)

From the above diagram, we can intuitively visualize a network of perceptrons as one which can filter previous layer consecutively i.e. weigh and decide the previous layer’s decisions and so on until we reach the final layer which is the output. Sometimes inputs can be represented as a perceptron just with an inherent or constant bias with no weights or inputs that point at it. Earlier when perceptron models were created, it was thought that NAND or XNOR can’t be simulated and this kept the pace of innovation dormant and later it was found that all bitwise operation can be simulated or matter of fact a function exhibiting any pattern. As networks with multiple layers became apparent and deep learning finally smashed its dormancy and is fast picking up its slack.

Workhorses inside Nodes – Activation Functions

Top

Modelling any process has been a fundamental activity of measuring output versus given input to gain understanding and more to guarantee repeatability of outcome for a given input with that established model in place. Essentially this model is always to me in simplified terms is ‘Curve Fitting’. With an input and model, you can always determine what is the output. Activation functions are ‘Curve Fitters’ where given an input, they provide an output based on the model defined. Again the variations in models gives rise to different activation functions.

Stepper Function

A basic activation function is a stepper function that’s shown in the graph below:

Equation : output=\left\{\begin{array}{c}0\;\;if\;\;x\;\leq\;0\\1\;\;if\;\;x\;>\;0\end{array}\right.

Stepper or Binary is a function whose output is either 1 or 0 (Yes or No). Alternatively its also called a href=”https://en.wikipedia.org/wiki/Heaviside_step_function”>Heaviside step function or Unit step function. This function emits 0 for all negative inputs and 1 for all positives. Perceptrons use this kind of activation to signal whether to proceed or not based on the input.

Linear Function

Linear function as the name suggests will ramp up the given input by a multiple. Mathematically a linear function is a function whose graph is a straight line, that is a polynomial function of degree one or zero.

Equation : \mathrm f(\mathrm x)\;=\;\mathrm{ax} The graph represents a linear function f(x) = 4x, essentially having a multiplier effect on the input by 4 times. This will be useful in neurons whenever we want to proportionally scale the input but its use case is rare.

Sigmoid Function

Sigmoid is a S curve, where output is = 0.5 for all positives and output never exceeds 1. For higher values in x and -x axis, y value is an asymptote. Of the ‘S’ shape, top portion is concave and the bottom being convex. The definition of sigmoidal functions are very broad, but uniquely it’s one that has real value output, one inflection point and has a bell shaped first derivative.

Equation : \mathrm f(\mathrm x)\;=\;\mathrm{ax}

In general, Sigmoid function is obtained from a Logistic function whose common shape is ‘S’ with the equation f(x)=\frac{L}{1+e^{-k\left(x-x_{0}\right)}} where

  • e = the natural logarithm base (also known as Euler’s number),
  • x0 = the x-value of the sigmoid’s midpoint,
  • L = the curve’s maximum value, and
  • k = the logistic growth rate or steepness of the curve

In the case of standard logistic sigmoid function the values will be as follows : L = 1, k = 1, x0 = 0, giving the Sigmoid equation above. f(x)=\frac{1}{1+e^{-x}}=\frac{e^{x}}{e^{x}+1}=\frac{1}{2}+\frac{1}{2} \tanh \left(\frac{x}{2}\right).
One great advantage of this function is it’s non-linear. shows linearity within the range of -3 and 3 and tapers of to an asymptote line on the rest of the range. It can be used to transform a continuous space value into value between binary range and such discretization method is very useful to dampen the extremes – characterized by limiting value of zero at low activation and one at high activation. For now this function suffices to start our journey into neural networks but we’ll peek comprehensively into other functions and their suitability later.

Idea Behind Sigmoid Optimization

The interest in sigmoid is that it can manage linear and non-linear modeling in the same function. In the above diagram the inflection point is at 0.5 indicated by z. Sigmoid curve models a typical economy of scale problem. For example, if a manufacturer starts making a toy, initial investment is heavy and the rate of return is minimal or near zero,once the public starts buying the toy and sales increases gradually and takes on a linear path as demand sets in where the supply cannot fully meet the demand – the profit soars up to a point where supply catches up demand at the inflection point where the profit is ideal beyond which supply outstrips demand and the profitability simply drops and cannot increase beyond a max no matter how much influence exerted.

This surmises that any activity dependent on threshold is an ideal candidate for optimization by a sigmoid function. Let’s look at a couple of such candidates and relate it to how it can help in deep learning use cases. Election campaign spending is one where if campaign manager tried to optimize their advertising spending, then their objective function will model an outcome where they want to spend money on states that have a feasible change of voting in favor of their candidate. Whereas money spent on states eventually lost is completely wasted. Sigmoidal functions will behave in the same way, only recommending to spend money in a state if they can pass the threshold of getting a majority of voters. They will predict that it is better to spend a lot of money on one state that has a good chance to be influenced than spending half each on two states where they may end up losing both. In ‘deep learning use case’ of predicting a number from a given image, it will be useful that each pixel or group of pixels collectively crosses a threshold and if they are consecutively cross validated using the layered sigmiod functions, then it can perfectly predict the outcome with less errors and during this process, our model would have learned well.

A Gentle Detour on Basics – Differential Calculus

Top

Differentiation

Before we delve further, we have to recollect what we studied in our high school mathematics about calculus – the dreaded subject to our dilettante brains then, it’s a perfect time to reflect with fun and an utilitarian mindset rather than rote learning we did in school days.

Derivatives is all about rate of change. How fast do you travel to reach your destination from home? How quick population grew over years? how wealthy are people? All these talk about rate of change, one against another. In comparison to above examples, they are distance versus time, count versus years, dollar versus people count.

Incidentally rate of change is also slope, if one of the above use case data is plotted in a graph of x-y axes, distance in y axis and time in x axis, average speed between 2 time points is given by distance travelled over the time taken – which is essentially the slope. If you really want to find the speed at a particular point of time, it’s difficult to measure distance traversed at infinitesimal slice of time, hence there’s no way to compute rate of change. But differentiation helps here.

Inspired by lucid and simple explanations on complex mathematical concepts, let’s borrow the learning concept (thanks to MathIsFun site, by Ed. Rod Pierce) to explain the derivatives in a more clear, fun and concise manner. Please ensure that you go through each one of them to recollect/refresh back what you’ve studied then and internalize differential calculus which is paramount in understanding the functioning of neural net.

  1. Introduction to Derivatives – talks about rate of change, its relation to slope, how we arrive at slope for a given function
  2. Derivative Rules – covers rules that govern the derivatives of various types of functions.
  3. Power Rule – most easiest way to de-clutter power functions to its derivatives
  4. Second Derivative – fantastic way to understand 2nd derivative with speed, acceleration and jerk
  5. Partial Derivatives – practical use case to understand real world usage of partial derivatives using our familiar cylinder
  6. Differentiable – this test allows to check differentialibity in a curve to determine presence of slope
  7. Finding Maxima and Minima using Derivatives – fundamental in learning Gradient Descent and Back Propagation

Differentiation Ready Reckoner

Common Functions Function Derivative
Constant c 0
Line x 1
  ax a
Square x2 2x
Square Root √x (½)x
Exponential ex ex
  ax ln(a) ax
Logarithms ln(x) 1/x
  loga(x) 1 / (x ln(a))
Trigonometry (x is in radians) sin(x) cos(x)
  cos(x) −sin(x)
  tan(x) sec2(x)
Inverse Trigonometry sin-1(x) 1/√(1−x2)
  cos-1(x) −1/√(1−x2)
  tan-1(x) 1/(1+x2)
     
Rules Function Derivative
Multiplication by constant cf cf’
Power Rule xn nxn−1
Sum Rule f + g f’ + g’
Difference Rule f – g f’ − g’
Product Rule fg f g’ + f’ g
Quotient Rule f/g (f’ g − g’ f )/g2
Reciprocal Rule 1/f −f’/f2
     
Chain Rule
(as "Composition of Functions")
f º g (f’ º g) × g’
Chain Rule (using ’ ) f(g(x)) f’(g(x))g’(x)
Chain Rule (using \frac d{dx} ) \frac{dy}{dx}=\frac{dy}{du}\frac{du}{dx}

How To Understand Derivatives: The Product, Power & Chain Rules – provides an intuitive way to understanding the product, power and chain rules we always memorized in high school
With a firm understanding of differential calculus, we take our next step!

The Underpinnings – Essential Statistics and Loss Reduction

Top

Statistics

In statistics, you should be familiar with estimates of location and variability which is central to how data is represented and varies and how data can be scaled to normalize and speed up analysis. A quick summary of essential concepts connected to this ‘Deep Learning’ journey:

Estimates of Location

Variables with measured data might have thousands of distinct values. A basic step in exploring your data is getting a “typical value” for each feature (variable): an estimate of where most of the data are located (i.e. their central tendency).

  • Mean
    • The sum of all values divided by the number of values
    • Synonyms
      • average
    • Equation
      • \overline{x}=\frac{1}{n}\left(\sum_{i=1}^{n} x_{i}\right)=\frac{x_{1}+x_{2}+\cdots+x_{n}}{n}
    • Weighted Mean
      • The sum of all values times a weight divided by the sum of the weights.
      • Synonyms
        • weighted average
      • Equation
        • \overline{x}=\frac{\sum_{i=1}^{n} w_{i} x_{i}}{\sum_{i=1}^{n} w_{i}}
  • Median
    • The value such that one-half of the data lies above and below.
    • Synonyms
      • 50th percentile
    • Weighted Median
      • The value such that one-half of the sum of the weights lies above and below the sorted data
  • Robust
    • Not sensitive to extreme values.
    • Synonyms
      • resistant
  • Outlier
    • A data value that is very different from most of the data.
    • Synonyms
      • extreme value

Estimates of Variability

Location is just one dimension in summarizing a feature. another dimension being variability, also referred to as dispersion, measures whether data values are tightly clustered or spread out. At the center of statistics lies variability: measuring it, reducing it, distinguishing random from real variability, identifying the various sources of real variability and making decisions in the presence of it.

  • Deviations
    • The difference between the observed values and the estimate of location.
    • Synonyms
      • errors, residuals.
  • Variance
    • The sum of squared deviations from the mean divided by N-1 where N is the number of data values.
    • Synonyms
      • mean-squared-error.
    • Equation
      • s^{2}=\frac{\sum(x-\overline{x})^{2}}{n-1}
  • Standard Deviation
    • The square root of the variance.
    • Synonyms
      • l2-norm, Euclidean norm
  • Mean Absolute Deviation
    • The mean of the absolute value of the deviations from the mean.
    • Synonyms
      • l1-norm, Manhattan norm
    • Median Absolute Deviation from the Median
      • The median of the absolute value of the deviations from the median.
  • Range
    • The difference between the largest and the smallest value in a data set.

  • Order Statistics
    • Metrics based on the data values sorted from smallest to biggest.
    • Synonyms
      • ranks
  • Percentile
    • The value such that P percent of the values take on this value or less and (100-P) percent take on this value or more.
    • Synonyms
      • quantile
  • Inter-quartile Range
    • The difference between the 75th percentile and the 25th percentile
    • Synonyms
      • IQR

Data Scaling

From Wikipedia: Since the range of values of raw data varies widely, in some machine learning algorithms, objective/loss functions will not work properly without normalization. For example, the majority of classifiers calculate the distance between two points by the Euclidean distance. If one of the features has a broad range of values, the distance will be governed by this particular feature. Therefore, the range of all features should be normalized so that each feature contributes approximately proportionately to the final distance.

Another reason why feature scaling is applied is that gradient descent (we’re going to see this in more detail later) converges much faster with feature scaling than without it.

  • Min-Max Scaling
    • Formula for mix-max:
      \tilde{x}=\frac{x-\min (x)}{\max (x)-\min (x)}
    • This method re-scales individual data points within a range say [0,1]. Let’s say we have a dataset with 100 data points and x be an individual data point,min(x) being the minimum among 100 data points and max(x) being the largest among them, min-max scaling squeezes (or stretches) every x value to be within the range of [0,1]. Sometimes the range could be set to [-1,1]. The target range depends on the nature of data and algorithm employed in analysis. Below figure gives you an intuitive illustration.
  • Unit Scaling or L2 Normalization
    • Formula for unit-scaling:
      \tilde{x}=\frac{x}{\|x\|_{2}}
    • This is obtained by dividing each data point by the Euclidean length of the vector from a centre or mean point. The L2 norm measures the length of the vector in coordinate space. Pythagorean theorem gives the distance \|x\|_{2}=\sqrt{x_{1}^{2}+x_{2}^{2}+\ldots+x_{m}^{2}}. Due to the division, data normalizes within -1 to 1 range.
  • Variance Scaling or Standardization
    • Formula for Standardization:
      \widetilde x=\frac{x-mean(x)}{sd(x)}
    • Average or mean intuitively signifies the absolute spread of data whereas Standard Deviation (SD/sd) gives an idea of weighted spread. A measure of spread is most helpful when the distribution of your data is symmetric around the mean and has a variance relatively close to that of the Normal distribution. (This means that it is approximately Normal). In the case where data is approximately Normal, the standard deviation has a canonical interpretation: 68% data will within 1 SD, 95% within 2 SDs and 99% within 3 SDs. The question is whether the distribution is wide or tight? This gives an idea of data and possible outliers that may appear.
      Here for every data point, mean value is stripped and further weighted by their standard deviation. Essentially scales every data point centered to mean by their overall spread.
      The resulting scaled data point is standardized to have a mean of 0 and a variance of 1. If the original feature has a Gaussian distribution, then the scaled feature is a standard Gaussian, a.k.a. standard normal. Below figure intuitively illustrates this notion.
    • Image Credit : Mastering Feature Engineering by Alice Zheng

Loss Functions

Variability inspires how deviation in prediction is measured, similar to standard deviation, loss can be measured, scaled and normalized. At the core, loss function measures how well your prediction performed compared to original value. Now you have a tool to track model performance and at the same time tune its performance so that loss is reduced to a minimum. Very basic loss function is to find the absolute difference between original and predicted values. This provides a good judgement on individual prediction (Absolute Error – AE) and again this can be averaged on all predictions to get Mean Absolute Error (MAE / L1 Loss) and averaging their squared errors gives Mean Squared Error (MSE / L2 Loss / Quadratic Loss).

MSE & MAE – An Analysis

Let’s take a simple function y = f(x) = 4(x), for which a graph is drawn as below. The corresponding predicted values yp is also drawn to show the deviation from actual values along with loss, absolute loss and squared loss. Also given are the statistical values and MAE and MSE.

The formulae are:
MSE=\frac{\sum_{i=1}^{n}\left(y_{i}-y_{i}^{p}\right)^{2}}{n}

MAE=\frac{\sum_{i=1}^{n}\left|y_{i}-y_{i}^{p}\right|}{n}

yp = forecasts (predicted),
y = observed values (actuals).

If formulas is not your staple, you can find the MSE & RMSE by:
MSE

  • Squaring the residuals/loss.
  • Finding the average of the residuals/loss.

RMSE

  • Taking the square root of the result.

Both MAE and MSE signify average model prediction errors slightly in different units of interest. Both metrics can range from 0 to ∞ and are indifferent to directional errors. They are negatively-calibrated scores, which means lower values are better. Taking the square root of the average squared errors or plain simple squared errors has some interesting implications – since the errors are squared and before they are averaged, MSE/RMSE gives a relatively high weight to large errors. This means the MSE/RMSE should be very useful when large errors are particularly undesirable.

Looking at the graph above, we can see that the MAE and MSE (RMSE is just square root of MSE) values for actual and predicted values. Most of the prediction are inline except 2 outliers where the error is very high. Compared to MAE,MSE/RMSE will weight the outlier heavily as you notice the Squared Loss is way above their nearer loss counterparts. Intuitively, we can interpret prediction performance from a loss function perspective where if try to minimize MSE, then that prediction should be the mean of all target values. But if we try to minimize MAE, that prediction would be the median of all observations. Median is more robust to outliers than mean, which consequently makes MAE more robust to outliers than MSE.

One big problem in using MAE loss (for neural nets especially) is that its gradient is the same throughout, which means the gradient will be large even for small loss values. This isn’t good for learning. To fix this, we can use dynamic learning rate which decreases as we move closer to the minima. MSE behaves nicely in this case and will converge even with a fixed learning rate. The gradient of MSE loss is high for larger loss values and decreases as loss approaches 0, making it more precise at the end of training (see figure below.)

  • MAE LOSS

  • MSE LOSS

Above is a plot of an MAE & MSE function where the true target value is 100, and the predicted values range between -10,000 to 10,000. The MSE loss (Y-axis) reaches its minimum value at prediction (X-axis) = 100. The range is 0 to ∞.

Some Loss Functions in Neural Networks

Loss functions play a vital role in stabilizing and improving prediction accuracy of a given model be in ML/DL. They are used in a way to minimize difference between actual and predicted values. A summary of loss functions used in deep learning is listed below – name, it’s equation and a gist of what it does.

Note that y is the actual point or node data and yhat is the predicted or functional output of y. C is the overall error across all data points.

Name Equation
Mean Squared Error \mathcal{C}=\frac{1}{n} \sum_{i=1}^{n}\left(y_{(i)}-\hat{y}_{(i)}\right)^{2}
Widely used in regression, the resultant fitting line for data points should be a line which minimizes the sum of distance of each point to the fitted regression line. Square difference between actual and predicted is named ‘residual’ and the target of loss function is to minimize the residual sum of squares in DL.
Half Mean Squared Error \mathcal{C}=\frac{1}{2n} \sum_{i=1}^{n}\left(y_{(i)}-\hat{y}_{(i)}\right)^{2}
As the name implies, its half the value of MSE and it’s for calculus convenience which we’ll come across later.
Squared/Quadratic Error – L2 \mathcal{C}=\sum_{i=1}^{n}\left(y_{(i)}-\hat{y}_{(i)}\right)^{2}
Simply a sum, without averaging the MSE. Also called L2 loss
Mean Squared Lograthmic Error \mathcal{C}=\frac{1}{n} \sum_{i=1}^{n}\left(\log \left(y_{(i)}+1\right)-\log                                 \left(\hat{y}_{(i)}+1\right)\right)^{2}
It measures the variance difference in log values of actual and predicted. It useful when predicted and actual are huge values – using the power of log to reduce the exponential to progressive values. MSLE penalizes under-estimates more than over-estimates.
Mean Absolute Error \mathcal{C}=\frac{1}{n} \sum_{i=1}^{n}\left|y_{(i)}-\hat{y}_{(i)}\right|
Absolute difference between actual and predicted value. Computing a nice gradient in MSE is easier whereas MAE gives only a constant.
Mean Absolute Percentage Error \mathcal{C}=\frac{1}{n}                                 \sum_{i=1}^{n}\left|\frac{y_{(i)}-\hat{y}_{(i)}}{y_{(i)}}\right| \cdot 100
Obviously a percentage error of actual vs predicted over actual. Cannot be used when actual is zero. When the difference between actual and predicted is wide, it’ll exceed 100%. This function has a tendency to choose a model whose predictions are too low.
Absolute Error -L1 \mathcal{C}=\frac{1}{n} \sum_{i=1}^{n}\left|y_{(i)}-\hat{y}_{(i)}\right|
MAE without the mean i.e. not divided by n
Cross Entropy Loss \mathcal{C}=-\frac{1}{n} \sum_{i=1}^{n}\left[y_{(i)} \log                                 \left(\hat{y}_{(i)}\right)+\left(1-y_{(i)}\right) \log                                 \left(1-\hat{y}_{(i)}\right)\right]
Typically in binary classification and neural networks, the output is a real value signifying the probability of expected result instead of a 0 or 1 for given event. Or it can be computed easily – given a set of data points, Log Loss measures the divergence of probability distribution of actual versus predicted. We can also assume that the actual and predicted data points represent a probability distribution. From a binary classification perspective, y being the either/or label for an event A or B (1 if A happens, 0 if doesn’t or B happens) and p(y) is its probability. Left hand side represents, label times probability of label being A whereas right side 1-label value times probability of being B. Further intuitive explanation here and refer this for a nice derivation from softmax derivative.
Negative Logarithmic Likelihood \mathcal{C}=-\frac{1}{n} \sum_{i=1}^{n} \log                                 \left(\hat{y}_{(i)}\right)
When the model output is probability of each class rather than most likely class – wherein it deals with multi-class i.e. each output will signify for example whether its a cat, dog, rat per output. A great read on soft-max and NLL from Understanding softmax and the negative log-likelihood in Lj Miranda Blog. I’m totally smitten by this fantastic visual from that article.

KL Divergence \mathcal{C}=-\frac{1}{n} \sum_{i=1}^{n} \log                                 \left(\hat{y}_{(i)}\right)
Links the above two and head to an intuitive article for a neat explanation.
Cosine Proximity \mathcal{C}=-\frac{\mathbf{y} \cdot \hat{\mathbf{y}}}{\|\mathbf{y}\|_{2}                                 \cdot\|\hat{\mathbf{y}}\|_{2}}=-\frac{\sum_{i=1}^{n} y_{(i)} \cdot                                 \hat{y}_{(i)}}{\sqrt{\sum_{i=1}^{n}\left(y_{(i)}\right)^{2}} \cdot                                 \sqrt{\sum_{i=1}^{n}\left(\hat{y}_{(i)}\right)^{2}}}
Form high school, we know Cos 0 is 1 and Cos 90 is 0, which kind of indicates that a given set of points whose cosine proximity being near zero are similar than near 1 where they are totally divergent. This can ignore magnitude differences and only take into account angle convergence.

Losses and their impact in error reduction and in comparison to gradient descent is discussed in detail here. Next we are going to tackle Gradient Descent more intuitively.

The Grand Optimization – Gradient Descent

Top

Gradient Descent

Oft refereed algorithm in any mention of ML/DL is Gradient Descent. As it is so fundamental that tomes have been covered illustrating, explaining and debunking it but here were are going to understand it more intuitively with 2 examples rather than the fancy 3D graphs and animations you come across.

Before we delve proper into an example, Gradient Descent literally is ‘a way to traverse’ such that we reach the bottom of a valley. Figuratively we have to step down and check are we at the lowest, if not look for the next best direction of descend and repeat it iteratively. One caveat is if we are stuck in one of the low minimums but not the lowest (assume we have a laser vision that can see through rocks), then retrace back and start in direction which will lead to lowest minimum, the global minimum which is deep down,the nadir point in the valley.

Assuming my valley is without contortions and is single dimension and a look-a-like of curve representing y = x2, then we can draw that valley in below graph. Recall back of finding the maxima and minima using 1st and 2nd derivatives earlier. Here the slope increases to the right and decreases to the left. Refer to the 2 curves, first is a curve y = x2 and next one is y = 1 + 2x – x2. Below diagram illustrates the idea.

Understanding Slope & Descend on Basic Curves

As noticed, in the left image above, slope is positive and increases as we move to right direction and reverses as we move left. In order to find the global minimum, which is simple operation – visually look where the slope reaches zero in the slope graph or mathematically if you equate the slope equation to zero, we get x = 0, the point where slope is minimum and is the global minimum. You can apply the second derivative test at this point of x (=0), 2nd derivative = 2 and when its greater than zero, it is local minimum and single its one dimensional, it’s also global minimum. Also note there’s no iterative process and hence single dimension functions don’t have gradient descent per se. Similarly right image gives an idea what happens to a parabola equation when we venture to find global maximum. Here the 2nd derivative = 2 – 2x and slope will be zero when x = 1. Applying 2nd derivative test, it’ll be -2 at this point and we know that it’s the local max and also global max for a 1-dimensional function.

Slope on Linear Function

There’s no chance for the 2nd derivative being a zero except for y = x (as shown above) where the slope is constant and there’s neither a minimum or maximum point rather technically an undecided point called saddle point. Saddle points happen in multi-dimensional functions.

Another important intuition, in single dimensional functions, we move left or right but from another point of view, if you move against when the slope is positive, you’ll reach the minimum i.e. when slope is steepest at a point then move in the opposite direction, the steepness reduces and reaches zero. This comes handy in gradient descent in multi-dimensional functions. Why it’s important – we can simply multiply the steepest slope at that point by a small factor and subtract/add to move the x-axis point in the opposite direction and re-evaluate the function’s slope and see it has become less steeper and reached the minimum and to repeat this process until we reach it.

Intuitive Examples to the Rescue – Descent Demystified

Top

Gradient Descent on a simple Function – f(x,y) = x + y

Time to recall the partial derivatives that we worked up earlier. If f(x,y,z) = x4 – 3xyz, then using curly dee notation:

  • \frac{\partial f}{\partial x}=4 x^{3}-3 y z

  • \frac{\partial f}{\partial y}=-3 x z

  • \frac{\partial f}{\partial z}=-3 x y

Now lets consider a function f(x,y) = x + y. Initialize the values of x and y as 2 and 4 respectively, output will be 6.

What beckons here is, how to adjust the values of x and y to bring the functional output to the minimum? The formula is simple:

  • new x position = old x position – (small constant) * slope of fn w.r.t x
  • new y position = old y position – (small constant) * slope of fn w.r.t y
  • Formally: \Theta_{n+1}=\Theta_{n}-\alpha\left(\frac{\partial}{\partial \Theta_{n}}                             J\left(\Theta_{n}\right)\right)
    • \theta_{n+1} – new position
    • \theta_{n} – old position
    • \alpha – small constant or learning rate
    • \frac{\partial}{\partial \Theta_{n}} \boldsymbol{J}\left(\Theta_{n}\right) – slope of fn w.r.t a variable or Gradient of function J for \theta

Let’s differentiate f(x,y) w.r.t its components

\begin{array}{c}{x_{\text {gradient}}=\frac{\partial f(x, y)}{\partial x}}                     {=\frac{\partial(x+y)}{\partial x}=1}\end{array}

\begin{array}{c}{y_{\text {gradient}}=\frac{\partial f(x, y)}{\partial y}}                     {=\frac{\partial(x+y)}{\partial y}=1}\end{array}

These partial derivatives can be thought of as forces that act on x and y inputs that produce the output. As we produce the output, the objective is to reach the minimum or zero output (incidentally slope at this point for x and y is zero, with 2nd derivative being +ve – recall from our earlier minima/maxima exploration, and also the output and loss function is same here). Let see how to apply the above formula and we assume \alpha = 0.01.

\begin{aligned} x_{\text {new}} &=x_{\text {old}}-\alpha * x_{\text {gradient}} \\ &=2-0.01^{*} 1                     \\ &=1.99 \end{aligned}

\begin{aligned} y_{\text { new }} &=y_{\text { old }}-\alpha * y \text { gradient } \\ &=4-0.01 *                     1 \\ &=3.99 \end{aligned}

This new positions for x and y leads to f(x,y) = x + y to be valued @ 1.99 + 3.99 = 5.98. Compared to our starting value of 6, it’s slightly less. In order to reach to a value of 0, we’ll utilize the important characteristics of gradient descent, that is to iterate steps until we reach convergence. Process is to repeat \theta_{j} \leftarrow \theta_{j}-\alpha \frac{\partial}{\partial\theta_{j}}J(\theta) till x+y reaches zero.

So if we repeat the above steps 300 times, we’ll get x = -1 and y = 1 and final output x+y = 0 and we have ultimately achieved the values of x and y that will give minimum value of 0.

Gradient Descent on Function f(x) = mx + C

Lets take a simple example to immerse on the inner workings and hence can ascend gradient descend with clarity.

From a survey conducted on city’s Central Business District (CBD), office floor space rental prices were collected and tabulated in column A and column B which respectively identifies floor space area (X) in sq.ft. versus the rental yield price per month (Y) in thousands of dollars it fetches.

Problem Statement : For a new office space in CBD, given the floor area, what’ll be the rental yield per month in dollar amount?

The green line plots the current yield price versus floor cover. We’ll use a simple linear model to predict the new rental yield (YPred) given its floor area (X). Amber line depicts the model which predict the yield given the area

YPred = mx + C

Green line gives the actual rental yield (YActual) based on the current survey.

The green dots and amber dots with interconnecting lines between them shows the difference between YActual and (YPred) which is the prediction loss or error (E).

The onus is on us to find optimal values for C and m (called weights) so that it best fits the prediction line governed by YPred = mX + C to reduce prediction error and improves accuracy. We’ll employ quadratic loss function RMSE as loss function which need to be minimized taking into account the actual and predicted values.

Simply put HMSE =,
\frac1{2n}\overset n{\underset1{\cdot\sum}}\left({\mathrm Y}_{\mathrm i}-{\mathrm{Ypred}}_{\mathrm i}\right)^2

Let’s utilize Gradient Descent algorithm to find optimal value for weights C and m to minimize HMSE.

Steps involved enumerated here:

  1. Normalize data in Column A and C using Min-Max Normalization into Column D & F espectively and will be used as X and Y
  2. Initialize C and m to random values and compute Squared Error (SE) per data point
  3. Compute individually ‘Change in HMSE’ per line/data point when C and m are changed by very small value from original random value, overall change constitutes summing up all these for all data points.
    After summing up, we could divide the summed value by number of data points to average but this isn’t necessary as it simply scales down the total change and moreover affects the change we introduce subsequently – as in the next statement. Essentially this may slow down convergence, so we can stop ar summing and no need to average. UIn dynamic cases, where the number of input data points are not known upfront, then we can’t also average and hence it’s best left summed up
  4. Adjust the corresponding weights C and m subtracting total ‘Change in HMSE’ w.r.t C and m that’s multiplied by a learning rate
  5. Use the new weights to compute total of SE and check it’s reducing
  6. Repeat Steps 3 and 4 till no more significant reductions in total SE

Steps in Detail:

  1. Refer to Columns D, E & F – which are min-max normalized rental data corresponding to Columns A, B & C
  2. To fit the line YPred = C + mx, start off with a random seed value and compute prediction loss HMSE. HMSE is over all data points, when applied to a single data point, n from the formula is stripped off – just work out only half squared error for each data point and sum all. SE and HSME refers the same in this context.
  3. Compute the error gradients w.r.t to the weights, in our case C and m \mathrm{SE}\;=\;\frac12(\mathrm Y-(\mathrm{mx}+\mathrm C))^2\;
    \frac{\partial\mathrm{SE}}{\partial\mathrm C}\;=\;-(\mathrm Y-{\mathrm Y}_{\mathrm{Pred}})
    \frac{\partial\mathrm{SE}}{\partial\mathrm m}\;=\;-(\mathrm Y-{\mathrm Y}_{\mathrm{Pred}})\mathrm x

    We’ll work out the calculus for the above equations separately later. The above 2 equations give the changes of ‘C and m’ and are simply gradients that points the direction of movement w.r.t to SE. Refer to last 2 columns of computation

  4. Change the weights C and m with the gradients to reach optimal SE value which is the minimum.

    We’ll adjust the initial random values of C and m, so that we move in the direction of optimal C and m.

    The update rules:
    \mathrm{New}\;\mathrm C\;=\;\mathrm C\;-\;\mathrm{lr}\;\ast\;\sum_1^{\mathrm n}\frac{\partial\mathrm{SE}}{\partial\mathrm C}
    \mathrm{New}\;\mathrm m\;=\;\mathrm m\;-\;\mathrm{lr}\;\ast\;\sum_1^{\mathrm n}\frac{\partial\mathrm{SE}}{\partial\mathrm m}

    here lr is the learning rate, hence:
    New C = 0.52 – 0.01 * 5.040 = 0.47
    New m = 0.81 = 0.01 * 2.279 = 0.79

  5. Use the new weights to compute total of SE and check it’s reducing – check the summed HSME value id decreasing in the graphic above
  6. As Steps 3 and 4 are repeated till no more significant reductions in total SE consecutively, then we have reached the optimal C and m with best predication rate

Calculus behind Gradient Descent

Squared Loss Function

HMSE per data point ends up applying   \mathrm{SE}\;=\;\frac12(\mathrm Y-(\mathrm{mx}+\mathrm C))^2\;
First derivative is w.r.t ‘C’ is
\frac{\partial\mathrm{SE}}{\partial\mathrm C}\;=\;-(\mathrm Y-{\mathrm Y}_{\mathrm{Pred}})

Let’s work out the partial derivatives now:

  • Equation as such
    =\;\frac12(\mathrm Y-(\mathrm{mx}+\mathrm C))^2\;
  • Expanding first
    =\;\frac12(\mathrm Y^2-(\mathrm{mx}+\mathrm C)^2\;-\;2Y(mx+C))\;
  • Expanding further
    =\;\frac12(\mathrm Y^2-\;(m^2x^2+\mathrm C^2+2\mathrm{mxC})\;-\;2Ymx\;-\;2YC)\;
  • removing all non C terms as they become zero upon differentiating on C
    =\;\frac12(-\;(2\mathrm C+2\mathrm{mx})\;-\;2Y)
  • =-\frac22(\;{Y-(\mathrm C+\mathrm{mx}))}
  • =-(\;{Y-\;Y_{Pred}))}

Second derivative is w.r.t ‘m’ is
\frac{\partial\mathrm{SE}}{\partial\mathrm m}\;=\;-(\mathrm Y-{\mathrm Y}_{\mathrm{Pred}})\mathrm X

  • From step 3 previously, fully expanding, equation becomes =\;\frac12(\mathrm Y^2-\;(m^2x^2+\mathrm C^2+2\mathrm{mxC})\;-\;2Ymx\;-\;2YC)\;
  • removing all non m terms as they become zero upon differentiating on m =\frac12(-\;(2\mathrm{mx}\;+\:2\mathrm{xC})\;-\;2Yx
  • taking 2 aside to left and x within separated,
    =\frac22(\;(\mathrm m\;+\:\mathrm C)\mathrm x\;-\;Yx)
  • removing minus and x aside
    =-(Y\;-\;Y_{pred})x

Applying a Different Loss Function

If you decide to apply a different loss function in the above endeavour, we can try Log-Cosh function which supports 1st and 2nd derivatives as well. If you applied MAE or RMSE, both yields a constant and so conducive from a learning and convergence perspective in neural network gradient descent process though!
Log-Cosh per data point ends up applying   \log(\cosh({\mathrm Y}_{\mathrm{Pred}}\;-\;\mathrm Y))
First derivative is w.r.t ‘C’ is
=\;\sinh({\mathrm Y}_{\mathrm{Pred}}\;-\;\mathrm Y)/\cosh({\mathrm Y}_{\mathrm{Pred}}\;-\;\mathrm Y)

Let’s work out the partial derivatives now:

  • Equation as such
    =\;\log(\cosh({\mathrm Y}_{\mathrm{Pred}}\;-\;\mathrm Y))
  • Applying Chain rule for y = u(v(w)), where u = log(), v = cosh(), w=mx+C:
    \;\frac{\mathrm{dy}}{\mathrm{dw}}=\;\;\frac{\mathrm{dy}}{\mathrm{du}}.\;\;\frac{\mathrm{du}}{\mathrm{dv}}.\;\;\frac{\mathrm{dv}}{\mathrm{dw}}
  • differentiating each term
    =\;\;\frac1{\cosh(\mathrm{mx}+\mathrm C-\mathrm Y)}\;.\;\sinh(\mathrm{mx}+\mathrm C-\mathrm Y)\;.\;1
  • simplifying each term
    =\;\;\frac{\;\sinh(\mathrm{mx}+\mathrm C-\mathrm Y)\;}{\cosh(\mathrm{mx}+\mathrm C-\mathrm Y)}\;
  • in terms of Y
    =\;\;\frac{\;\sinh({\mathrm Y}_{\mathrm{Pred}}-\mathrm Y)\;}{\cosh({\mathrm Y}_{\mathrm{Pred}}-\mathrm Y)}\;

Second derivative is w.r.t m simple multiplies the above term by x

  • =\;\;\frac{\;\sinh({\mathrm Y}_{\mathrm{Pred}}-\mathrm Y)}{\cosh({\mathrm Y}_{\mathrm{Pred}}-\mathrm Y)}\;.\;\mathrm x

Gradient Descent Variants

Batch – Vanilla Gradient Descent

As in the immersive example before, there were 12 data points and model parameters m & C were updated only at the end i.e. after computing all derivatives and loss values going through each data point for entire record set of 12 data points. Essentially this is Vanilla Batch GD where model parameters are updated only after all data point computations are complete – Processing all data points is refereed as an epoch.

Mini Batch – Gradient Descent

If in the above example, among 12 data points, we can process first 4 and update the model parameters, followed by next 4 data points to update the parameters and so on until all data points are exhausted. Essentially after every mini batch (4 data points in our case) are processed, model parameters are updated and hence within an epoch, parameters are updated by n times where n = total data points/mini batch data points.

Stochastic Gradient Descent – SGD

Generally SGD computes the gradient using a single data point but most applications of SGD actually use a mini-batch of several samples. Hence prior to every epoch start, all data points are randomly shuffled and mini batch process is followed. SGD works well (compared to batch gradient descent and there are advanced methods now) for error manifolds that have lots of local maxima/minima. With this, the somewhat noisier gradient calculated using the reduced number of data points tends to jerk the model out of local minima into a region that hopefully is more optimal. Single data points are really noisy, while mini-batches tend to average a little of the noise out. Hence the amount of jerk is reduced while using mini-batches. Possibly a good balance can be reached when the mini-batch size can avoid some of the poor local minima but large enough to include global minima or better performing local minima and ideally leading us to easier fall into larger and deeper basin of attraction wherein best minima are placed.

SGD is akin to sampling a cross-section of voters to get a forecast than running an actual election. The learning rate is spectacular compared to Batch, let’s say our total data points being a million (1M) and if the mini-batch size is 10K, then we get a speed-up factor of 1M/1K = 1000. The speed-up factor may not be accurate and includes statistical fluctuations but good enough to lead on a path of gradual reduction in cost function. One benefit of SGD is that it’s computationally faster and can be parallelized.

When building the Gradient Descent algorithm, maximum epoch count (iteration count) can be set along with loss factor difference which work in tandem to arrive at a convergence or stop endless processing whichever happens first. Explore the variants further here

Ensemble directed Back & Forth – Feed Forward & Back Propagation

Top

Chain Rule

In our earlier Gradient Descent example, we were looking at simple function y = mx + C. This can be imagined as a single node activation function which simply takes an input to provide an output – which essentially scaled the input by m and adds a bias C.
Now lets imagine a bunch of such activation functions interconnected – one to another – such that the output of one is input to next. This is a simple linear network which may have few nodes. In our example we’ll consider 4 node linear network as shown below. Again the same idea applies here as well – we need to find the right weights and biases such that the overall function’s output will be optimized so that any error difference between the original output and output from this ‘interconnected functions’ is minimum.
To demonstrate the idea, look at the diagram below – which is 2 hidden layered neural network with an input node and output node.

Each node is a function of the previous one connected to it. If we were to change either the input or the weight value in the input node, this change will cascade down – as each node will change the input they receive and finally reflects in output. As such the output change may seem complex but intuitively every node amplifies the cascading change. Since each node in interdependent and with notion of functional dependencies, we can mathematically formulate a composite function to represent the behavior of this network:

\begin{array}{l}output\;(O)\:=\;act(w3\;\ast\:h2)\;\\hidden2\;(h2)\;=\;act(w2\;\ast\;h1)\\hidden1\;(h1)\;=\;act(w1\;\ast\;input)\end{array}

Consolidating all equations in to one: output\;(O)\;=\;act(w3\ast act(w2\ast act\;(w1\ast input)))

Since the intermediary nodes are not working solo and dependent on input from the previous nodes, if we would like to understand the overall change in output for a small change in w1, then we have to apply the derivatives in tandem from last node to first for the output w.r.t w1. The resultant composite derivative will be:

\frac{\partial O}{\partial w1}\;=\;\frac{\partial O}{\partial h2}\cdot\;\frac{\partial h2}{\partial h1}\cdot\;\frac{\partial h1}{\partial w1}

Let’s pause here and look into a detailed example on how chain rule can be leveraged here.

Chain Rule in Action – Understanding With an Example

Chain rule of differentiation is fundamental to back propagation and getting an intuitive understanding goes a long way to comprehend neural networks in its entirety. With the example below, it allows to decipher how chain rule works and see how changes in factors attribute to final output change. In the diagram below, we have 3 activation functions connected in linear manner and essentially any input cascades through them to give an output following the rules stated. It’s very easy to follow along:

  • i/p signifies the input to this overall linear network
  • any i/p gets weighted by w1 and supplemented with a bias b1, collectively notated as z1, is fed to activation function #1
  • this function scales down any input by 1/8 times to give h1 as output
  • same pattern repeats until output h3
  • finally we have an error function wherein if we know the actual h3 and predicted h3, error/cost can be computed but for now we ignore this part

Let’s enumerate the equations for the above linear network from output all the way to input

  • h1 = f1(w1 * x + b1)
      = (w1 * x + b1)/8
  • h2 = f2(w2 * h1 + b2)
      = (w2 * h1 + b2)/7
  • O = h3 = f3(w3 * h2 + b3)
      = (w3 * h2 + b3)/9

As shown above essentially any wiggle that you make in x gets scaled down as per our example before a slight bump up due to weights and bias in each layer. One more thing to note is ‘h’ doesn’t know what has happened to x but only alters change from ‘g’ and ‘g’ also has no idea of change in x but only wiggles ‘f’ and so on. Now any change in x will gets amplified by delta x * rate of change in f, followed by its next layer and so on.
O = h(g(f(x))) and the change in h w.r.t. x will be   \frac{dh}{dx}=\frac{dh}{dg}\cdot\frac{dg}{df}\cdot\frac{df}{dx}

  1. x changes by dx
  2. f changes by df, so
    df=d x \cdot \frac{df}{dx}
    the derivative of f(df/dx) is how much to scale the initial wiggle
  3. g wiggles by dg, so
    dg=df\cdot\frac{dg}{df}
    the derivative of f(dg/df) is how much to scale the input wiggle
  4. h wiggles by dh, so
    dh=dg\cdot\frac{dh}{dg}
    the derivative of f(dh/dg) is how much to scale the input wiggle
  5. Rewriting
    dh=\left(dx\cdot\frac{df}{dx}\right)\cdot\frac{dg}{df}\cdot\frac{dh}{dg}
    \frac{dh}{dx}=\frac{df}{dx}\cdot\frac{dg}{df}\cdot\frac{dh}{dg}

The chain rule isn’t just unit cancellation of denominator and numerator — it’s cascading of a wiggle, which gets scaled up or down or adjusted at each step. The chain rule works for any number of variables (f depends on g depends on h and so on), just cascades the wiggle. Think of it as zooming into different variable’s point of view – beginning from dx and gazing up, you can visualize the entire chain of adjustments needed before the wiggle reaches h.

Any change in O w.r.t w1 (again w1 is scaled up value of x above) can be written using chain rule as follows:

  • Equation as such
    \frac{\partial O}{\partial w1}=\;\frac{\partial O}{\partial h2}\cdot\frac{\partial h2}{\partial h1}\cdot\frac{\partial h1}{\partial w1}
  • Expanding first
    \frac{\partial O}{\partial w1}=\;\frac{\partial\;(w3.h2+b3)/9}{\partial h2}\cdot\frac{\partial\;(w2.h1+b2)/7}{\partial h1}\cdot\frac{\partial\;(w1.x+b1)/8}{\partial w1}
  • Simplifying
    =\;\frac{w3}9\cdot\frac{w2}7\cdot\frac x8
  • replacing w3 and w2 values into equation, gives
    =\;\frac29\cdot\frac77\cdot\frac x8
  • Simplifying
    =\;\frac x{36}
  • To compute h3, after a change in w1, the formula would be new\;h3\;or\;output\;=\;\triangle w1\cdot\frac x{36}+\;old\;h3

Above tabulation sums up the idea behind the chain rule with numerical data

  1. Column A : consider few inputs to this linear network
  2. Column B : computes h1
  3. Column C : computes h2
  4. Column D : computes h3 – final output
  5. Column E : computes new h3 based on a change in w1 – just by calculating rate of change of h3 w.r.t w1 and multiplying with change in w1 and finally adding it to h3 to get the new values
  6. the repeated rows compute h1, h2 & h3 by incorporating the new w1 value and doing all end to end calculation to arrive at h3 – note column E data and data under heading ‘new h3’ in column D (on modified w1) are the same
  7. Column F : computes derivative using approximation – as given by the formula:
    \frac{\partial O}{\partial w1}\;=\;\frac{O(w1+\triangle w1)\;-\;O(w1)}x\cdot\triangle w1
    here O is nothing but computing h3
    The value computed by approximation and chain rule are the same as shown

Back Propagation – Chain Rule in Action in Reverse

In our previous linear network of nodes, let’s add a final function called cost function. This takes the final output (computed y value aka predicted value using the given model – our linear network) and calculates the difference with respect to given original output (original y values). Since this is another functional dependency, now our cost is also affected by input and weights and biases, and hence, is a function of them.

Now let’s attach a ‘Cost function’ at the end of this network as shown above, it’s obvious that this function’s output is also dependent opn the input and intermediary weights. Hence we can extend the same derivative formula to our cost function as well:

\frac{\partial C}{\partial w1}\;=\;\frac{\partial C}{\partial O}\cdot\frac{\partial O}{\partial h2}\cdot\frac{\partial h2}{\partial h1}\cdot\frac{\partial h1}{\partial w1}

In section ‘Understanding Chain Rule With an Example’, we were able to find the composite derivative using approximation. i.e by changing w1 by a small value and computing the new output and hence \frac{\partial O}{\partial w1} value.
Now a question would arise how about applying the approximation principle to compute changes in cost w.r.t weights in a real neural network. It’s quite feasible to take this path but leads to total slowdown in computation. Let’s assume a large network of weights, in order to find the change in Cost w.r.t w1, we need to find C and C(with delta w1) at all weights bearing nodes and sum up the get their final amounts and then apply the approximation formula. If the nodes are in millions, approximation algorithm screeches to a halt, whereas chain rule method comes handy in propagating the derivative back into the network to find the weights quickly by gradient descent to reduce overall cost – as depicted below.

Neural Network

Network

Time to delve in to the real neural network and work out the equations based on the understating so far and apply it to python code. The notation & code is based on the seminal book by Michael Nielsen. I had to spend quite a lot of time understating the python code with respect to notation, equations and deductions. What I’ve done here is to simply it further with better visualizations so that it can assimilated quickly and the book can be used as an reference resource.

Imagine a network of 4 layers (as above):

  1. Layer 0 – Input Layer of 3 nodes
    No transformation of inputs and they are just passed on
  2. Layer 1 – comprising 4 neurons
    For each neuron in this layer, layer 0 inputs are weighted and finally a bias is added and fed to activation function
  3. Layer 2 – comprising 3 neurons
    For each neuron in this layer, layer 1 inputs are weighted and finally a bias is added and fed to activation function
  4. Layer 3 – comprising 2 neurons
    For each neuron in this layer, layer 2 inputs are weighted and finally a bias is added and fed to activation function

Let’s assign proper notation to what we said now and properly express them mathematically and it’s very easy to follow. It’s also good to consider an input use case for this network but need not exactly match the inputs and outputs but provides a context and intuitive understanding of how the network is going to handle data overall.
MNIST is standard torch bearer dataset for deep learning – this contains 60K images of hand written numbers and annotated correct numeric value for each image.

  1. Input is an image of 28×28 pixels
    translates to a array 784 inputs storing the grey scale value within 0 to 1
  2. Output is indicate numeral between 0 – 9
    translates to an array of 10 outputs, each a value between 0 to 1 where a value greater than .5 indicates that numeral in order. In a way the output can alsi be considered as a probability for that numeral to occur in the i/p image

Mathematically expressed Network

A full network with complete notations is given here and to get a full scape click on the image.

Basic Definitions

Now let’s define all basics – inputs, weights, biases – step by step, Just refer to below diagram

  • Input (at 1st Layer)

    \boldsymbol {input\;=\;x}
    typically numerical data raw or normalized and is represented by x in this layer Also note that in this layer, there is no activation, weights and biases and hence all these are zero and we could also say that in this layer:
    x=z=a,\;\;given\;w=0,\;b=0
    We’ll see what’s z and a subsequently

  • Rows and Layers (Columns)

    Rows = \boldsymbol j and/or \boldsymbol k
    Layers = \boldsymbol l or \boldsymbol L
    j always being associated to the current layer and k refers to (l-1)th layer. ‘L‘ denotes either the last layer or collectives of a layer.
    Also note in the above diagram, neurons are not equal-spaced from top but nicely placed in their corresponding row and layer to match the matrix representation which we’ll see in a while.

  • Weights

    Weights = \boldsymbol {{w_{jk}^{l}}}
    Every neuron gets their multitude of inputs from previous layer neurons output ‘weighed in’ before a bias is combined to be taken in as input. j and k notation is as before and as illustrated in the below diagram

  • Bias

    Bias = \boldsymbol {{b_{j}^{l}}}
    all weighed in inputs from previous layer are summed up and added to bias in node at layer l in row j is fed to a neuron @ layer l and row j.

Other Definitions

Now let’s define next level – combined inputs, activations, outputs, cost function – step by step, Just refer to below diagram.

  • Input at a Neuron

    \boldsymbol {z_j^l=\sum_kw_{jk}^la_k^{l-1}+b_j^l}
    compact form (vectorized):
    \boldsymbol {z^{l} \equiv w^{l} a^{l-1}+b^{l}}
    all weights times inputs from previous layer (their nodes’ activation’s output) + bias in the given node

  • Activation

    \boldsymbol {a_{j}^{l}=\sigma\left(\sum_{k} w_{j k}^{l} a_{k}^{l-1}+b_{j}^{l}\right)}
    compact form (vectorized):
    \boldsymbol {a^{l}=\sigma\left(z^{l}\right)}
    \boldsymbol {a^{l}=\sigma\left(w^{l} a^{l-1}+b^{l}\right)}
    activation function is typically a sigmoid function.

  • Cost

    \boldsymbol {C=\frac{1}{2 n} \sum_{x}\left\|y(x)-a^{L}(x)\right\|^{2}}
    — for a single training sample at an output row in the last layer:
    \boldsymbol {C=\frac{1}{2}\left\|y-a^{L}\right\|^{2}=\frac{1}{2} \sum_{j}\left(y_{j}-a_{j}^{L}\right)^{2}}

    • \boldsymbol x
      number of inputs, from a MNIST dataset perspective: 50k training numeral images each represented individually as 784×1 matrix of real number between 0 to 1 which is a grey-scaled value of a pixel
    • \boldsymbol {a^{L}(x)}
      output, from a MNIST dataset perspective: after going through our neural net, is numeral output – array of 10 – representing each digit between 0 to 9 – whose value is again between 0 to 1 which is probability value for that particular numeral
      Number of outputs (for a given image input) = 10 outputs
    • \boldsymbol y(x)
      original output, from a MNIST dataset perspective: annotated & verified numeric value of a given image – which we can convert into an array of 10 values each representing numeral 0 to 9 where corresponding numeral is 1 and rest are zeros
      Number of original outputs (for a given image input – represented as an array of 10 – 1 for each digit) = 10 outputs
    • \boldsymbol n
      input/sample count, from a MNIST dataset perspective: 50K training samples

    as cost is a function of output activations, we can write:
    \boldsymbol {C=C\left(a^{L}\right)}

Backprop Equations Derivation

Following the chain rule example, we can say that any wiggle in weight/bias will affect combined input \boldsymbol {z^l} which in turn affects output \boldsymbol {a^l} and finally affecting Cost/Error \boldsymbol {C}. As a corollary, any change in input will affect output and cost and also any change in output will affect cost, notwithstanding or apart from changes in weight/bias i.e. every intermediate component affects it’s neighbor and all others down the lane. Conversely any change in final cost depends on output, input and weights and bias in that order. Enumerating w.r.t to notation described above in lth layer node j:

Possible Combinations of Changes & Affects

Changes and Affects Partial Derivative
Change in w & b affecting z, a & C
Change in weight/bias affects Cost \frac{\partial C}{\partial w_{jk}^l}, \frac{\partial C}{\partial b_j^l}
Change in weight/bias affects output \frac{\partial a_j^l}{\partial w_{jk}^l}, \frac{\partial a_j^l}{\partial b_j^l}
Change in weight/bias affects input \frac{\partial z_j^l}{\partial w_{jk}^l}, \frac{\partial z_j^l}{\partial b_j^l}
Change in z (due to change in w and /or b) affecting a & C
Change in total input affects corresponding Output \frac{\partial a_j^l}{\partial z_j^l}
Change in total input at a node affects Cost \frac{\partial C}{\partial z_j^l}
Change in a (due to change in z) affecting C
Change in output affects Cost \frac{\partial C}{\partial a_j^l}
Change in previous a affecting next a (handy in deriving BP2)
Change in previous output affects next output \frac{\partial a_j^l}{\partial a_j^{l-1}}   or   \frac{\partial a_j^{l+1}}{\partial a_j^l}

While finding change in cost w.r.t change in weight/bias yields 4 important equations that comprises the Backprop algorithm – BP1…BP4.

Change in Cost w.r.t Weight

Referring to above diagram, let’s workout the first and foremost equation – change in cost w.r.t. change in weight

  • using partial derivative notation and ‘chain rule’ to expand to its components, we can write:
    \frac{\partial C}{\partial w_{jk}^l}=\frac{\partial C}{\partial a_j^l}\cdot\frac{\partial a_j^l}{\partial z_j^l}\cdot\frac{\partial z_j^l}{\partial w_{jk}^l}
  • let’s simplify the first 2 terms in the right side as \delta^l
    \delta^l=\frac{\partial C}{\partial a_j^l}\cdot\frac{\partial a_j^l}{\partial z_j^l}
  • hence we can write the above equation as:
    \frac{\partial C}{\partial w_{jk}^l}=\delta^l\cdot\frac{\partial z_j^l}{\partial w_{jk}^l}
  • expanding for the right most expression
    \frac{\partial C}{\partial w_{jk}^l}=\delta_j^l\;\cdot\frac{\partial\left(\sum_kw_{jk}^la_k^{l-1}+b_j^l\right)}{\partial w_{jk}^l}
  • solving for the right most expression i.e. differentiating w.r.t w_{jk}^l
    \frac{\partial C}{\partial w_{jk}^l}=\delta_j^l\;\cdot a_k^{l-1}\;\;\;\;\;\;\cdots\cdots\cdots\;(BP4)

Change in Cost w.r.t Bias

  • hence we can write the above equation as:
    \frac{\partial C}{\partial b_j^l}=\delta^l\cdot\frac{\partial z_j^l}{\partial b_j^l}
  • expanding for the right most expression
    \frac{\partial C}{\partial b_j^l}=\delta_j^l\;\cdot\frac{\partial\left(\sum_kw_{jk}^la_k^{l-1}+b_j^l\right)}{\partial b_j^l}
  • solving for the right most expression i.e. differentiating w.r.t b_{j}, yields 1 for right most side
    \frac{\partial C}{\partial w_{jk}^l}=\delta_j^l\;\;\;\;\;\;\;\cdots\cdots\cdots\;(BP3)

Change in Cost w.r.t Input in Last Layer

  • using partial derivative notation and ‘chain rule’ to expand to its components, we can write:
    \frac{\partial C}{\partial z_j^l}=\frac{\partial C}{\partial a_j^l}\cdot\frac{\partial a_j^l}{\partial z_j^l}
  • we know that the output a is represented as activation function w.r.t input z as:
    a_j^l=\sigma\left(z_j^l\right)=\left(\frac1{1+e^{-z_j^l}}\right)
    • differentiating:
      \frac{\partial a_j^l}{\partial z_j^l}=\sigma'\left(z_j^l\right)=\frac{e^{-z_j^l}}{\left(1+e^{-z_j^l}\right)^2}
    • adding 1 and subtracting 1 in the numerator:
      =\frac{1+e^{-z_j^l}\;-\;1}{\left(1+e^{-z_j^l}\right)^2}
    • expanding and taking the numerators separately
      =\frac{1+e^{-z_j^l}\;}{\left(1+e^{-z_j^l}\right)^2}-\frac1{\left(1+e^{-z_j^l}\right)^2}
    • simplifying both fractions
      =\frac1{\left(1+e^{-z_j^l}\right)}-\frac1{\left(1+e^{-z_j^l}\right)^2}
    • re-writing in sigma notation
      =\sigma\left(z_j^j\right)\;-\sigma\left(z_j^j\right)^2\;
    • re-writing in vectorized form
      =\sigma\left(z^L\right)\;-\sigma\left(z^L\right)^2\;
      =\sigma\left(z^L\right)\left(1-\sigma\left(z^L\right)\right)
      =\sigma'\left(z^L\right)
  • expanding for the left most expression (in vector form)
    \frac{\partial C}{\partial a^L}=\frac{\partial C}{\partial y}=\frac{\partial\frac12\left(a^L-y\right)^2}{\partial y}
  • solving for the right most expression i.e. differentiating w.r.t y
    \left(a^L-y\right)
  • combining the above 2, we can write top equation (under this section) in vector form as:
    \delta^L=\left(a^L-y\right)\cdot\sigma'\left(z^L\right)\;\;\;\;\cdots\cdots\cdots\cdots\left(BP1\right)

Change in Cost w.r.t Input in mid Layer

We could use total derivative but it’s tedious and cumbersome and we’ll see how to derive this by applying the chain rule on single line multi-layered network and pattern it for a multi-row multi-layered network and understand the intuition behind it but for now the equation is as below:

  • vectored form:
    \delta^{l}=\left(\left(w^{l+1}\right)^{T} \delta^{l+1}\right) \odot \sigma^{\prime}\left(z^{l}\right)\;\;\;\;\cdots\cdots\cdots\cdots\left(BP2\right)
  • on a per row basis in the network:
    \delta_{j}^{l}=\sum_{k} w_{k j}^{l+1} \delta_{k}^{l+1} \sigma^{\prime}\left(z_{j}^{l}\right)

BP2 – Intuitive Deduction – Single Row Network

Let’s consider a linear network below with four neurons each with weights, biases, inputs/activations and output as shown. Superscripts represent the layer they are in.

Using the sigmoid neuron model, the output of each neuron in the 4-neuron linear network is defined as follows i.e. \boldsymbol{f} is a collective function of input comprising \boldsymbol{w^l\;a^{l-1}\;b^l} where input is notated as \boldsymbol{z^l} which is further applied to an sigmoid activation function referred byu sigma function. Superscript \boldsymbol{l} signifies the layer they belong to.

Forward Pass : Activations

  • Input at any node is given by:
    z^l\;=\;w^l\cdot a^{l-1}+b^l
  • Output at any node is given by:
    f(w^l\;a^{l-1}\;b^l)\;=\;\sigma(z^l)\;= \frac1{(1+e^{-z^1})}=\;\frac1{(1+e^{-(w^l.a^{l-1}+b^l)})}
  • At layer zero, where input
    \boldsymbol{x} is fed, a^0\;=\;x

For an input \boldsymbol{x}, the forward pass computation of the activations \boldsymbol{a^l} is as follows

  1. a^1=\sigma(z^1)\;=\;\sigma(w^1\cdot a^0+b^1)= \frac1{(1+e^{-z^1})}=\;\frac1{(1+e^{-(w^1.x+b^1)})}
  2. a^2=\sigma(z^2)\;=\;\sigma(w^2\cdot a^2+b^2)= \frac1{(1+e^{-z^2})}=\;\frac1{(1+e^{-(w^2.a^1+b^2)})}
  3. a^3=\sigma(z^3)\;=\;\sigma(w^3\cdot a^2+b^3)= \frac1{(1+e^{-z^3})}=\;\frac1{(1+e^{-(w^3.a^2+b^3)})}
  4. a^4=\sigma(z^4)\;=\;\sigma(w^4\cdot a^3+b^4)= \frac1{(1+e^{-z^4})}=\;\frac1{(1+e^{-(w^4.a^3+b^4)})}

Reverse Pass : optimize Weights for Cost Reduction

From gradient descend and chain rule, we know that we need to compute derivatives and hence compute the change to be effected in ‘w’ & ‘b’ so that cost is minimum i.e. – Recall that any wiggle in w will affect all intermediaries down the lane and we originally ventured to find change in Cost C w.r.t w for last layer. With the single row neuron multi-layer network, it’s easier to compute the intermediaries. By Chain Rule:

    • As previously computed, for last layer 4:
      \delta^4\;=\;\frac{\partial C}{\partial a^4}\;=\;a^4-y
    • computing change quantity for {w^4}
      {\triangle w^4=} -\eta\cdot\frac{\partial C}{\partial w^4}
    • by chain rule
      \frac{\partial C}{\partial a^4}\cdot\frac{\partial a^4}{\partial z^4}\cdot\frac{\partial z4}{\partial w^4}
    • expanding the constituents
      =\frac{\partial C}{\partial a^4}\cdot\frac{\partial\sigma(z^4)}{\partial z^4}\cdot\frac{\partial(w^4a^3+b^4)}{\partial w^4}
    • replacing first term from above,
      differentiating 2nd term yields \sigma'(z^4)
      differentiating last term w^4a^3 yields a^3 and b^4 becomes 0.
      hence
      =\delta^4\cdot\sigma'(z^4)\cdot a^3
    • computing change quantity for {w^3}
      \triangle w^3\;=\;-\eta\cdot\frac{\partial C}{\partial w^3}
    • by chain rule
      \frac{\partial C}{\partial w^3}=\frac{\partial c}{\partial a^4}\cdot\frac{\partial a^4}{\partial a^3}\cdot\frac{\partial a^3}{\partial w^3}
    • by chain rule
      =\frac{\partial c}{\partial a^4}\cdot\frac{\partial a^4}{\partial a^3}\cdot\frac{\partial a^3}{\partial z^3}\cdot\frac{\partial z^3}{\partial w^3}
    • solving 2nd term
      \frac{\partial a^4}{\partial a^3}=\frac{\partial(w^4a^3+b^4)}{\partial a^3}=\partial'(w^4a^3+b^4).w^4\;=\;\sigma'(z^4).w^4 \;\delta^4.\sigma'(z^4).w^4
    • replacing first & second term from above,
      \delta^3=\frac{\partial C}{\partial a^4}\frac{\partial a^4}{\partial a^3}=\;\delta^4.\sigma'(z^4).w^4
    • replacing first & second term from above,
      giving similar treatment to last 2 terms as in 1. above
      hence
      \triangle w^3=\;\delta^3.\sigma'(z^3).a^2

Now we can summarize ‘Backward Pass’ for weights and similarly you can workout for biases (encourage reader to work out in detail and left out for brevity) as follows:

based on weights based on biases
1.
  • \delta^4\;=\;a^4-y
  • \triangle w^4=-\eta\cdot\delta^4\cdot\sigma'(z^4)\cdot a^3
  • \delta^4\;=\;a^4-y
  • \triangle b^4=-\eta\cdot\delta^4\cdot\sigma'(z^4)
2.
  • \delta^3=\delta^4.\sigma'(z^4).w^4
  • \triangle w^3=\;\delta^3.\sigma'(z^3).a^2
  • \delta^3=\delta^4.\sigma'(z^4).w^4
  • \triangle b^3=\;\delta^3.\sigma'(z^3)
3.
  • \delta^2=\delta^3.\sigma'(z^3).w^3
  • \triangle w^2=\;\delta^2.\sigma'(z^2).a^1
  • \delta^2=\delta^3.\sigma'(z^3).w^3
  • \triangle b^2=\;\delta^2.\sigma'(z^2)
3.
  • \delta^1=\delta^2.\sigma'(z^2).w^2
  • \triangle w^1=\;\delta^1.\sigma'(z^1).x
  • \delta^1=\delta^2.\sigma'(z^2).w^2
  • \triangle b^1=\;\delta^1.\sigma'(z^1)

BP2 – Intuitive Deduction – 2-Row 4-layer Network

Previous section gives an idea of \delta^l computation across each layer for a single row multi-layer neural network. But a 2-row multi-layer network will provide the final answer and the previous section was a preamble and moreover we’ll not work the details of the derivation but provide the summary and for details refer to Mike Gordon’s work-out – some of which is used and given below.

Let’s focus on layer 4, from all our understanding before, we can work out the activations as given below:

\begin{array}{l}a_1^4=\sigma\left(w_{11}^4a_1^3+w_{12}^4a_2^3+b_1^4\right)\\a_2^4=\sigma\left(w_{21}^4a_1^3+w_{22}^4a_2^3+b_2^4\right)\end{array} \begin{bmatrix}a_1^4\\a_2^4\end{bmatrix}=\sigma\left(\begin{bmatrix}w_{11}^4&w_{12}^4\\w_{21}^4&w_{22}^4\end{bmatrix}\times\begin{bmatrix}a_1^3\\a_2^3\end{bmatrix}+\begin{bmatrix}b_1^4\\b_2^4\end{bmatrix}\right)
\begin{array}{l}{a_{1}^{l}=w_{11}^{l} a_{1}^{(l-1)}+w_{12}^{l} a_{2}^{(l-1)}+b_{1}^{l}} \\ {a_{2}^{l}=w_{21}^{l} a_{1}^{(l-1)}+w_{22}^{l} a_{2}^{(l-1)}+b_{2}^{l}}\end{array} \begin{bmatrix}a_1^l\\a_2^l\end{bmatrix}=\sigma\left(\begin{bmatrix}w_{11}^l&w_{12}^l\\w_{21}^l&w_{22}^{l2}\end{bmatrix}\times\begin{bmatrix}a_1^{(l-1)}\\a_2^{(l-1)}\end{bmatrix}+\begin{bmatrix}b_1^l\\b_2^l\end{bmatrix}\right)

In vector equation form: a^{l}=\sigma\left(w^{l} \times a^{(l-1)}+b^{l}\right) where

a^{l}=\left[ \begin{array}{l}{a_{1}^{l}} \\ {a_{2}^{l}}\end{array}\right]   and   w^{l}=\left[ \begin{array}{ll}{w_{11}^{l}} & {w_{12}^{l}} \\ {w_{21}^{l}} & {w_{22}^{l}}\end{array}\right]   and   a^{(l-1)}=\left[ \begin{array}{l}{a_{1}^{(l-1)}} \\ {a_{2}^{(l-1)}}\end{array}\right]   and   b^{l}=\left[ \begin{array}{ll}{b_{1}^{l}} \\ {b_{2}^{l}}\end{array}\right]

Working out the backward pass from layer 4 to layer 3 and summarized as below:

\delta Derivations Weight Partials Bias Partials
\delta_{1}^{3}=\left(\delta_{1}^{4} w_{11}^{4}+\delta_{2}^{4} w_{21}^{4}\right) \sigma^{\prime}\left(z_{1}^{3}\right) \frac{\partial c}{\partial w_{11}^{3}}=\delta_{1}^{3} a_{1}^{2} \quad \frac{\partial c}{\partial w_{12}^{3}}=\delta_{1}^{3} a_{2}^{2} \frac{\partial c}{\partial b_{1}^{3}}=\delta_{1}^{3}
\delta_{2}^{3}=\left(\delta_{1}^{4} w_{12}^{4}+\delta_{2}^{4} w_{22}^{4}\right) \sigma^{\prime}\left(z_{2}^{3}\right) \frac{\partial c}{\partial w_{21}^{3}}=\delta_{2}^{3} a_{1}^{2} \quad \frac{\partial c}{\partial w_{22}^{3}}=\delta_{2}^{3} a_{2}^{2} \frac{\partial c}{\partial b_{2}^{3}}=\delta_{2}^{3}

The above two \delta equations can be written as the single equation using matrices and vectors as
\left[ \begin{array}{l}{\delta_{1}^{3}} \\ {\delta_{2}^{3}}\end{array}\right]=\left(\left[ \begin{array}{cc}{w_{11}^{4}} & {w_{21}^{4}} \\ {w_{12}^{4}} & {w_{22}^{4}}\end{array}\right] \times \left[ \begin{array}{c}{\delta_{1}^{4}} \\ {\delta_{2}^{4}}\end{array}\right]\right) \odot \sigma^{\prime}\left(\left[ \begin{array}{c}{z_{1}^{3}} \\ {z_{2}^{3}}\end{array}\right]\right)

Let w^{l^{\top}} be the transpose of w^{l}, z^{l} the vector with components z_{j}^{l}, \delta^{l} \delta_{j}^{l} ,hence w^{l^{\top}}=\left[ \begin{array}{ll}{w_{11}^{l}} & {w_{21}^{l}} \\ {w_{12}^{l}} & {w_{22}^{l}}\end{array}\right] and z^{l}=\left[ \begin{array}{ll}{z_{1}} \\ {z_{2}^{l}}\end{array}\right] and \delta^{l}=\left[ \begin{array}{c}{\delta_{1}^{l}} \\ {\delta_{2}^{l}}\end{array}\right]

considering that multiplication is commutative, hence \delta_{j}^{l} w_{j k}^{l}=w_{j k}^{l} \delta_{j}^{l}, we can write finally as a single vector equation:
\delta^{3}=\left(\left(w^{4}\right)^{\top} \delta^{4}\right) \odot \sigma^{\prime}\left(z^{3}\right)
Which gets us, the generics BP2
\delta^{l}=\left(\left(w^{l+1}\right)^{\top} \delta^{l+1}\right) \odot \sigma^{\prime}\left(z^{l}\right)

Backprop – A Different Take

As in the above diagram, Activation w.r.t layer 3 (from 3 to 4): is nothing but sum of layer 4 weights multiplied by layer 3 outputs and added to layer 4 bias fed to activation functions gives the output for Layer 4 i.e. a^L\;=\;\{a_1^4,\;a_2^4\} whereas Back propagation w.r.t layer 3 (from 4 to 3): is layer 4 weights transpose times layer 4 change in cost w.r.t output hadmard product of layer 3 activation differential. This is illustrated in the below diagram.

Inner Workings of Bare NeuralNet – Matrices matched to Code

Top

MNIST Dataset on Our Neural Net – Nodes and Layers

Similar to ‘hello world’ program being defacto for any language, MNIST dataset is the defacto for NeuralNet learning. It contains 60K training images and 10K testing images. These images are normalized to fit a 28×28 pixel resolution and anti-aliased to give greyscale. We’ll go ahead to construct a neural net whose:

Generated with NN-SVG Tool
  1. First Layer – Input Layer:
    Will be a 784 node – featuring the 28×28 pixel bound box image as a linear 784×1 array matrix
  2. 2nd Layer – Hidden Layer
    Will be a hidden layer of 100 nodes – based on the notation, w_{jk}^l where j is the current layer (lth) and k is the previous layer ((l-1)th) and the weight matrix will be (100×784), bias matrix will be (100×1)
  3. 3rd Layer – Hidden Layer
    Will be a hidden layer of 40 nodes – and the weight matrix will be (40×100), bias matrix will be (40×1)
  4. 4th/Last Layer – Output Layer
    Will be a output (pseudo hidden) layer of 10 nodes – and the weight matrix will be (10×40), bias matrix will be (10×1). The last layer is so chosen, so that each will represent one digit (either one among the numerals 0…9. Since the value is real, the output signifies a probability of that particular numeric value, say above 0.5 is sure thing and less than 0.5 is a non-occurrence.

We will use above neural net architecture to understand the python code given in Michael Nielsen’s book. This provides a compact, simple bare minimum neural net implementation in python.

  1. Clone/copy below repo to a local directory
    https://github.com/mnielsen/neural-networks-and-deep-learning.git
  2. Move to your local directory where you cloned and launch python 2.6 or 2.7 from here
  3. Run this snippet to get going:
                                
                                import sys
                                # ---- set local directory path where repo was cloned /copied
                                sys.path.append("your local directory path/src")
                                
                                # ---- import the modules for data-prep and neural net respectively
                                import mnist_loader 
                                import network
                                
                                # ----- prep the data from source MNIST dataset 
                                training_data, validation_data, test_data = mnist_loader.load_data_wrapper()
                                # ----- create the network to our specs -- layers with requisite nodes 
                                net = network.Network([784, 100, 40, 10])
                                # ------ run the stochastic gradient descend routine with set configurations
                                net.SGD(training_data, 30, 10, 3.0, test_data=test_data)
                                

Overall NeuralNet in Python

Peek into source of network.py and follow along the functions mentioned below to get a better sense of what the python logic does!

  1. SGD Function
    (loop around number of epochs)
    1. Parameters:
      • training_data – training data matrices formatted accordingly
      • epochs – number of iterations on the entire dataset to go through
      • mini_batch_size – obvious as the name indicates
      • eta – learning rate multiplier
      • test_data – test data matrices formatted accordingly – used for evaluation of the learned weights and biases
    2. Shuffles all 50K training recirds (each containing 784×1 greyscale data) into mini-batches of size 30 records
    3. Performs mini-batch operation (which computes delta weights and biases, updates original weights and biases)
  2. update_mini_batch Function
    (loop around for each mini-batch count) where mini-batch count = total training samples divided by mini-batch size
    1. Initialize weight and biases matrices to zero
    2. Perform backprop (includes feed forward and back propogation). Loop around for each training record + its output record in mini-batch
    3. Get the delta matrices and add to previous delta matrices (for both weights and biases)
    4. Compute GD and subtract from weights and biases
  3. backprop Function
    1. Perform feed-forward – loop around each layer of weights/biases computing input & activation matrices for each layer
    2. Perform backward pass
    3. first compute last layer delta
    4. loop around from last-but-one layer till 1st layer & compute delta matrices at each layer for weights and biases

    Feed Forward & BackProp – Matrices Computed in Python

  4. evaluate Function
    1. With computed weights and biases by SGD, perform feed-forward for the training input
    2. Get the final output and compare it to actual output
    3. Report the compared correct result count


Hope it helps in your Deep Learning journey a bit and if you’d like to reach out, drop me an email

Learning Curve Retraced – References & Acknowledgements

Top

These are the books, blogs and code, I read, pondered and assimilated as I tried to understand deep learning and hope by stating it here, might help referring them in your journey and at same time acknowledging all who have similarly helped forging deep learning.

World Happiness Report 2019

On 20 Mar 2019, World Happiness Report 2019 was released and here are the ranks:

Rank – Country

  • 1 – Finland (First)
  • 9 – Canada
  • 15 – UK
  • 19 – US
  • 34 – Singapore
  • 67 – Pakistan
  • 93 – China
  • 125 – Bangladesh
  • 130 – Sri Lanka
  • 140 – India
  • 156 – South Sudan (Last)

Courtesy & Source: https://bit.ly/2Jq3RGH

People responsible apart from their routine power grabbing, hoarding money, unceasing corruption, mindless activities – should think seriously about this report and include it in their election manifesto about raising “Happiness Index” at least few notches up lest a leap ahead…if by any miracle they could do it, that’ll be their finest achievement and victory in service for the billions inhabiting the world!

Will they ever do it? There’s a proverb, can a dog’s tail be untangled? I think in the world where surgical strikes are possible to rectify certain things immediately, why not for “Happiness Index”? Think methodically and tackle consistently with surgical strike precision, few misses are inevitable, but overall work for the billions of humans in mind. They’re vexed and need solutions.

The factors that have to change (which are also taken up in rating the above countries) are:

  1. GDP per capita
  2. Social Support
  3. Healthy Life Expectancy
  4. Freedom to make Life Choices
  5. Generosity
  6. Perceptions of Corruption

Certain countries are pathetic and miserable in every count; their neighbors are far better. Would the power brokers finally take notice? Awareness is more important in their ballot decisions and in asking pertinent and daring questions to our leaders!