A collection of fragments of understanding in the pursuit of deeper questions.
Lecturer: Giacomo Indiveri
What is this course about?
The Human Brain The average human brain: 1.5kg weight, 1.1 - 1.2l volume.
We have a brain because we can move and interact with the world.
For example, the "Sea squirt" to the right has a brain in the early stages (few hours) of its life, but once it finds a rock to stick to, it digest its brain as it is no more needed for movements and would represent a high-consuming energy organ.
Producing Behaviour Our brain is fundamentally needed to interpret sensory data, take decisions and actions. It is needed to act in the environment.
Neurons and Brains Human brains are large, but by far not the largest (elephant, whales, ...). The cells (neurons) that make up brains are very similar between species.
![]() |
![]() |
|---|
Brain Evolution The size and density of the brain have been constantly increasing over time during evolution.
A very brief History of Neuroscience
Synapses In 1897 Charles Sherrington introduced the term synapse to describe the specialized structure at the zone of contact between neurons as the point in which one neuron communicates with another.
EPSC and EPSP
Spike Generating Mechanism If the membrane voltage increases above a certain threshold, a spike-generating mechanism is activated and an action potential is initiated.
Spike Properties and the F-I Curve
![]() |
![]() |
|---|
The First Models of Neurons - Warren McCulloch and Walter Pitts (1943)
![]() |
![]() |
|---|
"A logical calculus of the ideas immanent in nervous activity" The McCulloch&Pitts model quickly became extremely popular, and dominated the Artificial Neural Network scene for decades. Why? Isomorphism with calculus of logical propositions. In the hand of John von Neumann, the McCulloch & Pitts model became the basis for the logical design of digital computers.
The Turing Machine It is an universal computing machine, through a binary state machine all the logical operations could be performed.
Artificial vs Natural Intelligence How is the brain different from a computer?
Hard and Easy Problems "The main lesson of thirty-five years of AI research is that the hard problems are easy and the easy problems are hard" - Steve Pinker "It is comparatively easy to make computer exhibit adult level performance on intelligence tests or playing checkers, and difficult or impossible to give them the skills of a one-year-old when it comes to perception and mobility" - Hans Moravec
Artificial vs Real Neural Networks
Artificial Intelligence: The Deep Network Revolution
Deep Networks Galore
Deep Networks Computing Power Demands
Problems and Limitations of Artificial Intelligence
Neuromorphic Intelligence
Design Principles for Emulating Natural Intelligence
Brain-Inspired Principles: a Radical Paradigm Shift Exploit Physical Space
The Neuromorphic Computing Approach
Edge-Computing Application Specific Tasks Technology Transfer and Applications We are now entering the era of neuromorphic intelligence in which dedicated cognitive chiplets will be used to provide intelligence to a multitude of extreme edge-computing use cases.
On-going research: Mind-Brain-Body Iterative Refinement
Lecturer: Wolfger Von Der Behrens
Types of Neurons Morphology and electrical properties are what characterize neurons.
![]() |
![]() |
![]() |
|---|
Terminology
Axon
Vescicles Lyposomes release the contained neurotransmitter when the action potential arrives.
![]() |
![]() |
![]() |
|---|
Neurotransmitter Many different types, the most commons are the Glutamate (excitatory) and GABA (inhibitory). One neuron typically releases just one type of neurotransmitter.
![]() |
![]() |
|---|
Post-Synaptic Receptor There are two major classes of receptors: Ligand-Gated Ion Channels and G-Protein-Coupled Receptors (neurotransmitters bind to the receptor, which activates the G-Protein that modulates Effector protein) then depending on the ions flowing in, it can depolarize or hyperpolarize.
Dendrites The neuron depolarizes when Glutamate is released, such that it reaches the threshold to generate an action potential. On the other side, if the neuron is hyperpolarized by the release of GABA and Glutamate is released at the same time, we have a mediated effect (orange line) which is probably not reaching the threshold to generate an action potential.
![]() |
![]() |
|---|
Sequence of Events
Central Nervous System (CNS) vs Peripheral Nervous System (PNS) + Enteric Nervous System (ENS). Among the CNS we subdivide in Sympathetic ("Fight or Flight") and Parasympathetic Nervous System ("Feed & Breed").
Gross Anatomy: Protection and Sustenance of the Brain The brain is floating in the cerebrospinal fluid inside the skull.
The Meninges are three layers of tissues: Dura Mater (Hard structure, mechanical protection), Arachnoid Mater (Soft structure) and the Pia Mater.
Ventricular System
Navigating the Central Nervous System
![]() |
![]() |
![]() |
|---|
Direction of Orientation in the CNS
Neural Development

Major Divisions of the Brain
![]() |
![]() |
|---|
Cranial Nerves
![]() |
![]() |
|---|
The Limbic System It was named by Brocas, he hypothesized that it is involved in emotional processing and reactions. Hippocampus is one of the main memory forming structures.
Papez Circuit
Amygdala is strictly connected with the generation and processing of fear.
Hypothalamus and Thalamus On top we have the Thalamus, while the Hypothalamus sits below. The first one is the input structure into the cortex. Most of the sensory inputs are gated into the cortex by the Thalamus. It does not only relay information, but it also performs a processing of information. The Hypothalamus is more part of the endocrine nervous system, it has the control of the epituary gland.
Thalamus Upper Brain Stem: Diencephalon
Thalamus:
Hypothalamus:
Basal Ganglia
It sits below the forebrain, it is strongly interconnected with the cortex and the Thalamus. It is highly involved in movements. In the case of Parkinson's disease we assist at a loss of dopaminergic neurons.
Through Deep Brain Stimulation, i.e., an electrical stimulation of the deep brain tissues, it is possible to reduce the symptoms related to diseases affecting Basal Ganglia and Substantia Nigra.
Lobes of Cerebral Cortex
Cytoarchitecture The cells are organized in a very systematic and consistent structure made of 6 layers. In the first layer there are mostly projections, very few cells. Cell bodies are stained with Golgi technique even today and in the following picture cell bodies and their projections are shown in the layers 2 - 6. The 2 - 6 layers structure can change depending on the brain cortical area and the functions that they perform.
Cortical Areas
Brodmann's Areas (BA) The cytoarchitecture is not consistent throughout the cortex, Brodmann looked at the composition of layers in different areas in the brain. Hence, he defined a map of the brain based on the histological differences. The map has still some similarities with the actual knowledge of the brain areas, however it was not saying anything about the functions of the various regions.
Extracellular recordings Tungsten microelectrode for recording from single units: an electrode has been developed to fill the need for an easily made, sturdy device capable of resolving single-neuron action potentials at least as well as the commonly used micropipette.
Circuits The basic cortical circuits are either Feedforward or Feedback. In particular, layer 4 of the cortex is the "input layer" of the cortex.
Lecturer: Valerio Mante
Levels of Description Typical of biology to have multiple levels of description, and we don't know if each of them matter or not. Can we ignore some of these details and still get the same outcome when we try to replicate such a complex structure? No one can answer as of now.
Goal: create an artificial intelligent agent
Why study single neurons?
Neurons are Diverse Neurons are specialized to do some specific computations.
Two approaches to understanding neural computations:
Single Neuron Computations These kind of phenomena, such as transmission delays, dendritic computations and back-propagating action potentials, are not describable when the "point-neuron" model is used.
Neuromorphic Implementation
Experimental Procedures
The Resting Potential Why do cells have membrane potential? It is a way to store energy. Neurons invest energy to perform concentration differences.
The Basic Ingredients
A thought Experiments The previous 3 ingredients give rise to action potentials. We are going to perform a thought experiment, the box to the left simulates a cell environment. At t = 0, the molecules are only inside the cell. At t = 0, the concentration inside is higher than outside. After some time t inf, the concentration reaches an equilibrium at a macroscopic level, even though at a microscopic level small exchanges continue to happen through the channels.
![]() |
![]() |
|---|
Now, we are going to assume more complex molecules, i.e., ions with a charge. Furthermore, the ionic channels are going to be selective, which allow only the passage of positive ions. What do we expect now at t inf? In this setting we do not reach the same equilibrium as before, indeed the inside starts to turn negatively charged every time a positive charge io goes outside the cell. But then, for every positive ion that goes out, the inside becomes more negative and attracts more the remaining positive ion channels, reducing the chances that further positive ions "escape".
![]() |
![]() |
![]() |
|---|
The Cell Membrane This is essentially how things look like at the equilibrium.
Selective Ionic Channels & Ionic Flux There is asymmetry in the ionic flux, indeed, when one positive ion hits the channel from the outside, it will be dragged inside from the electric field. On the other side, a positive ion would need to have a kinetic energy bigger than to be able to cross the ionic channel from the inside.
![]() |
![]() |
|---|
The Boltzmann Factor The Boltzmann Factor is a statistical quantity that describes the probability of a system being in a certain energy state at thermal equilibrium. It is given by the following formula:
where , and .
On the y-axis, , represents the percentage of the ions having enough energy to cross the ionic channel.
The Nernst Equation The Nernst Equation is an equation that describes the relationship between the electrical potential of a cell and the concentration of ions in the cell. It describes dependencies in equilibrium potential, which also take the name of reversal potentials
![]() |
![]() |
|---|
What about the assumptions? We assumed fixed concentration.
The Reversal Potential
![]() |
![]() |
![]() |
|---|
Two Channel-Types The equilibrium is reached when the net current is 0.
Goldman-Hodgkin-Katz Equation
Energy Consumption in the brain
Lecturer: Valerio Mante
![]() |
![]() |
|---|
Ohmic Conductances
Current through the membrane passes through a particular channel type as a function of the voltage across the membrane. The slope in the following graph is the conductance through this channel. We say conductance is ohmic when we have passive properties, thus, . Different channel types have different resting potentials (as shown below). Example active conductances (mammals, approx. 37C degrees)
Synaptic Currents Synapses are injecting (external) current. Assume dendrite are passive cables that just conduct current. The dendrites tend to have much less active conductances respect to the axons. This way, we have three models:
![]() |
![]() |
|---|
Development of the Cable Equation The key formalism that was used successfully to justify what is happening is the cable equation. The cable equation was derived by Thomson and had practical relevance for transatlantic telegraph cable. In our discussion, we will consider cables as good approximations of dendrites.
We want to find an expression for , i.e., we want to derive the equation for the voltage of the cell as a function of time and spatial dimension. We will use two steps:
Single-Compartment Model It is a very simple model of a neuron.
![]() |
![]() |
|---|
The capacitive current might end on the surface of the membrane and increase/decrease the charges on the membrane. Otherwise, it could flow out as leak current. What would happen? In the beginning (if it is positive) will charge the inside of the membrane, so the potential will increase, as the potential increases and gets more different from the resting potential will eventually lead to a new equilibrium in which the cells is more depolarized, to the point that the current flowing out of the membrane is balanced with the current flowing in.
![]() |
![]() |
![]() |
![]() |
|---|
Deriving V(t)
Steady-State Solution We inject a current and the membrane potential will increase to an asymptote, which can be derived as shown below. Assuming that we will continue to inject the current constantly. If we double the current, we double the voltage difference that we get and the height of the asymptote, will depend on the properties of the resistance. Hence will increase if increases (less leak) or increases (more input).
Input Resistance Two neurons with the same concentration of channels, differing only regarding the size: the smaller neuron will have more resistance, thus, for the same amount of injected current it will have smaller voltage change. Furthermore, two neurons, one myelinated and another unmyelinated. The myelinated has less resistance thus it needs less current input to achieve the same voltage change. Also, you will need a larger current to create a certain amount of depolarization in a neuron with more channels compared with one of the same dimensions but with fewer channels.
General Solution is the memory of cell approx. 10 to 100 ms, Neuron forgets after . Longer memory requires few mechanisms, for example, charge in synapse or recurrency. In general, we will have a time dependency, so the potential as a function of time will be given by:
Where , and .
Implications Time constant :
Spatial and Temporal Summation
Integrate and Fire Neurons When an Integrate and Fire neuron achieve the threshold, it generates an action potential and right after, reset it.
Idealized synapse: If then leading to EPSP (depolarization). If then leading to IPSP (hyperpolarization).
Equivalent Circuits We can draw the electric circuit that captures the basic properties of neurons.
Then we add a synapse to the circuit:
Deriving the Cable Equation
So far, ions flow (in out) to achieve Erest. But, what if most of the time?
Longitudinal current In the cable equation we use the same variables as before but we need to express : longitudinal current. Now, we have a current that flows inside the membrane (for instance, from left to right).
![]() |
![]() |
|---|
![]() |
![]() |
![]() |
|---|
So cable equation derives from conservation of energy and conservation of charge.
![]() |
![]() |
|---|
Case 1: Infinite Cable & Constant Current
Case 2: Infinite Cable & Current Pulse
![]() |
![]() |
|---|
Passive Currents in a Branching Neuron
The Big Picture
Lecturer: Valerio Mante
An Action Potential is a depolarization that starts driving down the axon, which is then followed by hyperpolarization.
How to explain these properties of the AP
Answer: g = g(V, t): voltage-dependent channels in the axon. They open and close dependent on the voltage that they experience on the membrane. Hodgkin-Huxley: Nobel Prize Medicine Physiology in 1963. Everything they did was before the existence of ion-channels (membrane channels) were known.
During an AP we see channels opening and pulling V towards E. The hypotheses is that in the rising phase of AP the sodium and calcium conductances increase (gNa and gCa), and in the decaying phase of AP the sodium and calcium conductances decrease or potassium and chloride conductances increase. All as function of V. For testing these hypotheses, we need to measure , , etc... We can use the IV-relation: measuring , , etc... for different V then infer , , etc... To do this, we can use a voltage clamp.
Squid Giant Axon
Voltage Clamp A new technique invented by Hodgkin and Huxley. Previously, current was injected and voltage V was measured, now we set V and measure required to keep = . It measures the current required to clamp the membrane voltage. Fast feedback system to fix V and measure I. has the opposite sign, i.e., is positive if from outside to inside. But to keep constant, it is necessary to inject a current opposite to the ionic current. In the end, the current injected can be read as the ionic current (in the ionic current convention).
Space Clamp
It makes the axon isopotential, do not have an AP but it is the same mechanism. The giant axon in squid has approximately 1mm of diameter and it is like a long wire, making the axon isopotential.
Voltage Clamp Experiment
Identifying the Currents
![]() |
![]() |
|---|
Voltage and time-dependent conductances for , : increases quickly (fast activation), but then inactivation kicks in and it decreases again (fast inactivation). increases more slowly (slow activation), and only decreases once the voltage has decreased (no inactivation).
Towards a Mechanistic Model They proposed an hypothesis of what might be causing voltage and time-dependence, which is going to be formalized in the lines of the previous image. The white dots is what they measured and the models estimate the lines. How to explain voltage and time-dependence in and ?
Two possibilities:
Today we know that the second possibility is correct:
Single Channel Current
There are two types of voltage-dependent conductances:
Hodgkin & Huxley formalism is used for active conductances in general: where is the overall conductance of channels of type i; is the maximal conductance (if all channels were open); and is the probability of the channel to be open (or the fraction of channels that are open).
Persistent Conductances Assuming that k events (independents and identical) are necessary to open a single channel, then . n is a gating/activation variable: the probability of a subunit gate to be open, and it is voltage and time-dependent. k is the number of subunits necessary to open each channel. According to Hodgkin & Huxley (it is necessary 4 subunits to open the channel). When k was fitted to data it leaded to corrected predictions for K+ channels.
Gating-Variables: Time-Dependence
Persistent and Transient Conductances Transient conductance includes inactivation:
Where is the activation variable and is the inactivation variable, which also represents the probability that the channel is not blocked by the inactivation gate.
Gating-Variables: Voltage-Dependence
![]() |
![]() |
![]() |
|---|
The Hodgkin and Huxley Model It is the model that describes how action potentials in neurons are initiated and propagated.
![]() |
![]() |
|---|
Fitting the Hodgkin and Huxley Model
![]() |
![]() |
![]() |
|---|
{width="3.884027777777778in"
height="0.4263888888888889in"}
Model Predictions
Single Neuron Computations
Lecturer: Daniel Kiper
Discovery of Synaptic Transmission
"So far as our present knowledge goes, we are led to think that the tip of a twig of the arborescence is not continuous with but merely in contact with the substance of the dendrite or cell body on which it impinges. Such a special connection of one nerve cell with another might be called a synapse". "Such a surface might restrain diffusion, bank up osmotic pressure, restrict the movement of ions, accumulate electric charges, support a double electric layer, alter in shape and surface tension with changes in difference of potential ... or intervene as a membrane between dilute solutions of electrolytes of different concentration or colloidal suspensions with different sign of charge".
Soup vs Spark Controversy about Synaptic Transmission - Chemical or Electrical? Is Synaptic Transmission mediated chemically or by direct electrical transfer of charge? Evidence for chemical transmission at the Neuro Muscular Junction (NMJ) was widely accepted by Neuropharmacologists. Some of the physiologists thought that certain aspects were too fast to be mediated chemically. Chemical synapses is the predominant way of communication between neurons, but there are some electrical synapses. In the retina, we have large networks of photoreceptors: rods and cones. Rods are connected to each other via electrical synapses.
Otto Loewi, Chemical Transmitter
The picture above represents Otto Loewi vagus nerve experiment. Stimulating the vagus nerve slows down the heart beat, it has an inhibitory function. In the experiment, two solutions are connected, one heart in each. After stimulation of one of them, the other one, after a short time, has the same effect. Ringers solution is a mixture of chemicals in which the heart can continue beating. When switching the solution with one that has been used with an activated vagus nerve, the heart will slow down. It was found that the "Vagusstoff" is acetylcholine (Ach). The synapses are receptive for nicotine, muscarine and acetylcholine, because of Ach-receptors. This makes certain substances very addictive. Residual Ach has to be cleared and removed immediately. This happens with Ach esterase enzymes.
Chemical Synaptic Transmission
Communication between cells which involves the rapid release and diffusion of a substance to another cell where it binds to a receptor (at a localized site) resulting in a change in the postsynaptic cell properties.
A hall-mark of chemical transmission is a delay between presynaptic Ca2+ elevation and secretion. The delay can be as short as 0.2 ms, but is usually longer due to a variety of factors.
Steps to Chemical Synaptic Transmission
![]() |
![]() |
|---|
Criteria that Define a Neurotransmitter
Neurotransmitters may be either small molecules or peptides.
Mechanisms and Sites of synthesis are different
Model of Synaptic Transmission (Standard Katz Quantal Model) This theory has been developed by looking at the amplitude of EPP. Neurotransmitters are released in discrete packages, or quanta.
Failure analysis reveals that neurons release many quanta of neurotransmitter when stimulated, that all contribute to the response.
If the probability of a single unit responding is "p", and if each unit has an independent and equal "p", then the mean number of units responding to each stimulus is given by: "np" where n is the total number of available quanta.
Then, the probability that x-units successfully contributing is given by the binomial distribution: .
Quanta correspond to release of individual synaptic vesicles. EM images and biochemistry suggest that a MEPP could be caused by a single vesicle. EM studies revealed correlation between fusion of vesicles with plasma membrane and size of postsynaptic response.
CNS Synapses and Quanta At CNS there are fewer release sites in the synapses than in the NMJ, hence if we have only one synapse it is unlikely that it will lead above-threshold the postsynaptic site. So, we will need many synapses to be active at the same time. That implies that multiple cells connected to the neuron have to release simultaneously to drive a postsynaptic AP.
CNS Synapses and Miniature Release
Docked Synaptic Vesicles It is an expression that indicates the population of vesicles in the presynaptic site that are ready to release. It defines the number of readily releasable vesicles a synapse has available. A consequence of having a limited number is depletion at high stimulus frequency, CNS synapses may have only a small number of docked vesicles on the order of 5-10 vesicles for a hippocampal CA1 synapse.
Calcium influx is Necessary and Sufficient for Neurotransmitter release. In the following pictures, Calcium is artificially injected in the cell which results in the postsynaptic membrane potential to show a response. In this experiment, the injection results in an increase in the postsynaptic potential, while using a calcium buffer (absorbs the calcium) we can show a reduction in the postsynaptic potential.
![]() |
![]() |
![]() |
|---|
Summing Up:
The Synaptic Vesicle Cycle
Synaptic Vesicle Release Consists of Three Principal Steps
Priming Vesicles in the reserve pool undergo priming to enter the readily-releasable pool. At a molecular level, priming corresponds to the assembly of the SNARE complex.
The SNARE Complex
Endocytosis retrieves synaptic vesicles membrane and protein from the plasma membrane following fusion
Lecturer: Daniel Kiper
Post-Synaptic Receptors Neurotransmitters cross synaptic-cleft and can bind to two types of receptors:
![]() |
![]() |
|---|
Ionotropic Receptors
Metabotropic
![]() |
![]() |
|---|
NMDA Receptor One of the best known receptor because it is involved in synaptic plasticity. It is voltage-dependent and this peculiarity allows it to have a high flexibility in terms of behaviour to different conditions of current and neurotransmitters.
Two Principal Kinds of Synapses: Electrical and Chemical
![]() |
![]() |
![]() |
|---|
Gap Junctions are Formed where Hexameric Pores called Connexons Connect with one Between Cells.
Electrical Synapses are Built for Speed The delay between the onset of the depolarization in the pre- and post- synaptic neurons is very small. Electrical Coupling is a Way to Synchronize Neurons with One Another Mechanisms for how the neurons code information.
![]() |
![]() |
|---|
Electrical vs. Chemical Synapse
Modelling Synapses The following plots depict the results of an experiment where they patch-clamped a piece of membrane containing a synapse. In the left picture, we see the post-synaptic potential that is elicited when the pre-synaptic neuron was stimulated. Since a voltage clamp is negative when we have an excitatory post-synaptic potential. Depending when and at what level we clamp the membrane, it will produce different sizes of post-synaptic potentials. On the right, the plotting of the previously measured post-synaptic potentials shows the proportionality between potential and current. Hence, we can conclude that synaptic input is well-captured by Ohm's Law.
![]() |
![]() |
|---|
Equivalent Circuit of a Fast Chemical Synapse
Modified membrane patch equation with a synapse:
Rewriting, we get:
Alpha Function Synaptic input is usually approximated by an "Alpha Function" of the form:
You will need to add synapses in parallel with the RC circuit to create additional synaptic components. In the following way:
![]() |
![]() |
|---|
Plasticity
Imagine two neurons embedded in a complex network and they have a synapse that connects neuron A to neuron B, if such synapse occur with high probability, as soon as neuron A is active then also neuron B is also active, then this synapse will become stronger.
Spike-Time Dependent Plasticity
To the left of the dashed line the presynaptic neuron fired before the postsynaptic, while to the right we have the opposite. When the presynaptic neuron reliably fires before the postsynaptic one, we have LTP. However, the firing of the pre- and post-synaptic neurons have to be correlated and happen with a very small delay, indeed, if the time-span between the two AP is too large it either induces LTD (the synapse becomes weaker) or no LTP happens at all.
Lecturer: Benjamin Grewe
From temporal to rate coding, and single neurons to networks. How does the brain represent what we perceive? Perkel & Bullock (1968): The problem of neural coding is to elucidate "the representation and transformation of information in the nervous system".
Representation and Transformation of Information The simplest organism that uses spikes is the Paramecium (a "Swimming Neuron"), which represents some of the experience it has about the world. It can use such Action Potentials for movements.
The Coding Metaphor Considering three elements: Correspondence, Representation & Causality:
Encoding and Decoding of Information
In general:
Finding the Stimulus-Response Relation
Encoding Motor Output in Primates One of the first experiments investigating how is a motor command in arm reachment encoded by neurons's in the motor cortex. (Georgopoulus et. al., 1982).
Recording Neuronal Responses in Cat VI
Hubel & Wiesel wanted initially to find a neuron that was responsive to the black dot, they accidentally found the reaction to the edge of the paper onto which the black dot was depicted.
Orientation and Direction Selective Neurons in VI Neurons have a receptive field and they show direction selectivity to the stimuli.
Edge Filters in Primate Visual Cortex Edge filters constitute a way to represent in low-dimensional manner natural images, indeed with just a couple hundreds neurons you can reconstruct complex images through edges.
Paper: "Spatial Structure of Neuronal Receptive Field in Awake Monkey Secondary Visual Cortex (V2)".
![]() |
![]() |
|---|
Encoding Complex Stimuli in Primate V4 They proved the activity in V4 and figured out a way to design an experiment to understand what type of features maximally excite neurons in V4. Paper: "Neural Population Control Via Deep Image Synthesis".
Encoding Visual Stimuli in the Human Brain (Area MTL) They measured MTL neurons activity, they showed pics of people and measured that this patient had neurons responding to Jennifer Aniston's pictures. Hence, at MTL we have an high-representation of concepts, such as the Jennifer Aniston's character.
Encoding Spatial Information in Rata Hippocampus O'Keefe, M. B. Moser and E. Moser Nobel prize. It is possible to reconstruct the position of the mouse along the track based on the decoding of information encoded by spines. (Ziv and Schnitzer, 2013).
![]() |
![]() |
![]() |
![]() |
|---|
Which Features of the Spike Trains are the Signal? Rate Coding refers to information being carried by the firing rate. It is often argued, or assumed, that firing rate captures essentially all relevant information. (rate code means that I have a certain variable, which could be intensity or orientation, then I have a tuning curve and the more I tweak this variable the more I have a continuous reflection of my out-of-world variable and the spiking frequency of this neuron.) Temporal Coding may refer to several quite different ideas:
Temporal vs Rate Code
![]() |
![]() |
|---|
Phase Coding in Hippocampus Different cells responding to different stimuli encountered during the "trail". When the mouse is sleeping he replays the sequence faster but in the same order.
Hence, in the Hippocampus the information is mostly rate coded, but phase delay information (temporal coding) is also relevant. Indeed, the phase delay of spikes, with respect to the background oscillations, gives position cues that can be used to decode the position of the mouse.
Sound Localization by Measuring the Interaural Time Difference (ITD) The precise timing of spikes is directly used to hear and code the position of a prey (Barn Owl vs Mouse). The temporal delay between left and right ear is combined through delay lines. These neurons in the middle only activate when the stimuli arrive simultaneously.
How to Investigate the Stimulus Encoding of a Neuron? The same stimulus can be encoded very differently by different neurons. On the right we can see the factors that may cause such encoding differences.
![]() |
![]() |
|---|
In the cortex, most of the inputs for each neuron is not coming from the outside, but rather from neighboring neurons. In the cortex, approx. 4% of synaptic inputs are actually coming from the Thalamus and the Retina. Hence, the cortex is highly recurrent and the brain has a certain state that changes all the time, i.e., what we think. Depending on what we think, we might have different stimuli in the visual cortex.
What is the simplest possible relation between stimuli and encoded signals?
The Neuron as a Temporal Filter
Linear Temporal Filter:
The Running Average Filter In this filter we take N time points and we average them.
Linear Temporal Filter:
The Leaky Average Filter
Linear Temporal Filter:
Basic Model of Linear Spatial Filtering (against the previous temporal filtering) This filter is local in space. The center is weighted positively, while the surround is weighted negatively (On/Off).
![]() |
![]() |
|---|
![]() |
![]() |
![]() |
|---|
Combining Temporal and Spatial Filtering This is most likely what the brain is doing, i.e., integrate not only across space but also across time.
![]() |
![]() |
|---|
Combining Filtering with a Nonlinearity
Problems:
Linear Filter + Nonlinearity:
Taking into Account Spatio-temporal Features
![]() |
![]() |
![]() |
|---|
![]() |
![]() |
|---|
To Measure Population Activity in vivo it is possible to use Electrodes and Ca2+ Imaging. Now, we want to repeatedly sample the responses to a variety of stimuli so that we can characterize what feature combination triggers a spike or a behavior.
After collecting data, if we don't have any labels for the stimuli, we use an unsupervised/clustering approach, otherwise a supervised approach to identify the characteristics that trigger a behavior.
Population coding refers to information available from ensembles that goes beyond simple summation of individual signals. It is often associated with the method of Georgopoulos et. al. (1996), but many scientists have also asked what an "ideal observer" could learn from a population of neurons.
Finding the Single Neuron Response Vector & Projecting Stimuli in the direction of Neuronal Response (Encoding/Filtering)
![]() |
![]() |
|---|
Finding the I/O Function for a Single Neuron
The I/O function is:
Where as identified by our linear filter.
The I/O function can be found from data using the Bayes' rule:
![]() |
![]() |
![]() |
![]() |
![]() |
|---|
Population Distance Metrics
Lecturer: Benjamin Grewe
Content of this lecture
Plasticity in Neuronal Networks
Substrates of Neural Plasticity
The Hippocampus as a Model System to Study Neural Plasticity
ETH/ETH - Introduction to Neuroscience/Extracted Topics - ETH Introduction to Neuroscience/Learning in Artificial & Biological Neural Networks/The Perceptron
Lecturer: Matthew Cook
When we look inside the brain, we see a bunch of neurons and a mess of connectivities. It's really hard to identify the connections among neurons, and with this information to understand how the brain works. We don't learn how the brain works by studying neurons, the same way that just by studying transistors we do not know how a computer works. We know the brain does processing but we don't know how it works. The bottleneck to understand brain is probably that we do not have the right abstractions to understand it. McCulloch and Pitts developed a computational model of a biological neuron in 1943. The McCulloch-Pitts neural model is also known as linear threshold unit/gate. It models a neuron with a set of inputs and one output.
McCulloch-Pitts Neurons vs Biological Neurons
(Basic) Digital Logic Gates are processing units. They are functions that evaluate the inputs. Each input has a 0/1 value that can be seen as a false/true in digital logic. There are different types of gates, each type is represented by a shape, but we can also just use their names to refer to them. We use gates to help us to understand the computation that might be happening in the brain. Neurons and axons do not behave as wires, but this is the tool we have available. And while neurons do not have 0s and 1s, they can be active/inactive, what gives us a good approximation. Circuits are a combination of gates.
Linear Threshold Unit/Gates (Perceptron) The picture represents a neuron with inputs x's and one output y. Weights w determine the influence of the inputs. f is a function determining the output: if the influence of all the inputs combined crosses a threshold, then the neuron becomes active. Active state: . Otherwise, the neuron is inactive.
We add a bias input as so that activates the neuron. As a digital abstraction, we consider that the activity of a neuron can be described as 0 or 1. Where 0 is inactive and 1 is active. Neurons can learn by adjusting their weights, they can have thousand of inputs. For a desired behavior there are many possible weight vectors that can work (if any can work).
This model can create AND/OR/NOT-gates. The function that a McCulloch-Pitts neuron can represent are only the linearly separable functions. Hence, this model is not capable of computing XOR and Equality Gates. Indeed, XOR is not a linear combination of the inputs. And the perceptron uses a line to separate classes, i.e., all the points in the same side of the line will have the same output. Although these units can't model the XOR, they are still powerful. You can't calculate it with a single unit, but with a combination of them, it is possible. The same way one cannot compute the XOR function using one single OR, NOT or AND gate. In fact, we need three gates to compute the XOR function. In the end, these model of neuron units are more efficient than digital logic gates, however, these new analog units run into precision problems. For implementing the XOR function, McCulloch and Pitts units allow exponentially smaller circuits.
We can find the weights that make the unit produce certain desired outputs through the Perceptron Learning Algorithm.
Perceptron Learning Algorithm This algorithm is called a learning algorithm because we use it to the define the weights of the neuron. This learning is just a definition of parameters, and can be seen as an optimization, where we have some guess and we want to improve it.
Supervised Error-Correcting Rules
We start with an initial guess of weights, we compare the output in response to the input with the desired output and then we change the weights to improve the performance. Consider the mapping of the table shown before. Can we set w0, w1, w2 (bias term, and the weights of each input) so that this unit computes the f from the table? Learning happens by changing synaptic weights. How can we change the synaptic weights of a unit to make it behave as desired? Let's start with random weights, let's say all zero. With these weights, doesn't matter the values of x1 and x2, the result will always be 1. And for the second case ( we produce a wrong output. So, we reduce and because they contribute to the sum (. We reduce the weights if the sum should go down and we increase the weights if the sum should go up. We iterate this step until convergence. We can consider the bias as a weight with input value always one (, thus, we can write:
Convergence These single units can separate two classes if they are linearly separable (i.e., if a solution exists). The Perceptron Learning Algorithm must converge, i.e., it will update the weights a finite number of times.
Class Discussion: If there is some weight vector that implements the desired (partial) function, then the Perceptron learning Algorithm (PLA) will terminate, with the unit correctly implementing the (partial) function. Why? Because of two facts:
Proof 1 Suppose there is a solution , i.e., the data is linearly separable. The weights of the perceptron units can be seen as the components of a normal vector to the hyperplane that separates the classes (1 and 0). In fact, the length of the normal vector doesn't matter, we are looking for its direction. Pick any solution , for instance, starting with all weights as zero, if this solution already satisfies our conditions, we are done. Otherwise, we pick an arbitrary misclassified point and update . Each step makes progress in the direction, because additions to are always The magnitude of increases linearly. Since doesn't change, there is a maximum growth from to achieve the solution. By contradiction in the limit of infinite steps, we can say that the algorithm converges, i.e., if there is a solution and it takes infinite steps to achieve it, this is a contradiction.
Algorithm
Lecturer: Matthew Cook
The Hopfield Network, or Hopfield Model, was proposed by John Hopfield in the early 80s. He was trying to understand what neural networks do. When we look into the brain and how neurons are connected to each other, we do not see always a clear pathway, i.e., it is hard to think about the connections in the brain in terms of input-output. He thought about connecting units to each other in an all-to-all pattern and see what would happen.
Characteristics of Hopfield Networks
Hopfield and Memory We are used to think about computation as a process that receives an input, does something and generates an output. However, this is not what we "see" in the memory process, for instance. It seems our memory works with Pattern Completion, also known as Content Addressable Memory or Associative Memory. This memory has no input-output relation: given any piece of it, we can recover the rest. An example is when certain smell can bring us a memory, even though the smell is only part of the memory. Hopfield Networks can be used to give us insights about how the memory works by having a highly connected network that computes with no input-output, but with states. The Hopfield Network is dynamic and moves from one state to another until it arrives to a stable state (also known as an attractor). Having stable states help the network to recover information giving partial inputs.
In the figure: Model of a Hopfield Network. Red units are inactive, green units are active, blue unit is the bias node that is always active.
Updates and State Dynamics Hopfield found that these networks always converge to a stable state. Idea: Consider the sum of weights between active units (Q). The update rule is equivalent to always increase Q i.e., we turn an unit active if , where is the value of the state of each unit, is the weight between the units and is the threshold, thus maximizing Q (weight between active units) by updating the output of one unit at time. In this setup, the update algorithm of Hopfield Networks can be seen as a greedy algorithm to find the MAX-CLIQUE. We update all the units but the bias node (bias node are always active and do not change their state). Remember, inactive neurons with zero threshold don't send inhibitory signal, instead they do not take part in the activation of other neurons.
Consider the network presented in the above figure: states are in the table on the right. In the beginning A and D are active, the sum of the active weights (Q) is 1. Now, we look at B and we sum the weights of the active units linked to B (D and A). Q is (). If we put the unit on the top part (active unit), otherwise we put the unit on the bottom (inactive unit). If B becomes active, Q increases (. If B becomes inactive, Q does not change (sum = 1). After updating C, the network achieves a stable state, i.e., updating any unit does not change its output.
When we update a unit and change its value (active or inactive), then Q increases or stays the same (if we are making units active). When we make units inactive, Q doesn't change. We can also consider active units as having a value of "+1" and inactive units as "-1". In this new representation, the Hopfield Network dynamic is equivalent to a graph min-cut, i.e., we want to minimize the sum of weights that link active and inactive units.
Asynchronous Updates When we update one unit at time, we are using an asynchronous method. Asynchronous Hopfield Networks always converge to a stable state, regardless of the update order, however, to which stable state is dependent on the update order.
Considering a Hopfield Network with two units connected with weight = 2, threshold zero and a bias unit of -1. Starting in a state with 1 active and 2 inactive, we can see how the order of update can lead to different stable states. By updating first unit 1 then unit 2, the stable state is with both units inactive. However, updating unit 2 then unit, we arrive to a stable state where both units are active.
Synchronous Updates Let's now consider a synchronous case, i.e., when we update all units at the same time: for this, active units for time are computed based on active units at time t. In this case, the system becomes a deterministic system because it doesn't depend on the order of updates anymore. Consider the network in the figure to the right, starting with units 1 and 2 active, in the next step of update (after all units), unit 3 and 4 should be active. It may be confusing to see this by updating one unit per time, let's say we update units in the order 1,2,3,4: after updating unit 3 it will become active but, when updating unit 4, we shouldn't consider unit 3, yet.
Figure: Starting with 1 and 2 actives, in the next step 3 and 4 will be active. This network doesn't have a unique stable state but converges to a cycle between two states.
Trick for analysis: make a larger asynchronous network based on the network we want to analyze. Duplicate the units in two columns and only use non-zero weights between them.
Figure: Starting with 1 and 2 actives, we will end up with 3 and 4 actives in the second column. First column represents t = 0, second t + 1. With synchronous updates, a Hopfield Network converges to a cycle of length 2 or to a stable state.
Summary
Lecturer: Matthew Cook
Feed-Forward Networks (FFNs) are not like the networks in the brain. In the feed-forward networks, the information moves in only one direction (forward) from the input nodes, through the hidden nodes (if they exist), and to the output nodes. There are no cycles in this network. Usually, people are referring to feed-forward networks when they talk about Artificial Neural Networks (ANNs). General structure:
A single unit, like a perceptron, can be seen as a feed-forward network. We can write down the connections of a FNN as a matrix of weights, so is the weight from i to j. Why FNN are nice? Because we can think about functions that receive inputs and generate outputs. When we use FNN we know what we want to compute. We need to set the weights of the network in order to compute the function we want. The process of defining the weights is called learning or training. In feed-forward networks it is easy to evaluate each unit. The outputs are continuous functions of the input, which facilitates the optimization in case of wrong outputs. The training can be done using "training data": input/output pairs . Where is the input value and is the desired output. Then, we can define the error , where is the output of the network. Differently of Hopfield Networks, we don't need continuous updates and we do not reevaluate units. FNN have the idea of a pipeline (unlike the brain). If a node on layer n in a FNN is connected to layer n + i with i > 1, this is still a FNN, however the most common structure is to connect nodes on layer n to nodes on layer n + 1. How can we change the weights to reduce the error? We can use gradient descent. We calculate all the in the network, easily, by starting at the end and then walking backwards. This is known as Backpropagation. Training Process
Backpropagation and Error Function Backpropagation is the process of calculating the derivatives, using the chain rule, from the last layer (directly connected with the output, thus, with the loss function) to the first layer (connected with the inputs). This process can be seen as walking through the network in a backward manner.
Gradient Descent Consider , we want to adjust (weights of the network) to minimize E. Gradient descent is the process of descending through the gradients (using the derivatives calculated with backpropagation), in this algorithm we try to reach the minimum of the loss function. This is an iterative process.
Generally, the iterative process is given by , where .
In the above picture: the left-most figure: is not a good threshold function. To know in what direction we should move to find out minima, we need to use a threshold function that is continuous and differentiable, like the one in the center figure. Right figure: .
We haven't yet found biological mechanisms that would be similar to gradient descent in the brain.
Boltzmann Machines Boltzmann Machines were invented in 1985 but not by Boltzmann. The name is given because these units use a Boltzmann distribution in their sampling function. These units are similar to Hopfield Networks, however, they have a probability of being active. When updating a unit, we set its value to zero or one probabilistically, following a sampling function (see figure below). Boltzmann machines do not converge, they do not reach a stable state. It can be seen as a system for sampling. It's a way to do a random walking in the state space.
Sampling There are many types of distributions. When we want to get examples of these distributions, we need to sample from it. And giving some samples we can recover the distribution of the data.
Figure: Representations of activation function for Hopfield Networks (HN) and Boltzmann Machines.
We want the units to forget previous states so the sampling is not biased, i.e., not similar to previous ones, thus really "random". While this sampling is nice, it is not useful as a memory, i.e., we don't want to sample things randomly from our memory.
We can define a network where the units have a real-valued activity level , and also we can make time continuous, so . Using units like this, we can make a feed-forward network.
Lecturer: Matthew Cook
We know how to think about units and their computation, but we don't exactly know how to think about how the information is processed. To try to understand it, we use engineering tools, but we don't know how the brain is doing it.
Neural Encoding of Information How the neurons encoding/process information? We would like to understand it in a mechanistic level. To understand the brain, we need to understand its structure and functionality (processes). Different areas in the brain are highly connected, practically from any area you can reach, virtually, any other area. Humans are capable to learn things we were not evolved for, for instance, to fly a drone. This is an inspiration for us. It seems to exist a general solution for solving problems: the brain. Neurons respond to combinations of properties, features, aspects of the situation, etc. Neurons are tuned to a set of values for the parameters they care about. For example, a neuron can be "tuned" to a moving bar at a certain angle in its receptive field. It also responds to a certain velocity, position (x,y), bar width, etc. Neurons response is called firing rates. Although neurons are tuned, they aren't super picky about the exact values. And, experimentally, neurons do not code a single attribute but a combination of them.
Population Code Information is encoded by a group (population) of neurons. A group includes all neurons in that area. The values are encoded by the pattern of activity. This code of information by population is our understanding of how neurons represent information, but this is not how we do it, typically, in an engineering process. Given a parameter, and looking at the neurons tuned to some value of this parameter, we can define tuning curves. In the figure (a) below, we see in the x-axis the parameter we are evaluating, and in the y-axis, the activity of one single neuron (firing response). Doing this process to many neurons, we can sort them regarding to the parameter we are looking for. In figure (b) all neurons were ordered accordingly to the response to a specific parameters. The read out of information given a population code is an easy way to figure out what the population is doing, it is robust to noise, however, require a lot of units.
Consider now an example were we want to correlate the angle of the eye (E), the position of an object on retina (R) and the head-centered direction (A), as showed in figure (a) below. This is an example of processing that is not feed-forward. How can a relation like A = E + R be represented? Say R, A and E are encoded by population codes, i.e., by population of units, each tuned to a particular value of that variable. Considering these three variables, we can use units that are tuned to combinations of A, E and R. The set of (A, E, R) triple that satisfy the relation will be active. This can be seen in Figure (b): when certain neuron fires for a parameter in R, another one can fire for a parameter in E and this can result in the firing of neurons in E, after learning. Given any two information in this relation, the third one can be recovered. There is no obligatory order. R, A and E have the following relations:
Summary
Lecturer: Giacomo Indiveri
Content of the Lecture
VLSI
The computer hardware had a radical paradigm shift when we looked to real brains. For instance, a bee brain is much smaller and consumes much less power (using neurons in a slow way), and offer real time interaction with the environment and complex behavior.
The Term "Neuromorphic" The term neuromorphic was coined by Carver Mead in the late '80s to describe VLSI systems containing electronic analog/digital circuits that exploit the physics of silicon to reproduce the bio-physics of neural circuits present in the nervous system. It is a discipline characterized by two main goals.
New hardware different from conventional computers: radically different from Von Neumann architectures. Now, there are parallel elements with memory and computation co-localized, with continuous streaming data driven computation, no clock. The co-localization of memory and computation allows to have no I/O bottleneck and no memory bottleneck.
The INI Neuromorphic Engineering Mission Learn to build artificial neural processing systems that can interact intelligently with the physical world.
Neuromorphic Computing vs Engineering Neuromorphic computing uses a dedicated VLSI hardware, high performance computing, it is application driven and uses conservative approaches. Neuromorphic engineering is a fundamental research, deeply rooted in biology. It emulates neural function in subthreshold analog and asynchronous digital.
Neuromorphic Electronic Circuits
Circuits Digital transistors operating only in the minimum and the maximum. Analog transistors use also intermediate amounts, thus transistors can emulate physical proteic channels. In biology, at high voltages, the fraction of the channels that are open approaches unity, causing a saturation. The same can be seen in a subthreshold regime. In subthreshold, the current is smaller than 1V, it increases exponentially, and after threshold currents change quadratically. It changes from pico to nano amps.
![]() |
![]() |
|---|
In Complementary Metal-Oxide Semiconductor (CMOS) technology, there are two types of MOS-FETs: n-FETs and p-FETs. There is no current going to transistors. In traditional CMOS circuits, all n-FETs have the common bulk potential connected to ground (GND) and all p-FETs have a common bulk potential connected to the power supply rail .
![]() |
![]() |
|---|
![]() |
![]() |
|---|
The current is defined to be positive if it flows from the drain to the source.
Diffusion and Saturation
![]() |
![]() |
|---|
Output Current versus and and Current Source
![]() |
![]() |
|---|
The Current Mirror
The output current is a mirrored copy of the input current. If both MOSFETs are of the same size and have the same source voltage, they source the same current, which is why the device is called current mirror. The input current through the diode-connected transistor sets the common gate voltage and hence the output current of the second transistor .
The output current can be scaled by choosing different transistor sizes, or by choosing different source potentials and for the two MOSFETs. If is in saturation:
The Differential - Pair
![]() |
![]() |
|---|
Note that in equation at the numerator is a mistake, should instead be .
To implement the difference of currents ( - ) we can use the current-mirror circuit.
The Transconductance Amplifier For small differential voltages (e.g., ), the tanh() relationship is approximately linear and the equation can be reduced to: where .
![]() |
![]() |
|---|
Spike Generating Mechanism If the membrane voltage increases above a certain threshold, a spike-generating mechanism is activated and an action potential is initiated.
![]() |
![]() |
![]() |
|---|
A Conductance-Based Silicon Neuron In 1991 Misha Mahowald and Rodney Douglas proposed a conductance-based silicon neuron and showed that it had properties remarkably similar to those of real cortical neurons.
Neuron Models Traditionally there have been two main classes of neuron models:
![]() |
![]() |
![]() |
|---|
Current-mode CMOS circuits operated in the subthreshold, or weak-inversion regime can be used to implement log-domain filters. An example of classical log-domain integrator is presented in the picture below. This circuit's linear transfer function can be easily derived by applying the translinear principle on the loop highlighted by the arrows: given the exponential relationship between the subthreshold currents of the p-FETs and their voltages, we can write: . In subthreshold, the output n-FET produces a current that changes exponentially with its gate voltage . Differentiating with respect to and combining the result with the capacitor equation we obtain: .
The DPI is a CMOS current-mode circuit that operates in the subthreshold regime integrating voltage pulses. However, rather than using a single p-FET to generate the appropriate current, via the triangular principle (Gilbert, 1975), it uses a differential pair in negative feedback configuration. This allows the circuit to achieve LPF functionality with tunable dynamic conductances: Input voltage pulses are integrated to produce an output current that has maximum amplitude set by , and . (Silicon neuron circuits) It has additional advantages of providing a compact layout, better matching properties and lower power consumption. The differential - pair integrator is used to model synaptic dynamics. It comprises only 3 n-FETs, 2 p-FETs and 1 capacitor. The two current sources are implemented using two subthreshold MOSFETs: one n-FET for the current and on p-FET for the current. Following a similar derivation to the one used in the classical log-domain integrator, the characteristic equation is obtained as observed in the picture below.
Additional circuits can be attached to the DPI synapse to extend the model with extra features typical of biological synapses and implement various types of plasticity. For example, by adding two extra transistors, we can implement voltage-gated channels that model NMDA synapse behavior. Similarly, by using two more transistors, we can extend the synaptic model to be conductance based. Furthermore, the DPI circuit is compatible with previously proposed circuits for implementing synaptic plasticity, on both short timescales with models of short-term depression (STD) and on larger timescales with spike-based learning mechanisms, such as spike timing-dependent plasticity (STDP). The DPI neuron circuit is a variant of the generalized IF neuron and is depicted in the following picture. The input DPI low-pass filter (yellow, ML1 - ML3) models the neuron's leak conductance. A spike event generation amplifier (red, MA1 - MA6) implements current-based positive feedback (modeling both sodium activation and inactivation conductances) and produces address-events at extremely low-power. The reset block (blue, MR1 - MR6) resets the neuron and keeps it in a reset state for a refractory period, set by the bias voltage. An additional DPI filter integrates the spikes and produces a slow after hyper-polarizing current responsible for spike-frequency adaptation (green, MG1 - MG6). By applying a current-mode analysis to both the input and the spike-frequency adaptation DPI circuits, it is possible to derive a simplified analytical solution:
The state of the art version of this neuron circuit consumes one order of magnitude less power than the circuit described in the following figure and two orders of magnitude less power than the digital implementation of the I&F neuron. Given the exponential nature of the generalized IF neuro's non-linear term , the DPI-neuron implements an adaptive exponential IF model. This IF model has been shown to be able to reproduce a wide range of spiking behaviors, and explain a wide set of experimental measurements from pyramidal neurons.
Neuromorphic Processors Typical spiking neural network chips have the elements described in the figure below. Multiple instances of these elements can be integrated onto single chips and connected among each other either with on-chip hard-wired connections or via off-chip reconfigurable connectivity infrastructures. The most relevant characteristics of processors build based on analog circuits working in subthreshold are:
Neuromorphic vs Conventional Processors
Winner - Take - All Networks in Neuromorphic Hardware "Winner take all" (WTA) refers to a type of neural network architecture that is commonly used in neuromorphic hardware. In a WTA network, each neuron competes with other neurons to be the "winner" and output the highest value. This can be useful in situations where you want to identify the most active or strongest signal among a group of neurons. This architecture is inspired by the way the brain works where many neurons compete to fire.
Learning and Winner - Take - All Networks Memories can be formed in neuromorphic hardware using attractor networks, which are a type of recurrent neural network. An attractor network is composed of a group of neurons that are connected to each other through synapses. These synapses can be either excitatory, which increases the likelihood that a neuron will fire, or inhibitory, which decreases the likelihood that a neuron will fire. The network's dynamics are determined by the strengths of these connections, which can be modified through a process called synaptic plasticity. When the network is exposed to a specific input pattern, the neurons that are active will strengthen their connections to other active neurons, and inhibitory connections will form between active neurons and inactive neurons.This process creates a stable state, or an attractor, in the network. The attractor corresponds to the input pattern that was presented to the network, and the network will continue to settle to this attractor state even after the input pattern is removed. This behavior allows the network to "remember" the input pattern, and this process is called memory formation. The network can then be used to recall the stored information by providing a partial or noisy version of the input pattern.
Mechanisms operating at the network level can allow neural processing systems to form short-term memories, consolidate long-term ones, and carry out nonlinear processing functions such as selective amplification (e.g., to implement attention and decision making). An example of such a network-level mechanism is provided by "attractor networks". These are networks of neurons that are recurrently connected via excitatory synapses, and that can settle into stable patterns of firing even after the external stimulus is removed. Different stimuli can elicit different stable patterns, which consist of specific subsets of neurons firing at high rates. Each of the high-firing rate attractor states can represent a different memory. To make an analogy with conventional logic structures, a small attractor network with two stable states would be equivalent to a flip-flop gate in CMOS.
A particularly interesting class of attractor networks is the one of soft winner-take-all (sWTA) neural networks. In these networks, groups of neurons both cooperate and compete with each other. Cooperation takes place between groups of neurons spatially close to each other, while competition is typically achieved through global recurrent patterns of inhibitory connections. When stimulated by external inputs, the neurons excite their neighbors and the ones with the highest response suppress all other neurons to win the competition. Thanks to these competition and cooperation mechanisms, the outputs of individual neurons depend on the activity of the whole network and not just on their individual inputs.
Conclusion